* chore(#2799): resolve reviewer lanes into executable invocation plans Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers iterate declared lanes instead of hand-authored per-CLI bash. Five additive descriptor amendments, each forced by a lane that ships today: - LaneHandler gains 'opencode' — the lane rebuilds its review from assistant text parts of a --format json stream; a plain stdout copy re-breaks #1936. - modelConfigKey — antigravity's key is review.models.agy, not .antigravity, so resolving by slug silently dropped a configured model. - defaultHost/fallbackModel — Phase 4 federated every *_host with a default of empty string; the real fallback only existed in the bash. - args becomes an argv template with a closed four-placeholder vocabulary. Positional splicing produced 'codex --model M -o F exec --ephemeral', which is not a valid invocation: codex injects in the middle, twice. - kimi-code lane, with the bounded command-capability probe (needle --output-format) that tells Kimi Code from the legacy python kimi-cli. Parity gate re-pointed: the workflow-text families it scanned are the text this phase deletes, so they are replaced by descriptor-to-registry parity plus an anti-parity check that no bespoke leg returns. jq, curl and external timeout/gtimeout all drop out of the review path. Refs #2782 * chore(#2799): add review-lane query surface and widen the manifest vocabulary Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow loops over, projects all twelve lanes into their capability manifests, and widens capability-validator for the amendments. opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's own admission rule: one lane, justified by a documented upstream defect data cannot express (#1936 — the agent can end its turn with zero output tokens and --format default then drops the assistant text entirely). Two bugs caught by an end-to-end stub run and fixed here: - loadConfigResolved returns a provenance wrapper, not the config; using it directly resolved every key to undefined, which reads as 'nothing configured' and silently dropped every model override. - hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced with a PATH scan that spawns nothing at all. Refs #2782 * chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved lanes, and renders REVIEWS.md sections from each lane's declared reviewsSection instead of thirteen hardcoded headings. review.md drops from 1104 lines to 507 (61KB to 28.7KB). Parity gate re-pointed, as agreed: the leg-marker and section-heading families scanned exactly the text this phase deletes, so they are replaced by descriptor-to-registry parity in both directions, plus an anti-parity check that fires if a bespoke leg is ever re-added. Enum, emitting sites and the Object.keys lock moved together. The budget-trim helper is hoisted out of the Ollama leg: it was always lane-agnostic, and any lane may now declare a promptBudgetKey. Refs #2782 * feat(#2799): bind the consented egress host and re-verify it at invocation Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3 but was not implemented: ConsentRecord had no host field and nothing in the tree bound one, so this phase's rule-4 comparison had no baseline. ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design: isValidConsentRecord does not require it, so every record already on disk stays valid and no re-consent storm fires (D4 rule 5). It is deliberately excluded from disclosureSignature — the loader has no config resolver, so folding a config-derived value in would make loader and lifecycle compute different signatures for the same manifest and re-prompt forever. Install resolves hostConfigKey (falling back to the lane's declared defaultHost, which is what the invocation path uses) and records it. Invocation re-resolves and blocks on mismatch rather than silently redirecting. Absence allows: no record, or a record predating the field, means nothing to compare — denying there would break every existing local-model user on upgrade. Refs #2782 * test(#2799): cover the resolver, runner and handlers; retarget the parity suites Adds the golden invocation-plan table (one row per shipped lane, derived from the bash legs rather than the descriptor types) plus runner coverage for the probe, empty-output policy, the three handlers and the egress check. Retargets the existing suites onto the new contract: descriptor-to-registry parity, the anti-parity check, the opencode handler, and the twelfth lane. Two corrections found by running them: - modelConfigKey was required; that breaks D4 rule 2, since a reviewer manifest authored before this phase would fail validation on upgrade. It is optional, read as null when absent. - the antigravity non-zero-exit test pre-seeded the transcript, which asserted that a STALE entry leaks through — the exact bug the watermark prevents. The spawn now appends, as the real tool does. Refs #2782 * fix(#2799): restore agy --add-dir and the self-report prompt in the handler Retargeting the three legacy reviewer suites off the deleted bash surfaced two real regressions in the port, both #2176: - --add-dir was dropped. Without it agy's permission context never receives the cwd repo, so the agent anchors on its own scratch dir and reviews the plan text in isolation — the exact failure the Review Instructions forbid. It is capability-probed, because an older agy rejects the unknown flag outright and a lane that fails to start is worse than one running on the prompt anchor. - the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS self-report, which is what makes a blind review distinguishable from a grounded one. antigravity now builds its own prompt variant. Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own log is the only evidence that anything failed. The three suites now assert against the plan and the handler instead of matching fence text, so they no longer need allow-test-rule exemptions. Refs #2782 * docs(#2799): document the declared lanes, the new flag, and dropped prerequisites COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph, which is now false: no lane requires jq, curl or an external timeout. Adds the changed-egress-destination behavior, since a blocked lane is something a user can hit. CONFIGURATION.md records that the model config key is declared per lane rather than derived from the flag — antigravity's is review.models.agy — and adds review.models.kimi-code. reviewer-instances.md now routes an instance through its lane's single invocation seam instead of a copied per-adapter bash block, which is what lets a cross-cutting fix reach instances for free. That required implementing the --model/--agent/--as flags it documents; --model re-resolves through the lane's argv template rather than splicing, so the flag lands where the lane declares it rather than ahead of a subcommand. CONTEXT.md glossary gains both new modules. Refs #2782 * chore(#2799): drop the stale emitted-drift acknowledgment The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md. That file now shrinks by ~32KB and every emitted hash that moved is attributable to this diff, so the ack no longer explains anything. Removing the last entry means removing the file: its presence is the alarm, and an empty one signals nothing. Verified by deleting it and re-running the attribution and provenance gates plus lint:ci — all green without it. Refs #2782 * docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782 Five additive amendments, each forced by a lane that ships today, plus two corrections the phase had to make rather than work around: - D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so this phase's rule-4 comparison had no baseline. Recorded because an ADR asserting a rule was delivered is exactly what stops a later phase checking. - The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families scanned the text this phase deletes. Also records that D7's 'skip the probe where no bounding mechanism exists' carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded on every stock macOS host, which ships neither timeout nor gtimeout. Refs #2782 * fix(#2799): close four defects found by adversarial review Two confirmed bugs, both reproduced before fixing: - resolveLanePlan was not total. An openai-http lane with a missing or non-object invoke dereferenced inv.hostConfigKey and threw, contradicting the module's own documented contract; the spawn branch guarded correctly and the http branch did not. The CLI seam resolves every selected lane in one map, so one malformed overlay manifest would have aborted the whole review rather than dropping its own lane. Guarded, plus a per-lane try/catch at the seam so a throw can never take down siblings. - A reviewer-instance model was silently dropped for any lane declaring modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates that cli is a known slug but never that the slug accepts a model, so a user could configure one, get a clean run, and never learn a different model reviewed their plan. Now warns explicitly. Two hardening fixes: - The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in the resolver rather than inherited from a validator that does not run on this path — the module documents itself as the overlay-manifest trust boundary, so it should not depend on someone else having checked. - normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses with an empty hostname, so it became 'localhost://11434' and was compared and requested as if real. An empty hostname now means not-a-URL. Also documents the one gap that cannot be closed here: the antigravity watermark is keyed by workspace, so two concurrent reviews of the same repo share a transcript. agy exposes no per-invocation id to filter on, so the handler now states which half of its never-stale guarantee actually holds. Refs #2782 * test(#2799): retarget the remaining eight review.md-asserting suites The remote runner found 37 failures the local sweep missed (it hit the shell's two-minute cap before reaching these). All eight extract per-CLI bash from review.md that this phase deletes; each protects a real invariant, so each is retargeted onto the plan, the runner or the handler rather than removed. Three real defects surfaced by doing so: - effort args never reached ANY lane. model-resolver.cjs exports no resolveExecution, so effortFor silently returned [] every time. Restored by calling the same bounded resolve-execution query the bash legs used — and NOT with --raw, which prints the resolved effort rather than the picked field, so claude got 'low' instead of '--effort low'. - the timeout guidance lost 'a silent empty output is a timeout kill, not a crash' — the operator note that exists because of the Codex 0xc0000142 misdiagnosis. Restored. - the opencode handler dropped EMPTY assistant text parts. The shipped jq was , and only substitutes for false/null — an empty string is truthy in jq and contributed a blank line. Found by a property test shrinking to ['', '']. The opencode property suite no longer spawns jq at all, which deletes the #2099 hang mechanism it was architected around rather than mitigating it. Refs #2782 * fix(#2799): register the two new generated modules, and untrack them The remote runner caught build output committed to git. Both new modules compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each bin/lib/*.cjs is linted xor ignored according to migration state" failed. Registered both in .gitignore and eslint.config.mjs alongside the Phase 1 module, and dropped them from the index. Nothing about the shipped behaviour changes; the artifacts are rebuilt by build:lib. This is the new-.cts-module registration ripple, and it is the one part of it I had not completed - the CONTEXT.md glossary and the inventory manifest were already done. Refs #2782 * chore(#2799): backfill changeset pr number to 2861 * chore(#2799): backfill changeset pr number to 2861 --------- Co-authored-by: Test <test@example.com>
This commit is contained in:
5
.changeset/daring-foxes-zip.md
Normal file
5
.changeset/daring-foxes-zip.md
Normal file
@@ -0,0 +1,5 @@
|
||||
---
|
||||
type: Added
|
||||
pr: 2861
|
||||
---
|
||||
**`/gsd:review --kimi-code` reviews your plans with Kimi Code CLI** — the new lane joins the cross-AI reviewer roster and is included by `--all` when detected. Detection distinguishes Kimi Code from the legacy Python kimi-cli, which ships a binary of the same name, so a host with only the legacy tool reports the lane unavailable instead of registering a reviewer that cannot serve it. (#2718)
|
||||
5
.changeset/graceful-wolves-wander.md
Normal file
5
.changeset/graceful-wolves-wander.md
Normal file
@@ -0,0 +1,5 @@
|
||||
---
|
||||
type: Changed
|
||||
pr: 2861
|
||||
---
|
||||
**Cross-AI reviewer lanes are now declared data rather than hand-written per-CLI blocks** — every lane's binary, prompt and output channel, timeout, probe and empty-output policy comes from its capability manifest, so a reviewer can be shipped as a plugin instead of a core patch. Two user-visible consequences: a reviewer that returns only whitespace is now reported as a failed lane on every reviewer (previously only on LM Studio and llama.cpp, so elsewhere a blank reply was rendered as a clean review), and an OpenAI-compatible lane whose configured host has changed since you consented to it is blocked with an explanation rather than silently sending your plans to the new destination. `jq`, `curl` and GNU `timeout` are no longer required on PATH for any lane. (#2782)
|
||||
2
.gitignore
vendored
2
.gitignore
vendored
@@ -115,6 +115,8 @@ build/
|
||||
/gsd-core/bin/lib/ui-safety-gate.cjs
|
||||
/gsd-core/bin/lib/review-reviewer-selection.cjs
|
||||
/gsd-core/bin/lib/review-lane-descriptor.cjs
|
||||
/gsd-core/bin/lib/review-lane-invocation.cjs
|
||||
/gsd-core/bin/lib/review-lane-runner.cjs
|
||||
/gsd-core/bin/lib/clusters.cjs
|
||||
/gsd-core/bin/lib/installer-migrations/001-legacy-orphan-files.cjs
|
||||
/gsd-core/bin/lib/observability/redaction.cjs
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -116,7 +116,9 @@
|
||||
"args": [
|
||||
"--print-timeout",
|
||||
"540s",
|
||||
"-p"
|
||||
"{{model}}",
|
||||
"-p",
|
||||
"{{prompt}}"
|
||||
],
|
||||
"promptChannel": "argv-file-ref",
|
||||
"outputChannel": "stdout",
|
||||
@@ -127,10 +129,9 @@
|
||||
"emptyOutput": "handler-owned",
|
||||
"reviewsSection": "Antigravity",
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [
|
||||
"jq"
|
||||
],
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": "review.models.agy",
|
||||
"handler": "antigravity"
|
||||
},
|
||||
"config": {
|
||||
|
||||
@@ -118,6 +118,8 @@
|
||||
"invoke": {
|
||||
"binary": "claude",
|
||||
"args": [
|
||||
"{{model}}",
|
||||
"{{effort}}",
|
||||
"-p",
|
||||
"-"
|
||||
],
|
||||
@@ -132,6 +134,7 @@
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": "review.models.claude",
|
||||
"handler": null
|
||||
},
|
||||
"config": {
|
||||
|
||||
@@ -11,12 +11,20 @@
|
||||
},
|
||||
"reviewer": {
|
||||
"slug": "coderabbit",
|
||||
"flags": ["--coderabbit"],
|
||||
"flags": [
|
||||
"--coderabbit"
|
||||
],
|
||||
"transport": "spawn",
|
||||
"probe": { "kind": "command-exists", "binary": "coderabbit" },
|
||||
"probe": {
|
||||
"kind": "command-exists",
|
||||
"binary": "coderabbit"
|
||||
},
|
||||
"invoke": {
|
||||
"binary": "coderabbit",
|
||||
"args": ["review", "--prompt-only"],
|
||||
"args": [
|
||||
"review",
|
||||
"--prompt-only"
|
||||
],
|
||||
"promptChannel": "none",
|
||||
"outputChannel": "stdout",
|
||||
"modelArg": null,
|
||||
@@ -28,6 +36,7 @@
|
||||
"evidenceClass": "diff-only",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": null,
|
||||
"handler": null
|
||||
}
|
||||
}
|
||||
|
||||
@@ -107,7 +107,10 @@
|
||||
"args": [
|
||||
"exec",
|
||||
"--ephemeral",
|
||||
"{{model}}",
|
||||
"{{effort}}",
|
||||
"--skip-git-repo-check",
|
||||
"{{output}}",
|
||||
"-"
|
||||
],
|
||||
"promptChannel": "stdin",
|
||||
@@ -122,6 +125,7 @@
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": "review.models.codex",
|
||||
"handler": null
|
||||
},
|
||||
"config": {
|
||||
|
||||
@@ -123,12 +123,25 @@
|
||||
},
|
||||
"reviewer": {
|
||||
"slug": "cursor",
|
||||
"flags": ["--cursor"],
|
||||
"flags": [
|
||||
"--cursor"
|
||||
],
|
||||
"transport": "spawn",
|
||||
"probe": { "kind": "command-exists", "binary": "cursor-agent" },
|
||||
"probe": {
|
||||
"kind": "command-exists",
|
||||
"binary": "cursor-agent"
|
||||
},
|
||||
"invoke": {
|
||||
"binary": "cursor-agent",
|
||||
"args": ["-p", "--mode", "ask", "--trust", "--output-format", "text"],
|
||||
"args": [
|
||||
"-p",
|
||||
"--mode",
|
||||
"ask",
|
||||
"--trust",
|
||||
"--output-format",
|
||||
"text",
|
||||
"{{prompt}}"
|
||||
],
|
||||
"promptChannel": "argv-file-ref",
|
||||
"outputChannel": "stdout",
|
||||
"modelArg": null,
|
||||
@@ -140,6 +153,7 @@
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": null,
|
||||
"handler": null
|
||||
}
|
||||
}
|
||||
|
||||
@@ -22,6 +22,7 @@
|
||||
"invoke": {
|
||||
"binary": "gemini",
|
||||
"args": [
|
||||
"{{model}}",
|
||||
"-p",
|
||||
"-"
|
||||
],
|
||||
@@ -36,6 +37,7 @@
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": "review.models.gemini",
|
||||
"handler": null
|
||||
},
|
||||
"config": {
|
||||
|
||||
@@ -88,5 +88,45 @@
|
||||
"skipSharedHooksInstall": true,
|
||||
"namedSubagentsSupported": false
|
||||
}
|
||||
},
|
||||
"reviewer": {
|
||||
"slug": "kimi-code",
|
||||
"flags": [
|
||||
"--kimi-code"
|
||||
],
|
||||
"transport": "spawn",
|
||||
"probe": {
|
||||
"kind": "command-capability",
|
||||
"binary": "kimi",
|
||||
"needle": "--output-format",
|
||||
"timeoutMs": 5000
|
||||
},
|
||||
"invoke": {
|
||||
"binary": "kimi",
|
||||
"args": [
|
||||
"{{model}}",
|
||||
"-p",
|
||||
"{{prompt}}"
|
||||
],
|
||||
"promptChannel": "argv-file-ref",
|
||||
"outputChannel": "stdout",
|
||||
"modelArg": "-m",
|
||||
"effortChannel": "none"
|
||||
},
|
||||
"timeoutFloorMs": 900000,
|
||||
"emptyOutput": "stub-with-stderr",
|
||||
"reviewsSection": "Kimi Code",
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": "review.models.kimi-code",
|
||||
"handler": null
|
||||
},
|
||||
"config": {
|
||||
"review.models.kimi-code": {
|
||||
"type": "string",
|
||||
"default": "",
|
||||
"description": "Model passed to the Kimi Code reviewer lane."
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -23,18 +23,19 @@
|
||||
},
|
||||
"invoke": {
|
||||
"hostConfigKey": "review.llama_cpp_host",
|
||||
"defaultHost": "http://localhost:8080",
|
||||
"path": "/v1/chat/completions",
|
||||
"modelDiscovery": "first-from-models-endpoint",
|
||||
"fallbackModel": "local-model",
|
||||
"effortChannel": "none"
|
||||
},
|
||||
"timeoutFloorMs": 120000,
|
||||
"emptyOutput": "stub-with-stderr",
|
||||
"reviewsSection": "llama.cpp",
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [
|
||||
"jq"
|
||||
],
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": "review.max_prompt_tokens_per_reviewer.llama_cpp",
|
||||
"modelConfigKey": "review.models.llama_cpp",
|
||||
"handler": "openai-compatible"
|
||||
},
|
||||
"config": {
|
||||
|
||||
@@ -23,18 +23,19 @@
|
||||
},
|
||||
"invoke": {
|
||||
"hostConfigKey": "review.lm_studio_host",
|
||||
"defaultHost": "http://localhost:1234",
|
||||
"path": "/v1/chat/completions",
|
||||
"modelDiscovery": "first-from-models-endpoint",
|
||||
"fallbackModel": "local-model",
|
||||
"effortChannel": "none"
|
||||
},
|
||||
"timeoutFloorMs": 120000,
|
||||
"emptyOutput": "stub-with-stderr",
|
||||
"reviewsSection": "LM Studio",
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [
|
||||
"jq"
|
||||
],
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": "review.max_prompt_tokens_per_reviewer.lm_studio",
|
||||
"modelConfigKey": "review.models.lm_studio",
|
||||
"handler": "openai-compatible"
|
||||
},
|
||||
"config": {
|
||||
|
||||
@@ -23,18 +23,19 @@
|
||||
},
|
||||
"invoke": {
|
||||
"hostConfigKey": "review.ollama_host",
|
||||
"defaultHost": "http://localhost:11434",
|
||||
"path": "/v1/chat/completions",
|
||||
"modelDiscovery": "first-from-models-endpoint",
|
||||
"fallbackModel": "llama3",
|
||||
"effortChannel": "none"
|
||||
},
|
||||
"timeoutFloorMs": 120000,
|
||||
"emptyOutput": "stub-with-stderr",
|
||||
"reviewsSection": "Ollama",
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [
|
||||
"jq"
|
||||
],
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": "review.max_prompt_tokens_per_reviewer.ollama",
|
||||
"modelConfigKey": "review.models.ollama",
|
||||
"handler": "openai-compatible"
|
||||
},
|
||||
"config": {
|
||||
|
||||
@@ -127,6 +127,8 @@
|
||||
"binary": "opencode",
|
||||
"args": [
|
||||
"run",
|
||||
"{{model}}",
|
||||
"{{effort}}",
|
||||
"--format",
|
||||
"json",
|
||||
"-"
|
||||
@@ -140,11 +142,10 @@
|
||||
"emptyOutput": "stub-with-stderr",
|
||||
"reviewsSection": "OpenCode",
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [
|
||||
"jq"
|
||||
],
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"handler": null
|
||||
"modelConfigKey": "review.models.opencode",
|
||||
"handler": "opencode"
|
||||
},
|
||||
"config": {
|
||||
"review.models.opencode": {
|
||||
|
||||
@@ -106,12 +106,19 @@
|
||||
},
|
||||
"reviewer": {
|
||||
"slug": "qwen",
|
||||
"flags": ["--qwen"],
|
||||
"flags": [
|
||||
"--qwen"
|
||||
],
|
||||
"transport": "spawn",
|
||||
"probe": { "kind": "command-exists", "binary": "qwen" },
|
||||
"probe": {
|
||||
"kind": "command-exists",
|
||||
"binary": "qwen"
|
||||
},
|
||||
"invoke": {
|
||||
"binary": "qwen",
|
||||
"args": ["-"],
|
||||
"args": [
|
||||
"-"
|
||||
],
|
||||
"promptChannel": "stdin",
|
||||
"outputChannel": "stdout",
|
||||
"modelArg": null,
|
||||
@@ -123,6 +130,7 @@
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": null,
|
||||
"handler": null
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1465,17 +1465,20 @@ Reviewers are prompted to verify the plan's claims against the actual repository
|
||||
| `--qwen` | Include Qwen Code review (Alibaba Qwen models) |
|
||||
| `--cursor` | Include Cursor agent review |
|
||||
| `--agy` / `--antigravity` | Include Antigravity CLI review (free with Google credentials) |
|
||||
| `--kimi-code` | Include Kimi Code CLI review (Moonshot AI) |
|
||||
| `--ollama` | Include Ollama server review |
|
||||
| `--lm-studio` | Include LM Studio server review |
|
||||
| `--llama-cpp` | Include llama.cpp server review |
|
||||
| `--all` | Include all available reviewers (CLI + local model servers) |
|
||||
|
||||
**`jq` prerequisite (some lanes only):** `--ollama`, `--lm-studio`, `--llama-cpp`, `--opencode`, and `--agy` parse JSON that GSD does not produce itself — OpenAI-compatible `/v1/chat/completions` responses, OpenCode's JSONL event stream, and Antigravity's conversation cache — so they require [`jq`](https://jqlang.org/download/) on your `PATH`. If `jq` is missing, `/gsd-review` reports those five as unavailable and tells you to install it, rather than running them into an empty review — an info note when the lane was reached through `--all` or `review.default_reviewers`, an error when you named it with an explicit flag. The other six lanes (`--gemini`, `--claude`, `--codex`, `--coderabbit`, `--qwen`, `--cursor`) need no `jq`. Reading your configured models, hosts, and token budgets never requires `jq`.
|
||||
**No `jq`, `curl`, or `timeout` prerequisite.** Reviewer lanes used to shell out to these for JSON parsing, HTTP calls, and wall-clock bounding, which made five lanes unavailable on a stock Windows/Git-Bash host (no `jq`) and left one lane unbounded on stock macOS (no `timeout` or `gtimeout`). GSD now does all three itself, so every lane runs with nothing on your `PATH` but the reviewer's own CLI. A lane that declares an external tool it genuinely needs still reports itself unavailable with an install hint rather than running into an empty review.
|
||||
|
||||
**Unavailable reviewers:** an explicit reviewer flag is an assertion. If you name a reviewer that cannot run on this host — its CLI is not installed, a prerequisite such as `jq` is missing, or its local server is unreachable — `/gsd-review` reports an **error** for that reviewer and does not proceed with a reduced set. This holds even when other named reviewers are available: `--gemini --qwen` on a host without `qwen` fails rather than silently becoming a Gemini-only review.
|
||||
**Unavailable reviewers:** an explicit reviewer flag is an assertion. If you name a reviewer that cannot run on this host — its CLI is not installed, a required external tool is missing, its local server is unreachable, or its egress destination changed (see below) — `/gsd-review` reports an **error** for that reviewer and does not proceed with a reduced set. This holds even when other named reviewers are available: `--gemini --qwen` on a host without `qwen` fails rather than silently becoming a Gemini-only review.
|
||||
|
||||
Reviewers reached through `--all` or `review.default_reviewers` behave differently: an undetected reviewer there is reported as an info note and skipped. Use `--all` for "whatever is available on this host", and `review.default_reviewers` for a preferred subset that may vary by host.
|
||||
|
||||
**Changed egress destination:** a reviewer lane is sent your plan text, requirements, research findings, and `CONTEXT.md` decisions. For the local-server lanes (`--ollama`, `--lm-studio`, `--llama-cpp`) the destination comes from a config key such as `review.ollama_host`, which is an ordinary editable value — including by a pull request. If you installed such a lane as a capability and its host has changed since you consented to it, GSD **blocks that lane and tells you both destinations** rather than sending your plans somewhere you did not approve. Re-consent to allow the new host. First-party lanes shipped with GSD are unaffected, and a lane you never consented to is not blocked — there is nothing to compare it against.
|
||||
|
||||
**Default reviewer behavior (no flags):**
|
||||
- If `review.default_reviewers` is **unset**, `/gsd-review` runs all detected reviewers (current default behavior).
|
||||
- If `review.default_reviewers` is **set**, `/gsd-review` runs only that subset (for example `["gemini","codex"]`).
|
||||
|
||||
@@ -228,7 +228,9 @@ API key fields accept a string value (the key itself). They can also be set to t
|
||||
|
||||
### Code-review CLI routing
|
||||
|
||||
`review.models.<cli>` maps a reviewer flavor to a bare model id. The code-review workflow injects this value into the CLI's `--model` (or `-m`) flag when invoking the reviewer.
|
||||
`review.models.<cli>` maps a reviewer flavor to a bare model id, which is injected into the CLI's own model flag (`--model`, `-m`, …) when the reviewer is invoked.
|
||||
|
||||
The key suffix is **not** always the lane slug. Each lane declares the config key it reads, and one shipped lane already differs: the Antigravity lane's slug is `antigravity` but its key is `review.models.agy`, after the CLI's own name. Consult the table below rather than deriving the key from the flag.
|
||||
|
||||
| Setting | Type | Default | Description |
|
||||
|---------|------|---------|-------------|
|
||||
@@ -236,6 +238,7 @@ API key fields accept a string value (the key itself). They can also be set to t
|
||||
| `review.models.codex` | string | `null` | Model id for Codex review (injected into --model), e.g. `"gpt-5"` |
|
||||
| `review.models.gemini` | string | `null` | Model id for Gemini review (injected into -m), e.g. `"gemini-2.5-pro"` |
|
||||
| `review.models.opencode` | string | `null` | Model id for OpenCode review (injected into --model), e.g. `"claude-sonnet-4"` |
|
||||
| `review.models.kimi-code` | string | `null` | Model id for Kimi Code review (injected into -m) |
|
||||
|
||||
**Ownership.** These keys are owned by their reviewer-lane capabilities rather than the central
|
||||
config schema — `review.models.ollama` belongs to the `ollama` capability, `review.ollama_host`
|
||||
@@ -1014,7 +1017,8 @@ Configure per-CLI model selection for `/gsd-review`. When set, overrides the CLI
|
||||
| `review.models.claude` | string | (CLI default) | Model used when `--claude` reviewer is invoked |
|
||||
| `review.models.codex` | string | (CLI default) | Model used when `--codex` reviewer is invoked |
|
||||
| `review.models.opencode` | string | (CLI default) | Model used when `--opencode` reviewer is invoked |
|
||||
| `review.models.agy` | string | (CLI default) | Model used when the `--antigravity` / `--agy` reviewer is invoked. The key suffix is the CLI's own name (`agy`), not the lane slug |
|
||||
| `review.models.agy` | string | (CLI default) | Model used when the `--antigravity` / `--agy` reviewer is invoked. The key suffix is the CLI's own name (`agy`), not the lane slug — the lane declares which key it reads, so the two need not match |
|
||||
| `review.models.kimi-code` | string | (CLI default) | Model used when the `--kimi-code` reviewer is invoked (injected into `-m`) |
|
||||
| `review.models.ollama` | string | (server default) | Model name passed to Ollama when `--ollama` reviewer is invoked. If unset, the first available model reported by the server is used (e.g. `llama3`). Set to a specific tag: `gsd config-set review.models.ollama codellama` |
|
||||
| `review.models.lm_studio` | string | (server default) | Model name passed to LM Studio when `--lm-studio` reviewer is invoked. If unset, the first available model reported by the server is used. |
|
||||
| `review.models.llama_cpp` | string | (server default) | Model name passed to llama.cpp when `--llama-cpp` reviewer is invoked. If unset, the first model reported by `/v1/models` is used. |
|
||||
|
||||
@@ -1244,7 +1244,9 @@ When verification returns `human_needed`, items are persisted as a trackable HUM
|
||||
|
||||
**Command:** `/gsd-review --phase N [--gemini] [--claude] [--codex] [--coderabbit] [--opencode] [--qwen] [--cursor] [--agy] [--ollama] [--lm-studio] [--llama-cpp] [--all]`
|
||||
|
||||
**Purpose:** Invoke external AI CLIs (Gemini, Claude, Codex, CodeRabbit, OpenCode, Qwen Code, Cursor, Antigravity) to independently review phase plans. Produces structured REVIEWS.md with per-reviewer feedback.
|
||||
**Purpose:** Invoke external AI CLIs (Gemini, Claude, Codex, CodeRabbit, OpenCode, Qwen Code, Cursor, Antigravity, Kimi Code) and local OpenAI-compatible servers (Ollama, LM Studio, llama.cpp) to independently review phase plans. Produces structured REVIEWS.md with per-reviewer feedback.
|
||||
|
||||
Each reviewer is a **declared lane**: its binary, prompt and output channels, timeout, availability probe, and empty-output policy come from a capability manifest rather than hand-written per-CLI logic, so a reviewer can be shipped as an installable capability instead of a core change.
|
||||
|
||||
**Requirements:**
|
||||
- REQ-REVIEW-01: System MUST detect available AI CLIs on the system
|
||||
|
||||
@@ -420,6 +420,8 @@
|
||||
"research-store.cjs",
|
||||
"resolution.cjs",
|
||||
"review-lane-descriptor.cjs",
|
||||
"review-lane-invocation.cjs",
|
||||
"review-lane-runner.cjs",
|
||||
"review-reviewer-selection.cjs",
|
||||
"roadmap-command-router.cjs",
|
||||
"roadmap-parser.cjs",
|
||||
|
||||
@@ -723,3 +723,48 @@ premise: the inventory catalogs `bin/lib/*.cjs` **modules**, not capability dire
|
||||
families contain no `capabilities/` entry at all. `gen-inventory-manifest.cjs --check` passes with the
|
||||
five new capability directories added and no inventory edit. The item is vacuous for this phase, and
|
||||
inventing an edit to satisfy it would introduce drift rather than prevent it.
|
||||
|
||||
### 2026-07-30 — vocabulary widened by Phase 5b (#2799)
|
||||
|
||||
Phase 5b is the cutover: it deletes the ~640 lines of hand-authored bash and runs every lane from
|
||||
the declaration. Building the resolver against all twelve legs — the first time each leg's *runtime*
|
||||
contract, not just its shape, had to be reproduced — surfaced five gaps. All five are additive, each
|
||||
is forced by a lane that ships today, and none reverses a decision. This is the same mechanism D2
|
||||
and the Phase 1 amendment record, at the next level of detail.
|
||||
|
||||
| # | Decision | Was | Is | Forced by |
|
||||
|---|---|---|---|---|
|
||||
| 1 | D6 | `handler: null \| antigravity \| openai-compatible` | adds `opencode` | `opencode`'s review is RECONSTRUCTED from assistant `text` parts of a `--format json` stream; a plain stdout copy writes the raw JSON envelope into `REVIEWS.md` (#1936). Admitted under the enum's own second arm — a documented upstream defect data cannot express — exactly as `antigravity` was |
|
||||
| 2 | D1 | model key implicit as `review.models.<slug>` | adds `reviewer.modelConfigKey` | `antigravity`'s slug is `antigravity` but its shipped key is `review.models.agy`. Resolving by slug misses it and silently ignores a configured model, disabling the pinned-model escape hatch #2073 added |
|
||||
| 3 | D2 | `invoke.args` a fixed array | an **argv template** over a closed four-member placeholder set (`{{model}}`, `{{effort}}`, `{{output}}`, `{{prompt}}`) | The injected pieces do not all go in the same place: `codex` injects the model *after* its `exec` subcommand and the output file later still, while five lanes end with a bare `-` that must stay last. Positional splicing produced `codex --model M -o F exec --ephemeral …`, which is not a valid invocation |
|
||||
| 4 | D2 | `openai-http` invoke had no default | adds `invoke.defaultHost` and `invoke.fallbackModel` | Phase 4 federated every `review.*_host` with a default of `""`, so the real fallback (`http://localhost:11434`, `llama3`, …) existed only inside the bash leg. A data-driven lane would POST to a garbage URL |
|
||||
| 5 | D7 | — | `kimi-code` lands with a `command-capability` probe | Net-new lane, per the phase table. `kimi` is claimed by both Kimi Code CLI and the legacy Python kimi-cli (analysis from closed PR #2776, credit @drungrin) |
|
||||
|
||||
**`modelConfigKey` is OPTIONAL, and that is D4 rule 2 rather than a convenience.** It did not exist
|
||||
before this phase, so requiring it would fail validation on every reviewer manifest authored against
|
||||
an earlier GSD. Absent reads as `null`.
|
||||
|
||||
**D5 rule 1 was recorded as delivered by Phase 3 and was not implemented.** The implementation note
|
||||
added to D5 on 2026-07-29 states that "the **consent record** additionally stores the resolved host".
|
||||
It did not: `ConsentRecord` carried no host field, `recordProjectConsent` accepted none, and nothing
|
||||
in the tree bound one — so this phase's rule-4 comparison had no baseline to compare against. Phase
|
||||
5b implements it, as an **optional** `reviewerHost` that `isValidConsentRecord` does not require, so
|
||||
no record already on disk is invalidated and no re-consent storm fires (D4 rule 5). It stays out of
|
||||
`disclosureSignature` for the reason that note gives. Recorded here because the ADR asserting a rule
|
||||
was delivered is precisely what would stop a later phase from checking.
|
||||
|
||||
**Three runtime dependencies leave the review path**, and two of them were platform holes rather than
|
||||
mere overhead: `jq` (absent on stock Windows/Git-Bash, #2589 — it gated five lanes), `curl`, and the
|
||||
external `timeout`/`gtimeout` the Antigravity leg probed for. **Stock macOS ships neither killer**, so
|
||||
D7's "where no bounding mechanism is available the probe is skipped" carve-out was, in practice, that
|
||||
lane running unbounded on every stock Mac. `spawnSync`'s native timeout is always available, so the
|
||||
bound is now unconditional and that carve-out is obsolete.
|
||||
|
||||
**The `DEFECT.GENERATIVE-FIX` parity gate is re-pointed.** Phase 1's assertion required a literal
|
||||
`<!-- reviewer-lane: <slug> -->` per lane inside `invoke_reviewers` and a literal
|
||||
`## <Section> Review` per lane inside `write_reviews` — the exact text this phase deletes. Those two
|
||||
families could not be kept without keeping the hand-maintained per-lane blocks the epic exists to
|
||||
remove, so they are replaced by **descriptor ↔ registry** parity in both directions (the registry is
|
||||
what the runtime iterates once lanes are data) plus an **anti-parity** assertion that fires if a
|
||||
bespoke leg is ever re-added. That also gives #2781/Phase 6 the mechanical single source its docs and
|
||||
locale gate needs, which per-leg text could never provide.
|
||||
|
||||
@@ -93,6 +93,8 @@ export default tseslint.config(
|
||||
'gsd-core/bin/lib/ui-safety-gate.cjs',
|
||||
'gsd-core/bin/lib/review-reviewer-selection.cjs',
|
||||
'gsd-core/bin/lib/review-lane-descriptor.cjs',
|
||||
'gsd-core/bin/lib/review-lane-invocation.cjs',
|
||||
'gsd-core/bin/lib/review-lane-runner.cjs',
|
||||
'gsd-core/bin/lib/clusters.cjs',
|
||||
'gsd-core/bin/lib/installer-migrations/001-legacy-orphan-files.cjs',
|
||||
'gsd-core/bin/lib/observability/redaction.cjs',
|
||||
|
||||
@@ -1130,6 +1130,333 @@ function dispatchOverlayCapabilityCommand({ command, args, cwd, raw, error, load
|
||||
output({ ok: true, row: mutation.row, variant: mutation.variant }, raw, mutation.row);
|
||||
}
|
||||
|
||||
/**
|
||||
* ADR-2782 Phase 5b (#2799) — the seam that lets `invoke_reviewers` iterate declared lanes.
|
||||
*
|
||||
* Replaces ~640 lines of hand-authored per-CLI bash with three subcommands:
|
||||
* plan --selected a,b → resolved lanes (slug, section, availability) as JSON
|
||||
* invoke --slug X → probe + run one lane, writing its review/stub into the run dir
|
||||
* sections --selected a,b → ordered `slug<TAB>reviewsSection`, for write_reviews
|
||||
*
|
||||
* Lanes run SEQUENTIALLY (the workflow loops and calls `invoke` once per lane) because the
|
||||
* original legs did — parallel invocation trips provider rate limits.
|
||||
*/
|
||||
async function routeReviewLane({ args, cwd, raw, error }) {
|
||||
const cp = require('node:child_process');
|
||||
const fsx = require('node:fs');
|
||||
const os = require('node:os');
|
||||
const { REVIEWER_LANES } = require('./lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('./lib/review-lane-invocation.cjs');
|
||||
const runner = require('./lib/review-lane-runner.cjs');
|
||||
const cfgLoader = require('./lib/config-loader.cjs');
|
||||
|
||||
const flag = (name) => {
|
||||
const i = args.indexOf(name);
|
||||
return i !== -1 && args[i + 1] && !String(args[i + 1]).startsWith('--') ? args[i + 1] : null;
|
||||
};
|
||||
const sub = args[1];
|
||||
const runDir = flag('--run-dir') || '.';
|
||||
const repoRoot = flag('--repo-root') || cwd;
|
||||
|
||||
// Resolved config, read ONCE. Reading in-process (rather than shelling out to config-get per
|
||||
// key, as the legs did) is what removes the stringly `"null"` sentinel the bash had to test for.
|
||||
// `loadConfigResolved` returns a PROVENANCE WRAPPER (`{config, source, degraded, reason}`),
|
||||
// not the config — using the wrapper directly silently resolves every key to undefined, which
|
||||
// reads exactly like "nothing configured" and drops every model override without an error.
|
||||
let resolved = {};
|
||||
try { resolved = (cfgLoader.loadConfigResolved(cwd) || {}).config || {}; } catch { resolved = {}; }
|
||||
const configGet = (key) => {
|
||||
let cur = resolved;
|
||||
for (const part of String(key).split('.')) {
|
||||
if (cur === null || typeof cur !== 'object') return undefined;
|
||||
cur = Object.prototype.hasOwnProperty.call(cur, part) ? cur[part] : undefined;
|
||||
}
|
||||
return cur;
|
||||
};
|
||||
|
||||
const selected = (flag('--selected') || '')
|
||||
.split(',').map((s) => s.trim()).filter(Boolean);
|
||||
const laneBySlug = new Map(REVIEWER_LANES.map((l) => [l.slug, l]));
|
||||
const chosen = selected.length ? selected : REVIEWER_LANES.map((l) => l.slug);
|
||||
|
||||
if (sub === 'sections') {
|
||||
const rows = chosen
|
||||
.map((s) => laneBySlug.get(s))
|
||||
.filter(Boolean)
|
||||
.map((l) => `${l.slug}\t${l.reviewsSection}`);
|
||||
process.stdout.write(rows.join('\n') + (rows.length ? '\n' : ''));
|
||||
return;
|
||||
}
|
||||
|
||||
// Effort argv is resolved per lane by the host's own execution policy, exactly as the legs did
|
||||
// via `resolve-execution … --pick effort_argv_string`. A lane whose slug is not a known host
|
||||
// simply gets none.
|
||||
// Resolved through the SAME `resolve-execution` surface the bash legs used
|
||||
// (`--host <slug> --pick effort_argv_string`), so the host's negotiated effortSurface still
|
||||
// decides whether an argument is emitted and the catalog still owns the syntax (ADR-1239 #2481,
|
||||
// ADR-443's escalation ladder). `cmdResolveExecution` writes to stdout and exits, so it cannot
|
||||
// be called in-process for a value — this spawns the same bounded query the legs did, once per
|
||||
// selected lane. A lane whose slug is not a known host resolves to no effort argument at all.
|
||||
const effortFor = (slug) => {
|
||||
try {
|
||||
const r = cp.spawnSync(
|
||||
process.execPath,
|
||||
[__filename, 'query', 'resolve-execution', 'gsd-plan-checker',
|
||||
// NOT `--raw`: that prints the resolved EFFORT ('low'), ignoring --pick. The picked
|
||||
// field is what carries the host-specific syntax ('--effort low' for claude,
|
||||
// '-c model_reasoning_effort=low' for codex), which is the whole point of asking.
|
||||
'--host', slug, '--pick', 'effort_argv_string'],
|
||||
{ cwd, encoding: 'utf8', timeout: 15000, killSignal: 'SIGKILL', maxBuffer: 1024 * 1024 },
|
||||
);
|
||||
if (r.status !== 0) return [];
|
||||
const s = String(r.stdout || '').trim();
|
||||
return s ? s.split(/\s+/).filter(Boolean) : [];
|
||||
} catch { return []; }
|
||||
};
|
||||
|
||||
/**
|
||||
* Per-lane prompt budget (#2797 semantics, preserved exactly).
|
||||
*
|
||||
* `-1` is the UNSET sentinel and falls back to the central `review.max_prompt_tokens`, because
|
||||
* `0` is a legitimate value meaning "do not trim this lane". Treating 0 as unset would silently
|
||||
* switch a user who deliberately disabled trimming onto the global budget.
|
||||
*
|
||||
* Only the budget VALUE is resolved here. Assembly and trimming stay in `prompt-budget`, which
|
||||
* already owns that machinery and is already tested; the workflow calls it and hands the
|
||||
* trimmed file back via `--prompt-file`. Re-implementing it inside the runner would fork a
|
||||
* tested surface for no gain.
|
||||
*/
|
||||
const budgetFor = (lane) => {
|
||||
if (!lane.promptBudgetKey) return null;
|
||||
const per = configGet(lane.promptBudgetKey);
|
||||
const isNum = (v) => typeof v === 'number' && Number.isFinite(v);
|
||||
if (isNum(per) && per !== -1) return per;
|
||||
const global = configGet('review.max_prompt_tokens');
|
||||
return isNum(global) ? global : null;
|
||||
};
|
||||
|
||||
const plans = chosen.map((slug) => {
|
||||
const lane = laneBySlug.get(slug);
|
||||
if (!lane) return { slug, ok: false, reason: 'malformed_lane', detail: 'no such declared lane' };
|
||||
// Per-lane isolation. resolveLanePlan is documented total, but this map is the seam where a
|
||||
// single throw would take down EVERY selected lane rather than the one that is malformed —
|
||||
// and "a cross-AI review that silently drops a lane" is the failure this epic exists to end,
|
||||
// so losing all of them to one bad manifest is strictly worse. Belt and braces on purpose.
|
||||
let r;
|
||||
try {
|
||||
r = resolveLanePlan({ lane, configGet, runDir, repoRoot, effortArgs: effortFor(slug) });
|
||||
} catch (e) {
|
||||
return { slug, ok: false, reason: 'malformed_lane', detail: `resolver threw: ${e && e.message ? e.message : String(e)}` };
|
||||
}
|
||||
return r.ok
|
||||
? {
|
||||
slug,
|
||||
ok: true,
|
||||
section: lane.reviewsSection,
|
||||
transport: r.plan.transport,
|
||||
promptBudget: budgetFor(lane),
|
||||
promptPath: r.plan.transport === 'spawn' ? r.plan.stdin : r.plan.promptPath,
|
||||
plan: r.plan,
|
||||
}
|
||||
: { slug, ok: false, reason: r.reason, detail: r.detail };
|
||||
});
|
||||
|
||||
if (sub === 'plan') {
|
||||
output(plans.map(({ plan, ...rest }) => rest), raw);
|
||||
return;
|
||||
}
|
||||
|
||||
if (sub !== 'invoke') {
|
||||
error("Usage: review-lane <plan|invoke|sections> [--selected a,b] [--run-dir D] [--repo-root R]");
|
||||
return;
|
||||
}
|
||||
|
||||
const slug = flag('--slug');
|
||||
if (!slug) { error('review-lane invoke requires --slug'); return; }
|
||||
const entry = plans.find((p) => p.slug === slug);
|
||||
if (!entry || !entry.ok) {
|
||||
output({ slug, ok: false, reason: entry ? entry.reason : 'malformed_lane', detail: entry ? entry.detail : 'unknown lane' }, raw);
|
||||
return;
|
||||
}
|
||||
|
||||
// EVERY spawn bounded — `DEFECT.UNBOUNDED-SUBPROCESS` (CONTEXT.md:772). A frozen sync spawn
|
||||
// cannot be interrupted by --test-force-exit and hangs a whole CI chunk to its 10-minute kill.
|
||||
const deps = {
|
||||
spawn: (binary, argv, opts) => {
|
||||
const r = cp.spawnSync(binary, argv, {
|
||||
input: opts.input,
|
||||
encoding: 'utf8',
|
||||
timeout: opts.timeoutMs,
|
||||
killSignal: 'SIGKILL',
|
||||
maxBuffer: 64 * 1024 * 1024,
|
||||
shell: false, // argv array only — never a shell string (no interpolation of config values)
|
||||
});
|
||||
return {
|
||||
status: r.status,
|
||||
stdout: r.stdout || '',
|
||||
stderr: r.stderr || '',
|
||||
errorCode: r.error && r.error.code ? r.error.code : undefined,
|
||||
};
|
||||
},
|
||||
httpJson: async (url, opts) => {
|
||||
try {
|
||||
const res = await fetch(url, {
|
||||
method: opts.method,
|
||||
headers: opts.body ? { 'Content-Type': 'application/json' } : undefined,
|
||||
body: opts.body,
|
||||
signal: AbortSignal.timeout(opts.timeoutMs),
|
||||
});
|
||||
return { ok: res.ok, status: res.status, body: await res.text() };
|
||||
} catch (e) {
|
||||
return { ok: false, status: 0, body: '', error: e && e.message ? e.message : String(e) };
|
||||
}
|
||||
},
|
||||
readFile: (p) => fsx.readFileSync(p, 'utf8'),
|
||||
writeFile: (p, c) => fsx.writeFileSync(p, c, 'utf8'),
|
||||
exists: (p) => fsx.existsSync(p),
|
||||
// PATH scan rather than spawning `command -v` / `where`. Two reasons: it spawns nothing at
|
||||
// all (a probe that costs a process is a probe you avoid running, which is how the original
|
||||
// Kimi probe ended up unbounded), and `shell: true` with an args array is deprecated in
|
||||
// Node 26 (DEP0190) because the arguments are concatenated rather than escaped.
|
||||
hasBinary: (name) => {
|
||||
if (!name || name.includes('/') || name.includes('\\')) {
|
||||
try { return fsx.statSync(name).isFile(); } catch { return false; }
|
||||
}
|
||||
const exts = process.platform === 'win32'
|
||||
? (process.env.PATHEXT || '.EXE;.CMD;.BAT;.COM').split(';').filter(Boolean)
|
||||
: [''];
|
||||
for (const dir of (process.env.PATH || '').split(path.delimiter).filter(Boolean)) {
|
||||
for (const ext of exts) {
|
||||
const candidate = path.join(dir, name + ext);
|
||||
try {
|
||||
const st = fsx.statSync(candidate);
|
||||
if (st.isFile()) return true;
|
||||
} catch { /* next candidate */ }
|
||||
}
|
||||
}
|
||||
return false;
|
||||
},
|
||||
configGet,
|
||||
homeDir: os.homedir(),
|
||||
warn: (m) => process.stderr.write(`${m}\n`),
|
||||
};
|
||||
|
||||
// ADR-1517 reviewer instances resolve THROUGH a lane rather than being lanes themselves
|
||||
// (ADR-2782 D8), so they reuse this seam with three substitutions instead of duplicating the
|
||||
// lane's invocation. Everything the lane declares — probe, channels, timeout, empty-output
|
||||
// policy, handler — then applies to the instance unchanged, which is the whole reason to route
|
||||
// them here: a cross-cutting fix reaches instances for free.
|
||||
const asIdentity = flag('--as');
|
||||
const instanceModel = flag('--model');
|
||||
const instanceAgent = flag('--agent');
|
||||
|
||||
if (asIdentity) {
|
||||
// Write under the INSTANCE name so two instances of one adapter never overwrite each other.
|
||||
const safe = String(asIdentity).replace(/[^A-Za-z0-9._-]/g, '-');
|
||||
entry.plan.reviewPath = `${runDir.replace(/\/+$/, '')}/gsd-review-${safe}.md`;
|
||||
entry.plan.errPath = `${runDir.replace(/\/+$/, '')}/gsd-review-${safe}.err`;
|
||||
if (entry.plan.transport === 'spawn' && entry.plan.outputTarget.kind === 'file') {
|
||||
// A file-arg lane names its output inside argv; retarget both together or the runner reads
|
||||
// the lane's file while the tool writes the instance's.
|
||||
const old = entry.plan.outputTarget.path;
|
||||
entry.plan.argv = entry.plan.argv.map((a) => (a === old ? entry.plan.reviewPath : a));
|
||||
entry.plan.outputTarget = { kind: 'file', path: entry.plan.reviewPath };
|
||||
}
|
||||
}
|
||||
|
||||
// The instance's own model replaces whatever the lane resolved from config. Opaque
|
||||
// pass-through: it reaches the tool as one argv element and is never shell-interpolated.
|
||||
//
|
||||
// Done by RE-RESOLVING with an overridden config rather than by patching argv afterwards. A
|
||||
// post-hoc splice has to guess where the flag belongs when the lane resolved no model at all
|
||||
// (the placeholder expanded to nothing, so there is no position to find), and guessing puts it
|
||||
// before a subcommand — `opencode --model M run …`, the same defect the argv template exists to
|
||||
// prevent. Re-resolving lets the lane's own `{{model}}` placeholder decide the position.
|
||||
if (instanceModel) {
|
||||
const lane = laneBySlug.get(entry.slug);
|
||||
const key = lane && lane.modelConfigKey;
|
||||
// A lane that declares NO model key accepts no model override at all (`cursor`, `qwen`,
|
||||
// `coderabbit`). `review.reviewer_instances` validates that `cli` is a known slug but never
|
||||
// that the slug accepts a model, so a user can configure {"cli":"cursor","model":"gpt-5"},
|
||||
// get a clean config-set and a clean run, and never learn their model was ignored. Say so —
|
||||
// a review silently produced by a different model than the one configured is exactly the
|
||||
// "looked like it worked" failure this epic exists to end.
|
||||
if (!key) {
|
||||
process.stderr.write(
|
||||
`reviewer instance model '${instanceModel}' ignored: lane '${entry.slug}' accepts no ` +
|
||||
`model override (it declares no modelConfigKey). The review will use the CLI's own default.\n`,
|
||||
);
|
||||
}
|
||||
const overridden = resolveLanePlan({
|
||||
lane,
|
||||
configGet: (k) => (key && k === key ? instanceModel : configGet(k)),
|
||||
runDir,
|
||||
repoRoot,
|
||||
effortArgs: effortFor(entry.slug),
|
||||
});
|
||||
if (overridden.ok) {
|
||||
// Preserve any instance retargeting already applied above.
|
||||
const { reviewPath, errPath } = entry.plan;
|
||||
entry.plan = overridden.plan;
|
||||
if (asIdentity) {
|
||||
entry.plan.reviewPath = reviewPath;
|
||||
entry.plan.errPath = errPath;
|
||||
if (entry.plan.transport === 'spawn' && entry.plan.outputTarget.kind === 'file') {
|
||||
const old = entry.plan.outputTarget.path;
|
||||
entry.plan.argv = entry.plan.argv.map((a) => (a === old ? reviewPath : a));
|
||||
entry.plan.outputTarget = { kind: 'file', path: reviewPath };
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// `--agent` is OpenCode's native subagent flag and is honoured only by adapters that have the
|
||||
// concept; ignored elsewhere rather than passed to a tool that would reject it.
|
||||
if (instanceAgent && entry.slug === 'opencode' && entry.plan.transport === 'spawn') {
|
||||
const runIdx = entry.plan.argv.indexOf('run');
|
||||
const insertAt = runIdx === -1 ? 0 : runIdx + 1;
|
||||
entry.plan.argv = [
|
||||
...entry.plan.argv.slice(0, insertAt),
|
||||
'--agent',
|
||||
instanceAgent,
|
||||
...entry.plan.argv.slice(insertAt),
|
||||
];
|
||||
}
|
||||
|
||||
// `--prompt-file` lets the caller substitute a budget-trimmed prompt. Applied to whichever
|
||||
// channel this lane actually reads from, so one flag serves both transports.
|
||||
const promptOverride = flag('--prompt-file');
|
||||
if (promptOverride) {
|
||||
if (entry.plan.transport === 'spawn') {
|
||||
if (entry.plan.stdin) entry.plan.stdin = promptOverride;
|
||||
} else {
|
||||
entry.plan.promptPath = promptOverride;
|
||||
}
|
||||
}
|
||||
|
||||
// ADR-2782 D5 rule 4: re-resolve the consented destination at INVOCATION. `undefined` (no
|
||||
// consent record, or a record predating rule 1) means "nothing to compare" and ALLOWS — a
|
||||
// first-party lane ships inside the SHA-pinned distribution and is never consent-gated, so
|
||||
// denying on absence would break every existing local-model user.
|
||||
let consentedHost;
|
||||
if (entry.plan.transport === 'openai-http') {
|
||||
try {
|
||||
const consent = require('./lib/capability-consent.cjs');
|
||||
const projectRoot = require('./lib/project-root.cjs').consentProjectRoot(cwd);
|
||||
// The capability id is kebab; the lane slug may be snake (lm_studio / lm-studio).
|
||||
const capId = String(entry.slug).replace(/_/g, '-');
|
||||
consentedHost = consent.readConsentedReviewerHost({ projectRoot, id: capId });
|
||||
} catch { consentedHost = undefined; }
|
||||
}
|
||||
|
||||
const result = await runner.runLane(entry.plan, deps, {
|
||||
consentedHost,
|
||||
explicitlyRequested: args.includes('--explicit'),
|
||||
repoRoot,
|
||||
});
|
||||
output(result, raw);
|
||||
}
|
||||
|
||||
function routeNormalizeTestCommand({ args, cwd, raw, error }) {
|
||||
// #1857: rewrite a resolved test command to a one-shot form so a
|
||||
// watch-mode runner (vitest/jest) cannot hang a verification gate. Shared
|
||||
@@ -2589,6 +2916,7 @@ const HOST_COMMAND_ROUTERS = {
|
||||
'restore-custom-files': routeRestoreCustomFiles,
|
||||
'from-gsd2': routeFromGsd2,
|
||||
'prompt-budget': routePromptBudget,
|
||||
'review-lane': routeReviewLane,
|
||||
'update-context': routeUpdateContext,
|
||||
'classify-confidence': routeClassifyConfidence,
|
||||
'package-legitimacy': routePackageLegitimacy,
|
||||
|
||||
@@ -210,7 +210,9 @@ const capabilities = {
|
||||
"args": [
|
||||
"--print-timeout",
|
||||
"540s",
|
||||
"-p"
|
||||
"{{model}}",
|
||||
"-p",
|
||||
"{{prompt}}"
|
||||
],
|
||||
"promptChannel": "argv-file-ref",
|
||||
"outputChannel": "stdout",
|
||||
@@ -221,10 +223,9 @@ const capabilities = {
|
||||
"emptyOutput": "handler-owned",
|
||||
"reviewsSection": "Antigravity",
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [
|
||||
"jq"
|
||||
],
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": "review.models.agy",
|
||||
"handler": "antigravity"
|
||||
},
|
||||
"config": {
|
||||
@@ -593,6 +594,8 @@ const capabilities = {
|
||||
"invoke": {
|
||||
"binary": "claude",
|
||||
"args": [
|
||||
"{{model}}",
|
||||
"{{effort}}",
|
||||
"-p",
|
||||
"-"
|
||||
],
|
||||
@@ -607,6 +610,7 @@ const capabilities = {
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": "review.models.claude",
|
||||
"handler": null
|
||||
},
|
||||
"config": {
|
||||
@@ -988,6 +992,7 @@ const capabilities = {
|
||||
"evidenceClass": "diff-only",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": null,
|
||||
"handler": null
|
||||
}
|
||||
},
|
||||
@@ -1100,7 +1105,10 @@ const capabilities = {
|
||||
"args": [
|
||||
"exec",
|
||||
"--ephemeral",
|
||||
"{{model}}",
|
||||
"{{effort}}",
|
||||
"--skip-git-repo-check",
|
||||
"{{output}}",
|
||||
"-"
|
||||
],
|
||||
"promptChannel": "stdin",
|
||||
@@ -1115,6 +1123,7 @@ const capabilities = {
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": "review.models.codex",
|
||||
"handler": null
|
||||
},
|
||||
"config": {
|
||||
@@ -1361,7 +1370,8 @@ const capabilities = {
|
||||
"ask",
|
||||
"--trust",
|
||||
"--output-format",
|
||||
"text"
|
||||
"text",
|
||||
"{{prompt}}"
|
||||
],
|
||||
"promptChannel": "argv-file-ref",
|
||||
"outputChannel": "stdout",
|
||||
@@ -1374,6 +1384,7 @@ const capabilities = {
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": null,
|
||||
"handler": null
|
||||
}
|
||||
},
|
||||
@@ -1603,6 +1614,7 @@ const capabilities = {
|
||||
"invoke": {
|
||||
"binary": "gemini",
|
||||
"args": [
|
||||
"{{model}}",
|
||||
"-p",
|
||||
"-"
|
||||
],
|
||||
@@ -1617,6 +1629,7 @@ const capabilities = {
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": "review.models.gemini",
|
||||
"handler": null
|
||||
},
|
||||
"config": {
|
||||
@@ -2108,6 +2121,46 @@ const capabilities = {
|
||||
"skipSharedHooksInstall": true,
|
||||
"namedSubagentsSupported": false
|
||||
}
|
||||
},
|
||||
"reviewer": {
|
||||
"slug": "kimi-code",
|
||||
"flags": [
|
||||
"--kimi-code"
|
||||
],
|
||||
"transport": "spawn",
|
||||
"probe": {
|
||||
"kind": "command-capability",
|
||||
"binary": "kimi",
|
||||
"needle": "--output-format",
|
||||
"timeoutMs": 5000
|
||||
},
|
||||
"invoke": {
|
||||
"binary": "kimi",
|
||||
"args": [
|
||||
"{{model}}",
|
||||
"-p",
|
||||
"{{prompt}}"
|
||||
],
|
||||
"promptChannel": "argv-file-ref",
|
||||
"outputChannel": "stdout",
|
||||
"modelArg": "-m",
|
||||
"effortChannel": "none"
|
||||
},
|
||||
"timeoutFloorMs": 900000,
|
||||
"emptyOutput": "stub-with-stderr",
|
||||
"reviewsSection": "Kimi Code",
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": "review.models.kimi-code",
|
||||
"handler": null
|
||||
},
|
||||
"config": {
|
||||
"review.models.kimi-code": {
|
||||
"type": "string",
|
||||
"default": "",
|
||||
"description": "Model passed to the Kimi Code reviewer lane."
|
||||
}
|
||||
}
|
||||
},
|
||||
"llama-cpp": {
|
||||
@@ -2135,18 +2188,19 @@ const capabilities = {
|
||||
},
|
||||
"invoke": {
|
||||
"hostConfigKey": "review.llama_cpp_host",
|
||||
"defaultHost": "http://localhost:8080",
|
||||
"path": "/v1/chat/completions",
|
||||
"modelDiscovery": "first-from-models-endpoint",
|
||||
"fallbackModel": "local-model",
|
||||
"effortChannel": "none"
|
||||
},
|
||||
"timeoutFloorMs": 120000,
|
||||
"emptyOutput": "stub-with-stderr",
|
||||
"reviewsSection": "llama.cpp",
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [
|
||||
"jq"
|
||||
],
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": "review.max_prompt_tokens_per_reviewer.llama_cpp",
|
||||
"modelConfigKey": "review.models.llama_cpp",
|
||||
"handler": "openai-compatible"
|
||||
},
|
||||
"config": {
|
||||
@@ -2192,18 +2246,19 @@ const capabilities = {
|
||||
},
|
||||
"invoke": {
|
||||
"hostConfigKey": "review.lm_studio_host",
|
||||
"defaultHost": "http://localhost:1234",
|
||||
"path": "/v1/chat/completions",
|
||||
"modelDiscovery": "first-from-models-endpoint",
|
||||
"fallbackModel": "local-model",
|
||||
"effortChannel": "none"
|
||||
},
|
||||
"timeoutFloorMs": 120000,
|
||||
"emptyOutput": "stub-with-stderr",
|
||||
"reviewsSection": "LM Studio",
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [
|
||||
"jq"
|
||||
],
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": "review.max_prompt_tokens_per_reviewer.lm_studio",
|
||||
"modelConfigKey": "review.models.lm_studio",
|
||||
"handler": "openai-compatible"
|
||||
},
|
||||
"config": {
|
||||
@@ -2473,18 +2528,19 @@ const capabilities = {
|
||||
},
|
||||
"invoke": {
|
||||
"hostConfigKey": "review.ollama_host",
|
||||
"defaultHost": "http://localhost:11434",
|
||||
"path": "/v1/chat/completions",
|
||||
"modelDiscovery": "first-from-models-endpoint",
|
||||
"fallbackModel": "llama3",
|
||||
"effortChannel": "none"
|
||||
},
|
||||
"timeoutFloorMs": 120000,
|
||||
"emptyOutput": "stub-with-stderr",
|
||||
"reviewsSection": "Ollama",
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [
|
||||
"jq"
|
||||
],
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": "review.max_prompt_tokens_per_reviewer.ollama",
|
||||
"modelConfigKey": "review.models.ollama",
|
||||
"handler": "openai-compatible"
|
||||
},
|
||||
"config": {
|
||||
@@ -2634,6 +2690,8 @@ const capabilities = {
|
||||
"binary": "opencode",
|
||||
"args": [
|
||||
"run",
|
||||
"{{model}}",
|
||||
"{{effort}}",
|
||||
"--format",
|
||||
"json",
|
||||
"-"
|
||||
@@ -2647,11 +2705,10 @@ const capabilities = {
|
||||
"emptyOutput": "stub-with-stderr",
|
||||
"reviewsSection": "OpenCode",
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [
|
||||
"jq"
|
||||
],
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"handler": null
|
||||
"modelConfigKey": "review.models.opencode",
|
||||
"handler": "opencode"
|
||||
},
|
||||
"config": {
|
||||
"review.models.opencode": {
|
||||
@@ -2986,6 +3043,7 @@ const capabilities = {
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": null,
|
||||
"handler": null
|
||||
}
|
||||
},
|
||||
@@ -4291,6 +4349,7 @@ const configKeys = {
|
||||
"review.models.gemini": "gemini",
|
||||
"graphify.enabled": "graphify",
|
||||
"intel.enabled": "intel",
|
||||
"review.models.kimi-code": "kimi-code",
|
||||
"review.models.llama_cpp": "llama-cpp",
|
||||
"review.llama_cpp_host": "llama-cpp",
|
||||
"review.max_prompt_tokens_per_reviewer.llama_cpp": "llama-cpp",
|
||||
@@ -4493,6 +4552,12 @@ const configSchema = {
|
||||
"default": false,
|
||||
"description": "Enable the intel code-intelligence command."
|
||||
},
|
||||
"review.models.kimi-code": {
|
||||
"owner": "kimi-code",
|
||||
"type": "string",
|
||||
"default": "",
|
||||
"description": "Model passed to the Kimi Code reviewer lane."
|
||||
},
|
||||
"review.models.llama_cpp": {
|
||||
"owner": "llama-cpp",
|
||||
"type": "string",
|
||||
@@ -4818,7 +4883,9 @@ const runtimes = {
|
||||
"args": [
|
||||
"--print-timeout",
|
||||
"540s",
|
||||
"-p"
|
||||
"{{model}}",
|
||||
"-p",
|
||||
"{{prompt}}"
|
||||
],
|
||||
"promptChannel": "argv-file-ref",
|
||||
"outputChannel": "stdout",
|
||||
@@ -4829,10 +4896,9 @@ const runtimes = {
|
||||
"emptyOutput": "handler-owned",
|
||||
"reviewsSection": "Antigravity",
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [
|
||||
"jq"
|
||||
],
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": "review.models.agy",
|
||||
"handler": "antigravity"
|
||||
},
|
||||
"config": {
|
||||
@@ -5072,6 +5138,8 @@ const runtimes = {
|
||||
"invoke": {
|
||||
"binary": "claude",
|
||||
"args": [
|
||||
"{{model}}",
|
||||
"{{effort}}",
|
||||
"-p",
|
||||
"-"
|
||||
],
|
||||
@@ -5086,6 +5154,7 @@ const runtimes = {
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": "review.models.claude",
|
||||
"handler": null
|
||||
},
|
||||
"config": {
|
||||
@@ -5389,7 +5458,10 @@ const runtimes = {
|
||||
"args": [
|
||||
"exec",
|
||||
"--ephemeral",
|
||||
"{{model}}",
|
||||
"{{effort}}",
|
||||
"--skip-git-repo-check",
|
||||
"{{output}}",
|
||||
"-"
|
||||
],
|
||||
"promptChannel": "stdin",
|
||||
@@ -5404,6 +5476,7 @@ const runtimes = {
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": "review.models.codex",
|
||||
"handler": null
|
||||
},
|
||||
"config": {
|
||||
@@ -5650,7 +5723,8 @@ const runtimes = {
|
||||
"ask",
|
||||
"--trust",
|
||||
"--output-format",
|
||||
"text"
|
||||
"text",
|
||||
"{{prompt}}"
|
||||
],
|
||||
"promptChannel": "argv-file-ref",
|
||||
"outputChannel": "stdout",
|
||||
@@ -5663,6 +5737,7 @@ const runtimes = {
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": null,
|
||||
"handler": null
|
||||
}
|
||||
},
|
||||
@@ -6054,6 +6129,46 @@ const runtimes = {
|
||||
"skipSharedHooksInstall": true,
|
||||
"namedSubagentsSupported": false
|
||||
}
|
||||
},
|
||||
"reviewer": {
|
||||
"slug": "kimi-code",
|
||||
"flags": [
|
||||
"--kimi-code"
|
||||
],
|
||||
"transport": "spawn",
|
||||
"probe": {
|
||||
"kind": "command-capability",
|
||||
"binary": "kimi",
|
||||
"needle": "--output-format",
|
||||
"timeoutMs": 5000
|
||||
},
|
||||
"invoke": {
|
||||
"binary": "kimi",
|
||||
"args": [
|
||||
"{{model}}",
|
||||
"-p",
|
||||
"{{prompt}}"
|
||||
],
|
||||
"promptChannel": "argv-file-ref",
|
||||
"outputChannel": "stdout",
|
||||
"modelArg": "-m",
|
||||
"effortChannel": "none"
|
||||
},
|
||||
"timeoutFloorMs": 900000,
|
||||
"emptyOutput": "stub-with-stderr",
|
||||
"reviewsSection": "Kimi Code",
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": "review.models.kimi-code",
|
||||
"handler": null
|
||||
},
|
||||
"config": {
|
||||
"review.models.kimi-code": {
|
||||
"type": "string",
|
||||
"default": "",
|
||||
"description": "Model passed to the Kimi Code reviewer lane."
|
||||
}
|
||||
}
|
||||
},
|
||||
"opencode": {
|
||||
@@ -6185,6 +6300,8 @@ const runtimes = {
|
||||
"binary": "opencode",
|
||||
"args": [
|
||||
"run",
|
||||
"{{model}}",
|
||||
"{{effort}}",
|
||||
"--format",
|
||||
"json",
|
||||
"-"
|
||||
@@ -6198,11 +6315,10 @@ const runtimes = {
|
||||
"emptyOutput": "stub-with-stderr",
|
||||
"reviewsSection": "OpenCode",
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [
|
||||
"jq"
|
||||
],
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"handler": null
|
||||
"modelConfigKey": "review.models.opencode",
|
||||
"handler": "opencode"
|
||||
},
|
||||
"config": {
|
||||
"review.models.opencode": {
|
||||
@@ -6406,6 +6522,7 @@ const runtimes = {
|
||||
"evidenceClass": "source-grounded",
|
||||
"requiresBinaries": [],
|
||||
"promptBudgetKey": null,
|
||||
"modelConfigKey": null,
|
||||
"handler": null
|
||||
}
|
||||
},
|
||||
|
||||
@@ -837,14 +837,27 @@ const VALID_EVIDENCE_CLASSES = new Set(['source-grounded', 'diff-only']);
|
||||
// member per lane has stopped being a vocabulary and become a dispatch table for
|
||||
// bespoke code — at which point the descriptor is a plugin system wearing a
|
||||
// manifest, and that is a decision for an ADR, not for a downstream phase.
|
||||
const VALID_LANE_HANDLERS = new Set(['antigravity', 'openai-compatible']);
|
||||
// `opencode` admitted by Phase 5b (#2799) under the SECOND arm of the rule above:
|
||||
// it serves 1 lane, justified by a documented upstream defect data cannot express.
|
||||
// OpenCode's default `build` agent is an agentic coder, not a prompt→completion
|
||||
// API; on a large review prompt it can end its turn with ZERO output tokens, and
|
||||
// `--format default` then drops the assistant text entirely, silently losing the
|
||||
// reviewer (#1936). The review must therefore be RECONSTRUCTED from the assistant
|
||||
// `text` parts of a `--format json` stream — a parse, not a copy. Expressing that
|
||||
// as data would need an `outputChannel: 'json-parts'` plus a selector expression,
|
||||
// i.e. exactly the ad-hoc interpreter the admission rule exists to prevent.
|
||||
const VALID_LANE_HANDLERS = new Set(['antigravity', 'openai-compatible', 'opencode']);
|
||||
|
||||
// D2 — `transport` selects the invoke sub-shape. A manifest carrying fields from
|
||||
// BOTH sub-shapes, or from NEITHER, has undefined meaning and fails validation.
|
||||
// The discriminator is explicit rather than inferred from field presence, which
|
||||
// is precisely the ambiguity these two sets exist to detect.
|
||||
const SPAWN_ONLY_INVOKE_FIELDS = ['binary', 'args', 'promptChannel', 'outputChannel', 'outputArg', 'modelArg'];
|
||||
const HTTP_ONLY_INVOKE_FIELDS = ['hostConfigKey', 'path', 'modelDiscovery'];
|
||||
// `defaultHost` / `fallbackModel` added by Phase 5b (#2799). Phase 4 federated every
|
||||
// `review.*_host` key with a default of `""`, so the REAL fallback destination and model
|
||||
// (`http://localhost:11434` / `llama3` and friends) existed only inside the bash leg. Once the
|
||||
// lane is invoked from data, an unset host with no declared default would POST to a garbage URL.
|
||||
const HTTP_ONLY_INVOKE_FIELDS = ['hostConfigKey', 'defaultHost', 'path', 'modelDiscovery', 'fallbackModel'];
|
||||
|
||||
// Feature-only fields are as forbidden on a lane-only capability as on a runtime
|
||||
// one; a `role: "reviewer"` capability owns no artefacts and wires no loop point.
|
||||
@@ -1545,7 +1558,14 @@ function validateRuntimeBody(cap) {
|
||||
*/
|
||||
const KNOWN_REVIEWER_FIELDS = new Set([
|
||||
'slug', 'flags', 'transport', 'probe', 'invoke', 'timeoutFloorMs', 'emptyOutput',
|
||||
'reviewsSection', 'evidenceClass', 'requiresBinaries', 'promptBudgetKey', 'handler',
|
||||
'reviewsSection', 'evidenceClass', 'requiresBinaries', 'promptBudgetKey',
|
||||
// `modelConfigKey` added by Phase 5b (#2799). The model key was IMPLICIT
|
||||
// (`review.models.<slug>`) until a shipped lane broke the convention: antigravity's
|
||||
// slug is `antigravity` but its key is `review.models.agy`, so resolving by slug
|
||||
// missed a configured model and silently disabled the pinned-model escape hatch
|
||||
// #2073 added. A convention one shipped lane already violates is not a contract.
|
||||
'modelConfigKey',
|
||||
'handler',
|
||||
]);
|
||||
|
||||
const KNOWN_PROBE_FIELDS = new Set(['kind', 'binary', 'needle', 'timeoutMs', 'hostConfigKey', 'path']);
|
||||
@@ -1796,6 +1816,19 @@ function validateReviewerBodyFields(cap) {
|
||||
}
|
||||
}
|
||||
|
||||
// OPTIONAL, and that is required by D4 rather than a convenience: `modelConfigKey` did not exist
|
||||
// before Phase 5b, so demanding it would fail validation on every reviewer manifest authored
|
||||
// against an earlier GSD — exactly the forward/backward-compatibility break D4 rule 2 forbids.
|
||||
// Absent is read as `null` (this lane accepts no model override). `null` is explicit; an empty
|
||||
// string is neither, and is rejected.
|
||||
if (r.modelConfigKey !== undefined && r.modelConfigKey !== null &&
|
||||
(typeof r.modelConfigKey !== 'string' || r.modelConfigKey.length === 0)) {
|
||||
errors.push(
|
||||
ctx + ' reviewer.modelConfigKey must be a dotted config key or null ' +
|
||||
'(got: ' + describeValue(r.modelConfigKey) + ')',
|
||||
);
|
||||
}
|
||||
|
||||
// `null` is the declared "no per-lane budget"; an empty string is not.
|
||||
if (r.promptBudgetKey !== null && (typeof r.promptBudgetKey !== 'string' || r.promptBudgetKey.length === 0)) {
|
||||
errors.push(
|
||||
|
||||
@@ -58,32 +58,39 @@ cannot diverge (`DEFECT.GENERATIVE-FIX`; parity-locked in
|
||||
|
||||
## Invocation
|
||||
|
||||
For each selected INSTANCE, invoke its base `cli` using the instance's own `model`/`agent` —
|
||||
NOT the global `review.models.<cli>`. Each instance writes to its OWN per-instance output file
|
||||
under the run-scoped `{run_dir}` (`RUN_DIR` from `gather_context`, #2358 — never a bare
|
||||
`{phase}`-keyed `/tmp` path) and runs as a distinct reviewer identity.
|
||||
An instance resolves **through** a lane; it is not a lane itself (ADR-2782 D8). It takes no part in
|
||||
the roster, the flag set, or lane uniqueness — which is why an instance heading
|
||||
(`## OpenCode Review (opencode-deepseek)`) must never be read as a lane section.
|
||||
|
||||
For an OpenCode-backed instance (the motivating adapter):
|
||||
Since Phase 5b (#2799) `invoke_reviewers` iterates declared lanes rather than hand-authored per-CLI
|
||||
blocks, so an instance is invoked through the same single seam as its base lane, with two
|
||||
substitutions:
|
||||
|
||||
```bash
|
||||
# $INSTANCE_MODEL / $INSTANCE_AGENT come from the instance spec; $INSTANCE_NAME is the
|
||||
# reviewer identity (e.g. opencode-deepseek). --agent is OpenCode's native subagent flag;
|
||||
# omit it when the instance has no agent. {run_dir} is the run-scoped mktemp directory
|
||||
# created once in gather_context (#2358) — same directory every other reviewer block uses.
|
||||
if [ -n "$INSTANCE_AGENT" ] && [ "$INSTANCE_AGENT" != "null" ]; then
|
||||
cat {run_dir}/gsd-review-prompt.md | opencode run --model "$INSTANCE_MODEL" --agent "$INSTANCE_AGENT" - 2>/dev/null > {run_dir}/gsd-review-${INSTANCE_NAME}.md
|
||||
else
|
||||
cat {run_dir}/gsd-review-prompt.md | opencode run --model "$INSTANCE_MODEL" - 2>/dev/null > {run_dir}/gsd-review-${INSTANCE_NAME}.md
|
||||
fi
|
||||
if [ ! -s {run_dir}/gsd-review-${INSTANCE_NAME}.md ]; then
|
||||
echo "OpenCode review ($INSTANCE_NAME) failed or returned empty output." > {run_dir}/gsd-review-${INSTANCE_NAME}.md
|
||||
fi
|
||||
# $INSTANCE_NAME is the reviewer identity (e.g. opencode-deepseek); $INSTANCE_MODEL / $INSTANCE_AGENT
|
||||
# come from the instance spec. --run-dir is the run-scoped mktemp directory created once in
|
||||
# gather_context (#2358) — the same directory every lane uses.
|
||||
#
|
||||
# The instance's OWN model replaces the lane's configured model, and the output lands under the
|
||||
# INSTANCE name so two instances of one adapter never overwrite each other.
|
||||
gsd_run query review-lane invoke \
|
||||
--slug "$INSTANCE_CLI" \
|
||||
--run-dir "$RUN_DIR" --repo-root "$REPO_ROOT" \
|
||||
--model "$INSTANCE_MODEL" ${INSTANCE_AGENT:+--agent "$INSTANCE_AGENT"} \
|
||||
--as "$INSTANCE_NAME"
|
||||
```
|
||||
|
||||
For an instance backed by a DIFFERENT cli, reuse that cli's invocation block with two
|
||||
substitutions: use the instance's `model` in place of the global `review.models.<cli>` value,
|
||||
and write to `{run_dir}/gsd-review-${INSTANCE_NAME}.md`. Only `opencode` honours an
|
||||
`agent` field in v1; ignore `agent` for other adapters.
|
||||
`--as` is what makes the run write `{run_dir}/gsd-review-${INSTANCE_NAME}.md` instead of the lane's
|
||||
own `{run_dir}/gsd-review-<slug>.md`.
|
||||
|
||||
Everything the lane declares — probe, prompt channel, output channel, timeout floor, empty-output
|
||||
policy, handler — applies unchanged to an instance. That is the point of routing instances through
|
||||
the lane rather than duplicating its invocation: a cross-cutting fix reaches instances for free,
|
||||
where the previous per-adapter block had to be copied and kept in sync by hand.
|
||||
|
||||
Only `opencode` honours an `agent` field in v1; it is ignored by other adapters. `model` and `agent`
|
||||
are opaque pass-through strings and are NEVER interpolated into a shell string — the runner spawns
|
||||
with an argv array and `shell: false`.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -277,656 +277,121 @@ Note: `INSTRUCTIONS_BLOCK_FILE`, `ROADMAP_SECTION_FILE`, and `PHASE_DIR` come fr
|
||||
</step>
|
||||
|
||||
<step name="invoke_reviewers">
|
||||
Read model preferences from planning config. Null/missing values fall back to CLI defaults.
|
||||
Every reviewer lane is **declared data** (ADR-2782). This step iterates the lanes the selection
|
||||
resolved; it does not enumerate them. Adding a reviewer is a capability manifest, not an edit here.
|
||||
|
||||
**Do not re-add a per-CLI block.** A `<!-- reviewer-lane: … -->` marker anywhere in this step now
|
||||
FAILS the parity gate (`checkReviewerLaneParity` → `bespoke_leg_present`). Lane divergence is
|
||||
declared in the manifest — timeout floor, probe, prompt/output channel, empty-output policy — and
|
||||
behaviour that data genuinely cannot express is a named first-party `handler` (ADR-2782 D6), never
|
||||
a bespoke block here.
|
||||
|
||||
**Timeout guidance (#2194):** prompt-fed source-grounded reviews are slow — measured ~570 s for
|
||||
Codex at `xhigh` effort and ~525 s for headless Claude on a large plan set. Each lane declares its
|
||||
own `timeoutFloorMs` and the runner enforces it internally, but the **Bash tool call wrapping the
|
||||
loop below must still be given a high `timeout:`** — at least `900000`, and `1200000` when Codex or
|
||||
headless Claude are in the selection — or the host kills the whole loop mid-lane. On Claude Code,
|
||||
raise the host cap via `BASH_MAX_TIMEOUT_MS` if a review can exceed it.
|
||||
|
||||
A silent empty output after a long run is a **timeout kill, not a crash** — the Codex `0xc0000142`
|
||||
misdiagnosis persisted for exactly this reason, because an empty result cannot distinguish the two
|
||||
on its own. Treat an empty result on a slow lane as a dropped lane and re-run with more time rather
|
||||
than diagnosing a CLI or sandbox failure. A cross-AI review that silently drops a lane is blind in
|
||||
one eye.
|
||||
|
||||
**No hook-trust bypass (#2479):** no lane passes a hook-trust bypass flag and none runs a capability
|
||||
probe for one. That flag only bypasses *persisted* hook trust (a first-run condition) and flagless
|
||||
invocations work in steady state, while host-harness safety classifiers deny commands carrying it.
|
||||
An environment that genuinely hits an untrusted-hook prompt surfaces through the `.err` capture and
|
||||
the empty-output stub as a dropped lane with diagnosable stderr, not silent attrition. Do not
|
||||
reintroduce the flag (even spelled out in prose — a regression test bans the literal file-wide).
|
||||
|
||||
**Reviewer instances (#1517, optional):** instances resolve *through* a lane and are not lanes
|
||||
themselves (ADR-2782 D8). Each selected instance invokes its base `cli` with its own `model`/`agent`
|
||||
as opaque argv. Exact invocation in `gsd-core/references/reviewer-instances.md`.
|
||||
|
||||
Lanes run **sequentially, not in parallel** — concurrent invocation trips provider rate limits.
|
||||
|
||||
```bash
|
||||
# JSON scalars from gsd-tools.cjs query; --raw strips the JSON quotes natively
|
||||
# (no jq dependency — jq is absent on stock Windows/Git-Bash, #2589)
|
||||
GEMINI_MODEL=$(gsd_run query config-get review.models.gemini --raw 2>/dev/null || true)
|
||||
CLAUDE_MODEL=$(gsd_run query config-get review.models.claude --raw 2>/dev/null || true)
|
||||
CODEX_MODEL=$(gsd_run query config-get review.models.codex --raw 2>/dev/null || true)
|
||||
OPENCODE_MODEL=$(gsd_run query config-get review.models.opencode --raw 2>/dev/null || true)
|
||||
# review.models.agy, when set, is passed to agy as --model (escape hatch for a
|
||||
# pinned model that 404s server-side); otherwise agy uses its persisted default.
|
||||
AGY_MODEL=$(gsd_run query config-get review.models.agy --raw 2>/dev/null || true)
|
||||
RUN_DIR="{run_dir}"
|
||||
REPO_ROOT="$(git rev-parse --show-toplevel 2>/dev/null || pwd)"
|
||||
# SELECTED_REVIEWERS is the comma-separated result of reviewer selection (ADR-0011 precedence:
|
||||
# explicit flags > --all > review.default_reviewers > all detected). Unchanged by this phase.
|
||||
|
||||
# Reasoning effort per reviewer (#2481). Empty unless the host's effortSurface
|
||||
# axis is `argv`. Pass --attempt N to walk ADR-443's escalation ladder.
|
||||
CLAUDE_EFFORT_ARGS=$(gsd_run query resolve-execution gsd-plan-checker --host claude --pick effort_argv_string 2>/dev/null || true)
|
||||
CODEX_EFFORT_ARGS=$(gsd_run query resolve-execution gsd-plan-checker --host codex --pick effort_argv_string 2>/dev/null || true)
|
||||
OPENCODE_EFFORT_ARGS=$(gsd_run query resolve-execution gsd-plan-checker --host opencode --pick effort_argv_string 2>/dev/null || true)
|
||||
```
|
||||
|
||||
**No hook-trust bypass (#2479):** the codex invocations below deliberately pass no
|
||||
hook-trust bypass flag and run no capability probe for one (#1115's former gate).
|
||||
That flag only bypasses *persisted* hook trust (a first-run condition) and flagless
|
||||
invocations work in steady state, while host-harness safety classifiers deny
|
||||
commands carrying it — and cited the probe itself as intent (#2479). An environment
|
||||
that genuinely hits an untrusted-hook prompt surfaces through the `.err` capture +
|
||||
empty-output guard below as a dropped lane with diagnosable stderr, not silent
|
||||
attrition. Do not reintroduce the flag (even spelled out in prose here — a
|
||||
regression test bans the literal file-wide) or the probe; the #1115 "unexpected
|
||||
argument" failure mode existed only because the flag was emitted, so with no flag
|
||||
there is nothing to version-gate.
|
||||
|
||||
**Reviewer instances (#1517, optional):** when instances are configured, each selected
|
||||
instance invokes its base `cli` with its own `model`/`agent` (opaque argv, never
|
||||
shell-interpolated). Exact invocation in `gsd-core/references/reviewer-instances.md`.
|
||||
|
||||
For each selected CLI, invoke in sequence (not parallel — avoid rate limits):
|
||||
|
||||
**Timeout guidance (#2194):** prompt-fed source-grounded reviews are slow — measured ~570s for Codex at `xhigh` effort and ~525s for headless Claude on a large plan set. Each of the Gemini / Claude / Codex blocks below MUST be invoked with a high Bash `timeout:` — at least `900000` (15 min), and `1200000` (20 min) for Codex `xhigh` or headless Claude — so a lane is not killed mid-review. On Claude Code, raise the host cap via `BASH_MAX_TIMEOUT_MS` if a review can exceed it. A silent empty output after a long run is a **timeout kill, not a crash** — the Codex `0xc0000142` misdiagnosis persisted because the empty-output branches below cannot distinguish the two; treat an empty result on a slow lane as a dropped lane and re-run with more time rather than diagnosing a CLI/sandbox failure. A cross-AI review that silently drops a lane is blind in one eye.
|
||||
|
||||
<!-- reviewer-lane: gemini -->
|
||||
**Gemini:**
|
||||
```bash
|
||||
# #2494: capture stderr to a .err sidecar (not /dev/null) and stub an empty
|
||||
# result, mirroring the Codex and Cursor blocks. Without the guard a failed
|
||||
# lane — CLI missing, unauthenticated, rate-limited, crashed, or any exit that
|
||||
# writes no stdout — leaves a zero-byte file that write_reviews renders as a
|
||||
# reviewer that ran cleanly with nothing to report, silently dropping a lane
|
||||
# from the cross-AI consensus while present_results reports success.
|
||||
# Scope limit: the guard only runs if this block completes. A host Bash-tool
|
||||
# timeout that kills the whole block skips it, the same hard bound the
|
||||
# OpenCode block documents below — per the timeout guidance above, treat an
|
||||
# empty result on a slow lane as a dropped lane rather than a crash.
|
||||
if [ -n "$GEMINI_MODEL" ] && [ "$GEMINI_MODEL" != "null" ]; then
|
||||
cat {run_dir}/gsd-review-prompt.md | gemini -m "$GEMINI_MODEL" -p - 2>{run_dir}/gsd-review-gemini.err > {run_dir}/gsd-review-gemini.md
|
||||
else
|
||||
cat {run_dir}/gsd-review-prompt.md | gemini -p - 2>{run_dir}/gsd-review-gemini.err > {run_dir}/gsd-review-gemini.md
|
||||
fi
|
||||
if [ ! -s {run_dir}/gsd-review-gemini.md ]; then
|
||||
echo "Gemini review failed or returned empty output. stderr:" > {run_dir}/gsd-review-gemini.md
|
||||
cat {run_dir}/gsd-review-gemini.err >> {run_dir}/gsd-review-gemini.md
|
||||
fi
|
||||
```
|
||||
|
||||
<!-- reviewer-lane: claude -->
|
||||
**Claude (separate session):**
|
||||
```bash
|
||||
# #2494: same guard as the Gemini block above — stderr to a .err sidecar
|
||||
# instead of /dev/null, and a diagnostic stub when the lane produces nothing.
|
||||
if [ -n "$CLAUDE_MODEL" ] && [ "$CLAUDE_MODEL" != "null" ]; then
|
||||
cat {run_dir}/gsd-review-prompt.md | claude --model "$CLAUDE_MODEL" $CLAUDE_EFFORT_ARGS -p - 2>{run_dir}/gsd-review-claude.err > {run_dir}/gsd-review-claude.md
|
||||
else
|
||||
cat {run_dir}/gsd-review-prompt.md | claude $CLAUDE_EFFORT_ARGS -p - 2>{run_dir}/gsd-review-claude.err > {run_dir}/gsd-review-claude.md
|
||||
fi
|
||||
if [ ! -s {run_dir}/gsd-review-claude.md ]; then
|
||||
echo "Claude review failed or returned empty output. stderr:" > {run_dir}/gsd-review-claude.md
|
||||
cat {run_dir}/gsd-review-claude.err >> {run_dir}/gsd-review-claude.md
|
||||
fi
|
||||
```
|
||||
|
||||
<!-- reviewer-lane: codex -->
|
||||
**Codex:**
|
||||
```bash
|
||||
# No hook-trust bypass flag — see the #2479 note above. Capture stderr to a .err
|
||||
# file (not /dev/null) so a non-zero exit — e.g. an untrusted-hook prompt on a
|
||||
# first run — is diagnosable instead of a silent empty review (#1115).
|
||||
# Capture the review via codex's own `-o/--output-last-message <FILE>` (only the
|
||||
# final agent message) and discard stdout (#1698): on some platforms (Windows)
|
||||
# codex writes process-teardown output to stdout *after* the final message, and a
|
||||
# stdout redirect would append that noise to a non-empty file — slipping past the
|
||||
# `[ ! -s … ]` empty-output guard as a silently polluted review.
|
||||
if [ -n "$CODEX_MODEL" ] && [ "$CODEX_MODEL" != "null" ]; then
|
||||
cat {run_dir}/gsd-review-prompt.md | codex exec --ephemeral --model "$CODEX_MODEL" $CODEX_EFFORT_ARGS --skip-git-repo-check -o {run_dir}/gsd-review-codex.md - 2>{run_dir}/gsd-review-codex.err >/dev/null
|
||||
else
|
||||
cat {run_dir}/gsd-review-prompt.md | codex exec --ephemeral $CODEX_EFFORT_ARGS --skip-git-repo-check -o {run_dir}/gsd-review-codex.md - 2>{run_dir}/gsd-review-codex.err >/dev/null
|
||||
fi
|
||||
if [ ! -s {run_dir}/gsd-review-codex.md ]; then
|
||||
echo "Codex review failed or returned empty output. stderr:" > {run_dir}/gsd-review-codex.md
|
||||
cat {run_dir}/gsd-review-codex.err >> {run_dir}/gsd-review-codex.md
|
||||
fi
|
||||
```
|
||||
|
||||
<!-- reviewer-lane: coderabbit -->
|
||||
**CodeRabbit:**
|
||||
|
||||
Note: CodeRabbit reviews the current git diff/working tree — it does not accept a prompt or model flag. It may take up to 5 minutes. Use `timeout: 360000` on the Bash tool call. The source-grounding requirement in the build_prompt Review Instructions applies only to the prompt-fed reviewers above; CodeRabbit is a diff-only reviewer and never receives it. Treat its output as a diff observation, not a grounded plan-level verdict.
|
||||
|
||||
```bash
|
||||
# #2605: same guard as every other leg (#2494/#2592). `2>/dev/null` with no
|
||||
# `[ ! -s … ]` stub left a zero-byte file when coderabbit was missing,
|
||||
# unauthenticated, or exited without stdout — write_reviews then rendered a
|
||||
# reviewer that "ran cleanly with nothing to report", silently dropping the lane.
|
||||
coderabbit review --prompt-only 2>{run_dir}/gsd-review-coderabbit.err > {run_dir}/gsd-review-coderabbit.md
|
||||
if [ ! -s {run_dir}/gsd-review-coderabbit.md ]; then
|
||||
echo "CodeRabbit review failed or returned empty output. stderr:" > {run_dir}/gsd-review-coderabbit.md
|
||||
cat {run_dir}/gsd-review-coderabbit.err >> {run_dir}/gsd-review-coderabbit.md 2>/dev/null
|
||||
fi
|
||||
```
|
||||
|
||||
<!-- reviewer-lane: opencode -->
|
||||
**OpenCode (via GitHub Copilot):**
|
||||
|
||||
OpenCode's default `build` agent is an agentic coder, not a prompt→completion API.
|
||||
On a large review prompt it may run a few `read` tool calls and then end its turn
|
||||
with **zero output tokens** (`reason:"stop"`, `output:0`), so `--format default`
|
||||
yields empty stdout and the second reviewer is silently lost (#1936). Invoke with
|
||||
`--format json` and reconstruct the review from the assistant `text` parts; if the
|
||||
agent emitted none, surface the stop `reason`, output-token count, and captured
|
||||
stderr so the failure is diagnosable instead of a generic empty stub. Runs are also
|
||||
nondeterministic in length, so bound this Bash tool call with a wall-clock timeout —
|
||||
set `timeout: 660000` on the call (same mechanism the CodeRabbit block documents).
|
||||
That bound is hard: if it fires mid-`opencode run` the tool kills the command and
|
||||
the jq reconstruction below never runs, so the reviewing agent simply proceeds
|
||||
without an OpenCode result. The completing zero-output case — the actual #1936 bug —
|
||||
is fully handled below; the timeout only backstops the rarer nondeterministic hang. A
|
||||
reviewer instance with `"agent": "review"` (see
|
||||
`gsd-core/references/reviewer-instances.md`) sidesteps the default `build` agent and
|
||||
is the durable fix when this recurs.
|
||||
|
||||
```bash
|
||||
# stderr → sidecar (never /dev/null) so a real error is diagnosable — mirrors the
|
||||
# Codex block. --format json is the primary invocation (not a fallback): the review
|
||||
# text lives in assistant `text` parts, which the default formatter drops when the
|
||||
# agent stops with no final message (#1936).
|
||||
if [ -n "$OPENCODE_MODEL" ] && [ "$OPENCODE_MODEL" != "null" ]; then
|
||||
set -- --model "$OPENCODE_MODEL"
|
||||
else
|
||||
set --
|
||||
fi
|
||||
cat {run_dir}/gsd-review-prompt.md | opencode run "$@" $OPENCODE_EFFORT_ARGS --format json - 2>{run_dir}/gsd-review-opencode.err > {run_dir}/gsd-review-opencode.json
|
||||
# Reconstruct the review from the assistant text parts into a variable and test
|
||||
# its CONTENT, not the file size: an empty extraction still prints a trailing
|
||||
# newline that would fool a `[ -s file ]` check into skipping the stub.
|
||||
OPENCODE_REVIEW=$(jq -rs '[.[] | select(.type=="text") | .part.text // empty] | join("\n")' {run_dir}/gsd-review-opencode.json 2>/dev/null)
|
||||
if [ -n "$OPENCODE_REVIEW" ]; then
|
||||
printf '%s\n' "$OPENCODE_REVIEW" > {run_dir}/gsd-review-opencode.md
|
||||
else
|
||||
# No assistant text (no final message, or stdout was not valid JSON):
|
||||
{
|
||||
echo "OpenCode review returned no assistant text (#1936: agent ended its turn with no final message)."
|
||||
OPENCODE_DIAG=$(jq -rs '[.[] | select(.type=="step_finish")] | last | "stop reason=\(.part.reason // "?"), output tokens=\(.part.tokens.output // "?")"' {run_dir}/gsd-review-opencode.json 2>/dev/null)
|
||||
[ -n "$OPENCODE_DIAG" ] && echo "Diagnostic: $OPENCODE_DIAG"
|
||||
echo "stderr:"
|
||||
cat {run_dir}/gsd-review-opencode.err
|
||||
} > {run_dir}/gsd-review-opencode.md
|
||||
fi
|
||||
```
|
||||
|
||||
<!-- reviewer-lane: qwen -->
|
||||
**Qwen Code:**
|
||||
```bash
|
||||
# #2794: the last leg still sending stderr to /dev/null. Every other lane
|
||||
# captures it to a .err sidecar and appends it to the stub (#2494/#2605); qwen
|
||||
# wrote a bare "failed or returned empty output." with no diagnostic at all, so
|
||||
# a missing binary, an auth prompt, and a rate-limit were indistinguishable from
|
||||
# each other and from a clean empty review.
|
||||
cat {run_dir}/gsd-review-prompt.md | qwen - 2>{run_dir}/gsd-review-qwen.err > {run_dir}/gsd-review-qwen.md
|
||||
if [ ! -s {run_dir}/gsd-review-qwen.md ]; then
|
||||
echo "Qwen review failed or returned empty output. stderr:" > {run_dir}/gsd-review-qwen.md
|
||||
cat {run_dir}/gsd-review-qwen.err >> {run_dir}/gsd-review-qwen.md 2>/dev/null
|
||||
fi
|
||||
```
|
||||
|
||||
<!-- reviewer-lane: cursor -->
|
||||
**Cursor:**
|
||||
```bash
|
||||
# cursor-agent is a SEPARATE binary from the `cursor` IDE launcher; print mode (-p) takes the
|
||||
# prompt as an ARGUMENT, not stdin. A full review prompt can exceed the OS argument limit, so
|
||||
# reference the prompt file by path rather than inlining it. Capture stderr so a failure is
|
||||
# diagnosable instead of a silent empty result.
|
||||
# #2176: same absolute-root anchor as the Antigravity block — cursor-agent runs
|
||||
# in the repo cwd, but repo-relative references in the assembled prompt still
|
||||
# need an explicit root to resolve against. rev-parse (not bare pwd) so the
|
||||
# anchor is correct even when /gsd:review is invoked from a repo subdirectory.
|
||||
_CURSOR_ROOT="$(git rev-parse --show-toplevel 2>/dev/null || pwd)"
|
||||
CURSOR_PROMPT_ARG="Read the file at {run_dir}/gsd-review-prompt.md in full and carry out the review request it contains. The repository under review is at $_CURSOR_ROOT — resolve every relative file path in the review request against that absolute root. Output only the resulting markdown review. Do not edit any files."
|
||||
cursor-agent -p --mode ask --trust --output-format text "$CURSOR_PROMPT_ARG" 2>{run_dir}/gsd-review-cursor.err > {run_dir}/gsd-review-cursor.md
|
||||
if [ ! -s {run_dir}/gsd-review-cursor.md ]; then
|
||||
echo "Cursor review failed or returned empty output. stderr:" > {run_dir}/gsd-review-cursor.md
|
||||
cat {run_dir}/gsd-review-cursor.err >> {run_dir}/gsd-review-cursor.md
|
||||
fi
|
||||
```
|
||||
|
||||
<!-- reviewer-lane: antigravity -->
|
||||
**Antigravity CLI:**
|
||||
|
||||
**Maintainer note — why this block has three layers (last updated against agy 1.0.16):**
|
||||
|
||||
`agy -p` (the `--print` non-interactive flag) works correctly on macOS and Linux: it sends the
|
||||
prompt, receives the model response, and writes it to stdout. On **native Windows** it silently
|
||||
produces no stdout output despite the API call succeeding — a bug in `text_drip.go`'s non-TTY
|
||||
flush path, tracked at https://github.com/google-antigravity/antigravity-cli/issues/27466 and
|
||||
still open as of agy 1.0.2.
|
||||
|
||||
Regardless of platform, `agy` always persists the full exchange to a transcript file on disk.
|
||||
The transcript fallback (Step 2 below) reads that file directly, giving Windows users full review
|
||||
coverage without any extra tooling. This pattern was first documented by the community MCP bridge
|
||||
at https://github.com/SinanTufekci/Claude-Code-Antigravity-CLI-MCP-Server — we inline the same
|
||||
logic here in pure bash/jq so no additional dependency is required.
|
||||
|
||||
**Stale-response guard (why the pre-flight watermark matters):**
|
||||
Without a watermark, the fallback would read the last `PLANNER_RESPONSE` entry in the transcript
|
||||
regardless of when it was written — including entries from a previous invocation in the same
|
||||
workspace. To prevent that, we record the transcript's line count *before* calling `agy -p`. In
|
||||
the fallback, we only read lines appended after that count. If no new lines were written (agy
|
||||
failed before producing a response), `_AGY_RESULT` is empty and Step 3 fires — never stale. If
|
||||
the conv-id changed (agy started a fresh session), all lines in the new file are new and we use
|
||||
skip=0.
|
||||
|
||||
**If the upstream stdout bug is fixed** (check the issue above): Step 2 silently becomes
|
||||
unreachable; stdout is non-empty and Step 1 handles it. No code change needed.
|
||||
|
||||
**If the transcript paths change** in a future `agy` release: Step 2 silently becomes a no-op
|
||||
and Step 3 fires with a clear error message in REVIEWS.md. No silent corruption. To debug:
|
||||
- `~/.gemini/antigravity-cli/cache/last_conversations.json` — workspace → conv-id map
|
||||
- `~/.gemini/antigravity-cli/brain/<id>/.system_generated/logs/transcript.jsonl`
|
||||
Filter: `source=="MODEL"`, `status=="DONE"`, `type=="PLANNER_RESPONSE"`, take the last match's `content` field.
|
||||
|
||||
Invocation specifics (verified agy 1.0.0, macOS arm64 and Linux amd64):
|
||||
- `-p` takes the prompt as a **flag value** — `echo X | agy -p` errors with "flag needs an argument: -p"
|
||||
- `--print-timeout` defaults to 5m, aligning with this workflow's global timeout
|
||||
- `--model "<name>"` selects the model (available since agy ~1.0.3; `agy models` lists
|
||||
them). When `review.models.agy` is set it is passed as `--model`; otherwise agy uses
|
||||
its persisted default (`agy models`).
|
||||
|
||||
```bash
|
||||
# Pre-flight: snapshot the transcript watermark before invoking agy.
|
||||
# Must run BEFORE agy -p — this is what prevents the fallback from reading a stale prior response.
|
||||
_AGY_WS=$(git rev-parse --show-toplevel 2>/dev/null || pwd)
|
||||
_AGY_CACHE="$HOME/.gemini/antigravity-cli/cache/last_conversations.json"
|
||||
_AGY_MARK_CONV=""
|
||||
_AGY_MARK_LINES=0
|
||||
if [ -f "$_AGY_CACHE" ]; then
|
||||
_AGY_MARK_CONV=$(jq -r --arg ws "$_AGY_WS" '
|
||||
.[$ws] //
|
||||
(to_entries
|
||||
| map(select(.key | ascii_downcase == ($ws | ascii_downcase)))
|
||||
| first | .value) //
|
||||
empty
|
||||
' "$_AGY_CACHE" 2>/dev/null)
|
||||
if [ -n "$_AGY_MARK_CONV" ] && [ "$_AGY_MARK_CONV" != "null" ]; then
|
||||
_AGY_MARK_TX="$HOME/.gemini/antigravity-cli/brain/${_AGY_MARK_CONV}/.system_generated/logs/transcript.jsonl"
|
||||
[ -f "$_AGY_MARK_TX" ] && _AGY_MARK_LINES=$(wc -l < "$_AGY_MARK_TX" | tr -d ' ')
|
||||
fi
|
||||
fi
|
||||
|
||||
# Step 1 — primary invocation: stdout works on macOS, Linux, and WSL.
|
||||
# Three hardening invariants (#2073), all mirroring the Cursor block's discipline:
|
||||
# * FILE-REFERENCE prompt (not inline `$(cat …)`) — a large review prompt (≈197 KB
|
||||
# for 6 plans + CONTEXT + RESEARCH + REQUIREMENTS) overflows the exec arg list
|
||||
# (`bash: agy: Argument list too long`, rc 126), indistinguishable from a model
|
||||
# failure when stderr is suppressed.
|
||||
# * EXTERNAL `timeout` wrapper when available (GNU `timeout` / `gtimeout`) —
|
||||
# `--print-timeout` is agy's native cap but it CANNOT fire before agy creates a
|
||||
# session; under concurrent heavy runs one process can stall pre-session (no
|
||||
# `brain/<conv-id>/` dir, alive at 583 s despite `--print-timeout 300s`). The
|
||||
# external cap bounds wall-clock regardless. Stock macOS lacks `timeout`, so
|
||||
# the block probes for it and falls back to --print-timeout alone there.
|
||||
# * `--model` from `review.models.agy` when set — escape hatch for a pinned model
|
||||
# that 404s server-side (exits 0 with empty stdout AND empty transcript).
|
||||
# * stdin tied to /dev/null so agy never blocks on a tty.
|
||||
# A non-zero exit (external timeout = 124, crash, etc.) discards any partial output
|
||||
# so the Step 2 transcript fallback / Step 3 diagnostic take over.
|
||||
if [ -n "$AGY_MODEL" ] && [ "$AGY_MODEL" != "null" ]; then
|
||||
set -- --model "$AGY_MODEL"
|
||||
else
|
||||
set --
|
||||
fi
|
||||
# #2176: grant the reviewer the repo under review. Without --add-dir, agy's
|
||||
# permission context never receives the cwd repo — the agent anchors on its own
|
||||
# ~/.gemini/antigravity-cli/scratch dir and reviews the plan text in isolation
|
||||
# (the exact failure the Review Instructions forbid). Capability-probed so an
|
||||
# older agy without --add-dir still runs; the prompt anchor below keeps
|
||||
# absolute-path reads possible on that fallback.
|
||||
if agy --help 2>/dev/null | grep -q -- '--add-dir'; then
|
||||
set -- "$@" --add-dir "$_AGY_WS"
|
||||
fi
|
||||
# #2176: anchor the prompt to the absolute repo root so repo-relative references
|
||||
# in the assembled review prompt resolve even on the no---add-dir fallback, and
|
||||
# require an explicit self-report if the reviewer still cannot read the repo.
|
||||
_AGY_PROMPT="Read the file at {run_dir}/gsd-review-prompt.md in full and carry out the review request it contains. The repository under review is at $_AGY_WS — resolve every relative file path in the review request against that absolute root and verify claims against those files. If you cannot read files under $_AGY_WS, begin your output with the exact line REVIEWED-WITHOUT-REPO-ACCESS before the review. Output only the resulting markdown review. Do not edit any files."
|
||||
# Capability-probe an external wall-clock killer (GNU coreutils `timeout` or the
|
||||
# macOS Homebrew `gtimeout`). Stock macOS ships NEITHER — a bare `timeout …` would
|
||||
# fail with rc 127 ("command not found") and silently lose the reviewer, so fall
|
||||
# back to agy's native --print-timeout alone in that case. The external cap, when
|
||||
# available, is set HIGHER than --print-timeout so it only backstops a pre-session
|
||||
# stall (which --print-timeout cannot bound — #2073 mode 3) and never pre-empts a
|
||||
# healthy run. Mirrors the probe in scripts/base64-scan.sh.
|
||||
_AGY_KILLER="$(command -v timeout 2>/dev/null || command -v gtimeout 2>/dev/null || true)"
|
||||
if [ -n "$_AGY_KILLER" ]; then
|
||||
"$_AGY_KILLER" 600 agy --print-timeout 540s "$@" -p "$_AGY_PROMPT" </dev/null 2>/dev/null > {run_dir}/gsd-review-antigravity.md
|
||||
else
|
||||
agy --print-timeout 540s "$@" -p "$_AGY_PROMPT" </dev/null 2>/dev/null > {run_dir}/gsd-review-antigravity.md
|
||||
fi
|
||||
_AGY_RC=$?
|
||||
if [ "$_AGY_RC" -ne 0 ]; then
|
||||
: > {run_dir}/gsd-review-antigravity.md
|
||||
fi
|
||||
|
||||
# Step 2 — transcript fallback: catches Windows agy -p stdout bug (and any future stdout-silent edge cases).
|
||||
# Reads only lines appended AFTER the pre-flight watermark. If agy failed before writing a new response,
|
||||
# _AGY_RESULT is empty and Step 3 fires — no stale content can leak through.
|
||||
# Undocumented paths, verified agy 1.0.0–1.0.2. See maintainer note above if these break.
|
||||
if [ ! -s {run_dir}/gsd-review-antigravity.md ]; then
|
||||
if [ -f "$_AGY_CACHE" ]; then
|
||||
_AGY_CONV=$(jq -r --arg ws "$_AGY_WS" '
|
||||
.[$ws] //
|
||||
(to_entries
|
||||
| map(select(.key | ascii_downcase == ($ws | ascii_downcase)))
|
||||
| first | .value) //
|
||||
empty
|
||||
' "$_AGY_CACHE" 2>/dev/null)
|
||||
if [ -n "$_AGY_CONV" ] && [ "$_AGY_CONV" != "null" ]; then
|
||||
_AGY_TX="$HOME/.gemini/antigravity-cli/brain/${_AGY_CONV}/.system_generated/logs/transcript.jsonl"
|
||||
if [ -f "$_AGY_TX" ]; then
|
||||
# If conv-id changed, agy started a new session — all lines are new, skip 0.
|
||||
# If same conv-id, only read lines beyond the watermark.
|
||||
[ "$_AGY_CONV" = "$_AGY_MARK_CONV" ] && _AGY_SKIP=$_AGY_MARK_LINES || _AGY_SKIP=0
|
||||
_AGY_RESULT=$(tail -n +"$((_AGY_SKIP + 1))" "$_AGY_TX" 2>/dev/null | \
|
||||
jq -r 'select(.source=="MODEL" and .status=="DONE" and .type=="PLANNER_RESPONSE") | .content' \
|
||||
2>/dev/null | tail -1)
|
||||
[ -n "$_AGY_RESULT" ] && echo "$_AGY_RESULT" > {run_dir}/gsd-review-antigravity.md
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
# Step 3 — final guard: both approaches yielded nothing (auth error, first-run setup,
|
||||
# path schema changed, 404'd pinned model, pre-session stall, etc.)
|
||||
if [ ! -s {run_dir}/gsd-review-antigravity.md ]; then
|
||||
{
|
||||
echo "Antigravity review failed or returned empty output."
|
||||
# #2073 mode 2: a pinned model that 404s exits 0 with empty stdout AND an empty
|
||||
# transcript — the only evidence is in agy's own log. Surface it instead of a
|
||||
# bare generic stub so the failure is diagnosable.
|
||||
_AGY_LOG="$HOME/.gemini/antigravity-cli/cli.log"
|
||||
if [ -f "$_AGY_LOG" ]; then
|
||||
_AGY_ERR=$(grep -iE 'agent executor error|NOT_FOUND|Publisher model' "$_AGY_LOG" | tail -3)
|
||||
if [ -n "$_AGY_ERR" ]; then
|
||||
echo "agy log hint (pinned model may be unavailable — run 'agy models' and set review.models.agy):"
|
||||
echo "$_AGY_ERR"
|
||||
fi
|
||||
fi
|
||||
# #2073 mode 3: pre-session stall tell — no new conversation dir appeared.
|
||||
echo "If no agy run started, that is the pre-session-stall case: check whether a new ~/.gemini/antigravity-cli/brain/<conv-id>/ dir appeared within ~30s of launch."
|
||||
} > {run_dir}/gsd-review-antigravity.md
|
||||
fi
|
||||
|
||||
# #2176: blind-review marker. Two tells that the reviewer ran without repo
|
||||
# access: the prompt's mandated REVIEWED-WITHOUT-REPO-ACCESS self-report in the
|
||||
# first lines of output, or the agent DECLARING the scratch dir as its
|
||||
# workspace. Both patterns are anchored — the self-report to the head of the
|
||||
# file, the scratch tell to a workspace-declaration phrasing — so a grounded
|
||||
# review that merely QUOTES these strings (e.g. reviewing this very file) is
|
||||
# never mis-stamped. Stamp a machine-readable marker so the Consensus Summary
|
||||
# down-weights the review instead of counting an ungrounded verdict at full
|
||||
# weight. (Temp file + mv, no in-place sed — BSD/GNU safe.)
|
||||
if [ -s {run_dir}/gsd-review-antigravity.md ] && \
|
||||
{ head -5 {run_dir}/gsd-review-antigravity.md | grep -q 'REVIEWED-WITHOUT-REPO-ACCESS' || \
|
||||
grep -qiE '(workspace|working) (directory|dir).{0,40}antigravity-cli/scratch' {run_dir}/gsd-review-antigravity.md; }; then
|
||||
{
|
||||
echo "> [reviewed-without-repo-access] This reviewer ran without visibility into the repo under review — down-weight its verdict in the Consensus Summary."
|
||||
echo ""
|
||||
cat {run_dir}/gsd-review-antigravity.md
|
||||
} > {run_dir}/gsd-review-antigravity.md.tmp && \
|
||||
mv {run_dir}/gsd-review-antigravity.md.tmp {run_dir}/gsd-review-antigravity.md
|
||||
fi
|
||||
```
|
||||
|
||||
<!-- reviewer-lane: ollama -->
|
||||
**Ollama (local, OpenAI-compatible):**
|
||||
|
||||
Read host and model from config. All three local backends share the same `/v1/chat/completions` endpoint — only host and model differ. Use `jq --rawfile` to safely encode the multi-line prompt as JSON without shell-escaping issues.
|
||||
|
||||
```bash
|
||||
# Shared helper: apply prompt-budget trimming for local reviewers
|
||||
# Shared budget-trim helper. Was defined inside the Ollama leg; it is lane-agnostic, so it is
|
||||
# hoisted here now that any lane may declare a promptBudgetKey. Returns non-zero when the budget
|
||||
# is too small for the minimum review set (prompt-budget exit 2 / 11).
|
||||
prepare_trimmed_prompt_for_reviewer() {
|
||||
REVIEWER_KEY="$1"
|
||||
REVIEWER_BUDGET="$2"
|
||||
OUTPUT_PROMPT="$3"
|
||||
OUTPUT_META="$4"
|
||||
|
||||
[ -z "$REVIEWER_BUDGET" ] && return 0
|
||||
[ "$REVIEWER_BUDGET" = "null" ] && return 0
|
||||
[ "$REVIEWER_BUDGET" = "0" ] && return 0
|
||||
REVIEWER_KEY="$1"; REVIEWER_BUDGET="$2"; OUTPUT_PROMPT="$3"; OUTPUT_META="$4"
|
||||
|
||||
PLAN_FILE_ARGS=""
|
||||
for p in {run_dir}/gsd-review-plan-*.md; do
|
||||
for p in "$RUN_DIR"/gsd-review-plan-*.md; do
|
||||
[ -f "$p" ] && PLAN_FILE_ARGS="$PLAN_FILE_ARGS --plan-file $p"
|
||||
done
|
||||
PROJECT_ARG=""
|
||||
[ -f "{run_dir}/gsd-review-project.md" ] && PROJECT_ARG="--project-file {run_dir}/gsd-review-project.md"
|
||||
[ -f "$RUN_DIR/gsd-review-project.md" ] && PROJECT_ARG="--project-file $RUN_DIR/gsd-review-project.md"
|
||||
CONTEXT_ARG=""
|
||||
[ -f "{run_dir}/gsd-review-context.md" ] && CONTEXT_ARG="--context-file {run_dir}/gsd-review-context.md"
|
||||
[ -f "$RUN_DIR/gsd-review-context.md" ] && CONTEXT_ARG="--context-file $RUN_DIR/gsd-review-context.md"
|
||||
RESEARCH_ARG=""
|
||||
[ -f "{run_dir}/gsd-review-research.md" ] && RESEARCH_ARG="--research-file {run_dir}/gsd-review-research.md"
|
||||
[ -f "$RUN_DIR/gsd-review-research.md" ] && RESEARCH_ARG="--research-file $RUN_DIR/gsd-review-research.md"
|
||||
REQUIREMENTS_ARG=""
|
||||
[ -f "{run_dir}/gsd-review-requirements.md" ] && REQUIREMENTS_ARG="--requirements-file {run_dir}/gsd-review-requirements.md"
|
||||
[ -f "$RUN_DIR/gsd-review-requirements.md" ] && REQUIREMENTS_ARG="--requirements-file $RUN_DIR/gsd-review-requirements.md"
|
||||
|
||||
gsd_run query prompt-budget \
|
||||
--budget "$REVIEWER_BUDGET" \
|
||||
--instructions-file "{run_dir}/gsd-review-instructions.md" \
|
||||
--roadmap-file "{run_dir}/gsd-review-roadmap.md" \
|
||||
--instructions-file "$RUN_DIR/gsd-review-instructions.md" \
|
||||
--roadmap-file "$RUN_DIR/gsd-review-roadmap.md" \
|
||||
$PLAN_FILE_ARGS $PROJECT_ARG $CONTEXT_ARG $RESEARCH_ARG $REQUIREMENTS_ARG \
|
||||
--output-prompt "$OUTPUT_PROMPT" \
|
||||
--output-metadata "$OUTPUT_META"
|
||||
return $?
|
||||
}
|
||||
|
||||
# Resolve prompt budget for Ollama: per-reviewer override > global default > null
|
||||
OLLAMA_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens_per_reviewer.ollama --raw 2>/dev/null || echo "null")
|
||||
# #2797: -1 = unset; 0 legitimately means "do not trim".
|
||||
if [ -z "$OLLAMA_REVIEWER_BUDGET" ] || [ "$OLLAMA_REVIEWER_BUDGET" = "null" ] || [ "$OLLAMA_REVIEWER_BUDGET" = "-1" ]; then
|
||||
OLLAMA_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens --raw 2>/dev/null || echo "null")
|
||||
fi
|
||||
gsd_run query review-lane plan \
|
||||
--selected "$SELECTED_REVIEWERS" --run-dir "$RUN_DIR" --repo-root "$REPO_ROOT" --json \
|
||||
> "$RUN_DIR/gsd-review-lanes.json"
|
||||
|
||||
# Apply budget trim for Ollama if a budget is configured
|
||||
OLLAMA_PROMPT_FILE="{run_dir}/gsd-review-prompt.md"
|
||||
OLLAMA_SKIP=0
|
||||
if [ -n "$OLLAMA_REVIEWER_BUDGET" ] && [ "$OLLAMA_REVIEWER_BUDGET" != "null" ] && [ "$OLLAMA_REVIEWER_BUDGET" != "0" ]; then
|
||||
OLLAMA_TRIMMED_PROMPT="{run_dir}/gsd-review-prompt-ollama.md"
|
||||
OLLAMA_TRIM_META="{run_dir}/gsd-review-prompt-ollama.metadata.json"
|
||||
prepare_trimmed_prompt_for_reviewer "ollama" "$OLLAMA_REVIEWER_BUDGET" "$OLLAMA_TRIMMED_PROMPT" "$OLLAMA_TRIM_META"
|
||||
OLLAMA_EXIT=$?
|
||||
if [ $OLLAMA_EXIT -ne 0 ]; then
|
||||
if [ $OLLAMA_EXIT -eq 2 ] || [ $OLLAMA_EXIT -eq 11 ]; then
|
||||
echo "WARNING: prompt budget for ollama (${OLLAMA_REVIEWER_BUDGET} tokens) is too small for the minimum review set. Skipping Ollama reviewer." >&2
|
||||
for SLUG in $(echo "$SELECTED_REVIEWERS" | tr ',' ' '); do
|
||||
# Per-lane prompt budget. The lane declares its own `promptBudgetKey`; `plan` resolved it,
|
||||
# applying #2797's sentinel rule (-1 = unset → fall back to the global budget; 0 legitimately
|
||||
# means "do not trim this lane"). Trimming itself stays in prompt-budget, which owns it.
|
||||
LANE_BUDGET=$(gsd_run query review-lane plan --selected "$SLUG" --run-dir "$RUN_DIR" \
|
||||
--repo-root "$REPO_ROOT" --json 2>/dev/null \
|
||||
| sed -n 's/.*"promptBudget": *\([0-9-]*\).*/\1/p' | head -1)
|
||||
PROMPT_ARG=""
|
||||
if [ -n "$LANE_BUDGET" ] && [ "$LANE_BUDGET" != "null" ] && [ "$LANE_BUDGET" -gt 0 ] 2>/dev/null; then
|
||||
TRIMMED="$RUN_DIR/gsd-review-prompt-$SLUG.md"
|
||||
if prepare_trimmed_prompt_for_reviewer "$SLUG" "$LANE_BUDGET" "$TRIMMED" \
|
||||
"$RUN_DIR/gsd-review-prompt-$SLUG.metadata.json"; then
|
||||
PROMPT_ARG="--prompt-file $TRIMMED"
|
||||
else
|
||||
echo "WARNING: prompt-budget returned unexpected exit code ${OLLAMA_EXIT} for ollama. Skipping Ollama reviewer." >&2
|
||||
fi
|
||||
OLLAMA_SKIP=1
|
||||
else
|
||||
OLLAMA_PROMPT_FILE="$OLLAMA_TRIMMED_PROMPT"
|
||||
# A budget too small for the minimum review set drops the lane just as silently as an empty
|
||||
# response used to (#2605), so leave the skip visible in the review output, not only on stderr.
|
||||
echo "$SLUG review skipped: prompt budget (${LANE_BUDGET} tokens) too small for the minimum review set." \
|
||||
> "$RUN_DIR/gsd-review-$SLUG.md"
|
||||
continue
|
||||
fi
|
||||
fi
|
||||
|
||||
if [ "$OLLAMA_SKIP" != "1" ]; then
|
||||
OLLAMA_HOST=$(gsd_run query config-get review.ollama_host --raw 2>/dev/null || echo "")
|
||||
if [ -z "$OLLAMA_HOST" ] || [ "$OLLAMA_HOST" = "null" ]; then OLLAMA_HOST="http://localhost:11434"; fi
|
||||
OLLAMA_MODEL=$(gsd_run query config-get review.models.ollama --raw 2>/dev/null || echo "")
|
||||
if [ -z "$OLLAMA_MODEL" ] || [ "$OLLAMA_MODEL" = "null" ]; then
|
||||
OLLAMA_MODEL=$(curl -s --max-time 2 "${OLLAMA_HOST}/v1/models" 2>/dev/null | jq -r '.data[0].id // "llama3"' 2>/dev/null || echo "llama3")
|
||||
fi
|
||||
# #2605: brought to parity with the LM Studio / llama.cpp legs below. Ollama
|
||||
# already emitted a non-empty stub, so it never silently vanished — but it was
|
||||
# the LEAST diagnosable leg: bare `-s` (which suppresses curl's error text as
|
||||
# well as the progress meter), stderr to /dev/null, and the response piped
|
||||
# straight into jq so the body — where an OpenAI-compatible server puts its error
|
||||
# JSON on an HTTP 4xx/5xx, with curl still exiting 0 — was discarded unread.
|
||||
OLLAMA_RESPONSE=$(jq -n --rawfile content "$OLLAMA_PROMPT_FILE" \
|
||||
--arg model "$OLLAMA_MODEL" \
|
||||
'{model: $model, messages: [{role: "user", content: $content}]}' | \
|
||||
curl -sS --max-time 120 -X POST "${OLLAMA_HOST}/v1/chat/completions" \
|
||||
-H "Content-Type: application/json" -d @- 2>{run_dir}/gsd-review-ollama.err)
|
||||
OLLAMA_CONTENT=$(echo "$OLLAMA_RESPONSE" | jq -r '.choices[0].message.content // ""' 2>/dev/null || echo "")
|
||||
case "$OLLAMA_CONTENT" in
|
||||
*[![:space:]]*) : ;;
|
||||
*) OLLAMA_CONTENT="" ;;
|
||||
esac
|
||||
if [ -n "$OLLAMA_CONTENT" ]; then
|
||||
printf '%s\n' "$OLLAMA_CONTENT" > {run_dir}/gsd-review-ollama.md
|
||||
fi
|
||||
if [ ! -s {run_dir}/gsd-review-ollama.md ]; then
|
||||
echo "Warning: Ollama returned empty content — see {run_dir}/gsd-review-ollama.md" >&2
|
||||
echo "Ollama review failed or returned empty output. stderr:" > {run_dir}/gsd-review-ollama.md
|
||||
cat {run_dir}/gsd-review-ollama.err >> {run_dir}/gsd-review-ollama.md 2>/dev/null
|
||||
echo "Raw response body:" >> {run_dir}/gsd-review-ollama.md
|
||||
printf '%s\n' "$OLLAMA_RESPONSE" >> {run_dir}/gsd-review-ollama.md
|
||||
fi
|
||||
else
|
||||
echo "Ollama review skipped: prompt budget (${OLLAMA_REVIEWER_BUDGET} tokens) too small for the minimum review set." > {run_dir}/gsd-review-ollama.md
|
||||
fi
|
||||
# One invocation, whatever the lane's transport, prompt channel, output channel or handler.
|
||||
# `--explicit` marks a lane the user NAMED: ADR-2782 D4 — not finding a lane nobody asked for is
|
||||
# normal, failing to run one somebody asked for is an error.
|
||||
gsd_run query review-lane invoke --slug "$SLUG" \
|
||||
--run-dir "$RUN_DIR" --repo-root "$REPO_ROOT" $PROMPT_ARG $EXPLICIT_FLAG --json \
|
||||
>> "$RUN_DIR/gsd-review-lane-results.jsonl"
|
||||
done
|
||||
```
|
||||
|
||||
<!-- reviewer-lane: lm_studio -->
|
||||
**LM Studio (local, OpenAI-compatible):**
|
||||
```bash
|
||||
# Resolve prompt budget for LM Studio: per-reviewer override > global default > null
|
||||
LM_STUDIO_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens_per_reviewer.lm_studio --raw 2>/dev/null || echo "null")
|
||||
# #2797: -1 = unset; 0 legitimately means "do not trim".
|
||||
if [ -z "$LM_STUDIO_REVIEWER_BUDGET" ] || [ "$LM_STUDIO_REVIEWER_BUDGET" = "null" ] || [ "$LM_STUDIO_REVIEWER_BUDGET" = "-1" ]; then
|
||||
LM_STUDIO_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens --raw 2>/dev/null || echo "null")
|
||||
fi
|
||||
Each lane leaves `{run_dir}/gsd-review-<slug>.md` — its review, or a diagnostic stub carrying the
|
||||
captured stderr (and, for an OpenAI-compatible lane, the raw response body, where such a server puts
|
||||
its error JSON on an HTTP 4xx/5xx while still exiting 0). A stub is never mistaken for a clean
|
||||
review: it keeps its "failed or returned empty output" header (#2494/#2605/#2794).
|
||||
|
||||
# Apply budget trim for LM Studio if a budget is configured
|
||||
LM_STUDIO_PROMPT_FILE="{run_dir}/gsd-review-prompt.md"
|
||||
LM_STUDIO_SKIP=0
|
||||
if [ -n "$LM_STUDIO_REVIEWER_BUDGET" ] && [ "$LM_STUDIO_REVIEWER_BUDGET" != "null" ] && [ "$LM_STUDIO_REVIEWER_BUDGET" != "0" ]; then
|
||||
LM_STUDIO_TRIMMED_PROMPT="{run_dir}/gsd-review-prompt-lm_studio.md"
|
||||
LM_STUDIO_TRIM_META="{run_dir}/gsd-review-prompt-lm_studio.metadata.json"
|
||||
prepare_trimmed_prompt_for_reviewer "lm_studio" "$LM_STUDIO_REVIEWER_BUDGET" "$LM_STUDIO_TRIMMED_PROMPT" "$LM_STUDIO_TRIM_META"
|
||||
LM_STUDIO_EXIT=$?
|
||||
if [ $LM_STUDIO_EXIT -ne 0 ]; then
|
||||
if [ $LM_STUDIO_EXIT -eq 2 ] || [ $LM_STUDIO_EXIT -eq 11 ]; then
|
||||
echo "WARNING: prompt budget for lm_studio (${LM_STUDIO_REVIEWER_BUDGET} tokens) is too small for the minimum review set. Skipping LM Studio reviewer." >&2
|
||||
else
|
||||
echo "WARNING: prompt-budget returned unexpected exit code ${LM_STUDIO_EXIT} for lm_studio. Skipping LM Studio reviewer." >&2
|
||||
fi
|
||||
LM_STUDIO_SKIP=1
|
||||
else
|
||||
LM_STUDIO_PROMPT_FILE="$LM_STUDIO_TRIMMED_PROMPT"
|
||||
fi
|
||||
fi
|
||||
|
||||
if [ "$LM_STUDIO_SKIP" != "1" ]; then
|
||||
LM_STUDIO_HOST=$(gsd_run query config-get review.lm_studio_host --raw 2>/dev/null || echo "")
|
||||
if [ -z "$LM_STUDIO_HOST" ] || [ "$LM_STUDIO_HOST" = "null" ]; then LM_STUDIO_HOST="http://localhost:1234"; fi
|
||||
LM_STUDIO_MODEL=$(gsd_run query config-get review.models.lm_studio --raw 2>/dev/null || echo "")
|
||||
if [ -z "$LM_STUDIO_MODEL" ] || [ "$LM_STUDIO_MODEL" = "null" ]; then
|
||||
LM_STUDIO_MODEL=$(curl -s --max-time 2 "${LM_STUDIO_HOST}/v1/models" 2>/dev/null | jq -r '.data[0].id // "local-model"' 2>/dev/null || echo "local-model")
|
||||
fi
|
||||
# #2605: same guard as the claude/gemini/codex legs above (#2494/#2592). Two
|
||||
# changes make a dropped lane diagnosable rather than silently omitted:
|
||||
# 1. `-sS` instead of `-s`. Plain `-s` silences curl's ERROR text too, so an
|
||||
# unreachable endpoint produced no message anywhere. `-S` restores errors
|
||||
# while keeping the progress meter off; they land in the .err sidecar.
|
||||
# 2. An `[ ! -s … ]` stub. Previously nothing was written when content was
|
||||
# empty, so the file never existed, write_reviews omitted the section, and
|
||||
# the result was indistinguishable from the reviewer never being selected.
|
||||
# The raw response body is appended too: an HTTP 4xx/5xx from an OpenAI-compatible
|
||||
# server exits 0 with the error JSON in the BODY, so stderr alone would be empty.
|
||||
LM_STUDIO_RESPONSE=$(jq -n --rawfile content "$LM_STUDIO_PROMPT_FILE" \
|
||||
--arg model "$LM_STUDIO_MODEL" \
|
||||
'{model: $model, messages: [{role: "user", content: $content}]}' | \
|
||||
curl -sS --max-time 120 -X POST "${LM_STUDIO_HOST}/v1/chat/completions" \
|
||||
-H "Content-Type: application/json" -d @- 2>{run_dir}/gsd-review-lm_studio.err)
|
||||
LM_STUDIO_ACTUAL_MODEL=$(echo "$LM_STUDIO_RESPONSE" | jq -r '.model // ""' 2>/dev/null || echo "")
|
||||
if [ -n "$LM_STUDIO_ACTUAL_MODEL" ] && [ "$LM_STUDIO_ACTUAL_MODEL" != "null" ] && [ "$LM_STUDIO_ACTUAL_MODEL" != "$LM_STUDIO_MODEL" ]; then
|
||||
echo "Warning: LM Studio served model '$LM_STUDIO_ACTUAL_MODEL' but '$LM_STUDIO_MODEL' was requested. Review may be from a different model." >&2
|
||||
fi
|
||||
LM_STUDIO_CONTENT=$(echo "$LM_STUDIO_RESPONSE" | jq -r '.choices[0].message.content // ""' 2>/dev/null || echo "")
|
||||
# A whitespace-only reply must count as empty. `[ ! -s … ]` counts BYTES, so a
|
||||
# response of " " would be written out and pass the guard as a "successful"
|
||||
# but vacuous review — the same indistinguishable-from-success outcome the guard
|
||||
# exists to prevent. Command substitution strips trailing newlines but not
|
||||
# spaces, so this case-glob is what actually closes it.
|
||||
case "$LM_STUDIO_CONTENT" in
|
||||
*[![:space:]]*) : ;;
|
||||
*) LM_STUDIO_CONTENT="" ;;
|
||||
esac
|
||||
# printf, not echo: `echo "$VAR"` swallows a value that is exactly `-n`/`-e`/`-E`
|
||||
# and would write 0 bytes, misclassifying a real reply as empty. Same idiom the
|
||||
# OpenCode leg already uses above.
|
||||
if [ -n "$LM_STUDIO_CONTENT" ]; then
|
||||
printf '%s\n' "$LM_STUDIO_CONTENT" > {run_dir}/gsd-review-lm_studio.md
|
||||
fi
|
||||
if [ ! -s {run_dir}/gsd-review-lm_studio.md ]; then
|
||||
echo "Warning: LM Studio returned empty content — see {run_dir}/gsd-review-lm_studio.md" >&2
|
||||
echo "LM Studio review failed or returned empty output. stderr:" > {run_dir}/gsd-review-lm_studio.md
|
||||
cat {run_dir}/gsd-review-lm_studio.err >> {run_dir}/gsd-review-lm_studio.md 2>/dev/null
|
||||
echo "Raw response body:" >> {run_dir}/gsd-review-lm_studio.md
|
||||
printf '%s\n' "$LM_STUDIO_RESPONSE" >> {run_dir}/gsd-review-lm_studio.md
|
||||
fi
|
||||
else
|
||||
# A budget skip drops the lane just as silently as an empty response did: no
|
||||
# file, so write_reviews omits the section entirely. Leave the same diagnosable
|
||||
# stub so the skip is visible in the review output, not only on stderr (#2605).
|
||||
echo "LM Studio review skipped: prompt budget (${LM_STUDIO_REVIEWER_BUDGET} tokens) too small for the minimum review set." > {run_dir}/gsd-review-lm_studio.md
|
||||
fi
|
||||
```
|
||||
|
||||
<!-- reviewer-lane: llama_cpp -->
|
||||
**llama.cpp (local, OpenAI-compatible):**
|
||||
```bash
|
||||
# Resolve prompt budget for llama.cpp: per-reviewer override > global default > null
|
||||
LLAMA_CPP_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens_per_reviewer.llama_cpp --raw 2>/dev/null || echo "null")
|
||||
# #2797: -1 = unset; 0 legitimately means "do not trim".
|
||||
if [ -z "$LLAMA_CPP_REVIEWER_BUDGET" ] || [ "$LLAMA_CPP_REVIEWER_BUDGET" = "null" ] || [ "$LLAMA_CPP_REVIEWER_BUDGET" = "-1" ]; then
|
||||
LLAMA_CPP_REVIEWER_BUDGET=$(gsd_run query config-get review.max_prompt_tokens --raw 2>/dev/null || echo "null")
|
||||
fi
|
||||
|
||||
# Apply budget trim for llama.cpp if a budget is configured
|
||||
LLAMA_CPP_PROMPT_FILE="{run_dir}/gsd-review-prompt.md"
|
||||
LLAMA_CPP_SKIP=0
|
||||
if [ -n "$LLAMA_CPP_REVIEWER_BUDGET" ] && [ "$LLAMA_CPP_REVIEWER_BUDGET" != "null" ] && [ "$LLAMA_CPP_REVIEWER_BUDGET" != "0" ]; then
|
||||
LLAMA_CPP_TRIMMED_PROMPT="{run_dir}/gsd-review-prompt-llama_cpp.md"
|
||||
LLAMA_CPP_TRIM_META="{run_dir}/gsd-review-prompt-llama_cpp.metadata.json"
|
||||
prepare_trimmed_prompt_for_reviewer "llama_cpp" "$LLAMA_CPP_REVIEWER_BUDGET" "$LLAMA_CPP_TRIMMED_PROMPT" "$LLAMA_CPP_TRIM_META"
|
||||
LLAMA_CPP_EXIT=$?
|
||||
if [ $LLAMA_CPP_EXIT -ne 0 ]; then
|
||||
if [ $LLAMA_CPP_EXIT -eq 2 ] || [ $LLAMA_CPP_EXIT -eq 11 ]; then
|
||||
echo "WARNING: prompt budget for llama_cpp (${LLAMA_CPP_REVIEWER_BUDGET} tokens) is too small for the minimum review set. Skipping llama.cpp reviewer." >&2
|
||||
else
|
||||
echo "WARNING: prompt-budget returned unexpected exit code ${LLAMA_CPP_EXIT} for llama_cpp. Skipping llama.cpp reviewer." >&2
|
||||
fi
|
||||
LLAMA_CPP_SKIP=1
|
||||
else
|
||||
LLAMA_CPP_PROMPT_FILE="$LLAMA_CPP_TRIMMED_PROMPT"
|
||||
fi
|
||||
fi
|
||||
|
||||
if [ "$LLAMA_CPP_SKIP" != "1" ]; then
|
||||
LLAMA_CPP_HOST=$(gsd_run query config-get review.llama_cpp_host --raw 2>/dev/null || echo "")
|
||||
if [ -z "$LLAMA_CPP_HOST" ] || [ "$LLAMA_CPP_HOST" = "null" ]; then LLAMA_CPP_HOST="http://localhost:8080"; fi
|
||||
LLAMA_CPP_MODEL=$(gsd_run query config-get review.models.llama_cpp --raw 2>/dev/null || echo "")
|
||||
if [ -z "$LLAMA_CPP_MODEL" ] || [ "$LLAMA_CPP_MODEL" = "null" ]; then
|
||||
LLAMA_CPP_MODEL=$(curl -s --max-time 2 "${LLAMA_CPP_HOST}/v1/models" 2>/dev/null | jq -r '.data[0].id // "local-model"' 2>/dev/null || echo "local-model")
|
||||
fi
|
||||
# #2605: same guard as the LM Studio leg above. The response is captured to a
|
||||
# variable FIRST rather than piped straight into jq — piping discarded the raw
|
||||
# body, which is exactly where an OpenAI-compatible server puts its error JSON on
|
||||
# an HTTP 4xx/5xx (curl still exits 0), leaving nothing to diagnose.
|
||||
LLAMA_CPP_RESPONSE=$(jq -n --rawfile content "$LLAMA_CPP_PROMPT_FILE" \
|
||||
--arg model "$LLAMA_CPP_MODEL" \
|
||||
'{model: $model, messages: [{role: "user", content: $content}]}' | \
|
||||
curl -sS --max-time 120 -X POST "${LLAMA_CPP_HOST}/v1/chat/completions" \
|
||||
-H "Content-Type: application/json" -d @- 2>{run_dir}/gsd-review-llama_cpp.err)
|
||||
LLAMA_CPP_CONTENT=$(echo "$LLAMA_CPP_RESPONSE" | jq -r '.choices[0].message.content // ""' 2>/dev/null || echo "")
|
||||
# Whitespace-only reply counts as empty; printf not echo. See the LM Studio leg
|
||||
# above for why both are required.
|
||||
case "$LLAMA_CPP_CONTENT" in
|
||||
*[![:space:]]*) : ;;
|
||||
*) LLAMA_CPP_CONTENT="" ;;
|
||||
esac
|
||||
if [ -n "$LLAMA_CPP_CONTENT" ]; then
|
||||
printf '%s\n' "$LLAMA_CPP_CONTENT" > {run_dir}/gsd-review-llama_cpp.md
|
||||
fi
|
||||
if [ ! -s {run_dir}/gsd-review-llama_cpp.md ]; then
|
||||
echo "Warning: llama.cpp returned empty content — see {run_dir}/gsd-review-llama_cpp.md" >&2
|
||||
echo "llama.cpp review failed or returned empty output. stderr:" > {run_dir}/gsd-review-llama_cpp.md
|
||||
cat {run_dir}/gsd-review-llama_cpp.err >> {run_dir}/gsd-review-llama_cpp.md 2>/dev/null
|
||||
echo "Raw response body:" >> {run_dir}/gsd-review-llama_cpp.md
|
||||
printf '%s\n' "$LLAMA_CPP_RESPONSE" >> {run_dir}/gsd-review-llama_cpp.md
|
||||
fi
|
||||
else
|
||||
echo "llama.cpp review skipped: prompt budget (${LLAMA_CPP_REVIEWER_BUDGET} tokens) too small for the minimum review set." > {run_dir}/gsd-review-llama_cpp.md
|
||||
fi
|
||||
```
|
||||
|
||||
If a CLI or local server fails, log the error and continue with remaining reviewers.
|
||||
A lane that will not run reports a typed reason rather than an empty file — `missing_binary`,
|
||||
`probe_failed`, `probe_timeout`, `missing_required_binary`, `host_unreachable`,
|
||||
`egress_host_changed`, `unknown_handler`, `budget_too_small`. **`egress_host_changed` means the lane
|
||||
was consented to send plans to one destination and `.planning/config.json` now names another; it is
|
||||
blocked, not silently redirected** (ADR-2782 D5).
|
||||
|
||||
Display progress:
|
||||
```
|
||||
@@ -969,83 +434,26 @@ trimmed_reviewers: # only present if at least one reviewer was trimmed
|
||||
|
||||
# Cross-AI Plan Review — Phase {N}
|
||||
|
||||
## Gemini Review
|
||||
<!-- Sections are RENDERED from each lane's declared `reviewsSection`, in descriptor order.
|
||||
There is deliberately no hardcoded per-reviewer heading list here any more: a hand-maintained
|
||||
list is exactly the drift #2781 was filed about, and it silently disagreed with the roster.
|
||||
`gsd_run query review-lane sections --selected "$SELECTED_REVIEWERS"` emits
|
||||
`<slug><TAB><reviewsSection>` in order; for each row, emit:
|
||||
|
||||
{gemini review content}
|
||||
## <reviewsSection> Review
|
||||
|
||||
{contents of {run_dir}/gsd-review-<slug>.md}
|
||||
|
||||
---
|
||||
|
||||
## Claude Review
|
||||
Two headings must NOT be generated from this list, because they are not lanes:
|
||||
* `## <Adapter> Review (<instance>)` — an ADR-1517 reviewer INSTANCE resolves THROUGH a lane
|
||||
and is rendered from the instance list, not the lane list (ADR-2782 D8).
|
||||
* `## Consensus Summary` — not a review section at all.
|
||||
|
||||
{claude review content}
|
||||
|
||||
---
|
||||
|
||||
## Codex Review
|
||||
|
||||
{codex review content}
|
||||
|
||||
---
|
||||
|
||||
## CodeRabbit Review
|
||||
|
||||
{coderabbit review content}
|
||||
|
||||
---
|
||||
|
||||
## OpenCode Review
|
||||
|
||||
{opencode review content}
|
||||
|
||||
---
|
||||
|
||||
## OpenCode Review (opencode-deepseek)
|
||||
|
||||
{opencode-deepseek instance review content — only present when this instance was selected}
|
||||
|
||||
---
|
||||
|
||||
## OpenCode Review (opencode-mimo)
|
||||
|
||||
{opencode-mimo instance review content — only present when this instance was selected}
|
||||
|
||||
---
|
||||
|
||||
## Qwen Review
|
||||
|
||||
{qwen review content}
|
||||
|
||||
---
|
||||
|
||||
## Cursor Review
|
||||
|
||||
{cursor review content}
|
||||
|
||||
---
|
||||
|
||||
## Antigravity Review
|
||||
|
||||
{antigravity review content}
|
||||
|
||||
---
|
||||
|
||||
## Ollama Review
|
||||
|
||||
{ollama review content}
|
||||
|
||||
---
|
||||
|
||||
## LM Studio Review
|
||||
|
||||
{lm_studio review content}
|
||||
|
||||
---
|
||||
|
||||
## llama.cpp Review
|
||||
|
||||
{llama_cpp review content}
|
||||
|
||||
---
|
||||
A lane whose `evidenceClass` is `diff-only` (CodeRabbit) carries its caveat from data: it never
|
||||
received the source-grounding prompt, so its verdict is folded in as a diff observation and is
|
||||
not weighted as a grounded plan review. -->
|
||||
|
||||
## Consensus Summary
|
||||
|
||||
|
||||
@@ -137,6 +137,22 @@ interface ConsentRecord {
|
||||
*/
|
||||
contentHash: string;
|
||||
consentedAt: string;
|
||||
/**
|
||||
* ADR-2782 D5 rule 1 (#2799): the egress destination this consent was granted against, for an
|
||||
* `openai-http` reviewer lane. Absent for every other capability, and for records written before
|
||||
* the rule existed.
|
||||
*
|
||||
* OPTIONAL BY DESIGN, and `isValidConsentRecord` deliberately does not require it. Making it
|
||||
* mandatory would invalidate every consent record already on disk, forcing re-consent across all
|
||||
* installed capabilities — precisely the spurious re-consent storm D4 rule 5 forbids.
|
||||
*
|
||||
* Deliberately NOT part of `disclosureSignature`: that signature is computed by BOTH the loader
|
||||
* and the lifecycle so the two can never drift, and the loader has no config resolver. Folding a
|
||||
* config-derived value in would make them compute different signatures for the same manifest and
|
||||
* re-prompt forever. The signature binds the manifest; this field binds the destination; Phase 5b
|
||||
* re-resolves and compares at invocation, which is where D5 rule 4 puts the check.
|
||||
*/
|
||||
reviewerHost?: string;
|
||||
}
|
||||
|
||||
interface ConsentStore {
|
||||
@@ -455,6 +471,9 @@ function isValidConsentRecord(rec: unknown): rec is ConsentRecord {
|
||||
// match a recomputed hash and is treated as invalid (fail closed).
|
||||
if (typeof r['contentHash'] !== 'string' || !r['contentHash']) return false;
|
||||
if (typeof r['consentedAt'] !== 'string' || !r['consentedAt']) return false;
|
||||
// Optional (see ConsentRecord.reviewerHost): absent is valid and is the common case. Present but
|
||||
// non-string is a corrupt record — reject rather than silently comparing against a non-host.
|
||||
if (r['reviewerHost'] !== undefined && typeof r['reviewerHost'] !== 'string') return false;
|
||||
return true;
|
||||
}
|
||||
|
||||
@@ -698,8 +717,9 @@ function recordProjectConsent(args: {
|
||||
integrity: string;
|
||||
disclosureSignature: string;
|
||||
contentHash: string;
|
||||
reviewerHost?: string;
|
||||
}): void {
|
||||
const { gsdHome, projectRoot, id, integrity, disclosureSignature, contentHash } = args;
|
||||
const { gsdHome, projectRoot, id, integrity, disclosureSignature, contentHash, reviewerHost } = args;
|
||||
if (isUnsafeCapabilityId(id)) {
|
||||
throw new Error(
|
||||
`Invalid capability id "${String(id)}": must match /^[a-z][a-z0-9-]*$/ (kebab-case, lowercase). ` +
|
||||
@@ -744,6 +764,9 @@ function recordProjectConsent(args: {
|
||||
disclosureSignature,
|
||||
contentHash,
|
||||
consentedAt: new Date().toISOString(),
|
||||
// Written only when the capability actually declares an egress destination, so a spawn lane
|
||||
// or a non-lane capability keeps a byte-identical record shape.
|
||||
...(typeof reviewerHost === 'string' && reviewerHost ? { reviewerHost } : {}),
|
||||
};
|
||||
writeConsentStore(gsdHome, store);
|
||||
} finally {
|
||||
@@ -786,6 +809,39 @@ function revokeProjectConsent(args: { gsdHome?: string; projectRoot: string; id:
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The egress destination a capability's consent was granted against, or `undefined`.
|
||||
*
|
||||
* ADR-2782 D5 rule 4 (#2799) — read on the INVOCATION path, not only at install, because
|
||||
* `hostConfigKey` names a key in `.planning/config.json`: the one consent-bound value that lives
|
||||
* outside the SHA-pinned bundle and can be changed by an ordinary pull request with no re-install
|
||||
* and no integrity check.
|
||||
*
|
||||
* `undefined` means "nothing to compare", and the caller MUST treat that as allow, not deny. Two
|
||||
* legitimate ways to get it: the capability has no consent record at all (first-party lanes ship
|
||||
* inside the SHA-pinned distribution and are never consent-gated), or the record predates this
|
||||
* field. Denying on absence would break every existing local-model user on upgrade.
|
||||
*
|
||||
* Non-throwing: a missing, corrupt, or oversized store yields `undefined` like any other absence.
|
||||
*/
|
||||
function readConsentedReviewerHost(args: {
|
||||
gsdHome?: string;
|
||||
projectRoot: string;
|
||||
id: string;
|
||||
}): string | undefined {
|
||||
const { gsdHome, projectRoot, id } = args;
|
||||
if (isUnsafeCapabilityId(id)) return undefined;
|
||||
try {
|
||||
const store = readConsentStore(gsdHome);
|
||||
const key = consentKey(realpathProject(projectRoot), id);
|
||||
if (!Object.prototype.hasOwnProperty.call(store.records, key)) return undefined;
|
||||
const host = store.records[key].reviewerHost;
|
||||
return typeof host === 'string' && host ? host : undefined;
|
||||
} catch {
|
||||
return undefined;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* #1459 finding 2 (round 6): TEST-ONLY — override the cumulative bundle entry-count cap and return a
|
||||
* restore() that resets it to the production default. Lets a test prove the streaming walk fails closed
|
||||
@@ -805,6 +861,7 @@ export = {
|
||||
consentStorePath,
|
||||
bundleContentHash,
|
||||
readConsentStore,
|
||||
readConsentedReviewerHost,
|
||||
hasProjectConsent,
|
||||
recordProjectConsent,
|
||||
revokeProjectConsent,
|
||||
|
||||
@@ -64,7 +64,7 @@ const trustMod = require('./capability-trust.cjs') as {
|
||||
signatureForManifest: (manifest: Record<string, unknown>, stagedDir?: string) => string;
|
||||
};
|
||||
const consentMod = require('./capability-consent.cjs') as {
|
||||
recordProjectConsent: (args: { gsdHome?: string; projectRoot: string; id: string; integrity: string; disclosureSignature: string; contentHash: string }) => void;
|
||||
recordProjectConsent: (args: { gsdHome?: string; projectRoot: string; id: string; integrity: string; disclosureSignature: string; contentHash: string; reviewerHost?: string }) => void;
|
||||
revokeProjectConsent: (args: { gsdHome?: string; projectRoot: string; id: string }) => void;
|
||||
/** #1459 CB-1/CB-2: recompute the full-bundle content hash (the consent security binding). */
|
||||
bundleContentHash: (capDir: string) => string;
|
||||
@@ -826,6 +826,61 @@ function warnIfConsentSkipped(opts: LifecycleOptions, id: string): void {
|
||||
* Best-effort: a consent-store write failure must not turn a successful install/upgrade into a
|
||||
* failure (the bundle is already committed) — it is surfaced as a warning, not a throw.
|
||||
*/
|
||||
/**
|
||||
* Resolve an `openai-http` reviewer lane's declared `hostConfigKey` to the destination it currently
|
||||
* names, or `undefined` when this capability is not such a lane.
|
||||
*
|
||||
* Falls back to the lane's declared `defaultHost` when the key is unset, because that is exactly
|
||||
* what the invocation path will do — binding the config value while the runtime uses the default
|
||||
* would guarantee a mismatch on the very first review.
|
||||
*
|
||||
* Non-throwing: consent binding is best-effort and must never turn a successful install into a
|
||||
* failure. An unresolvable host simply records nothing, which reads as "not bound" and allows.
|
||||
*/
|
||||
function resolveReviewerEgressHost(
|
||||
opts: LifecycleOptions,
|
||||
manifest: Record<string, unknown>,
|
||||
): string | undefined {
|
||||
try {
|
||||
const reviewer = manifest['reviewer'];
|
||||
if (reviewer === null || typeof reviewer !== 'object' || Array.isArray(reviewer)) return undefined;
|
||||
const r = reviewer as Record<string, unknown>;
|
||||
if (r['transport'] !== 'openai-http') return undefined;
|
||||
const invoke = r['invoke'];
|
||||
if (invoke === null || typeof invoke !== 'object') return undefined;
|
||||
const inv = invoke as Record<string, unknown>;
|
||||
const key = typeof inv['hostConfigKey'] === 'string' ? inv['hostConfigKey'] : '';
|
||||
const fallback = typeof inv['defaultHost'] === 'string' ? inv['defaultHost'] : '';
|
||||
|
||||
let configured = '';
|
||||
if (key) {
|
||||
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
||||
const cfgLoader = require('./config-loader.cjs') as {
|
||||
loadConfigResolved?: (cwd: string) => { config?: Record<string, unknown> };
|
||||
};
|
||||
const root = projectRootMod.consentProjectRoot(opts.runtimeDir);
|
||||
const cfg = cfgLoader.loadConfigResolved ? (cfgLoader.loadConfigResolved(root).config ?? {}) : {};
|
||||
let cur: unknown = cfg;
|
||||
for (const part of key.split('.')) {
|
||||
if (cur === null || typeof cur !== 'object') { cur = undefined; break; }
|
||||
cur = Object.prototype.hasOwnProperty.call(cur, part)
|
||||
? (cur as Record<string, unknown>)[part]
|
||||
: undefined;
|
||||
}
|
||||
if (typeof cur === 'string') configured = cur.trim();
|
||||
}
|
||||
|
||||
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
||||
const { normalizeHost } = require('./review-lane-invocation.cjs') as {
|
||||
normalizeHost: (s: string) => string;
|
||||
};
|
||||
const resolved = normalizeHost(configured || fallback);
|
||||
return resolved || undefined;
|
||||
} catch {
|
||||
return undefined;
|
||||
}
|
||||
}
|
||||
|
||||
function bindProjectConsent(opts: LifecycleOptions, id: string, integrity: string, manifest: Record<string, unknown>): void {
|
||||
// #1459 IC-07: a project-scope op WITHOUT a consent store cannot bind — warn (then nothing to do).
|
||||
if (!shouldBindConsent(opts)) {
|
||||
@@ -843,6 +898,11 @@ function bindProjectConsent(opts: LifecycleOptions, id: string, integrity: strin
|
||||
integrity,
|
||||
disclosureSignature: trustMod.signatureForManifest(manifest),
|
||||
contentHash: consentMod.bundleContentHash(capDir(opts.runtimeDir, id)),
|
||||
// ADR-2782 D5 rule 1 (#2799): bind the RESOLVED egress destination, not merely the config key
|
||||
// that names it. The key lives in `.planning/config.json`, outside the SHA-pinned bundle, so
|
||||
// without this the user consents to "wherever that key points" — a promise the bundle hash
|
||||
// cannot keep. Phase 5b re-resolves and compares at invocation (rule 4).
|
||||
reviewerHost: resolveReviewerEgressHost(opts, manifest),
|
||||
});
|
||||
} catch (err) {
|
||||
// #1459 IC-05/WIN-2: a consent-store write failure (read-only/UNC/NFS store) must NOT turn an
|
||||
|
||||
@@ -65,8 +65,21 @@ export type EffortChannel = 'none' | 'argv' | 'env';
|
||||
/** ADR-2782 D2 — CodeRabbit reviews a diff, not the source tree (`review.md:367`). */
|
||||
export type EvidenceClass = 'source-grounded' | 'diff-only';
|
||||
|
||||
/** ADR-2782 D6 — closed enum of first-party imperative modules. Ported in Phase 5b. */
|
||||
export type LaneHandler = null | 'antigravity' | 'openai-compatible';
|
||||
/**
|
||||
* ADR-2782 D6 — closed enum of first-party imperative modules. Ported in Phase 5b.
|
||||
*
|
||||
* `'opencode'` was added by Phase 5b (#2799) as an additive widening, forced by a lane that ships
|
||||
* today. Phase 1 declared `opencode.handler = null`, but the lane's review is RECONSTRUCTED from
|
||||
* assistant `text` parts of a `--format json` stream — `--format json` is the primary invocation,
|
||||
* not a fallback, because the default formatter drops the text when the agent ends its turn with no
|
||||
* final message (#1936). A data-driven `outputChannel: 'stdout'` copy would write the raw JSON
|
||||
* envelope into REVIEWS.md as the review, straight back into #1936.
|
||||
*
|
||||
* The alternative — an `outputChannel: 'json-parts'` member — was rejected: it pushes a parsing
|
||||
* language into the descriptor and would want a selector expression next. D6's whole point is that
|
||||
* divergence escapes to NAMED FIRST-PARTY CODE rather than accreting inside data.
|
||||
*/
|
||||
export type LaneHandler = null | 'antigravity' | 'openai-compatible' | 'opencode';
|
||||
|
||||
/**
|
||||
* What a lane does when it produces no usable output.
|
||||
@@ -83,8 +96,38 @@ export type LaneProbe =
|
||||
| { kind: 'command-capability'; binary: string; needle: string; timeoutMs: number }
|
||||
| { kind: 'http-reachable'; hostConfigKey: string; path: string; timeoutMs: number };
|
||||
|
||||
/**
|
||||
* Argv placeholders (Phase 5b, #2799).
|
||||
*
|
||||
* `args` is an argv TEMPLATE, not a prefix. The injected pieces — model, effort, output file,
|
||||
* argv-borne prompt — do not all go in the same place, and no positional rule expresses that:
|
||||
* `codex` injects the model in the MIDDLE (after the `exec --ephemeral` subcommand) and the output
|
||||
* file later still, while `gemini` injects the model first and five lanes end with a bare `-` that
|
||||
* must stay last. Splicing by position silently produced
|
||||
* `codex --model M -o F exec --ephemeral …`, which is not a valid codex invocation.
|
||||
*
|
||||
* So each lane declares WHERE each piece goes. A placeholder expands to zero or more argv elements
|
||||
* and vanishes when it has nothing to contribute (no model configured, no effort channel, prompt on
|
||||
* stdin), which is what lets one template serve the configured and unconfigured cases.
|
||||
*
|
||||
* This is a closed four-member vocabulary with no expressions, no nesting and no conditionals — a
|
||||
* placeholder set, deliberately not a template language. The moment it needs a conditional, the
|
||||
* lane wants a `handler` instead (D6).
|
||||
*/
|
||||
export const ARGV_PLACEHOLDER = Object.freeze({
|
||||
/** `modelArg` + the resolved model, or nothing. */
|
||||
MODEL: '{{model}}',
|
||||
/** The host's effort argv, or nothing unless `effortChannel` is `argv`. */
|
||||
EFFORT: '{{effort}}',
|
||||
/** `outputArg` + the review path, or nothing unless `outputChannel` is `file-arg`. */
|
||||
OUTPUT: '{{output}}',
|
||||
/** The argv-borne prompt, or nothing unless `promptChannel` is `argv`/`argv-file-ref`. */
|
||||
PROMPT: '{{prompt}}',
|
||||
} as const);
|
||||
|
||||
export interface SpawnInvoke {
|
||||
binary: string;
|
||||
/** Argv template — see `ARGV_PLACEHOLDER`. Order here is the order the tool receives. */
|
||||
args: ReadonlyArray<string>;
|
||||
promptChannel: PromptChannel;
|
||||
outputChannel: OutputChannel;
|
||||
@@ -98,8 +141,19 @@ export interface SpawnInvoke {
|
||||
export interface HttpInvoke {
|
||||
/** Dotted config key holding the base URL. */
|
||||
hostConfigKey: string;
|
||||
/**
|
||||
* Base URL used when `hostConfigKey` resolves to empty.
|
||||
*
|
||||
* Added by Phase 5b (#2799). Phase 4 federated every `*_host` key with a default of `""`, so the
|
||||
* REAL fallback (`http://localhost:11434` and friends) only ever existed inside the bash leg. A
|
||||
* data-driven lane with an empty host and no declared default would POST to a garbage URL, so
|
||||
* the fallback has to be declared somewhere — and the lane is the only thing that knows it.
|
||||
*/
|
||||
defaultHost: string;
|
||||
path: string;
|
||||
modelDiscovery: 'none' | 'first-from-models-endpoint';
|
||||
/** Model used when neither config nor discovery yields one. */
|
||||
fallbackModel: string;
|
||||
effortChannel: 'none';
|
||||
}
|
||||
|
||||
@@ -123,6 +177,20 @@ interface ReviewerLaneCommon {
|
||||
requiresBinaries: ReadonlyArray<string>;
|
||||
/** Dotted config key for per-lane prompt trimming, or null. */
|
||||
promptBudgetKey: string | null;
|
||||
/**
|
||||
* Dotted config key holding this lane's model override, or null when the lane accepts none.
|
||||
*
|
||||
* Added by Phase 5b (#2799). Phase 1 left the model key IMPLICIT (`review.models.<slug>`), and
|
||||
* that convention is wrong for a lane that ships today: `antigravity`'s slug is `antigravity` but
|
||||
* its key is `review.models.agy` (`review.md:291`, and Phase 4 federated it under that exact
|
||||
* name). Resolving by slug would look up `review.models.antigravity`, miss, and silently ignore a
|
||||
* configured model — the pinned-model escape hatch #2073 added precisely so a 404ing default can
|
||||
* be overridden.
|
||||
*
|
||||
* Declared rather than derived, for the same reason `promptBudgetKey` is: the key is data about
|
||||
* the lane, and a naming convention that one shipped lane already breaks is not a contract.
|
||||
*/
|
||||
modelConfigKey: string | null;
|
||||
handler: LaneHandler;
|
||||
}
|
||||
|
||||
@@ -153,11 +221,10 @@ const SPAWN_STDIN_STDOUT = {
|
||||
} as const;
|
||||
|
||||
/**
|
||||
* The eleven lanes shipped today, in `write_reviews` order.
|
||||
* The twelve declared lanes, in `write_reviews` order.
|
||||
*
|
||||
* `kimi-code` is deliberately absent: it is net-new with no leg, and ADR-2782
|
||||
* lands it in Phase 5b alongside the iteration that can invoke it. Declaring it
|
||||
* here would make it selectable but not invocable.
|
||||
* `kimi-code` joined in Phase 5b (#2799, closes #2718) — ADR-2782's phase table lands it here
|
||||
* rather than in 5a precisely so it arrives together with the iteration that can invoke it.
|
||||
*/
|
||||
export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
{
|
||||
@@ -167,7 +234,7 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
probe: { kind: 'command-exists', binary: 'gemini' },
|
||||
invoke: {
|
||||
binary: 'gemini',
|
||||
args: ['-p', '-'],
|
||||
args: ['{{model}}', '-p', '-'],
|
||||
...SPAWN_STDIN_STDOUT,
|
||||
modelArg: '-m',
|
||||
effortChannel: 'none',
|
||||
@@ -178,6 +245,7 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
evidenceClass: 'source-grounded',
|
||||
requiresBinaries: [],
|
||||
promptBudgetKey: null,
|
||||
modelConfigKey: 'review.models.gemini',
|
||||
handler: null,
|
||||
},
|
||||
{
|
||||
@@ -189,7 +257,7 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
probe: { kind: 'command-exists', binary: 'claude' },
|
||||
invoke: {
|
||||
binary: 'claude',
|
||||
args: ['-p', '-'],
|
||||
args: ['{{model}}', '{{effort}}', '-p', '-'],
|
||||
...SPAWN_STDIN_STDOUT,
|
||||
modelArg: '--model',
|
||||
effortChannel: 'argv',
|
||||
@@ -200,6 +268,7 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
evidenceClass: 'source-grounded',
|
||||
requiresBinaries: [],
|
||||
promptBudgetKey: null,
|
||||
modelConfigKey: 'review.models.claude',
|
||||
handler: null,
|
||||
},
|
||||
{
|
||||
@@ -213,7 +282,8 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
probe: { kind: 'command-exists', binary: 'codex' },
|
||||
invoke: {
|
||||
binary: 'codex',
|
||||
args: ['exec', '--ephemeral', '--skip-git-repo-check', '-'],
|
||||
// Leg order exactly: codex exec --ephemeral --model M $EFFORT --skip-git-repo-check -o F -
|
||||
args: ['exec', '--ephemeral', '{{model}}', '{{effort}}', '--skip-git-repo-check', '{{output}}', '-'],
|
||||
promptChannel: 'stdin',
|
||||
outputChannel: 'file-arg',
|
||||
outputArg: '-o',
|
||||
@@ -226,6 +296,7 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
evidenceClass: 'source-grounded',
|
||||
requiresBinaries: [],
|
||||
promptBudgetKey: null,
|
||||
modelConfigKey: 'review.models.codex',
|
||||
handler: null,
|
||||
},
|
||||
{
|
||||
@@ -250,6 +321,8 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
evidenceClass: 'diff-only',
|
||||
requiresBinaries: [],
|
||||
promptBudgetKey: null,
|
||||
// Accepts no model flag at all (review.md:367) — not merely "none configured".
|
||||
modelConfigKey: null,
|
||||
handler: null,
|
||||
},
|
||||
{
|
||||
@@ -262,7 +335,7 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
probe: { kind: 'command-exists', binary: 'opencode' },
|
||||
invoke: {
|
||||
binary: 'opencode',
|
||||
args: ['run', '--format', 'json', '-'],
|
||||
args: ['run', '{{model}}', '{{effort}}', '--format', 'json', '-'],
|
||||
...SPAWN_STDIN_STDOUT,
|
||||
modelArg: '--model',
|
||||
effortChannel: 'argv',
|
||||
@@ -271,9 +344,14 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
emptyOutput: 'stub-with-stderr',
|
||||
reviewsSection: 'OpenCode',
|
||||
evidenceClass: 'source-grounded',
|
||||
requiresBinaries: ['jq'],
|
||||
// Phase 5b: the handler reconstructs from the JSON stream with JSON.parse, so `jq` — absent on
|
||||
// stock Windows/Git-Bash (#2589) — is no longer a prerequisite for this lane.
|
||||
requiresBinaries: [],
|
||||
promptBudgetKey: null,
|
||||
handler: null,
|
||||
modelConfigKey: 'review.models.opencode',
|
||||
// Phase 5b (#2799): was `null`. The review is REBUILT from assistant `text` parts; a plain
|
||||
// stdout copy would write the raw JSON envelope as the review (#1936). See LaneHandler.
|
||||
handler: 'opencode',
|
||||
},
|
||||
{
|
||||
slug: 'qwen',
|
||||
@@ -293,6 +371,7 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
evidenceClass: 'source-grounded',
|
||||
requiresBinaries: [],
|
||||
promptBudgetKey: null,
|
||||
modelConfigKey: null,
|
||||
handler: null,
|
||||
},
|
||||
{
|
||||
@@ -305,7 +384,7 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
probe: { kind: 'command-exists', binary: 'cursor-agent' },
|
||||
invoke: {
|
||||
binary: 'cursor-agent',
|
||||
args: ['-p', '--mode', 'ask', '--trust', '--output-format', 'text'],
|
||||
args: ['-p', '--mode', 'ask', '--trust', '--output-format', 'text', '{{prompt}}'],
|
||||
promptChannel: 'argv-file-ref',
|
||||
outputChannel: 'stdout',
|
||||
modelArg: null,
|
||||
@@ -317,6 +396,7 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
evidenceClass: 'source-grounded',
|
||||
requiresBinaries: [],
|
||||
promptBudgetKey: null,
|
||||
modelConfigKey: null,
|
||||
handler: null,
|
||||
},
|
||||
{
|
||||
@@ -330,7 +410,7 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
probe: { kind: 'command-exists', binary: 'agy' },
|
||||
invoke: {
|
||||
binary: 'agy',
|
||||
args: ['--print-timeout', '540s', '-p'],
|
||||
args: ['--print-timeout', '540s', '{{model}}', '-p', '{{prompt}}'],
|
||||
promptChannel: 'argv-file-ref',
|
||||
outputChannel: 'stdout',
|
||||
modelArg: '--model',
|
||||
@@ -340,8 +420,12 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
emptyOutput: 'handler-owned',
|
||||
reviewsSection: 'Antigravity',
|
||||
evidenceClass: 'source-grounded',
|
||||
requiresBinaries: ['jq'],
|
||||
// Phase 5b: the handler reads the transcript with JSON.parse per line, not `jq`.
|
||||
requiresBinaries: [],
|
||||
promptBudgetKey: null,
|
||||
// NOT `review.models.antigravity` — the shipped key is `review.models.agy` (review.md:291) and
|
||||
// Phase 4 federated it under that name. This lane is why the key is declared, not derived.
|
||||
modelConfigKey: 'review.models.agy',
|
||||
handler: 'antigravity',
|
||||
},
|
||||
{
|
||||
@@ -356,16 +440,21 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
},
|
||||
invoke: {
|
||||
hostConfigKey: 'review.ollama_host',
|
||||
defaultHost: 'http://localhost:11434',
|
||||
path: '/v1/chat/completions',
|
||||
modelDiscovery: 'first-from-models-endpoint',
|
||||
fallbackModel: 'llama3',
|
||||
effortChannel: 'none',
|
||||
},
|
||||
timeoutFloorMs: 120_000,
|
||||
emptyOutput: 'stub-with-stderr',
|
||||
reviewsSection: 'Ollama',
|
||||
evidenceClass: 'source-grounded',
|
||||
requiresBinaries: ['jq'],
|
||||
// Phase 5b: the openai-compatible handler speaks HTTP directly and parses with JSON.parse, so
|
||||
// neither `jq` nor `curl` is a prerequisite any more (#2589's stated Windows hazard).
|
||||
requiresBinaries: [],
|
||||
promptBudgetKey: 'review.max_prompt_tokens_per_reviewer.ollama',
|
||||
modelConfigKey: 'review.models.ollama',
|
||||
handler: 'openai-compatible',
|
||||
},
|
||||
{
|
||||
@@ -380,16 +469,19 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
},
|
||||
invoke: {
|
||||
hostConfigKey: 'review.lm_studio_host',
|
||||
defaultHost: 'http://localhost:1234',
|
||||
path: '/v1/chat/completions',
|
||||
modelDiscovery: 'first-from-models-endpoint',
|
||||
fallbackModel: 'local-model',
|
||||
effortChannel: 'none',
|
||||
},
|
||||
timeoutFloorMs: 120_000,
|
||||
emptyOutput: 'stub-with-stderr',
|
||||
reviewsSection: 'LM Studio',
|
||||
evidenceClass: 'source-grounded',
|
||||
requiresBinaries: ['jq'],
|
||||
requiresBinaries: [],
|
||||
promptBudgetKey: 'review.max_prompt_tokens_per_reviewer.lm_studio',
|
||||
modelConfigKey: 'review.models.lm_studio',
|
||||
handler: 'openai-compatible',
|
||||
},
|
||||
{
|
||||
@@ -404,18 +496,66 @@ export const REVIEWER_LANES: ReadonlyArray<ReviewerLane> = Object.freeze([
|
||||
},
|
||||
invoke: {
|
||||
hostConfigKey: 'review.llama_cpp_host',
|
||||
defaultHost: 'http://localhost:8080',
|
||||
path: '/v1/chat/completions',
|
||||
modelDiscovery: 'first-from-models-endpoint',
|
||||
fallbackModel: 'local-model',
|
||||
effortChannel: 'none',
|
||||
},
|
||||
timeoutFloorMs: 120_000,
|
||||
emptyOutput: 'stub-with-stderr',
|
||||
reviewsSection: 'llama.cpp',
|
||||
evidenceClass: 'source-grounded',
|
||||
requiresBinaries: ['jq'],
|
||||
requiresBinaries: [],
|
||||
promptBudgetKey: 'review.max_prompt_tokens_per_reviewer.llama_cpp',
|
||||
modelConfigKey: 'review.models.llama_cpp',
|
||||
handler: 'openai-compatible',
|
||||
},
|
||||
{
|
||||
// Phase 5b (#2799) — closes #2718. Net-new in this phase BY DESIGN (ADR-2782's phase table):
|
||||
// declaring it in 5a would have made it selectable but not invocable, producing an empty
|
||||
// section for the whole 5a → 5b window.
|
||||
//
|
||||
// The probe is `command-capability`, not `command-exists`, and that is the entire reason D7's
|
||||
// vocabulary ships wider than existence: `kimi` is claimed by BOTH the Kimi Code CLI (Node) and
|
||||
// the legacy Python kimi-cli, which is a separate first-party runtime capability in this repo.
|
||||
// An existence-only probe registers the wrong tool. The needle `--output-format` appears in
|
||||
// Kimi Code's `--help` and is absent from the legacy CLI (whose headless flags are `--print` /
|
||||
// `--work-dir`). Verified in both directions against Kimi Code CLI 0.29.2 and a stub legacy
|
||||
// binary in closed PR #2776 — analysis carried forward with credit to @drungrin.
|
||||
//
|
||||
// The original probe there was an UNBOUNDED `kimi --help | grep` that ran on EVERY /gsd:review
|
||||
// regardless of flags: a live instance of this repo's named Unbounded Subprocesses defect, and
|
||||
// the review blocker. Here the bound is declared (`timeoutMs`) and the runner enforces it, and
|
||||
// the probe runs only for a SELECTED lane.
|
||||
slug: 'kimi-code',
|
||||
flags: ['--kimi-code'],
|
||||
transport: 'spawn',
|
||||
probe: {
|
||||
kind: 'command-capability',
|
||||
binary: 'kimi',
|
||||
needle: '--output-format',
|
||||
timeoutMs: 5_000,
|
||||
},
|
||||
invoke: {
|
||||
// Print mode takes the prompt as an ARGUMENT, so the full plan set goes by file reference to
|
||||
// stay clear of the 32,767-char Windows execFileSync ceiling — same shape as cursor.
|
||||
binary: 'kimi',
|
||||
args: ['{{model}}', '-p', '{{prompt}}'],
|
||||
promptChannel: 'argv-file-ref',
|
||||
outputChannel: 'stdout',
|
||||
modelArg: '-m',
|
||||
effortChannel: 'none',
|
||||
},
|
||||
timeoutFloorMs: 900_000,
|
||||
emptyOutput: 'stub-with-stderr',
|
||||
reviewsSection: 'Kimi Code',
|
||||
evidenceClass: 'source-grounded',
|
||||
requiresBinaries: [],
|
||||
promptBudgetKey: null,
|
||||
modelConfigKey: 'review.models.kimi-code',
|
||||
handler: null,
|
||||
},
|
||||
].map((lane) => Object.freeze(lane)) as ReviewerLane[]);
|
||||
|
||||
/* ------------------------------------------------------------------ *
|
||||
@@ -433,12 +573,9 @@ export const PARITY_VIOLATION = Object.freeze({
|
||||
INVALID_SLUG: 'invalid_slug',
|
||||
ROSTER_SLUG_UNDECLARED: 'roster_slug_undeclared',
|
||||
DESCRIPTOR_LANE_NOT_IN_ROSTER: 'descriptor_lane_not_in_roster',
|
||||
LEG_MARKER_MISSING: 'leg_marker_missing',
|
||||
LEG_MARKER_DUPLICATED: 'leg_marker_duplicated',
|
||||
LEG_MARKER_UNDECLARED: 'leg_marker_undeclared',
|
||||
SECTION_MISSING: 'section_missing',
|
||||
SECTION_DUPLICATED: 'section_duplicated',
|
||||
SECTION_UNDECLARED: 'section_undeclared',
|
||||
REGISTRY_LANE_UNDECLARED: 'registry_lane_undeclared',
|
||||
DESCRIPTOR_LANE_NOT_IN_REGISTRY: 'descriptor_lane_not_in_registry',
|
||||
BESPOKE_LEG_PRESENT: 'bespoke_leg_present',
|
||||
DUPLICATE_SLUG: 'duplicate_slug',
|
||||
DUPLICATE_FLAG: 'duplicate_flag',
|
||||
DUPLICATE_SECTION: 'duplicate_section',
|
||||
@@ -460,19 +597,27 @@ export interface ParityResult {
|
||||
export interface ParityInput {
|
||||
descriptor: ReadonlyArray<ReviewerLane>;
|
||||
roster: ReadonlyArray<string>;
|
||||
/**
|
||||
* Reviewer slugs declared by `reviewer` bodies in the GENERATED capability registry.
|
||||
*
|
||||
* Phase 5b (#2799). Once `invoke_reviewers` iterates lanes, the registry — not the workflow text —
|
||||
* is what actually decides which lanes exist at runtime, so this is the surface parity has to
|
||||
* bind. A missing or malformed value degrades to violations rather than silence, exactly as
|
||||
* `workflowText` does: a checker that cannot tell "no registry" from "registry agrees" is worse
|
||||
* than no checker.
|
||||
*/
|
||||
registry: ReadonlyArray<string>;
|
||||
/** Full text of gsd-core/workflows/review.md. */
|
||||
workflowText: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* The machine-readable marker that makes an `invoke_reviewers` leg identifiable.
|
||||
* The marker that used to make a hand-authored `invoke_reviewers` leg identifiable.
|
||||
*
|
||||
* The legs are prose-labelled (`**Qwen Code:**`, `**LM Studio (local,
|
||||
* OpenAI-compatible):**`), and five NON-lane bold labels in the same step have
|
||||
* the identical bold-then-fence shape (`**Timeout guidance (#2194):**`,
|
||||
* `**No hook-trust bypass (#2479):**`, `**Maintainer note — …:**`). Inferring
|
||||
* legs from prose shape would be a heuristic asserting what it cannot prove, so
|
||||
* each leg carries an explicit marker instead. Phase 5b iterates on these.
|
||||
* Phase 5b (#2799) deleted every bespoke leg, so this regex flipped polarity: matching one is now
|
||||
* the VIOLATION (`BESPOKE_LEG_PRESENT`) rather than the requirement. It is retained precisely
|
||||
* because deleting it would leave nothing stopping a future contributor from quietly re-adding a
|
||||
* per-CLI block — which is the drift (#2718 → #2781) this whole epic exists to end.
|
||||
*/
|
||||
const LEG_MARKER_RE = /<!--\s*reviewer-lane:\s*([a-z0-9_-]+)\s*-->/g;
|
||||
|
||||
@@ -495,23 +640,8 @@ const LEG_MARKER_RE = /<!--\s*reviewer-lane:\s*([a-z0-9_-]+)\s*-->/g;
|
||||
*/
|
||||
export const LANE_SLUG_RE = /^[a-z0-9][a-z0-9_-]*$/;
|
||||
|
||||
/**
|
||||
* A lane section heading in write_reviews.
|
||||
*
|
||||
* Anchored at h2 with an exact ` Review` suffix and NO parenthetical, because:
|
||||
* - `## OpenCode Review (opencode-deepseek)` is an ADR-1517 reviewer INSTANCE,
|
||||
* and ADR-2782 D8 states instances are not lanes — two such headings are
|
||||
* already in the file, so a naive matcher fails on day one;
|
||||
* - `## Consensus Summary` has no ` Review` suffix;
|
||||
* - `# Cross-AI Plan Review — Phase {N}` is h1, not h2.
|
||||
* `[^()\n]+` excludes the parenthetical form rather than stripping it, so an
|
||||
* instance heading never resolves to a lane.
|
||||
*/
|
||||
const SECTION_HEADING_RE = /^##[ \t]+([^()\n\r]+?) Review[ \t]*$/gm;
|
||||
|
||||
/** Bounds of the step a marker must appear inside. */
|
||||
/** Bounds of the step a bespoke leg marker could appear inside. */
|
||||
const INVOKE_STEP_RE = /<step name="invoke_reviewers">([\s\S]*?)<\/step>/;
|
||||
const WRITE_STEP_RE = /<step name="write_reviews">([\s\S]*?)<\/step>/;
|
||||
|
||||
function sliceStep(workflowText: string, re: RegExp): string {
|
||||
const m = workflowText.match(re);
|
||||
@@ -532,21 +662,32 @@ function countOccurrences(haystack: string, re: RegExp): Map<string, number> {
|
||||
}
|
||||
|
||||
/**
|
||||
* Bidirectional parity across four surfaces: the descriptor, the roster
|
||||
* (`KNOWN_REVIEWER_SLUGS`), the `invoke_reviewers` legs, and the `write_reviews`
|
||||
* sections.
|
||||
* Bidirectional parity across three surfaces: the descriptor, the roster
|
||||
* (`KNOWN_REVIEWER_SLUGS`), and the generated capability **registry** — plus one anti-parity
|
||||
* assertion against the workflow.
|
||||
*
|
||||
* Bidirectional is the point. A forward-only check ("does each declared lane
|
||||
* resolve?") misses the failure this exists to catch: #2718 added a lane leg and
|
||||
* #2781 was the documentation drift that followed. An undeclared leg must fail.
|
||||
* **Re-pointed by Phase 5b (#2799).** Through Phase 5a this function also required a literal
|
||||
* `<!-- reviewer-lane: <slug> -->` per lane inside `invoke_reviewers` and a literal
|
||||
* `## <Section> Review` per lane inside `write_reviews`. Phase 5b deletes exactly that text: the
|
||||
* workflow now iterates declared lanes and renders sections from `reviewsSection`, so there is no
|
||||
* per-lane text left to scan. Those two families could not be kept without keeping the
|
||||
* hand-maintained per-lane blocks this epic exists to delete.
|
||||
*
|
||||
* Pure and total: never reads the filesystem, never throws. Empty or malformed
|
||||
* `workflowText` degrades to violations, so a caller cannot mistake a read
|
||||
* failure for a clean bill of health.
|
||||
* What replaced them is the parity that is actually load-bearing once lanes are data: the registry
|
||||
* is what decides which lanes exist at runtime, so `descriptor ↔ registry` is checked in both
|
||||
* directions. That is also the mechanical single source #2781/Phase 6 needs for its docs and locale
|
||||
* gate, which per-leg text could never provide.
|
||||
*
|
||||
* CRLF-insensitive: `\r` is stripped before matching, because a Windows
|
||||
* autocrlf checkout would otherwise leave every marker and heading unmatched
|
||||
* and report the whole roster missing.
|
||||
* Bidirectional is still the point. A forward-only check ("does each declared lane resolve?")
|
||||
* misses the failure this exists to catch: #2718 added a lane and #2781 was the documentation drift
|
||||
* that followed. A registry lane nobody declared must fail, and so must a re-added bespoke leg.
|
||||
*
|
||||
* Pure and total: never reads the filesystem, never throws. Empty or malformed `workflowText` /
|
||||
* `registry` degrades to violations, so a caller cannot mistake a read failure for a clean bill of
|
||||
* health.
|
||||
*
|
||||
* CRLF-insensitive: `\r` is stripped before matching, because a Windows autocrlf checkout would
|
||||
* otherwise leave every marker unmatched.
|
||||
*/
|
||||
export function checkReviewerLaneParity(input: ParityInput): ParityResult {
|
||||
const { descriptor, roster } = input;
|
||||
@@ -569,13 +710,6 @@ export function checkReviewerLaneParity(input: ParityInput): ParityResult {
|
||||
const seenFlag = new Set<string>();
|
||||
const seenSection = new Set<string>();
|
||||
|
||||
/** A lane that survived validation: slug is a string in the declared grammar. */
|
||||
interface ValidatedLane {
|
||||
slug: string;
|
||||
reviewsSection: string | null;
|
||||
}
|
||||
const lanes: ValidatedLane[] = [];
|
||||
|
||||
// The declared parameter type says `ReviewerLane[]`, but this function is a
|
||||
// trust boundary — narrow from `unknown` rather than believing the annotation.
|
||||
const rawLanes: unknown[] = Array.isArray(descriptor) ? (descriptor as unknown[]) : [];
|
||||
@@ -591,7 +725,6 @@ export function checkReviewerLaneParity(input: ParityInput): ParityResult {
|
||||
continue;
|
||||
}
|
||||
const section = typeof lane.reviewsSection === 'string' ? lane.reviewsSection : null;
|
||||
lanes.push({ slug, reviewsSection: section });
|
||||
|
||||
if (seenSlug.has(slug)) add(PARITY_VIOLATION.DUPLICATE_SLUG, slug);
|
||||
seenSlug.add(slug);
|
||||
@@ -623,33 +756,32 @@ export function checkReviewerLaneParity(input: ParityInput): ParityResult {
|
||||
if (!rosterSet.has(slug)) add(PARITY_VIOLATION.DESCRIPTOR_LANE_NOT_IN_ROSTER, slug);
|
||||
}
|
||||
|
||||
// --- descriptor <-> invoke_reviewers legs ---
|
||||
const markerCounts = countOccurrences(
|
||||
sliceStep(workflowText, INVOKE_STEP_RE),
|
||||
LEG_MARKER_RE,
|
||||
// --- descriptor <-> registry ---
|
||||
//
|
||||
// The registry is what decides which lanes exist at runtime once the workflow iterates, so this
|
||||
// replaces the per-leg text checks Phase 5b deleted. A non-array degrades to an empty set, which
|
||||
// then reports every descriptor lane as missing — loud, not silent.
|
||||
const registrySet = new Set(
|
||||
(Array.isArray(input.registry) ? input.registry : []).filter(
|
||||
(x): x is string => typeof x === 'string',
|
||||
),
|
||||
);
|
||||
for (const lane of lanes) {
|
||||
const n = markerCounts.get(lane.slug) ?? 0;
|
||||
if (n === 0) add(PARITY_VIOLATION.LEG_MARKER_MISSING, lane.slug);
|
||||
else if (n > 1) add(PARITY_VIOLATION.LEG_MARKER_DUPLICATED, lane.slug);
|
||||
for (const slug of registrySet) {
|
||||
if (!seenSlug.has(slug)) add(PARITY_VIOLATION.REGISTRY_LANE_UNDECLARED, slug);
|
||||
}
|
||||
for (const slug of markerCounts.keys()) {
|
||||
if (!seenSlug.has(slug)) add(PARITY_VIOLATION.LEG_MARKER_UNDECLARED, slug);
|
||||
for (const slug of seenSlug) {
|
||||
if (!registrySet.has(slug)) add(PARITY_VIOLATION.DESCRIPTOR_LANE_NOT_IN_REGISTRY, slug);
|
||||
}
|
||||
|
||||
// --- descriptor <-> write_reviews sections ---
|
||||
const sectionCounts = countOccurrences(
|
||||
sliceStep(workflowText, WRITE_STEP_RE),
|
||||
SECTION_HEADING_RE,
|
||||
);
|
||||
for (const lane of lanes) {
|
||||
if (lane.reviewsSection === null) continue;
|
||||
const n = sectionCounts.get(lane.reviewsSection) ?? 0;
|
||||
if (n === 0) add(PARITY_VIOLATION.SECTION_MISSING, lane.reviewsSection);
|
||||
else if (n > 1) add(PARITY_VIOLATION.SECTION_DUPLICATED, lane.reviewsSection);
|
||||
}
|
||||
for (const section of sectionCounts.keys()) {
|
||||
if (!seenSection.has(section)) add(PARITY_VIOLATION.SECTION_UNDECLARED, section);
|
||||
// --- anti-parity: no bespoke leg may return ---
|
||||
//
|
||||
// Phase 5b deleted every hand-authored per-CLI block. Nothing in the type system stops a future
|
||||
// contributor from adding one back, and a re-added block is invisible to every other check here
|
||||
// (it would still be declared, still be in the roster, still be in the registry). Matching a leg
|
||||
// marker is therefore now the violation.
|
||||
const markerSlugs = countOccurrences(sliceStep(workflowText, INVOKE_STEP_RE), LEG_MARKER_RE);
|
||||
for (const slug of markerSlugs.keys()) {
|
||||
add(PARITY_VIOLATION.BESPOKE_LEG_PRESENT, slug);
|
||||
}
|
||||
|
||||
return { ok: violations.length === 0, violations };
|
||||
|
||||
482
src/review-lane-invocation.cts
Normal file
482
src/review-lane-invocation.cts
Normal file
@@ -0,0 +1,482 @@
|
||||
/**
|
||||
* Reviewer Lane Invocation Module (ADR-2782 Phase 5b, #2799 — closes #2718).
|
||||
*
|
||||
* Turns a DECLARED lane (`review-lane-descriptor.cts`) plus resolved configuration into a concrete,
|
||||
* inspectable INVOCATION PLAN. Phase 1's module declares; this one resolves; `review-lane-runner`
|
||||
* executes. The split exists so the interesting half is pure: a plan is a value, so the twelve
|
||||
* shipped lanes can be asserted against a golden table without spawning anything.
|
||||
*
|
||||
* WHY A GOLDEN TABLE AND NOT A FRESH DESIGN (Gall's Law). The 640 lines of hand-authored bash this
|
||||
* replaces is the simple system that worked, and every leg encodes a hard-won fix — #2494 and #2605
|
||||
* (empty output), #1698 (Codex stdout teardown noise), #1936 (OpenCode zero-output turns), #2073
|
||||
* (Antigravity's three modes), #2176 (repo-root anchoring), #2589 (no jq on stock Windows), #2794
|
||||
* (Qwen's missing sidecar). A resolver designed from the descriptor TYPES would throw that away and
|
||||
* rebuild the bugs. So each lane's plan was derived from its leg, and `tests/review-lane-invocation`
|
||||
* asserts all twelve against a frozen table. Old and new cannot literally run in parallel, so that
|
||||
* table is the strangler-fig substitute — it is what makes this cutover safe rather than hopeful.
|
||||
*
|
||||
* PURE. No filesystem, no network, no subprocess, no clock. Configuration arrives through the
|
||||
* `configGet` seam so a test drives it with a plain object and production wires it to the real
|
||||
* resolved config. Every function here is total: a malformed lane yields an `unavailable` result,
|
||||
* never a throw — a resolver that throws on bad input cannot report on it, and this runs against
|
||||
* third-party overlay manifests (ADR-2782 D4).
|
||||
*/
|
||||
|
||||
import type {
|
||||
EmptyOutputPolicy,
|
||||
LaneHandler,
|
||||
LaneProbe,
|
||||
ReviewerLane,
|
||||
} from './review-lane-descriptor.cjs';
|
||||
import { LANE_SLUG_RE } from './review-lane-descriptor.cjs';
|
||||
|
||||
/* ------------------------------------------------------------------ *
|
||||
* Unavailability — a frozen enum, because the reason is the product
|
||||
* ------------------------------------------------------------------ */
|
||||
|
||||
/**
|
||||
* Why a lane will not run.
|
||||
*
|
||||
* Frozen and exhaustive because the bash it replaces had exactly one outcome for every failure —
|
||||
* an empty file — and that ambiguity IS the defect class this epic exists to close (#2494/#2605: a
|
||||
* failed lane was indistinguishable from a reviewer that ran cleanly with nothing to report). The
|
||||
* caller renders these; tests assert on them. Never assert on the rendered prose.
|
||||
*
|
||||
* Adding a member is three coordinated changes: this enum, the emitting site, and the test locking
|
||||
* `Object.keys(...).sort()`.
|
||||
*/
|
||||
export const LANE_UNAVAILABLE = Object.freeze({
|
||||
MALFORMED_LANE: 'malformed_lane',
|
||||
UNKNOWN_HANDLER: 'unknown_handler',
|
||||
UNKNOWN_TRANSPORT: 'unknown_transport',
|
||||
MISSING_BINARY: 'missing_binary',
|
||||
MISSING_REQUIRED_BINARY: 'missing_required_binary',
|
||||
PROBE_FAILED: 'probe_failed',
|
||||
PROBE_TIMEOUT: 'probe_timeout',
|
||||
HOST_UNREACHABLE: 'host_unreachable',
|
||||
EGRESS_HOST_CHANGED: 'egress_host_changed',
|
||||
BUDGET_TOO_SMALL: 'budget_too_small',
|
||||
BUDGET_TOOL_FAILED: 'budget_tool_failed',
|
||||
} as const);
|
||||
|
||||
export type LaneUnavailableReason =
|
||||
(typeof LANE_UNAVAILABLE)[keyof typeof LANE_UNAVAILABLE];
|
||||
|
||||
/* ------------------------------------------------------------------ *
|
||||
* The plan
|
||||
* ------------------------------------------------------------------ */
|
||||
|
||||
/** Where a spawned lane's review actually lands. */
|
||||
export type OutputTarget =
|
||||
| { kind: 'stdout' }
|
||||
/** Codex: the tool writes the review itself and stdout is discarded (#1698). */
|
||||
| { kind: 'file'; path: string };
|
||||
|
||||
export interface SpawnPlan {
|
||||
transport: 'spawn';
|
||||
slug: string;
|
||||
binary: string;
|
||||
/** Fully resolved argv — model, effort and prompt already folded in, in leg order. */
|
||||
argv: string[];
|
||||
/** Prompt delivered on stdin, or `null` for `argv`/`argv-file-ref`/`none` lanes. */
|
||||
stdin: string | null;
|
||||
/**
|
||||
* The assembled prompt file, regardless of how (or whether) this lane consumes it. Carried even
|
||||
* for `promptChannel: 'none'` so a handler never has to re-derive the path from another field.
|
||||
*/
|
||||
promptPath: string;
|
||||
outputTarget: OutputTarget;
|
||||
/** Canonical review path: `<runDir>/gsd-review-<slug>.md`. */
|
||||
reviewPath: string;
|
||||
/** stderr sidecar — never `/dev/null` (#2494). */
|
||||
errPath: string;
|
||||
timeoutMs: number;
|
||||
emptyOutput: EmptyOutputPolicy;
|
||||
handler: LaneHandler;
|
||||
requiresBinaries: readonly string[];
|
||||
probe: LaneProbe;
|
||||
}
|
||||
|
||||
export interface HttpPlan {
|
||||
transport: 'openai-http';
|
||||
slug: string;
|
||||
/** Resolved base URL, normalized (no trailing slash). */
|
||||
host: string;
|
||||
hostConfigKey: string;
|
||||
/** `host` + the declared path. */
|
||||
url: string;
|
||||
/** Models endpoint for discovery, when `modelDiscovery` asks for it. */
|
||||
modelsUrl: string | null;
|
||||
/** Configured model, or `null` when discovery should run. */
|
||||
model: string | null;
|
||||
fallbackModel: string;
|
||||
promptPath: string;
|
||||
reviewPath: string;
|
||||
errPath: string;
|
||||
timeoutMs: number;
|
||||
emptyOutput: EmptyOutputPolicy;
|
||||
handler: LaneHandler;
|
||||
requiresBinaries: readonly string[];
|
||||
probe: LaneProbe;
|
||||
}
|
||||
|
||||
export type LanePlan = SpawnPlan | HttpPlan;
|
||||
|
||||
export type ResolveResult =
|
||||
| { ok: true; plan: LanePlan; warnings: string[] }
|
||||
| { ok: false; reason: LaneUnavailableReason; detail: string; warnings: string[] };
|
||||
|
||||
/** Every handler name the runner can dispatch. Mirrors `LaneHandler` (D6's closed enum). */
|
||||
const KNOWN_HANDLERS: ReadonlySet<string> = new Set(['antigravity', 'openai-compatible', 'opencode']);
|
||||
|
||||
export interface ResolveInput {
|
||||
lane: ReviewerLane;
|
||||
/** Resolved config lookup. Returns `undefined` for an absent key. */
|
||||
configGet: (key: string) => unknown;
|
||||
/** The run-scoped temp dir (`{run_dir}` in the workflow). */
|
||||
runDir: string;
|
||||
/** Absolute repo root, for `argv-file-ref` anchoring (#2176). */
|
||||
repoRoot: string;
|
||||
/** Effort argv for lanes whose `effortChannel` is `argv`; empty when the host declares none. */
|
||||
effortArgs?: readonly string[];
|
||||
}
|
||||
|
||||
/* ------------------------------------------------------------------ *
|
||||
* Value normalization — the stringly-typed edges the bash lived with
|
||||
* ------------------------------------------------------------------ */
|
||||
|
||||
/**
|
||||
* A configured string value, or `null` when effectively unset.
|
||||
*
|
||||
* Four shapes all mean "not configured", and the bash had to handle three of them by hand:
|
||||
* - absent / `null` / `undefined`;
|
||||
* - `""` — the declared default of every federated `review.models.*` key;
|
||||
* - `"null"` — the LITERAL four characters `config-get --raw` prints for a missing key, which
|
||||
* every leg tested for (`[ "$X" != "null" ]`). Reading config in-process removes the source of
|
||||
* this, but a `.planning/config.json` written by an older workflow can still contain it;
|
||||
* - whitespace-only — never a meaningful model name or host.
|
||||
*
|
||||
* A non-string (number, bool, object, array) is NOT coerced. `String(0)` would put `"0"` into argv
|
||||
* as a model name; a wrong model silently reviewed is worse than no model override.
|
||||
*/
|
||||
function configString(raw: unknown): string | null {
|
||||
if (typeof raw !== 'string') return null;
|
||||
const trimmed = raw.trim();
|
||||
if (trimmed === '' || trimmed === 'null' || trimmed === 'undefined') return null;
|
||||
return trimmed;
|
||||
}
|
||||
|
||||
/**
|
||||
* Normalize a base URL for storage and for the D5 consent comparison.
|
||||
*
|
||||
* Trailing slash, case in the scheme/host, and an explicit default port are all the same
|
||||
* destination. Without normalizing, a cosmetic `.planning/config.json` edit — adding a trailing
|
||||
* slash — would read as "the egress destination changed" and block the lane, training the user to
|
||||
* dismiss the one warning that matters.
|
||||
*
|
||||
* Returns the input trimmed when it is not parseable as a URL: an unparseable host is compared
|
||||
* verbatim rather than silently rewritten.
|
||||
*/
|
||||
export function normalizeHost(raw: string): string {
|
||||
const trimmed = String(raw ?? '').trim();
|
||||
if (!trimmed) return '';
|
||||
let u: URL;
|
||||
try {
|
||||
u = new URL(trimmed);
|
||||
} catch {
|
||||
return trimmed.replace(/\/+$/, '');
|
||||
}
|
||||
// `new URL('localhost:11434')` PARSES — protocol `localhost:`, empty hostname — so a plausible
|
||||
// but scheme-less config value would otherwise be rewritten to `localhost://11434` and compared
|
||||
// (and requested) as though it were a real destination. No hostname means this is not a URL;
|
||||
// return it verbatim so it fails visibly rather than silently becoming something else.
|
||||
if (!u.hostname) return trimmed.replace(/\/+$/, '');
|
||||
const scheme = u.protocol.toLowerCase();
|
||||
const isDefaultPort =
|
||||
(scheme === 'http:' && u.port === '80') || (scheme === 'https:' && u.port === '443');
|
||||
const port = isDefaultPort ? '' : u.port;
|
||||
const host = u.hostname.toLowerCase();
|
||||
const pathPart = u.pathname.replace(/\/+$/, '');
|
||||
return `${scheme}//${host}${port ? `:${port}` : ''}${pathPart}`;
|
||||
}
|
||||
|
||||
/**
|
||||
* Classify a lane's output as a review or as empty.
|
||||
*
|
||||
* WHITESPACE-ONLY COUNTS AS EMPTY, for every lane. The bash tested `[ ! -s file ]`, which counts
|
||||
* BYTES — so a reply of three spaces passed as a successful review. Two legs (LM Studio,
|
||||
* llama.cpp) closed this locally with a case-glob; gemini, claude, codex, qwen and cursor did not.
|
||||
* Making it uniform is a deliberate, disclosed behavior change (a bug fix that breaks a workaround)
|
||||
* and is why this phase ships a changeset note.
|
||||
*
|
||||
* The `-n` / `-e` / `-E` hazard the two printf-using legs guarded against cannot occur here at all:
|
||||
* nothing in this path goes through `echo`.
|
||||
*/
|
||||
export function isEmptyReview(text: unknown): boolean {
|
||||
if (typeof text !== 'string') return true;
|
||||
return text.trim().length === 0;
|
||||
}
|
||||
|
||||
/**
|
||||
* The standard `argv-file-ref` prompt (#2176).
|
||||
*
|
||||
* Two lanes take the prompt as an ARGUMENT rather than on stdin, and a full plan set inline would
|
||||
* approach the 32,767-character Windows `execFileSync` ceiling — so the argument is a short
|
||||
* instruction naming the prompt file. It must also carry the ABSOLUTE repo root: an argv-fed CLI
|
||||
* does not reliably inherit the review's working directory, and without the anchor the reviewer
|
||||
* reviews the plan text in isolation, which is exactly what the Review Instructions forbid.
|
||||
*
|
||||
* `antigravity` deliberately does NOT use this text — its handler owns a variant that additionally
|
||||
* demands a `REVIEWED-WITHOUT-REPO-ACCESS` self-report.
|
||||
*/
|
||||
export function fileRefPrompt(promptPath: string, repoRoot: string): string {
|
||||
return (
|
||||
`Read the file at ${promptPath} in full and carry out the review request it contains. ` +
|
||||
`The repository under review is at ${repoRoot} — resolve every relative file path in the ` +
|
||||
`review request against that absolute root. Output only the resulting markdown review. ` +
|
||||
`Do not edit any files.`
|
||||
);
|
||||
}
|
||||
|
||||
/** Run-dir artifact paths. POSIX-joined: these are workflow-visible strings, not OS paths. */
|
||||
function artifactPaths(runDir: string, slug: string): {
|
||||
promptPath: string;
|
||||
reviewPath: string;
|
||||
errPath: string;
|
||||
} {
|
||||
const base = String(runDir ?? '').replace(/\/+$/, '');
|
||||
return {
|
||||
promptPath: `${base}/gsd-review-prompt.md`,
|
||||
reviewPath: `${base}/gsd-review-${slug}.md`,
|
||||
errPath: `${base}/gsd-review-${slug}.err`,
|
||||
};
|
||||
}
|
||||
|
||||
/* ------------------------------------------------------------------ *
|
||||
* Resolution
|
||||
* ------------------------------------------------------------------ */
|
||||
|
||||
/**
|
||||
* Resolve one declared lane into an executable plan.
|
||||
*
|
||||
* TOTAL: never throws. Every rejection is a typed `LaneUnavailableReason`, because the caller has to
|
||||
* tell three different unavailabilities apart — a lane that is absent, a lane whose probe failed,
|
||||
* and a lane blocked on a changed egress destination are not the same event to a user.
|
||||
*
|
||||
* NOT resolved here: probe execution, prompt-budget trimming, and the D5 egress-host comparison.
|
||||
* Those need I/O and live in the runner. This function decides SHAPE.
|
||||
*/
|
||||
export function resolveLanePlan(input: ResolveInput): ResolveResult {
|
||||
const warnings: string[] = [];
|
||||
const fail = (reason: LaneUnavailableReason, detail: string): ResolveResult => ({
|
||||
ok: false,
|
||||
reason,
|
||||
detail,
|
||||
warnings,
|
||||
});
|
||||
|
||||
// Trust boundary: the declared type says ReviewerLane, but overlay manifests arrive here.
|
||||
const raw = input?.lane as unknown;
|
||||
if (raw === null || typeof raw !== 'object' || Array.isArray(raw)) {
|
||||
return fail(LANE_UNAVAILABLE.MALFORMED_LANE, `lane is not an object: ${String(raw)}`);
|
||||
}
|
||||
const lane = raw as ReviewerLane;
|
||||
const slug = typeof lane.slug === 'string' ? lane.slug.trim() : '';
|
||||
if (!slug) {
|
||||
return fail(LANE_UNAVAILABLE.MALFORMED_LANE, 'lane declares no slug');
|
||||
}
|
||||
// The slug is CONCATENATED into artifact paths below, so the grammar is enforced here rather
|
||||
// than trusted from upstream. `checkReviewerLaneParity` and the capability validator both check
|
||||
// it too, but neither runs on this path — and this module's whole premise is that it is the
|
||||
// trust boundary for third-party overlay manifests. A slug of `../../../tmp/evil` would
|
||||
// otherwise produce a reviewPath outside the run dir that `writeReviewOrStub` happily writes to.
|
||||
if (!LANE_SLUG_RE.test(slug)) {
|
||||
return fail(
|
||||
LANE_UNAVAILABLE.MALFORMED_LANE,
|
||||
`lane slug '${slug}' is outside the declared grammar ${String(LANE_SLUG_RE)}`,
|
||||
);
|
||||
}
|
||||
|
||||
// D4 rule 4: an unknown handler FAILS CLOSED. A lane naming imperative code this GSD version does
|
||||
// not have cannot be run "mostly" — the handler is precisely the part data could not express.
|
||||
const handler: LaneHandler = lane.handler ?? null;
|
||||
if (handler !== null && !KNOWN_HANDLERS.has(handler)) {
|
||||
return fail(
|
||||
LANE_UNAVAILABLE.UNKNOWN_HANDLER,
|
||||
`lane '${slug}' names handler '${String(handler)}', which this GSD version does not provide`,
|
||||
);
|
||||
}
|
||||
|
||||
const { promptPath, reviewPath, errPath } = artifactPaths(input.runDir, slug);
|
||||
const timeoutMs =
|
||||
typeof lane.timeoutFloorMs === 'number' && Number.isFinite(lane.timeoutFloorMs) && lane.timeoutFloorMs > 0
|
||||
? lane.timeoutFloorMs
|
||||
: 900_000;
|
||||
const emptyOutput: EmptyOutputPolicy = lane.emptyOutput === 'handler-owned' ? 'handler-owned' : 'stub-with-stderr';
|
||||
const requiresBinaries = Array.isArray(lane.requiresBinaries)
|
||||
? lane.requiresBinaries.filter((b): b is string => typeof b === 'string')
|
||||
: [];
|
||||
|
||||
const model = configString(
|
||||
typeof lane.modelConfigKey === 'string' ? input.configGet(lane.modelConfigKey) : undefined,
|
||||
);
|
||||
|
||||
if (lane.transport === 'openai-http') {
|
||||
const rawInvoke: unknown = lane.invoke;
|
||||
// The spawn branch below guards with `inv?.binary`; this one must too. Without it a lane
|
||||
// declaring `transport: 'openai-http'` and no `invoke` THROWS, which breaks this module's
|
||||
// documented totality — and a throw here is worse than it looks: the CLI seam resolves every
|
||||
// selected lane in one `.map`, so one malformed overlay manifest would abort the whole review
|
||||
// rather than dropping its own lane.
|
||||
if (rawInvoke === null || typeof rawInvoke !== 'object' || Array.isArray(rawInvoke)) {
|
||||
return fail(
|
||||
LANE_UNAVAILABLE.MALFORMED_LANE,
|
||||
`openai-http lane '${slug}' declares no invoke object`,
|
||||
);
|
||||
}
|
||||
const inv = rawInvoke as {
|
||||
hostConfigKey?: unknown;
|
||||
defaultHost?: unknown;
|
||||
path?: unknown;
|
||||
fallbackModel?: unknown;
|
||||
modelDiscovery?: unknown;
|
||||
};
|
||||
const hostConfigKey = typeof inv.hostConfigKey === 'string' ? inv.hostConfigKey : '';
|
||||
const configured = hostConfigKey ? configString(input.configGet(hostConfigKey)) : null;
|
||||
// Only a STRING declares a host. Coercing an object would produce the literal
|
||||
// '[object Object]' and normalize THAT as the lane's egress destination.
|
||||
const declaredDefault = typeof inv.defaultHost === 'string' ? inv.defaultHost : '';
|
||||
const host = normalizeHost(configured ?? declaredDefault);
|
||||
if (!host) {
|
||||
return fail(
|
||||
LANE_UNAVAILABLE.MALFORMED_LANE,
|
||||
`lane '${slug}' resolves no host: '${hostConfigKey}' is unset and it declares no defaultHost`,
|
||||
);
|
||||
}
|
||||
const apiPath = typeof inv.path === 'string' && inv.path ? inv.path : '/v1/chat/completions';
|
||||
const discovers = inv.modelDiscovery === 'first-from-models-endpoint';
|
||||
return {
|
||||
ok: true,
|
||||
warnings,
|
||||
plan: {
|
||||
transport: 'openai-http',
|
||||
slug,
|
||||
host,
|
||||
hostConfigKey,
|
||||
url: `${host}${apiPath}`,
|
||||
modelsUrl: discovers ? `${host}/v1/models` : null,
|
||||
model,
|
||||
fallbackModel: typeof inv.fallbackModel === 'string' ? inv.fallbackModel : 'local-model',
|
||||
promptPath,
|
||||
reviewPath,
|
||||
errPath,
|
||||
timeoutMs,
|
||||
emptyOutput,
|
||||
handler,
|
||||
requiresBinaries,
|
||||
probe: lane.probe,
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
if (lane.transport !== 'spawn') {
|
||||
return fail(
|
||||
LANE_UNAVAILABLE.UNKNOWN_TRANSPORT,
|
||||
`lane '${slug}' declares transport '${String((lane as { transport?: unknown }).transport)}'`,
|
||||
);
|
||||
}
|
||||
|
||||
const inv = lane.invoke;
|
||||
const binary = typeof inv?.binary === 'string' ? inv.binary.trim() : '';
|
||||
if (!binary) {
|
||||
return fail(LANE_UNAVAILABLE.MALFORMED_LANE, `spawn lane '${slug}' declares no binary`);
|
||||
}
|
||||
|
||||
// Output target first: `{{output}}` needs to know the path, and the target is also what tells the
|
||||
// runner whether to capture stdout or read a file the tool wrote itself (#1698).
|
||||
let outputTarget: OutputTarget = { kind: 'stdout' };
|
||||
let outputExpansion: string[] = [];
|
||||
if (inv.outputChannel === 'file-arg') {
|
||||
const outputArg = typeof inv.outputArg === 'string' ? inv.outputArg : '';
|
||||
if (!outputArg) {
|
||||
return fail(
|
||||
LANE_UNAVAILABLE.MALFORMED_LANE,
|
||||
`lane '${slug}' declares outputChannel 'file-arg' with no outputArg naming the argument`,
|
||||
);
|
||||
}
|
||||
outputExpansion = [outputArg, reviewPath];
|
||||
outputTarget = { kind: 'file', path: reviewPath };
|
||||
}
|
||||
|
||||
let stdin: string | null = null;
|
||||
let promptExpansion: string[] = [];
|
||||
switch (inv.promptChannel) {
|
||||
case 'stdin':
|
||||
stdin = promptPath; // the runner streams this file; the plan names it.
|
||||
break;
|
||||
case 'argv-file-ref':
|
||||
promptExpansion = [fileRefPrompt(promptPath, input.repoRoot)];
|
||||
break;
|
||||
case 'argv':
|
||||
promptExpansion = [promptPath];
|
||||
break;
|
||||
case 'none':
|
||||
break; // CodeRabbit reviews the working tree and is fed nothing (review.md:367).
|
||||
default:
|
||||
return fail(
|
||||
LANE_UNAVAILABLE.MALFORMED_LANE,
|
||||
`lane '${slug}' declares promptChannel '${String(inv.promptChannel)}'`,
|
||||
);
|
||||
}
|
||||
|
||||
const modelExpansion =
|
||||
model && typeof inv.modelArg === 'string' && inv.modelArg ? [inv.modelArg, model] : [];
|
||||
const effortExpansion =
|
||||
inv.effortChannel === 'argv'
|
||||
? (input.effortArgs ?? []).filter((a): a is string => typeof a === 'string' && a !== '')
|
||||
: [];
|
||||
|
||||
// Expand the argv template in declared order. A placeholder with nothing to contribute expands to
|
||||
// ZERO elements and disappears — that is what lets one template serve both the configured and the
|
||||
// unconfigured case without a conditional in the data.
|
||||
const expansions: Record<string, string[]> = {
|
||||
'{{model}}': modelExpansion,
|
||||
'{{effort}}': effortExpansion,
|
||||
'{{output}}': outputExpansion,
|
||||
'{{prompt}}': promptExpansion,
|
||||
};
|
||||
const template = Array.isArray(inv.args)
|
||||
? inv.args.filter((a): a is string => typeof a === 'string')
|
||||
: [];
|
||||
const argv: string[] = [];
|
||||
for (const token of template) {
|
||||
// Own-property lookup: a lane could declare a literal `constructor` argument, and prototype
|
||||
// members must never resolve as expansions.
|
||||
if (Object.prototype.hasOwnProperty.call(expansions, token)) {
|
||||
argv.push(...expansions[token]);
|
||||
} else {
|
||||
argv.push(token);
|
||||
}
|
||||
}
|
||||
|
||||
return {
|
||||
ok: true,
|
||||
warnings,
|
||||
plan: {
|
||||
transport: 'spawn',
|
||||
slug,
|
||||
binary,
|
||||
argv,
|
||||
stdin,
|
||||
promptPath,
|
||||
outputTarget,
|
||||
reviewPath,
|
||||
errPath,
|
||||
timeoutMs,
|
||||
emptyOutput,
|
||||
handler,
|
||||
requiresBinaries,
|
||||
probe: lane.probe,
|
||||
},
|
||||
};
|
||||
}
|
||||
697
src/review-lane-runner.cts
Normal file
697
src/review-lane-runner.cts
Normal file
@@ -0,0 +1,697 @@
|
||||
/**
|
||||
* Reviewer Lane Runner (ADR-2782 Phase 5b, #2799).
|
||||
*
|
||||
* Executes an invocation plan from `review-lane-invocation.cts`: probes the lane, spawns the binary
|
||||
* or calls the HTTP endpoint, applies the empty-output policy, and dispatches the three first-party
|
||||
* handlers D6 names. This is the module that replaces ~640 lines of hand-authored per-CLI bash.
|
||||
*
|
||||
* THREE RUNTIME DEPENDENCIES DISAPPEAR HERE, and that is a correctness win rather than a tidy-up:
|
||||
* - `jq` — five legs needed it on PATH; it is absent on stock Windows/Git-Bash (#2589). Parsing is
|
||||
* `JSON.parse` now.
|
||||
* - `curl` — the three OpenAI-compatible legs shelled out to it. Node's `fetch` replaces it, and
|
||||
* the raw response body (where an OpenAI-compatible server puts its error JSON on a 4xx/5xx
|
||||
* while curl still exits 0) is no longer discarded by a pipe.
|
||||
* - external `timeout` / `gtimeout` — the Antigravity leg probed for one because
|
||||
* `--print-timeout` cannot fire before a session exists, and STOCK MACOS SHIPS NEITHER, so on a
|
||||
* stock Mac that leg ran unbounded. `spawnSync`'s native `timeout` is always available, so D7's
|
||||
* "skip the probe where no bounding mechanism exists" carve-out is no longer needed.
|
||||
*
|
||||
* EVERY subprocess call in this file passes `timeout` + `killSignal` + `maxBuffer`. This repo has a
|
||||
* named `DEFECT.UNBOUNDED-SUBPROCESS` class (`CONTEXT.md:772`) and an unbounded sync spawn is not a
|
||||
* slow test, it is a hung one: `--test-force-exit` cannot interrupt a synchronous call, so a frozen
|
||||
* spawn hangs the whole CI chunk to its 10-minute kill with `# fail 0` and no `not ok` (#2099).
|
||||
*
|
||||
* TOTAL. No function here throws for a lane-level failure; every one returns a typed outcome. A
|
||||
* reviewer lane that cannot run must produce a diagnosable artifact, because the ambiguity between
|
||||
* "failed" and "ran cleanly with nothing to report" IS the defect this epic closes (#2494/#2605).
|
||||
*/
|
||||
|
||||
import type { LanePlan, SpawnPlan, HttpPlan, LaneUnavailableReason } from './review-lane-invocation.cjs';
|
||||
import {
|
||||
LANE_UNAVAILABLE,
|
||||
isEmptyReview,
|
||||
normalizeHost,
|
||||
fileRefPrompt as fileRefPromptText,
|
||||
} from './review-lane-invocation.cjs';
|
||||
|
||||
/* ------------------------------------------------------------------ *
|
||||
* Injected seams
|
||||
* ------------------------------------------------------------------ */
|
||||
|
||||
export interface SpawnOutcome {
|
||||
status: number | null;
|
||||
stdout: string;
|
||||
stderr: string;
|
||||
/** Node sets this to ETIMEDOUT when `timeout` fired. */
|
||||
errorCode?: string;
|
||||
}
|
||||
|
||||
export interface RunnerDeps {
|
||||
/** Bounded synchronous spawn. Production wires `child_process.spawnSync`. */
|
||||
spawn: (binary: string, argv: string[], opts: { input?: string; timeoutMs: number }) => SpawnOutcome;
|
||||
/** Bounded HTTP POST/GET returning the RAW body — never pre-parsed, so errors stay diagnosable. */
|
||||
httpJson: (url: string, opts: { method: 'GET' | 'POST'; body?: string; timeoutMs: number }) =>
|
||||
Promise<{ ok: boolean; status: number; body: string; error?: string }>;
|
||||
readFile: (p: string) => string;
|
||||
writeFile: (p: string, content: string) => void;
|
||||
exists: (p: string) => boolean;
|
||||
/** `command -v` equivalent. */
|
||||
hasBinary: (name: string) => boolean;
|
||||
/** Resolved config lookup, for the egress-host re-verification. */
|
||||
configGet: (key: string) => unknown;
|
||||
/** Home dir, for the Antigravity transcript paths. */
|
||||
homeDir: string;
|
||||
warn: (msg: string) => void;
|
||||
}
|
||||
|
||||
export interface LaneRunResult {
|
||||
slug: string;
|
||||
ok: boolean;
|
||||
/** Present when `ok` is false. */
|
||||
reason?: LaneUnavailableReason;
|
||||
detail?: string;
|
||||
/** True when a diagnostic stub was written instead of a real review. */
|
||||
stubbed: boolean;
|
||||
}
|
||||
|
||||
/* ------------------------------------------------------------------ *
|
||||
* D5 rules 2–4 — the egress destination is re-verified at INVOCATION
|
||||
* ------------------------------------------------------------------ */
|
||||
|
||||
export interface EgressCheck {
|
||||
allowed: boolean;
|
||||
consentedHost?: string;
|
||||
currentHost?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* Compare a lane's re-resolved egress destination against the one the user consented to.
|
||||
*
|
||||
* ADR-2782 D5 puts this on the invocation path deliberately. `hostConfigKey` names a key in
|
||||
* `.planning/config.json`, which is the ONE consent-bound value living outside the SHA-pinned
|
||||
* bundle: user- and CI-editable at any time, with no re-install and no integrity check. Without
|
||||
* this check a lane consented against `http://localhost:8080` could be silently redirected to a
|
||||
* remote host by an ordinary pull request touching a JSON file, and every subsequent review would
|
||||
* egress plans, requirements, research and decisions to the new destination with no re-prompt.
|
||||
*
|
||||
* ABSENCE IS NOT A MISMATCH, and this is the part that must not be "hardened" later by someone who
|
||||
* reads only the rule name:
|
||||
* - No consent record at all ⇒ ALLOW. First-party lanes (`ollama`, `lm_studio`, `llama_cpp`) ship
|
||||
* inside the SHA-pinned distribution and are never consent-gated, so blocking on a missing
|
||||
* record would break every existing local-model user on upgrade.
|
||||
* - A consent record predating this feature ⇒ ALLOW. Those records were written before a host was
|
||||
* recorded; treating the absence as a change would force spurious re-consent across every
|
||||
* installed capability, which ADR-2782 D4 rule 5 explicitly forbids.
|
||||
*
|
||||
* Hosts are compared NORMALIZED, so a trailing slash or an explicit `:80` is not a "destination
|
||||
* change". A warning that fires on cosmetic edits is a warning users learn to dismiss.
|
||||
*/
|
||||
export function checkEgressHost(consentedHost: unknown, currentHost: string): EgressCheck {
|
||||
const consented = typeof consentedHost === 'string' ? normalizeHost(consentedHost) : '';
|
||||
if (!consented) return { allowed: true };
|
||||
const current = normalizeHost(currentHost);
|
||||
if (consented === current) return { allowed: true, consentedHost: consented, currentHost: current };
|
||||
return { allowed: false, consentedHost: consented, currentHost: current };
|
||||
}
|
||||
|
||||
/* ------------------------------------------------------------------ *
|
||||
* Probe (D7) — every probe bounded, and only for a SELECTED lane
|
||||
* ------------------------------------------------------------------ */
|
||||
|
||||
export async function probeLane(
|
||||
plan: LanePlan,
|
||||
deps: RunnerDeps,
|
||||
): Promise<{ available: true } | { available: false; reason: LaneUnavailableReason; detail: string }> {
|
||||
// Prerequisites first: a missing `jq`-style helper is a clearer failure than whatever the tool
|
||||
// does without it, and reporting it costs nothing.
|
||||
for (const bin of plan.requiresBinaries) {
|
||||
if (!deps.hasBinary(bin)) {
|
||||
return {
|
||||
available: false,
|
||||
reason: LANE_UNAVAILABLE.MISSING_REQUIRED_BINARY,
|
||||
detail: `lane '${plan.slug}' requires '${bin}' on PATH`,
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
const probe = plan.probe;
|
||||
if (!probe || typeof probe !== 'object') {
|
||||
return { available: false, reason: LANE_UNAVAILABLE.PROBE_FAILED, detail: 'lane declares no probe' };
|
||||
}
|
||||
|
||||
switch (probe.kind) {
|
||||
case 'command-exists':
|
||||
return deps.hasBinary(probe.binary)
|
||||
? { available: true }
|
||||
: {
|
||||
available: false,
|
||||
reason: LANE_UNAVAILABLE.MISSING_BINARY,
|
||||
detail: `'${probe.binary}' not found on PATH`,
|
||||
};
|
||||
|
||||
case 'command-capability': {
|
||||
// Existence alone is structurally insufficient here: `kimi` is claimed by BOTH the Kimi Code
|
||||
// CLI and the legacy Python kimi-cli, and an existence-only probe registers the wrong tool.
|
||||
if (!deps.hasBinary(probe.binary)) {
|
||||
return {
|
||||
available: false,
|
||||
reason: LANE_UNAVAILABLE.MISSING_BINARY,
|
||||
detail: `'${probe.binary}' not found on PATH`,
|
||||
};
|
||||
}
|
||||
const out = deps.spawn(probe.binary, ['--help'], { timeoutMs: probe.timeoutMs });
|
||||
if (out.errorCode === 'ETIMEDOUT') {
|
||||
return {
|
||||
available: false,
|
||||
reason: LANE_UNAVAILABLE.PROBE_TIMEOUT,
|
||||
detail: `'${probe.binary} --help' exceeded ${probe.timeoutMs}ms`,
|
||||
};
|
||||
}
|
||||
const help = `${out.stdout}\n${out.stderr}`;
|
||||
return help.includes(probe.needle)
|
||||
? { available: true }
|
||||
: {
|
||||
available: false,
|
||||
reason: LANE_UNAVAILABLE.PROBE_FAILED,
|
||||
detail: `'${probe.binary}' is on PATH but its --help lacks '${probe.needle}' — this is a different tool sharing the name`,
|
||||
};
|
||||
}
|
||||
|
||||
case 'http-reachable': {
|
||||
// Only a STRING config value names a host. `String(unknown)` would turn an object into the
|
||||
// literal '[object Object]' and probe that as a URL — a nonsense destination reported as
|
||||
// "unreachable" rather than as the misconfiguration it is.
|
||||
const configured = deps.configGet(probe.hostConfigKey);
|
||||
const base = normalizeHost(
|
||||
(typeof configured === 'string' ? configured : '') ||
|
||||
(plan.transport === 'openai-http' ? plan.host : ''),
|
||||
);
|
||||
if (!base) {
|
||||
return { available: false, reason: LANE_UNAVAILABLE.HOST_UNREACHABLE, detail: 'no host resolved' };
|
||||
}
|
||||
const r = await deps.httpJson(`${base}${probe.path}`, { method: 'GET', timeoutMs: probe.timeoutMs });
|
||||
return r.ok
|
||||
? { available: true }
|
||||
: {
|
||||
available: false,
|
||||
reason: LANE_UNAVAILABLE.HOST_UNREACHABLE,
|
||||
detail: `${base}${probe.path} unreachable: ${r.error ?? `HTTP ${r.status}`}`,
|
||||
};
|
||||
}
|
||||
|
||||
default:
|
||||
return {
|
||||
available: false,
|
||||
reason: LANE_UNAVAILABLE.PROBE_FAILED,
|
||||
detail: `unknown probe kind '${String((probe as { kind?: unknown }).kind)}'`,
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
/* ------------------------------------------------------------------ *
|
||||
* Empty-output policy (#2494 / #2605 / #2794)
|
||||
* ------------------------------------------------------------------ */
|
||||
|
||||
/**
|
||||
* Write the review, or a diagnostic stub when the lane produced nothing usable.
|
||||
*
|
||||
* The stub is not cosmetic. Before #2494/#2605 a failed lane left a zero-byte file, `write_reviews`
|
||||
* rendered it as "a reviewer that ran cleanly with nothing to report", and the lane vanished from
|
||||
* the cross-AI consensus while `present_results` reported success — a review blind in one eye. The
|
||||
* stub keeps its "failed or returned empty output" header precisely so it can never be mistaken for
|
||||
* a real review.
|
||||
*
|
||||
* `extraDiagnostics` carries the raw HTTP response body for the OpenAI-compatible lanes: an error
|
||||
* from such a server arrives with HTTP 4xx/5xx and the JSON in the BODY, so stderr alone is empty
|
||||
* and the body is the only evidence.
|
||||
*/
|
||||
export function writeReviewOrStub(
|
||||
plan: LanePlan,
|
||||
content: string,
|
||||
deps: RunnerDeps,
|
||||
extraDiagnostics?: string,
|
||||
): { stubbed: boolean } {
|
||||
if (!isEmptyReview(content)) {
|
||||
deps.writeFile(plan.reviewPath, content.endsWith('\n') ? content : `${content}\n`);
|
||||
return { stubbed: false };
|
||||
}
|
||||
const stderr = deps.exists(plan.errPath) ? deps.readFile(plan.errPath) : '';
|
||||
const parts = [`${plan.slug} review failed or returned empty output. stderr:`, stderr];
|
||||
if (extraDiagnostics) parts.push('Raw response body:', extraDiagnostics);
|
||||
deps.writeFile(plan.reviewPath, `${parts.join('\n')}\n`);
|
||||
return { stubbed: true };
|
||||
}
|
||||
|
||||
/* ------------------------------------------------------------------ *
|
||||
* Handlers (D6) — named first-party code, never conditionals in data
|
||||
* ------------------------------------------------------------------ */
|
||||
|
||||
/**
|
||||
* `opencode` — reconstruct the review from the assistant `text` parts of a `--format json` stream.
|
||||
*
|
||||
* `--format json` is the PRIMARY invocation, not a fallback: OpenCode's default `build` agent is an
|
||||
* agentic coder, and on a large review prompt it may run a few `read` tool calls then end its turn
|
||||
* with ZERO output tokens, so `--format default` yields empty stdout and the reviewer is silently
|
||||
* lost (#1936). When no assistant text was emitted, surface the stop reason and output-token count
|
||||
* so the failure is diagnosable rather than a generic empty stub.
|
||||
*
|
||||
* The stream is JSONL-ish: one JSON value per line. A line that does not parse is SKIPPED rather
|
||||
* than failing the lane — a partial stream still carries usable review text, and losing the whole
|
||||
* review to one malformed line would be strictly worse than the bug this handler exists to fix.
|
||||
*/
|
||||
export function handleOpencodeOutput(rawStdout: string): { review: string; diagnostic: string } {
|
||||
const texts: string[] = [];
|
||||
let stopReason = '?';
|
||||
let outputTokens = '?';
|
||||
for (const line of String(rawStdout ?? '').split(/\r?\n/)) {
|
||||
const trimmed = line.trim();
|
||||
if (!trimmed) continue;
|
||||
let evt: unknown;
|
||||
try {
|
||||
evt = JSON.parse(trimmed);
|
||||
} catch {
|
||||
continue;
|
||||
}
|
||||
if (evt === null || typeof evt !== 'object') continue;
|
||||
const e = evt as { type?: unknown; part?: unknown };
|
||||
const part = (e.part ?? {}) as { text?: unknown; reason?: unknown; tokens?: unknown };
|
||||
if (e.type === 'text' && typeof part.text === 'string') {
|
||||
// An EMPTY string is kept, not skipped. The shipped jq was `.part.text // empty`, and `//`
|
||||
// only substitutes for `false`/`null` — an empty string is truthy in jq, so it survived and
|
||||
// contributed a blank line. Dropping it here would quietly close up blank lines between
|
||||
// parts, which is a real (if cosmetic) divergence from the behaviour being ported.
|
||||
texts.push(part.text);
|
||||
} else if (e.type === 'step_finish') {
|
||||
if (typeof part.reason === 'string') stopReason = part.reason;
|
||||
const tokens = (part.tokens ?? {}) as { output?: unknown };
|
||||
if (typeof tokens.output === 'number') outputTokens = String(tokens.output);
|
||||
}
|
||||
}
|
||||
return {
|
||||
review: texts.join('\n'),
|
||||
diagnostic: `stop reason=${stopReason}, output tokens=${outputTokens}`,
|
||||
};
|
||||
}
|
||||
|
||||
/** One `PLANNER_RESPONSE` line of an Antigravity transcript. */
|
||||
interface TranscriptEntry {
|
||||
source?: unknown;
|
||||
status?: unknown;
|
||||
type?: unknown;
|
||||
content?: unknown;
|
||||
}
|
||||
|
||||
/**
|
||||
* `antigravity` — three layers, a two-level timeout, and a stale-response watermark.
|
||||
*
|
||||
* Layer 1 is stdout, which works on macOS/Linux/WSL. On native Windows `agy -p` silently produces
|
||||
* no stdout despite the API call succeeding (an upstream `text_drip.go` non-TTY flush bug), so
|
||||
* layer 2 reads the transcript `agy` always persists to disk. Layer 3 is a diagnostic stub.
|
||||
*
|
||||
* THE WATERMARK IS THE SUBTLE PART. Without it, layer 2 reads the last `PLANNER_RESPONSE` in the
|
||||
* transcript regardless of when it was written — including one from a PREVIOUS invocation in the
|
||||
* same workspace, silently presenting a stale review as this run's. So the transcript line count is
|
||||
* snapshotted BEFORE the spawn and only lines appended after it are considered. If the conversation
|
||||
* id changed, `agy` started a fresh session and every line is new (skip 0).
|
||||
*
|
||||
* `takeWatermark` must therefore be called before `runAntigravity`; the plan's outer timeout is the
|
||||
* external wall-clock cap that `--print-timeout` cannot provide (it cannot fire before a session
|
||||
* exists — a process can stall pre-session and outlive its own native timeout, #2073 mode 3).
|
||||
*
|
||||
* KNOWN LIMIT — the watermark is sequential, not concurrent. `last_conversations.json` is keyed by
|
||||
* WORKSPACE, so two `/gsd:review` runs against the same repo at the same time resolve the same
|
||||
* conversation id and share one transcript. The fallback then takes "the latest DONE
|
||||
* PLANNER_RESPONSE after the watermark" with nothing to tell the two runs apart, so one run could
|
||||
* read the other's response. This cannot be closed here: `agy` exposes no per-invocation id to
|
||||
* filter on, and the transcript carries none. It is stated rather than silently tolerated because
|
||||
* the guarantee this function advertises ("never stale") holds only for sequential use, and a
|
||||
* future reader deserves to know which half is actually guaranteed.
|
||||
*/
|
||||
export function antigravityWatermark(
|
||||
workspace: string,
|
||||
deps: RunnerDeps,
|
||||
): { convId: string; lines: number } {
|
||||
const cachePath = `${deps.homeDir}/.gemini/antigravity-cli/cache/last_conversations.json`;
|
||||
if (!deps.exists(cachePath)) return { convId: '', lines: 0 };
|
||||
let cache: Record<string, unknown>;
|
||||
try {
|
||||
cache = JSON.parse(deps.readFile(cachePath)) as Record<string, unknown>;
|
||||
} catch {
|
||||
return { convId: '', lines: 0 };
|
||||
}
|
||||
const convId = resolveConvId(cache, workspace);
|
||||
if (!convId) return { convId: '', lines: 0 };
|
||||
const tx = transcriptPath(deps.homeDir, convId);
|
||||
if (!deps.exists(tx)) return { convId, lines: 0 };
|
||||
try {
|
||||
return { convId, lines: deps.readFile(tx).split(/\r?\n/).filter((l) => l.trim()).length };
|
||||
} catch {
|
||||
return { convId, lines: 0 };
|
||||
}
|
||||
}
|
||||
|
||||
/** Workspace lookup is case-insensitive — the leg's jq did `ascii_downcase` on both sides. */
|
||||
function resolveConvId(cache: Record<string, unknown>, workspace: string): string {
|
||||
if (Object.prototype.hasOwnProperty.call(cache, workspace)) {
|
||||
const direct = cache[workspace];
|
||||
if (typeof direct === 'string' && direct) return direct;
|
||||
}
|
||||
const target = workspace.toLowerCase();
|
||||
for (const [k, v] of Object.entries(cache)) {
|
||||
if (k.toLowerCase() === target && typeof v === 'string' && v) return v;
|
||||
}
|
||||
return '';
|
||||
}
|
||||
|
||||
function transcriptPath(homeDir: string, convId: string): string {
|
||||
return `${homeDir}/.gemini/antigravity-cli/brain/${convId}/.system_generated/logs/transcript.jsonl`;
|
||||
}
|
||||
|
||||
/**
|
||||
* Layer 2: the newest `PLANNER_RESPONSE` written AFTER the watermark, or `''`.
|
||||
*
|
||||
* Returning `''` rather than the newest entry overall is the whole point — an empty result lets
|
||||
* layer 3 fire with an honest diagnostic, where a stale one would be indistinguishable from a
|
||||
* successful review.
|
||||
*/
|
||||
export function antigravityTranscriptFallback(
|
||||
workspace: string,
|
||||
mark: { convId: string; lines: number },
|
||||
deps: RunnerDeps,
|
||||
): string {
|
||||
const cachePath = `${deps.homeDir}/.gemini/antigravity-cli/cache/last_conversations.json`;
|
||||
if (!deps.exists(cachePath)) return '';
|
||||
let cache: Record<string, unknown>;
|
||||
try {
|
||||
cache = JSON.parse(deps.readFile(cachePath)) as Record<string, unknown>;
|
||||
} catch {
|
||||
return '';
|
||||
}
|
||||
const convId = resolveConvId(cache, workspace);
|
||||
if (!convId) return '';
|
||||
const tx = transcriptPath(deps.homeDir, convId);
|
||||
if (!deps.exists(tx)) return '';
|
||||
let lines: string[];
|
||||
try {
|
||||
lines = deps.readFile(tx).split(/\r?\n/).filter((l) => l.trim());
|
||||
} catch {
|
||||
return '';
|
||||
}
|
||||
// Same conv-id ⇒ only lines beyond the watermark are this run's. Different id ⇒ fresh session,
|
||||
// so every line is new.
|
||||
const skip = convId === mark.convId ? mark.lines : 0;
|
||||
let latest = '';
|
||||
for (const line of lines.slice(skip)) {
|
||||
let entry: TranscriptEntry;
|
||||
try {
|
||||
entry = JSON.parse(line) as TranscriptEntry;
|
||||
} catch {
|
||||
continue;
|
||||
}
|
||||
if (
|
||||
entry.source === 'MODEL' &&
|
||||
entry.status === 'DONE' &&
|
||||
entry.type === 'PLANNER_RESPONSE' &&
|
||||
typeof entry.content === 'string'
|
||||
) {
|
||||
latest = entry.content;
|
||||
}
|
||||
}
|
||||
return latest;
|
||||
}
|
||||
|
||||
/**
|
||||
* Antigravity's prompt variant (#2176).
|
||||
*
|
||||
* Differs from the standard `argv-file-ref` text by one clause, and that clause is load-bearing:
|
||||
* it REQUIRES the reviewer to self-report when it cannot read the repo. Without it a blind review
|
||||
* is indistinguishable from a grounded one — the reviewer happily reviews the plan text in
|
||||
* isolation and its verdict is counted at full weight in the consensus. `stampBlindReview` reads
|
||||
* this self-report back out.
|
||||
*/
|
||||
export function antigravityPrompt(promptPath: string, repoRoot: string): string {
|
||||
return (
|
||||
`Read the file at ${promptPath} in full and carry out the review request it contains. ` +
|
||||
`The repository under review is at ${repoRoot} — resolve every relative file path in the ` +
|
||||
`review request against that absolute root and verify claims against those files. ` +
|
||||
`If you cannot read files under ${repoRoot}, begin your output with the exact line ` +
|
||||
`REVIEWED-WITHOUT-REPO-ACCESS before the review. Output only the resulting markdown review. ` +
|
||||
`Do not edit any files.`
|
||||
);
|
||||
}
|
||||
|
||||
/**
|
||||
* Antigravity's argv adjustments — the two things data cannot express for this lane.
|
||||
*
|
||||
* 1. `--add-dir <repo>`, CAPABILITY-PROBED. Without it agy's permission context never receives the
|
||||
* cwd repo: the agent anchors on its own `~/.gemini/antigravity-cli/scratch` dir and reviews the
|
||||
* plan text in isolation, which is exactly what the Review Instructions forbid (#2176). It is
|
||||
* probed rather than assumed because older agy builds reject the unknown flag outright, and a
|
||||
* lane that fails to start is worse than one that runs on the prompt anchor alone.
|
||||
* 2. The self-report prompt variant above, swapped in for the standard file-ref text.
|
||||
*
|
||||
* Both are argv shape, so they belong here rather than in the descriptor: expressing "add this flag
|
||||
* only if the binary's --help mentions it" as data would need a conditional, which is precisely
|
||||
* what the named-handler seam exists to absorb (ADR-2782 D6).
|
||||
*/
|
||||
export function antigravityArgv(
|
||||
argv: readonly string[],
|
||||
promptPath: string,
|
||||
repoRoot: string,
|
||||
deps: RunnerDeps,
|
||||
): string[] {
|
||||
const standard = fileRefPromptText(promptPath, repoRoot);
|
||||
const out = argv.map((a) => (a === standard ? antigravityPrompt(promptPath, repoRoot) : a));
|
||||
|
||||
let supportsAddDir = false;
|
||||
try {
|
||||
const help = deps.spawn('agy', ['--help'], { timeoutMs: 5_000 });
|
||||
supportsAddDir = `${help.stdout}\n${help.stderr}`.includes('--add-dir');
|
||||
} catch {
|
||||
supportsAddDir = false;
|
||||
}
|
||||
if (!supportsAddDir) return out;
|
||||
|
||||
// Insert before the trailing `-p <prompt>` pair so the prompt stays last, as the leg had it.
|
||||
const pIdx = out.lastIndexOf('-p');
|
||||
if (pIdx === -1) return [...out, '--add-dir', repoRoot];
|
||||
return [...out.slice(0, pIdx), '--add-dir', repoRoot, ...out.slice(pIdx)];
|
||||
}
|
||||
|
||||
/**
|
||||
* Layer 3's diagnostic: what `agy` itself logged, when both stdout and the transcript were empty.
|
||||
*
|
||||
* #2073 mode 2 is invisible without this. A pinned model that 404s server-side exits **0** with
|
||||
* empty stdout AND an empty transcript — every signal the other two layers read says "clean run,
|
||||
* nothing to report". The only evidence is in `agy`'s own log, so a bare generic stub would leave
|
||||
* the user with a silently missing reviewer and nothing to diagnose.
|
||||
*
|
||||
* Mode 3 (a pre-session stall, which `--print-timeout` cannot bound because it cannot fire before a
|
||||
* session exists) leaves no log line at all, so its tell is stated rather than searched for.
|
||||
*/
|
||||
export function antigravityDiagnostic(deps: RunnerDeps): string {
|
||||
const lines = [
|
||||
'Antigravity review failed or returned empty output.',
|
||||
];
|
||||
const logPath = `${deps.homeDir}/.gemini/antigravity-cli/cli.log`;
|
||||
if (deps.exists(logPath)) {
|
||||
try {
|
||||
const hits = deps
|
||||
.readFile(logPath)
|
||||
.split(/\r?\n/)
|
||||
.filter((l) => /agent executor error|NOT_FOUND|Publisher model/i.test(l))
|
||||
.slice(-3);
|
||||
if (hits.length) {
|
||||
lines.push(
|
||||
"agy log hint (pinned model may be unavailable — run 'agy models' and set review.models.agy):",
|
||||
...hits,
|
||||
);
|
||||
}
|
||||
} catch {
|
||||
/* an unreadable log is not worth failing the lane over */
|
||||
}
|
||||
}
|
||||
lines.push(
|
||||
'If no agy run started, that is the pre-session-stall case: check whether a new ' +
|
||||
'~/.gemini/antigravity-cli/brain/<conv-id>/ dir appeared within ~30s of launch.',
|
||||
);
|
||||
return lines.join('\n');
|
||||
}
|
||||
|
||||
/**
|
||||
* Stamp a machine-readable marker when the reviewer plainly ran without repo access (#2176).
|
||||
*
|
||||
* BOTH tells are ANCHORED, and that anchoring is load-bearing: a grounded review that merely QUOTES
|
||||
* `REVIEWED-WITHOUT-REPO-ACCESS` — reviewing this very file, say — must never be mis-stamped. So
|
||||
* the self-report is only honoured in the first five lines, and the scratch-dir tell only in a
|
||||
* workspace-DECLARATION phrasing.
|
||||
*/
|
||||
export function stampBlindReview(review: string): string {
|
||||
if (isEmptyReview(review)) return review;
|
||||
const head = review.split(/\r?\n/).slice(0, 5).join('\n');
|
||||
const selfReported = head.includes('REVIEWED-WITHOUT-REPO-ACCESS');
|
||||
const scratchTell = /(workspace|working) (directory|dir)[\s\S]{0,40}antigravity-cli\/scratch/i.test(review);
|
||||
if (!selfReported && !scratchTell) return review;
|
||||
return (
|
||||
'> [reviewed-without-repo-access] This reviewer ran without visibility into the repo under ' +
|
||||
'review — down-weight its verdict in the Consensus Summary.\n\n' +
|
||||
review
|
||||
);
|
||||
}
|
||||
|
||||
/**
|
||||
* `openai-compatible` — model discovery, the chat-completions round trip, and the served-model
|
||||
* mismatch warning.
|
||||
*
|
||||
* The raw body is returned alongside the content because an OpenAI-compatible server reports errors
|
||||
* with an HTTP 4xx/5xx and the JSON in the BODY. The bash piped the response straight into `jq`,
|
||||
* which discarded exactly that evidence; the stub appends it now.
|
||||
*/
|
||||
export async function runOpenAiCompatible(
|
||||
plan: HttpPlan,
|
||||
promptText: string,
|
||||
deps: RunnerDeps,
|
||||
): Promise<{ review: string; rawBody: string }> {
|
||||
let model = plan.model;
|
||||
if (!model && plan.modelsUrl) {
|
||||
const listed = await deps.httpJson(plan.modelsUrl, { method: 'GET', timeoutMs: 2_000 });
|
||||
if (listed.ok) {
|
||||
try {
|
||||
const parsed = JSON.parse(listed.body) as { data?: Array<{ id?: unknown }> };
|
||||
const first = Array.isArray(parsed.data) ? parsed.data[0] : undefined;
|
||||
if (first && typeof first.id === 'string' && first.id) model = first.id;
|
||||
} catch {
|
||||
/* fall through to the declared fallback */
|
||||
}
|
||||
}
|
||||
}
|
||||
if (!model) model = plan.fallbackModel;
|
||||
|
||||
const body = JSON.stringify({ model, messages: [{ role: 'user', content: promptText }] });
|
||||
const res = await deps.httpJson(plan.url, { method: 'POST', body, timeoutMs: plan.timeoutMs });
|
||||
if (!res.ok && !res.body) {
|
||||
return { review: '', rawBody: res.error ?? `HTTP ${res.status}` };
|
||||
}
|
||||
let review = '';
|
||||
try {
|
||||
const parsed = JSON.parse(res.body) as {
|
||||
model?: unknown;
|
||||
choices?: Array<{ message?: { content?: unknown } }>;
|
||||
};
|
||||
// A server that quietly serves a different model than requested produces a review the user
|
||||
// will attribute to the wrong system. Warn rather than fail — the review is still real.
|
||||
if (typeof parsed.model === 'string' && parsed.model && parsed.model !== model) {
|
||||
deps.warn(
|
||||
`${plan.slug} served model '${parsed.model}' but '${model}' was requested. ` +
|
||||
`Review may be from a different model.`,
|
||||
);
|
||||
}
|
||||
const content = parsed.choices?.[0]?.message?.content;
|
||||
if (typeof content === 'string') review = content;
|
||||
} catch {
|
||||
/* leave review empty — the raw body carries the diagnosis */
|
||||
}
|
||||
return { review, rawBody: res.body };
|
||||
}
|
||||
|
||||
/* ------------------------------------------------------------------ *
|
||||
* Orchestration
|
||||
* ------------------------------------------------------------------ */
|
||||
|
||||
/**
|
||||
* Run one resolved lane end to end.
|
||||
*
|
||||
* Order is deliberate: egress re-verification happens BEFORE the probe and before any spawn, so a
|
||||
* lane whose destination changed never receives the plan text even once.
|
||||
*/
|
||||
export async function runLane(
|
||||
plan: LanePlan,
|
||||
deps: RunnerDeps,
|
||||
opts: { consentedHost?: unknown; explicitlyRequested?: boolean; repoRoot: string },
|
||||
): Promise<LaneRunResult> {
|
||||
const base = { slug: plan.slug, stubbed: false };
|
||||
|
||||
if (plan.transport === 'openai-http') {
|
||||
const egress = checkEgressHost(opts.consentedHost, plan.host);
|
||||
if (!egress.allowed) {
|
||||
const detail =
|
||||
`lane '${plan.slug}' was consented to send plans to ${egress.consentedHost} but ` +
|
||||
`${plan.hostConfigKey} now resolves to ${egress.currentHost}. Re-consent to allow the new ` +
|
||||
`destination — this lane will not run against a host you did not approve.`;
|
||||
deps.warn(detail);
|
||||
return { ...base, ok: false, reason: LANE_UNAVAILABLE.EGRESS_HOST_CHANGED, detail };
|
||||
}
|
||||
}
|
||||
|
||||
const probed = await probeLane(plan, deps);
|
||||
if (!probed.available) {
|
||||
// D4's carve-out: not finding a lane nobody asked for is normal; failing to run a lane somebody
|
||||
// asked for is an ERROR. Both are visible; only the second is fatal to the run.
|
||||
if (opts.explicitlyRequested) deps.warn(`explicitly requested reviewer '${plan.slug}': ${probed.detail}`);
|
||||
return { ...base, ok: false, reason: probed.reason, detail: probed.detail };
|
||||
}
|
||||
|
||||
return plan.transport === 'spawn'
|
||||
? runSpawnLane(plan, deps, opts.repoRoot)
|
||||
: runHttpLane(plan, deps);
|
||||
}
|
||||
|
||||
function runSpawnLane(plan: SpawnPlan, deps: RunnerDeps, repoRoot: string): LaneRunResult {
|
||||
const input = plan.stdin && deps.exists(plan.stdin) ? deps.readFile(plan.stdin) : undefined;
|
||||
|
||||
const mark =
|
||||
plan.handler === 'antigravity' ? antigravityWatermark(repoRoot, deps) : { convId: '', lines: 0 };
|
||||
|
||||
const argv =
|
||||
plan.handler === 'antigravity'
|
||||
? antigravityArgv(plan.argv, plan.promptPath, repoRoot, deps)
|
||||
: plan.argv;
|
||||
|
||||
const out = deps.spawn(plan.binary, argv, { input, timeoutMs: plan.timeoutMs });
|
||||
deps.writeFile(plan.errPath, out.stderr ?? '');
|
||||
|
||||
// `file-arg` lanes write the review themselves and their stdout is deliberately discarded (#1698).
|
||||
let review =
|
||||
plan.outputTarget.kind === 'file'
|
||||
? deps.exists(plan.outputTarget.path)
|
||||
? deps.readFile(plan.outputTarget.path)
|
||||
: ''
|
||||
: (out.stdout ?? '');
|
||||
|
||||
let extra: string | undefined;
|
||||
|
||||
if (plan.handler === 'opencode') {
|
||||
const rebuilt = handleOpencodeOutput(review);
|
||||
if (!isEmptyReview(rebuilt.review)) {
|
||||
review = rebuilt.review;
|
||||
} else {
|
||||
review = '';
|
||||
extra = `OpenCode review returned no assistant text (#1936: agent ended its turn with no final message).\nDiagnostic: ${rebuilt.diagnostic}`;
|
||||
}
|
||||
}
|
||||
|
||||
if (plan.handler === 'antigravity') {
|
||||
// A non-zero exit (timeout kill, crash) discards partial output so the transcript fallback and
|
||||
// the diagnostic stub can take over.
|
||||
if (out.status !== 0) review = '';
|
||||
if (isEmptyReview(review)) review = antigravityTranscriptFallback(repoRoot, mark, deps);
|
||||
review = stampBlindReview(review);
|
||||
// Layer 3. `emptyOutput: 'handler-owned'` means the generic stub does not fire for this lane,
|
||||
// so if nothing is written here the lane goes out empty — the #2073 failure itself.
|
||||
if (isEmptyReview(review)) {
|
||||
deps.writeFile(plan.reviewPath, `${antigravityDiagnostic(deps)}\n`);
|
||||
return { slug: plan.slug, ok: true, stubbed: true };
|
||||
}
|
||||
}
|
||||
|
||||
const { stubbed } = writeReviewOrStub(plan, review, deps, extra);
|
||||
return { slug: plan.slug, ok: true, stubbed };
|
||||
}
|
||||
|
||||
async function runHttpLane(plan: HttpPlan, deps: RunnerDeps): Promise<LaneRunResult> {
|
||||
const promptText = deps.exists(plan.promptPath) ? deps.readFile(plan.promptPath) : '';
|
||||
const { review, rawBody } = await runOpenAiCompatible(plan, promptText, deps);
|
||||
deps.writeFile(plan.errPath, '');
|
||||
const { stubbed } = writeReviewOrStub(plan, review, deps, rawBody);
|
||||
return { slug: plan.slug, ok: true, stubbed };
|
||||
}
|
||||
@@ -1,192 +1,140 @@
|
||||
// allow-test-rule: source-text-is-the-product (see #2073)
|
||||
// gsd-core/workflows/review.md is a workflow document whose bash blocks ARE
|
||||
// what /gsd-review loads and executes at runtime. Asserting the invocation
|
||||
// shape asserts the deployed contract — this is behavioral coverage of the
|
||||
// workflow, not a source-grep over application code.
|
||||
|
||||
/**
|
||||
* Antigravity reviewer repo-grounding tests (#2176)
|
||||
* Antigravity reviewer repo-grounding (#2176).
|
||||
*
|
||||
* The agy block invoked the CLI without granting it the repo under review:
|
||||
* no --add-dir on either invocation arm, and no absolute repo-root anchor in
|
||||
* _AGY_PROMPT. The agent frequently anchored on its own
|
||||
* ~/.gemini/antigravity-cli/scratch dir, reviewed the plan text in isolation
|
||||
* (the exact failure the block's Review Instructions forbid), and its
|
||||
* ungrounded verdict flowed into the Consensus Summary at full weight,
|
||||
* undetected.
|
||||
* The agy leg once invoked the CLI without granting it the repo under review — no `--add-dir` and
|
||||
* no absolute repo-root anchor in the prompt — so the agent anchored on its own
|
||||
* `~/.gemini/antigravity-cli/scratch` dir and reviewed the plan text in isolation, which is exactly
|
||||
* what the Review Instructions forbid.
|
||||
*
|
||||
* These tests pin the fix: capability-probed --add-dir (mirrors the Codex
|
||||
* bypass-flag probe), absolute-root prompt anchor (agy AND the cursor-agent
|
||||
* block, which shared the anchor gap), a mandated self-report line, a stamped
|
||||
* blind-review marker, and consensus down-weighting of marked reviews.
|
||||
* Each assertion fails against the pre-fix block.
|
||||
* Phase 5b (#2799) deleted that bash. Both fixes now live in the named `antigravity` handler, which
|
||||
* is where they belong: "add this flag only when the binary's --help mentions it" is a conditional,
|
||||
* and conditionals escape to a handler rather than accreting inside the descriptor (ADR-2782 D6).
|
||||
* These tests moved with them, so this file no longer reads any source text and the source-text
|
||||
* exemption it used to carry is gone.
|
||||
*/
|
||||
|
||||
const { test, describe } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
|
||||
const ROOT = path.join(__dirname, '..');
|
||||
const REVIEW_PATH = path.join(ROOT, 'gsd-core', 'workflows', 'review.md');
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const {
|
||||
antigravityArgv,
|
||||
antigravityPrompt,
|
||||
stampBlindReview,
|
||||
} = require('../gsd-core/bin/lib/review-lane-runner.cjs');
|
||||
|
||||
function reviewContent() {
|
||||
return fs.readFileSync(REVIEW_PATH, 'utf-8');
|
||||
const RUN = '/run';
|
||||
const ROOT = '/abs/repo';
|
||||
|
||||
function planFor(slug) {
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === slug);
|
||||
const r = resolveLanePlan({ lane, configGet: () => undefined, runDir: RUN, repoRoot: ROOT });
|
||||
assert.equal(r.ok, true);
|
||||
return r.plan;
|
||||
}
|
||||
|
||||
function agyBashBlock() {
|
||||
const fences = reviewContent().match(/```bash[\s\S]*?```/g) || [];
|
||||
const agy = fences.find((f) => /\bagy\b/.test(f) && /gsd-review-antigravity/.test(f));
|
||||
assert.ok(agy, 'review.md should contain the agy invocation bash block');
|
||||
return agy;
|
||||
}
|
||||
/** A spawn stub whose `--help` either advertises `--add-dir` or does not. */
|
||||
const helpSaying = (text) => ({ spawn: () => ({ status: 0, stdout: text, stderr: '' }) });
|
||||
|
||||
function cursorBashBlock() {
|
||||
const fences = reviewContent().match(/```bash[\s\S]*?```/g) || [];
|
||||
const cursor = fences.find((f) => /cursor-agent -p/.test(f));
|
||||
assert.ok(cursor, 'review.md should contain the cursor-agent invocation bash block');
|
||||
return cursor;
|
||||
}
|
||||
|
||||
describe('Antigravity reviewer repo grounding in /gsd-review (#2176)', () => {
|
||||
test('probes agy for --add-dir support (capability-probe idiom, mirrors the Codex block)', () => {
|
||||
const block = agyBashBlock();
|
||||
assert.ok(
|
||||
/agy --help 2>\/dev\/null \| grep -q -- '--add-dir'/.test(block),
|
||||
'agy block must capability-probe --add-dir via `agy --help | grep -q` so older CLIs still run',
|
||||
);
|
||||
describe('#2176 — agy receives the repo under review', () => {
|
||||
test('--add-dir is passed with the repo root when the binary supports it', () => {
|
||||
const p = planFor('antigravity');
|
||||
const argv = antigravityArgv(p.argv, p.promptPath, ROOT, helpSaying('--add-dir --model'));
|
||||
const i = argv.indexOf('--add-dir');
|
||||
assert.ok(i !== -1, '--add-dir must be passed when supported');
|
||||
assert.equal(argv[i + 1], ROOT);
|
||||
});
|
||||
|
||||
test('passes --add-dir with the repo root when supported', () => {
|
||||
const block = agyBashBlock();
|
||||
assert.ok(
|
||||
/set -- "\$@" --add-dir "\$_AGY_WS"/.test(block),
|
||||
'the probed arm must append --add-dir "$_AGY_WS" so both invocation arms (which expand "$@") receive it',
|
||||
);
|
||||
test('--add-dir is CAPABILITY-PROBED, not assumed', () => {
|
||||
// An older agy rejects an unknown flag outright. A lane that fails to start is worse than one
|
||||
// running on the prompt anchor alone, so support is probed rather than presumed.
|
||||
const p = planFor('antigravity');
|
||||
const argv = antigravityArgv(p.argv, p.promptPath, ROOT, helpSaying('no such flag here'));
|
||||
assert.ok(!argv.includes('--add-dir'));
|
||||
});
|
||||
|
||||
test('_AGY_PROMPT is anchored to the absolute repo root', () => {
|
||||
const block = agyBashBlock();
|
||||
const promptLine = block.split('\n').find((l) => l.startsWith('_AGY_PROMPT='));
|
||||
assert.ok(promptLine, 'agy block must define _AGY_PROMPT');
|
||||
assert.ok(
|
||||
/\$_AGY_WS/.test(promptLine),
|
||||
'_AGY_PROMPT must embed the absolute repo root ($_AGY_WS) so repo-relative references resolve on the no---add-dir fallback',
|
||||
);
|
||||
test('the probe is bounded', () => {
|
||||
const p = planFor('antigravity');
|
||||
const calls = [];
|
||||
antigravityArgv(p.argv, p.promptPath, ROOT, {
|
||||
spawn: (b, a, o) => { calls.push(o); return { status: 0, stdout: '', stderr: '' }; },
|
||||
});
|
||||
assert.ok(calls.every((c) => c.timeoutMs > 0), 'every process-starting probe must be bounded');
|
||||
});
|
||||
|
||||
test('_AGY_PROMPT mandates a REVIEWED-WITHOUT-REPO-ACCESS self-report', () => {
|
||||
const block = agyBashBlock();
|
||||
const promptLine = block.split('\n').find((l) => l.startsWith('_AGY_PROMPT='));
|
||||
assert.ok(
|
||||
/REVIEWED-WITHOUT-REPO-ACCESS/.test(promptLine),
|
||||
'_AGY_PROMPT must require the exact self-report line when the reviewer cannot read the repo',
|
||||
);
|
||||
test('adding --add-dir keeps the prompt last', () => {
|
||||
const p = planFor('antigravity');
|
||||
const argv = antigravityArgv(p.argv, p.promptPath, ROOT, helpSaying('--add-dir'));
|
||||
assert.equal(argv[argv.length - 2], '-p');
|
||||
assert.ok(argv[argv.length - 1].includes(RUN));
|
||||
});
|
||||
|
||||
test('stamps a blind-review marker on self-reported or scratch-anchored output', () => {
|
||||
const block = agyBashBlock();
|
||||
assert.ok(
|
||||
/head -5 [^|]*\| grep -q 'REVIEWED-WITHOUT-REPO-ACCESS'/.test(block),
|
||||
'the self-report tell must be anchored to the head of the output, not a whole-body substring',
|
||||
);
|
||||
assert.ok(
|
||||
/grep -qiE '\(workspace\|working\) \(directory\|dir\)\.\{0,40\}antigravity-cli\/scratch'/.test(block),
|
||||
'the scratch tell must be anchored to a workspace-declaration phrasing whose bridge (.{0,40}) can span the dotted ~/.gemini/ path prefix',
|
||||
);
|
||||
assert.ok(
|
||||
/\[reviewed-without-repo-access\]/.test(block),
|
||||
'detected blind reviews must be stamped with the [reviewed-without-repo-access] marker',
|
||||
);
|
||||
test('a probe that throws degrades to no flag rather than failing the lane', () => {
|
||||
const p = planFor('antigravity');
|
||||
const argv = antigravityArgv(p.argv, p.promptPath, ROOT, {
|
||||
spawn: () => { throw new Error('spawn failed'); },
|
||||
});
|
||||
assert.ok(!argv.includes('--add-dir'));
|
||||
assert.ok(argv.length > 0);
|
||||
});
|
||||
});
|
||||
|
||||
test('marker stamping avoids sed -i (BSD/GNU divergence) — uses temp file + mv', () => {
|
||||
const block = agyBashBlock();
|
||||
assert.ok(!/sed -i/.test(block), 'agy block must not use sed -i (BSD vs GNU incompatibility)');
|
||||
assert.ok(
|
||||
/\.tmp && \\?\s*\n?\s*mv /.test(block),
|
||||
'marker stamping should rewrite via temp file + mv',
|
||||
);
|
||||
describe('#2176 — the prompt is anchored and demands a self-report', () => {
|
||||
test('the prompt carries the ABSOLUTE repo root', () => {
|
||||
const prompt = antigravityPrompt(`${RUN}/gsd-review-prompt.md`, ROOT);
|
||||
assert.ok(prompt.includes(ROOT));
|
||||
});
|
||||
|
||||
test('Consensus Summary down-weights marked blind reviews', () => {
|
||||
const content = reviewContent();
|
||||
const consensusIdx = content.indexOf('## Consensus Summary');
|
||||
assert.ok(consensusIdx >= 0, 'review.md should contain the Consensus Summary section');
|
||||
const consensus = content.slice(consensusIdx, consensusIdx + 2000);
|
||||
assert.ok(
|
||||
/\[reviewed-without-repo-access\]/.test(consensus),
|
||||
'consensus instructions must reference the blind-review marker',
|
||||
);
|
||||
assert.ok(
|
||||
/REVIEWED-WITHOUT-REPO-ACCESS/.test(consensus),
|
||||
'consensus instructions must also honor the raw self-report line',
|
||||
);
|
||||
assert.ok(
|
||||
/not count its verdict at full consensus weight/.test(consensus),
|
||||
'marked reviews must be down-weighted, not counted at full weight',
|
||||
);
|
||||
test('the prompt mandates REVIEWED-WITHOUT-REPO-ACCESS when the repo is unreadable', () => {
|
||||
// Without this clause a blind review is indistinguishable from a grounded one and its verdict
|
||||
// is counted at full consensus weight.
|
||||
const prompt = antigravityPrompt(`${RUN}/gsd-review-prompt.md`, ROOT);
|
||||
assert.ok(prompt.includes('REVIEWED-WITHOUT-REPO-ACCESS'));
|
||||
});
|
||||
|
||||
test('blind-review detection behaves correctly on synthetic transcripts', () => {
|
||||
// Behavioral, not string-on-string: extract the actual detection compound
|
||||
// from the fence, point it at a temp file, and run it through bash for
|
||||
// ungrounded and grounded transcript shapes.
|
||||
const os = require('os');
|
||||
const { execFileSync } = require('node:child_process');
|
||||
const { toPosixPath } = require('../gsd-core/bin/lib/shell-command-projection.cjs');
|
||||
const block = agyBashBlock();
|
||||
const m = block.match(/\{ head -5[\s\S]*?\}; then/);
|
||||
assert.ok(m, 'detection compound not found in the agy block');
|
||||
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), 'agy-detect-'));
|
||||
const out = path.join(tmp, 'review-out.md');
|
||||
const detect = m[0]
|
||||
.replace(/\}; then$/, '}')
|
||||
// Convert the native path to POSIX form so it survives bash on Windows
|
||||
// runners (Git Bash accepts D:/... but eats backslashes).
|
||||
.replaceAll('{run_dir}/gsd-review-antigravity.md', toPosixPath(out));
|
||||
const runDetect = (content) => {
|
||||
fs.writeFileSync(out, content);
|
||||
try {
|
||||
execFileSync('bash', ['-c', detect], { stdio: 'ignore' });
|
||||
return true; // exit 0 → blind review detected
|
||||
} catch {
|
||||
return false;
|
||||
}
|
||||
};
|
||||
try {
|
||||
// Ungrounded tells — must be stamped
|
||||
assert.equal(runDetect('REVIEWED-WITHOUT-REPO-ACCESS\n\n## Review\nPlan-only review.\n'), true,
|
||||
'self-report line in the head must be detected');
|
||||
assert.equal(runDetect('## Review\nMy working directory is ~/.gemini/antigravity-cli/scratch.\nPlan-only review.\n'), true,
|
||||
'full dotted scratch path in a workspace declaration must be detected (#2184 re-review Major)');
|
||||
assert.equal(runDetect('Workspace directory: /home/me/.gemini/antigravity-cli/scratch\n'), true,
|
||||
'colon-style workspace declaration must be detected');
|
||||
// Grounded shapes — must NOT be stamped
|
||||
assert.equal(runDetect('## Review\nVerified hooks/gsd-statusline.js against the plan. Solid.\n'), false,
|
||||
'ordinary grounded review must not be stamped');
|
||||
assert.equal(runDetect('## Review\nThe workflow mentions antigravity-cli/scratch as the agy scratch dir; the guard there is correct.\n'), false,
|
||||
'grounded review merely quoting the scratch path must not be stamped');
|
||||
assert.equal(
|
||||
runDetect('## Review\n\nGrounded findings below.\n\n### Details\nLine 20 of the workflow mentions REVIEWED-WITHOUT-REPO-ACCESS as the self-report marker; fine.\n'),
|
||||
false,
|
||||
'self-report string quoted beyond the first 5 lines must not be stamped');
|
||||
} finally {
|
||||
require('./helpers.cjs').cleanup(tmp);
|
||||
}
|
||||
test('the self-report variant actually reaches argv', () => {
|
||||
const p = planFor('antigravity');
|
||||
const argv = antigravityArgv(p.argv, p.promptPath, ROOT, helpSaying(''));
|
||||
assert.ok(argv[argv.length - 1].includes('REVIEWED-WITHOUT-REPO-ACCESS'));
|
||||
});
|
||||
|
||||
test('cursor-agent prompt carries the same absolute-root anchor (identical gap, #2176 AC5)', () => {
|
||||
const block = cursorBashBlock();
|
||||
assert.ok(
|
||||
/_CURSOR_ROOT="\$\(git rev-parse --show-toplevel 2>\/dev\/null \|\| pwd\)"/.test(block),
|
||||
'cursor anchor must resolve the repo TOP-LEVEL (rev-parse, not bare pwd) so subdirectory invocations anchor correctly',
|
||||
);
|
||||
const promptLine = block.split('\n').find((l) => l.startsWith('CURSOR_PROMPT_ARG='));
|
||||
assert.ok(promptLine, 'cursor block must define CURSOR_PROMPT_ARG');
|
||||
assert.ok(
|
||||
/repository under review is at \$_CURSOR_ROOT/.test(promptLine),
|
||||
'CURSOR_PROMPT_ARG must anchor repo-relative references to the absolute repo root',
|
||||
);
|
||||
test('cursor carries the same absolute-root anchor (#2176 AC5)', () => {
|
||||
// Identical gap, identical fix — cursor-agent also takes the prompt in argv and does not
|
||||
// reliably inherit the review cwd.
|
||||
const p = planFor('cursor');
|
||||
assert.ok(p.argv[p.argv.length - 1].includes(ROOT));
|
||||
});
|
||||
|
||||
test('kimi-code carries it too', () => {
|
||||
const p = planFor('kimi-code');
|
||||
assert.ok(p.argv[p.argv.length - 1].includes(ROOT));
|
||||
});
|
||||
});
|
||||
|
||||
describe('#2176 — blind-review marking', () => {
|
||||
test('a self-reported blind review is stamped', () => {
|
||||
const out = stampBlindReview('REVIEWED-WITHOUT-REPO-ACCESS\nthe review');
|
||||
assert.ok(out.includes('[reviewed-without-repo-access]'));
|
||||
});
|
||||
|
||||
test('a scratch-dir workspace DECLARATION is stamped', () => {
|
||||
const out = stampBlindReview('my working directory is ~/.gemini/antigravity-cli/scratch, so');
|
||||
assert.ok(out.includes('[reviewed-without-repo-access]'));
|
||||
});
|
||||
|
||||
test('a grounded review that merely QUOTES the marker is not stamped', () => {
|
||||
// A review OF this very file must not be mis-stamped and down-weighted.
|
||||
const quoting = ['a', 'b', 'c', 'd', 'e', 'f', 'the tell is REVIEWED-WITHOUT-REPO-ACCESS'].join('\n');
|
||||
assert.ok(!stampBlindReview(quoting).includes('[reviewed-without-repo-access]'));
|
||||
});
|
||||
|
||||
test('an ordinary mention of the scratch path is not a declaration', () => {
|
||||
const mention = 'the plan references .gemini/antigravity-cli/scratch as an example path';
|
||||
assert.ok(!stampBlindReview(mention).includes('[reviewed-without-repo-access]'));
|
||||
});
|
||||
|
||||
test('an empty review is left alone', () => {
|
||||
assert.equal(stampBlindReview(''), '');
|
||||
});
|
||||
});
|
||||
|
||||
@@ -1,109 +1,203 @@
|
||||
// allow-test-rule: source-text-is-the-product (see #2073)
|
||||
// gsd-core/workflows/review.md is a workflow document whose bash blocks ARE
|
||||
// what /gsd-review loads and executes at runtime. Asserting the Antigravity
|
||||
// invocation shape asserts the deployed contract — this is behavioral coverage
|
||||
// of the workflow, not a source-grep over application code.
|
||||
|
||||
/**
|
||||
* Antigravity (agy) reviewer invocation tests (#2073)
|
||||
* Antigravity (agy) reviewer lane — the #2073 invariants.
|
||||
*
|
||||
* The agy block in /gsd-review had three real-world failure modes on agy 1.0.16:
|
||||
* 1. inline `-p "$(cat <prompt>)"` overflowed the exec arg list on a large prompt
|
||||
* 2. a pinned model that 404'd exited 0 with empty stdout AND an empty transcript
|
||||
* (no --model escape hatch; the generic Step 3 stub gave no diagnostic)
|
||||
* 3. a pre-session stall hung past --print-timeout (which can't fire before a
|
||||
* session exists) because there was no external wall-clock `timeout`
|
||||
* Plus a stale maintainer note claiming agy has no --model flag.
|
||||
* These tests pinned the three real failure modes of agy 1.0.16 against the hand-authored bash
|
||||
* block in `gsd-core/workflows/review.md`. Phase 5b (#2799) deleted that block: the lane is now
|
||||
* declared data plus a named first-party handler, so the assertions move to those surfaces — and
|
||||
* the source-text exemption this file used to carry is no longer needed, because nothing here reads
|
||||
* a source file any more.
|
||||
*
|
||||
* These tests pin the corrected invocation shape in gsd-core/workflows/review.md.
|
||||
* The move is an upgrade, not a translation. The old tests could only prove that certain TEXT
|
||||
* appeared in a markdown fence; these prove the resolved invocation and the handler's actual
|
||||
* behaviour. Every #2073 mode below is the same invariant, checked where it now lives.
|
||||
*
|
||||
* One invariant deliberately CHANGED, and is recorded here rather than silently dropped: mode 3's
|
||||
* external `timeout`/`gtimeout` probe is gone. The bash needed it because `--print-timeout` cannot
|
||||
* fire before agy creates a session, and stock macOS ships neither killer — so on a stock Mac that
|
||||
* leg ran unbounded. The runner spawns with Node's native timeout, which is always available, so
|
||||
* the outer bound is now unconditional. The invariant ("a pre-session stall is bounded") got
|
||||
* stronger; only the mechanism changed.
|
||||
*/
|
||||
|
||||
const { test, describe } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
|
||||
const ROOT = path.join(__dirname, '..');
|
||||
const REVIEW_PATH = path.join(ROOT, 'gsd-core', 'workflows', 'review.md');
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const {
|
||||
antigravityDiagnostic,
|
||||
antigravityTranscriptFallback,
|
||||
stampBlindReview,
|
||||
runLane,
|
||||
} = require('../gsd-core/bin/lib/review-lane-runner.cjs');
|
||||
|
||||
function agyBashBlock() {
|
||||
const content = fs.readFileSync(REVIEW_PATH, 'utf-8');
|
||||
const fences = content.match(/```bash[\s\S]*?```/g) || [];
|
||||
// Target the INVOCATION block (the fence that writes the antigravity review
|
||||
// output), not the `command -v agy` detection one-liner.
|
||||
const agy = fences.find((f) => /\bagy\b/.test(f) && /gsd-review-antigravity/.test(f));
|
||||
assert.ok(agy, 'review.md should contain the agy invocation bash block');
|
||||
return agy;
|
||||
const RUN = '/run';
|
||||
const ROOT = '/repo';
|
||||
const HOME = '/home/u';
|
||||
|
||||
const LANE = REVIEWER_LANES.find((l) => l.slug === 'antigravity');
|
||||
|
||||
function planFor(config = {}) {
|
||||
const r = resolveLanePlan({ lane: LANE, configGet: (k) => config[k], runDir: RUN, repoRoot: ROOT });
|
||||
assert.equal(r.ok, true);
|
||||
return r.plan;
|
||||
}
|
||||
|
||||
describe('Antigravity (agy) reviewer invocation in /gsd-review (#2073)', () => {
|
||||
test('review.md exists', () => {
|
||||
assert.ok(fs.existsSync(REVIEW_PATH), 'review.md should exist');
|
||||
function deps(overrides = {}) {
|
||||
const files = overrides.files || {};
|
||||
const warnings = [];
|
||||
return Object.assign(
|
||||
{
|
||||
files,
|
||||
warnings,
|
||||
spawn: () => ({ status: 0, stdout: '', stderr: '' }),
|
||||
httpJson: async () => ({ ok: true, status: 200, body: '{}' }),
|
||||
readFile: (p) => { if (!(p in files)) throw new Error(`ENOENT ${p}`); return files[p]; },
|
||||
writeFile: (p, c) => { files[p] = c; },
|
||||
exists: (p) => p in files,
|
||||
hasBinary: () => true,
|
||||
configGet: () => undefined,
|
||||
homeDir: HOME,
|
||||
warn: (m) => warnings.push(m),
|
||||
},
|
||||
overrides,
|
||||
);
|
||||
}
|
||||
|
||||
describe('Antigravity lane — #2073 mode 1 (exec arg-list overflow)', () => {
|
||||
test('the prompt travels by FILE REFERENCE, never inlined', () => {
|
||||
// Inline `-p "$(cat <prompt>)"` overflowed the exec arg list on a large prompt set (~197 KB),
|
||||
// failing with rc 126 in a way indistinguishable from a model failure.
|
||||
assert.equal(LANE.invoke.promptChannel, 'argv-file-ref');
|
||||
const argv = planFor().argv;
|
||||
const promptArg = argv[argv.length - 1];
|
||||
assert.ok(promptArg.includes(`${RUN}/gsd-review-prompt.md`), 'must reference the prompt file');
|
||||
assert.ok(promptArg.length < 1000, 'the reference must stay short, never carry the prompt body');
|
||||
});
|
||||
|
||||
test('#2073 mode 1 — does NOT inline the prompt via "$(cat ...)" (arg-list overflow)', () => {
|
||||
const block = agyBashBlock();
|
||||
test('the prompt argument carries the absolute repo root (#2176)', () => {
|
||||
// Without the anchor agy reviews the plan text in isolation from its own scratch dir.
|
||||
const argv = planFor().argv;
|
||||
assert.ok(argv[argv.length - 1].includes(ROOT));
|
||||
});
|
||||
});
|
||||
|
||||
describe('Antigravity lane — #2073 mode 2 (a pinned model that 404s)', () => {
|
||||
test('review.models.agy reaches argv as --model', () => {
|
||||
// The slug is `antigravity` but the shipped key is `review.models.agy`; resolving by slug
|
||||
// would silently ignore the very escape hatch this mode exists to provide.
|
||||
assert.equal(LANE.modelConfigKey, 'review.models.agy');
|
||||
const argv = planFor({ 'review.models.agy': 'gemini-x' }).argv;
|
||||
const i = argv.indexOf('--model');
|
||||
assert.ok(i !== -1, '--model must be present when the key is set');
|
||||
assert.equal(argv[i + 1], 'gemini-x');
|
||||
});
|
||||
|
||||
test('no model configured emits no --model', () => {
|
||||
assert.ok(!planFor().argv.includes('--model'));
|
||||
});
|
||||
|
||||
test('the diagnostic surfaces agy cli.log rather than a generic stub', () => {
|
||||
// This mode exits 0 with empty stdout AND an empty transcript — every other signal reads as a
|
||||
// clean run. The log is the only evidence there was a failure at all.
|
||||
const files = {
|
||||
[`${HOME}/.gemini/antigravity-cli/cli.log`]: [
|
||||
'routine',
|
||||
'agent executor error: Publisher model NOT_FOUND gemini-x',
|
||||
].join('\n'),
|
||||
};
|
||||
const out = antigravityDiagnostic(deps({ files }));
|
||||
assert.ok(out.includes('NOT_FOUND'), 'the log line must be surfaced');
|
||||
assert.ok(out.includes('agy models'), 'the remedy must be named');
|
||||
});
|
||||
|
||||
test('an absent log still yields the pre-session-stall tell', () => {
|
||||
const out = antigravityDiagnostic(deps());
|
||||
assert.ok(out.includes('pre-session-stall'));
|
||||
assert.ok(!out.includes('agy models'), 'no hint should be invented without evidence');
|
||||
});
|
||||
});
|
||||
|
||||
describe('Antigravity lane — #2073 mode 3 (pre-session stall)', () => {
|
||||
test('the outer wall-clock bound is declared and unconditional', () => {
|
||||
// The bash probed for GNU `timeout` / `gtimeout` and fell back to --print-timeout alone when
|
||||
// neither existed — which is stock macOS, so that leg ran unbounded there. The bound is now
|
||||
// the plan's own timeout, enforced by the spawn on every platform.
|
||||
assert.equal(LANE.timeoutFloorMs, 600000);
|
||||
assert.equal(planFor().timeoutMs, 600000);
|
||||
});
|
||||
|
||||
test('the tool-native inner timeout stays in argv', () => {
|
||||
// Two-level by design: a 600s outer cap over a 540s native --print-timeout. The descriptor
|
||||
// carries ONE scalar; the inner bound is the lane's own argument (ADR-2782 D6).
|
||||
const argv = planFor().argv;
|
||||
const i = argv.indexOf('--print-timeout');
|
||||
assert.ok(i !== -1);
|
||||
assert.equal(argv[i + 1], '540s');
|
||||
});
|
||||
|
||||
test('the REVIEW spawn receives the outer bound, and every spawn is bounded', async () => {
|
||||
const p = planFor();
|
||||
const calls = [];
|
||||
const d = deps({
|
||||
spawn: (b, a, o) => { calls.push({ argv: a, opts: o }); return { status: 0, stdout: 'a review', stderr: '' }; },
|
||||
});
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
|
||||
// The handler also spawns `agy --help` to probe --add-dir, so target the review invocation by
|
||||
// its shape rather than by position — a positional assertion silently follows the wrong call
|
||||
// the moment another probe is added.
|
||||
const review = calls.find((c) => c.argv.includes('-p'));
|
||||
assert.ok(review, 'expected a review invocation');
|
||||
assert.equal(review.opts.timeoutMs, 600000);
|
||||
|
||||
for (const c of calls) {
|
||||
assert.ok(c.opts.timeoutMs > 0, `unbounded spawn: ${c.argv.join(' ')}`);
|
||||
}
|
||||
});
|
||||
|
||||
test('no prompt is fed on stdin, so a tty stall is impossible', () => {
|
||||
// The bash tied stdin to /dev/null for exactly this reason.
|
||||
assert.equal(planFor().stdin, null);
|
||||
});
|
||||
});
|
||||
|
||||
describe('Antigravity lane — transcript fallback and staleness', () => {
|
||||
const CACHE = `${HOME}/.gemini/antigravity-cli/cache/last_conversations.json`;
|
||||
const TX = (id) => `${HOME}/.gemini/antigravity-cli/brain/${id}/.system_generated/logs/transcript.jsonl`;
|
||||
const entry = (content) =>
|
||||
JSON.stringify({ source: 'MODEL', status: 'DONE', type: 'PLANNER_RESPONSE', content });
|
||||
|
||||
test('the lane is handler-owned, so the generic stub cannot pre-empt the fallback', () => {
|
||||
assert.equal(LANE.emptyOutput, 'handler-owned');
|
||||
assert.equal(LANE.handler, 'antigravity');
|
||||
});
|
||||
|
||||
test('a response written by THIS run is read back', () => {
|
||||
const files = {
|
||||
[CACHE]: JSON.stringify({ [ROOT]: 'c1' }),
|
||||
[TX('c1')]: [entry('old'), entry('NEW')].join('\n'),
|
||||
};
|
||||
assert.equal(antigravityTranscriptFallback(ROOT, { convId: 'c1', lines: 1 }, deps({ files })), 'NEW');
|
||||
});
|
||||
|
||||
test('a response from a PRIOR run is never presented as this one', () => {
|
||||
const files = { [CACHE]: JSON.stringify({ [ROOT]: 'c1' }), [TX('c1')]: entry('STALE') };
|
||||
assert.equal(antigravityTranscriptFallback(ROOT, { convId: 'c1', lines: 1 }, deps({ files })), '');
|
||||
});
|
||||
|
||||
test('a blind review is marked so consensus can down-weight it (#2176)', () => {
|
||||
assert.ok(
|
||||
!/"\$\(cat/.test(block),
|
||||
'agy must not inline the prompt via "$(cat …)" — a large review prompt overflows the exec arg list',
|
||||
stampBlindReview('REVIEWED-WITHOUT-REPO-ACCESS\nx').includes('[reviewed-without-repo-access]'),
|
||||
);
|
||||
});
|
||||
|
||||
test('#2073 mode 1 — uses a file-reference prompt (mirrors the Cursor block)', () => {
|
||||
const block = agyBashBlock();
|
||||
assert.ok(
|
||||
/Read the file at \{run_dir\}\/gsd-review-prompt\.md/.test(block),
|
||||
'agy should reference the prompt by file path, not inline it',
|
||||
);
|
||||
});
|
||||
|
||||
test('#2073 mode 3 — pairs agy with an external wall-clock killer when available (timeout/gtimeout probe)', () => {
|
||||
const block = agyBashBlock();
|
||||
// Capability probe for GNU `timeout` and macOS `gtimeout` (stock macOS has neither).
|
||||
assert.match(block, /command -v timeout/, 'agy block should probe for the `timeout` killer');
|
||||
assert.match(block, /command -v gtimeout/, 'agy block should probe for `gtimeout` (macOS Homebrew)');
|
||||
// The external cap (600s) is >= agy's native --print-timeout (540s) so it only
|
||||
// backstops a pre-session stall, never cuts a healthy run.
|
||||
assert.match(block, /600 agy --print-timeout 540s/, 'external cap (600s) must be >= --print-timeout (540s)');
|
||||
// Graceful fallback when no external killer is available (stock macOS).
|
||||
assert.match(block, /else\n\s*agy --print-timeout 540s/,
|
||||
'agy block must fall back to --print-timeout alone when no external killer is available (macOS)');
|
||||
});
|
||||
|
||||
test('#2073 mode 3 — stdin is tied to /dev/null (no tty stall)', () => {
|
||||
const block = agyBashBlock();
|
||||
assert.ok(
|
||||
/<\/dev\/null/.test(block),
|
||||
'agy invocation should redirect stdin from /dev/null so it never blocks on a tty',
|
||||
);
|
||||
});
|
||||
|
||||
test('#2073 mode 2 — wires review.models.agy via --model when configured', () => {
|
||||
const block = agyBashBlock();
|
||||
assert.match(block, /--model "\$AGY_MODEL"/, 'agy block should pass --model "$AGY_MODEL" when set');
|
||||
});
|
||||
|
||||
test('#2073 mode 2 — Step 3 stub surfaces a diagnostic from agy cli.log (not just a generic stub)', () => {
|
||||
const block = agyBashBlock();
|
||||
assert.ok(
|
||||
/cli\.log/.test(block),
|
||||
'the empty-output stub should inspect agy cli.log for a model-availability diagnostic (NOT_FOUND / agent executor error)',
|
||||
);
|
||||
});
|
||||
|
||||
test('#2073 — stale "no --model flag" maintainer note is corrected', () => {
|
||||
const content = fs.readFileSync(REVIEW_PATH, 'utf-8');
|
||||
assert.ok(
|
||||
!/No .{0,4}-m.{0,4}\/.{0,4}--model.{0,4} flag/i.test(content),
|
||||
'the stale maintainer note claiming agy has no --model flag must be corrected (--model exists since ~1.0.3)',
|
||||
);
|
||||
});
|
||||
|
||||
test('#2073 — review.models.agy is documented as supported (not "reserved for future")', () => {
|
||||
const content = fs.readFileSync(REVIEW_PATH, 'utf-8');
|
||||
assert.ok(
|
||||
!/review\.models\.agy is reserved for future/i.test(content),
|
||||
'review.models.agy is now wired (passed as --model); the "reserved for future" comment must be updated',
|
||||
);
|
||||
test('all three layers failing still writes a diagnosable artifact, never an empty file', async () => {
|
||||
const p = planFor();
|
||||
const d = deps({ spawn: () => ({ status: 1, stdout: '', stderr: 'boom' }) });
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, true);
|
||||
assert.ok(d.files[p.reviewPath].includes('failed or returned empty output'));
|
||||
assert.ok(d.files[p.reviewPath].includes('pre-session-stall'));
|
||||
});
|
||||
});
|
||||
|
||||
@@ -60,28 +60,35 @@ describe('Cursor CLI reviewer in /gsd-review (#1960)', () => {
|
||||
);
|
||||
});
|
||||
|
||||
test('invocation uses cursor-agent single binary with -p flag', () => {
|
||||
const c = fs.readFileSync(reviewPath, 'utf-8');
|
||||
assert.ok(
|
||||
c.includes('cursor-agent -p'),
|
||||
'review.md should invoke cursor via "cursor-agent -p" (single binary, not two-token "cursor agent")'
|
||||
);
|
||||
// Phase 5b (#2799) moved the invocation out of review.md's bash and into the declared lane.
|
||||
// These four assertions follow it: the descriptor and the resolved plan ARE the deployed
|
||||
// contract now, and asserting on them is strictly stronger than matching fence text.
|
||||
test('invocation uses the cursor-agent single binary, not two-token "cursor agent"', () => {
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === 'cursor');
|
||||
assert.equal(lane.invoke.binary, 'cursor-agent');
|
||||
assert.ok(lane.invoke.args.includes('-p'));
|
||||
});
|
||||
|
||||
test('invocation includes --output-format text', () => {
|
||||
const c = fs.readFileSync(reviewPath, 'utf-8');
|
||||
assert.ok(
|
||||
c.includes('--output-format text'),
|
||||
'review.md cursor-agent invocation should include "--output-format text"'
|
||||
);
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === 'cursor');
|
||||
const i = lane.invoke.args.indexOf('--output-format');
|
||||
assert.ok(i !== -1);
|
||||
assert.equal(lane.invoke.args[i + 1], 'text');
|
||||
});
|
||||
|
||||
test('invocation passes prompt as a file-path argument (not via stdin pipe)', () => {
|
||||
const c = fs.readFileSync(reviewPath, 'utf-8');
|
||||
assert.ok(
|
||||
c.includes('Read the file at {run_dir}/gsd-review-prompt.md'),
|
||||
'review.md cursor-agent invocation should pass prompt by referencing the file path as an argument'
|
||||
);
|
||||
test('the prompt is a file-path ARGUMENT, never piped on stdin', () => {
|
||||
// Print mode takes the prompt as an argument, and a full plan set inline would approach the
|
||||
// 32,767-char Windows execFileSync ceiling — hence the file reference.
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === 'cursor');
|
||||
assert.equal(lane.invoke.promptChannel, 'argv-file-ref');
|
||||
const r = resolveLanePlan({ lane, configGet: () => undefined, runDir: '/rd', repoRoot: '/repo' });
|
||||
assert.equal(r.ok, true);
|
||||
assert.equal(r.plan.stdin, null, 'nothing may be fed on stdin');
|
||||
assert.ok(r.plan.argv[r.plan.argv.length - 1].includes('/rd/gsd-review-prompt.md'));
|
||||
});
|
||||
|
||||
test('does NOT use broken two-token "cursor agent " form', () => {
|
||||
@@ -104,11 +111,18 @@ describe('Cursor CLI reviewer in /gsd-review (#1960)', () => {
|
||||
);
|
||||
});
|
||||
|
||||
test('contains Cursor Review section in REVIEWS.md template', () => {
|
||||
const c = fs.readFileSync(reviewPath, 'utf-8');
|
||||
assert.ok(
|
||||
c.includes('Cursor Review'),
|
||||
'review.md should include a "Cursor Review" section in the REVIEWS.md template'
|
||||
test('the lane declares its REVIEWS.md section', () => {
|
||||
// The heading used to be a literal in review.md's write_reviews template. Phase 5b renders
|
||||
// sections from each lane's declared `reviewsSection`, so THAT is the contract now — and
|
||||
// uniqueness across lanes is enforced by the parity gate (ADR-2782 D8), which a hardcoded
|
||||
// list never was.
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === 'cursor');
|
||||
assert.equal(lane.reviewsSection, 'Cursor');
|
||||
const sections = REVIEWER_LANES.map((l) => l.reviewsSection);
|
||||
assert.equal(
|
||||
sections.filter((x) => x === 'Cursor').length, 1,
|
||||
'two lanes sharing a heading would silently merge their output in REVIEWS.md',
|
||||
);
|
||||
});
|
||||
|
||||
|
||||
@@ -393,16 +393,42 @@ describe('#2481 review workflow resolves effort per reviewer', () => {
|
||||
'utf8',
|
||||
);
|
||||
|
||||
test('review.md invokes resolve-execution — the grep ADR-443 said returned zero hits', () => {
|
||||
test('shipped orchestration invokes resolve-execution — the grep ADR-443 said returned zero hits', () => {
|
||||
// Phase 5b (#2799) moved the call out of review.md's per-lane bash and into the review-lane
|
||||
// route, which resolves effort once per selected lane through the SAME surface. ADR-443's
|
||||
// invariant is about shipped orchestration calling resolve-execution at all, not about which
|
||||
// file it lives in — so the assertion follows the call rather than pinning the old location.
|
||||
const toolsSrc = fs.readFileSync(
|
||||
path.join(__dirname, '..', 'gsd-core', 'bin', 'gsd-tools.cjs'), 'utf-8',
|
||||
);
|
||||
assert.ok(
|
||||
reviewMd.includes('resolve-execution'),
|
||||
toolsSrc.includes('resolve-execution'),
|
||||
'ADR-443 blocks on no shipped orchestration calling resolve-execution',
|
||||
);
|
||||
});
|
||||
|
||||
test('each argv reviewer receives its effort variable on the command line', () => {
|
||||
for (const [cli, varName] of [['claude', 'CLAUDE_EFFORT_ARGS'], ['codex', 'CODEX_EFFORT_ARGS'], ['opencode', 'OPENCODE_EFFORT_ARGS']]) {
|
||||
assert.ok(reviewMd.includes(`$${varName}`), `${cli} invocation must carry $${varName}`);
|
||||
test('each argv-effort reviewer places effort in its resolved command line', () => {
|
||||
// Stronger than the old shell-variable check: this asserts the effort actually lands in the
|
||||
// argv AT THE POSITION the lane declares, which a `$VAR` substring never proved. Lanes whose
|
||||
// effortChannel is not `argv` must receive nothing.
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const EFFORT = ['--effort', 'high'];
|
||||
for (const lane of REVIEWER_LANES.filter((l) => l.transport === 'spawn')) {
|
||||
const r = resolveLanePlan({
|
||||
lane, configGet: () => undefined, runDir: '/run', repoRoot: '/repo', effortArgs: EFFORT,
|
||||
});
|
||||
assert.equal(r.ok, true, `${lane.slug} failed to resolve`);
|
||||
const carries = r.plan.argv.includes('--effort');
|
||||
assert.equal(
|
||||
carries, lane.invoke.effortChannel === 'argv',
|
||||
`${lane.slug}: effortChannel=${lane.invoke.effortChannel} but argv ${carries ? 'carries' : 'omits'} effort`,
|
||||
);
|
||||
}
|
||||
// The three lanes ADR-1239 #2481 named must still be the argv-effort set.
|
||||
const argvEffort = REVIEWER_LANES
|
||||
.filter((l) => l.transport === 'spawn' && l.invoke.effortChannel === 'argv')
|
||||
.map((l) => l.slug).sort();
|
||||
assert.deepStrictEqual(argvEffort, ['claude', 'codex', 'opencode']);
|
||||
});
|
||||
});
|
||||
|
||||
@@ -1,8 +0,0 @@
|
||||
{
|
||||
"version": 1,
|
||||
"paths": {
|
||||
"review.md": {
|
||||
"reason": "#2797 / ADR-2782 D9 (config half). Growth is comment-only: the three per-lane prompt-budget fallback guards (ollama, lm_studio, llama_cpp) gain an explanation of why they test for the -1 sentinel rather than 0. A federated config key always resolves to its declared default rather than reporting not-found, so an unset per-lane budget needs a sentinel that cannot collide with a legitimate value — and 0 is legitimate: it already means \"do not trim this lane\" (the early-return guard in prepare_trimmed_prompt_for_reviewer). Treating 0 as unset would have silently switched a user who deliberately disabled trimming for one lane onto the global budget. No lane's invocation shape changed."
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -8,8 +8,13 @@
|
||||
* source-grounded review (~570s Codex xhigh, ~525s headless Claude) is killed
|
||||
* mid-review, its output is empty, and the cross-AI review silently proceeds
|
||||
* with fewer lanes. CodeRabbit and OpenCode already documented a timeout; the
|
||||
* four main lanes did not. review.md IS the product the runtime loads, so this
|
||||
* asserts the deployed text carries the guidance.
|
||||
* four main lanes did not. review.md IS the product the runtime loads, so this asserts the
|
||||
* deployed text carries the guidance.
|
||||
*
|
||||
* Phase 5b (#2799) replaced the per-lane blocks with a loop, so the anchor moved from the old
|
||||
* "invoke in sequence" prose to the step itself — but the guidance is MORE load-bearing now, not
|
||||
* less: one Bash call wraps the whole loop, so a host timeout kills every remaining lane rather
|
||||
* than one.
|
||||
*/
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
@@ -21,13 +26,19 @@ const REVIEW_MD = path.join(__dirname, '..', 'gsd-core', 'workflows', 'review.md
|
||||
|
||||
describe('#2194 review.md prompt-fed reviewer timeout guidance', () => {
|
||||
const content = fs.readFileSync(REVIEW_MD, 'utf-8');
|
||||
const sectionStart = content.indexOf('invoke in sequence');
|
||||
const section = sectionStart !== -1
|
||||
? content.slice(sectionStart, sectionStart + 3000)
|
||||
: '';
|
||||
const sectionStart = content.indexOf('<step name="invoke_reviewers">');
|
||||
const sectionEnd = content.indexOf('</step>', sectionStart);
|
||||
const section = sectionStart !== -1 ? content.slice(sectionStart, sectionEnd) : '';
|
||||
|
||||
test('review.md has the reviewer-invocation section', () => {
|
||||
assert.notEqual(sectionStart, -1, 'review.md must contain the "invoke in sequence" section');
|
||||
test('review.md has the reviewer-invocation step', () => {
|
||||
assert.notEqual(sectionStart, -1, 'review.md must contain the invoke_reviewers step');
|
||||
});
|
||||
|
||||
test('lanes are invoked sequentially, not in parallel', () => {
|
||||
// Concurrent invocation trips provider rate limits; the original prose said so and the loop
|
||||
// must keep saying so, since a future reader could otherwise "optimize" it into a fan-out.
|
||||
assert.ok(/sequential|in sequence|not in parallel/i.test(section),
|
||||
'the step must state that lanes run sequentially');
|
||||
});
|
||||
|
||||
test('the section carries Bash timeout guidance for the prompt-fed lanes', () => {
|
||||
|
||||
@@ -75,98 +75,49 @@ describe('#2358 review.md temp paths are run-scoped, not phase-only', () => {
|
||||
);
|
||||
});
|
||||
|
||||
test('the Antigravity reviewer prompt-instruction string references the run-scoped path', () => {
|
||||
const agyPromptMatch = content.match(/_AGY_PROMPT="Read the file at ([^ ]+)/);
|
||||
assert.ok(agyPromptMatch, '_AGY_PROMPT must contain a "Read the file at <path>" instruction');
|
||||
assert.equal(
|
||||
agyPromptMatch[1], '{run_dir}/gsd-review-prompt.md',
|
||||
'the Antigravity reviewer must be told to read the run-scoped prompt path, not a bare {phase}-only /tmp path — ' +
|
||||
'this is the exact instruction text the reporter forensically traced back to a stale cross-project read'
|
||||
);
|
||||
// Phase 5b (#2799) moved these strings out of review.md's bash and into the resolver and the
|
||||
// antigravity handler, so the assertions follow them. The invariant is unchanged and is what
|
||||
// #2358 was about: every reviewer artifact must live under the run-scoped mktemp directory, never
|
||||
// a bare `{phase}`-keyed /tmp path that a concurrent review could collide with.
|
||||
test('every lane anchors its prompt and artifacts under the run dir', () => {
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const RUN = '/run-scoped';
|
||||
for (const lane of REVIEWER_LANES) {
|
||||
const r = resolveLanePlan({
|
||||
lane, configGet: () => undefined, runDir: RUN, repoRoot: '/repo',
|
||||
});
|
||||
assert.equal(r.ok, true, `${lane.slug} failed to resolve`);
|
||||
const p = r.plan;
|
||||
assert.ok(p.reviewPath.startsWith(`${RUN}/`), `${lane.slug} review path escapes the run dir`);
|
||||
assert.ok(p.errPath.startsWith(`${RUN}/`), `${lane.slug} err path escapes the run dir`);
|
||||
assert.ok(p.promptPath.startsWith(`${RUN}/`), `${lane.slug} prompt path escapes the run dir`);
|
||||
}
|
||||
});
|
||||
|
||||
test('the Cursor reviewer prompt-instruction string references the run-scoped path', () => {
|
||||
const cursorPromptMatch = content.match(/CURSOR_PROMPT_ARG="Read the file at ([^ ]+)/);
|
||||
assert.ok(cursorPromptMatch, 'CURSOR_PROMPT_ARG must contain a "Read the file at <path>" instruction');
|
||||
assert.equal(
|
||||
cursorPromptMatch[1], '{run_dir}/gsd-review-prompt.md',
|
||||
'the Cursor reviewer must be told to read the run-scoped prompt path, not a bare {phase}-only /tmp path'
|
||||
test('the argv-borne prompt instruction references the run-scoped path', () => {
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const RUN = '/run-scoped';
|
||||
const fileRefLanes = REVIEWER_LANES.filter(
|
||||
(l) => l.transport === 'spawn' && l.invoke.promptChannel === 'argv-file-ref',
|
||||
);
|
||||
assert.ok(fileRefLanes.length > 0, 'expected at least one argv-file-ref lane');
|
||||
for (const lane of fileRefLanes) {
|
||||
const r = resolveLanePlan({ lane, configGet: () => undefined, runDir: RUN, repoRoot: '/repo' });
|
||||
const arg = r.plan.argv[r.plan.argv.length - 1];
|
||||
assert.ok(arg.includes(`${RUN}/gsd-review-prompt.md`), `${lane.slug} prompt not run-scoped`);
|
||||
}
|
||||
});
|
||||
|
||||
test('the run directory is cleaned up at the end of the review', () => {
|
||||
const presentResultsStart = content.indexOf('<step name="present_results">');
|
||||
assert.notEqual(presentResultsStart, -1, 'review.md must contain the present_results step');
|
||||
const section = content.slice(presentResultsStart, presentResultsStart + 1500);
|
||||
assert.ok(
|
||||
/rm -rf "\{run_dir\}"/.test(section),
|
||||
'present_results must remove the run-scoped temp directory once REVIEWS.md is written'
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
describe('#2358 ship.md external-review stderr capture is run-scoped', () => {
|
||||
const content = fs.readFileSync(SHIP_MD, 'utf-8');
|
||||
|
||||
test('no bare, unqualified /tmp/gsd-review-stderr.log path remains', () => {
|
||||
assert.ok(
|
||||
!content.includes('/tmp/gsd-review-stderr.log'),
|
||||
'ship.md must not write/read a shared, unqualified stderr log path — every ship run, phase, and project ' +
|
||||
'shares this exact path with zero disambiguator, which is strictly worse than review.md\'s phase-only keying'
|
||||
);
|
||||
});
|
||||
|
||||
test('stderr is captured to a per-run file via the portable ${TMPDIR:-/tmp} seam', () => {
|
||||
assert.ok(
|
||||
/REVIEW_STDERR_FILE=\$\(mktemp "\$\{TMPDIR:-\/tmp\}\/gsd-review-stderr-XXXXXX"\)/.test(content),
|
||||
'ship.md must create the stderr capture file via `mktemp "${TMPDIR:-/tmp}/gsd-review-stderr-XXXXXX"`'
|
||||
);
|
||||
assert.ok(
|
||||
/2>"\$\{REVIEW_STDERR_FILE\}"/.test(content),
|
||||
'the external review command invocation must redirect stderr to the per-run $REVIEW_STDERR_FILE, not a literal path'
|
||||
);
|
||||
assert.ok(
|
||||
/cat "\$\{REVIEW_STDERR_FILE\}"/.test(content),
|
||||
'the failure-handling block must read back the same per-run $REVIEW_STDERR_FILE'
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
describe('#2358 reviewer-instances.md (#1517, lazily loaded from invoke_reviewers) is run-scoped too', () => {
|
||||
// review.md's own invoke_reviewers step lazily loads this companion doc for the
|
||||
// review.reviewer_instances codepath — it was missed in the initial pass and still
|
||||
// pointed at the old, unscoped /tmp/gsd-review-*-{phase} paths, which both broke
|
||||
// reviewer-instances functionality (the prompt file build_prompt now writes lives
|
||||
// at {run_dir}/gsd-review-prompt.md, never the old path) and left the exact
|
||||
// cross-project collision bug open for that code path.
|
||||
const content = fs.readFileSync(REVIEWER_INSTANCES_MD, 'utf-8');
|
||||
|
||||
test('no bare, unscoped /tmp/gsd-review-* path remains', () => {
|
||||
assert.ok(
|
||||
!content.includes('/tmp/gsd-review'),
|
||||
'reviewer-instances.md must not contain any hardcoded /tmp/gsd-review* literal — ' +
|
||||
'every review temp path must be rooted under the run-scoped {run_dir} directory'
|
||||
);
|
||||
});
|
||||
|
||||
test('no temp path is still keyed on a bare {phase} placeholder', () => {
|
||||
assert.ok(
|
||||
!/gsd-review[^\r\n]*\{phase\}/.test(content),
|
||||
'no temp path may still be keyed on a bare {phase} placeholder'
|
||||
);
|
||||
});
|
||||
|
||||
test('the instance prompt read and output write are threaded through {run_dir}', () => {
|
||||
assert.ok(
|
||||
/\{run_dir\}\/gsd-review-prompt\.md/.test(content),
|
||||
'reviewer-instances.md must read the prompt from {run_dir}/gsd-review-prompt.md, ' +
|
||||
'the same run-scoped path build_prompt writes in review.md'
|
||||
);
|
||||
assert.ok(
|
||||
/\{run_dir\}\/gsd-review-\$\{INSTANCE_NAME\}\.md/.test(content),
|
||||
'reviewer-instances.md must write each instance\'s output to ' +
|
||||
'{run_dir}/gsd-review-${INSTANCE_NAME}.md, not a {phase}-keyed /tmp path'
|
||||
);
|
||||
test('an instance writes under the run dir, keyed by its own identity', () => {
|
||||
// Two instances of one adapter must not overwrite each other, and neither may escape the run
|
||||
// dir — the identity is sanitized to a flat filename.
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === 'opencode');
|
||||
const r = resolveLanePlan({ lane, configGet: () => undefined, runDir: '/run-scoped', repoRoot: '/repo' });
|
||||
assert.ok(r.plan.reviewPath.startsWith('/run-scoped/'));
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
@@ -1,204 +1,121 @@
|
||||
// allow-test-rule: source-text-is-the-product (see #2494)
|
||||
// The Gemini and Claude reviewer dispatch blocks in gsd-core/workflows/review.md
|
||||
// ARE the runtime contract — the workflow's text is what the reviewing agent
|
||||
// executes. This suite extracts those two shell blocks verbatim from the
|
||||
// workflow and runs them under a real bash against a failing CLI stub, so the
|
||||
// shipped guard is what gets exercised rather than a reimplementation of it.
|
||||
// The assertions on the produced review file are assertions on that guard's
|
||||
// documented output contract (the stub line the consensus step must be able to
|
||||
// tell apart from a clean empty review), not incidental string matching.
|
||||
|
||||
/**
|
||||
* Regression tests for #2494 — the claude and gemini reviewer legs discarded
|
||||
* stderr with no empty-output guard.
|
||||
* #2494 — a failed reviewer lane must be diagnosable, never a silent drop.
|
||||
*
|
||||
* Before the fix both blocks were `... 2>/dev/null > {run_dir}/gsd-review-<leg>.md`
|
||||
* with nothing after: a non-zero exit that wrote no stdout (CLI missing,
|
||||
* unauthenticated, rate-limited, timeout-killed, crashed) left a ZERO-BYTE
|
||||
* review file with the only diagnostic evidence — stderr — already discarded.
|
||||
* The write_reviews step then substituted that empty file into a
|
||||
* `## <Reviewer> Review` section indistinguishable from "ran cleanly, nothing
|
||||
* to report", silently degrading the advertised N-reviewer consensus to N-1.
|
||||
* Before the fix, gemini and claude sent stderr to `/dev/null` and wrote nothing on failure. A
|
||||
* failed lane — CLI missing, unauthenticated, rate-limited, crashed, any exit that writes no
|
||||
* stdout — left a zero-byte file that `write_reviews` rendered as "a reviewer that ran cleanly with
|
||||
* nothing to report", silently dropping a lane from the cross-AI consensus while `present_results`
|
||||
* reported success.
|
||||
*
|
||||
* These tests fail against pre-fix review.md: the produced file is zero-byte,
|
||||
* so both the non-empty and the diagnosable-message assertions trip.
|
||||
* The invariant is unchanged; what moved is where it lives. This suite used to extract the two
|
||||
* shell blocks verbatim from `review.md` and run them under a real bash against a failing stub.
|
||||
* Phase 5b (#2799) deleted those blocks, so the tests now drive the real runner with stubbed
|
||||
* dependencies — the same behavioural altitude (a lane is actually run and its artifacts
|
||||
* inspected), against the surface that ships today. Nothing here reads source text any more, so
|
||||
* the source-text exemption this file used to carry is gone.
|
||||
*
|
||||
* The guarantee also got STRONGER in one way worth locking: the policy is uniform across every
|
||||
* lane now rather than fixed per-leg, so these assertions run over the whole spawn roster instead
|
||||
* of the two legs the issue named.
|
||||
*/
|
||||
|
||||
'use strict';
|
||||
|
||||
const { describe, test, before, after } = require('node:test');
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const { spawnSync } = require('node:child_process');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
const { createTempDir, cleanup } = require('./helpers.cjs');
|
||||
|
||||
const ROOT = path.join(__dirname, '..');
|
||||
const REVIEW_PATH = path.join(ROOT, 'gsd-core', 'workflows', 'review.md');
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const { runLane } = require('../gsd-core/bin/lib/review-lane-runner.cjs');
|
||||
|
||||
// Normalize CRLF: on a Windows git-autocrlf checkout every line carries a
|
||||
// trailing \r, which would leave the extracted block's shebang/redirect tokens
|
||||
// mangled and defeat the fence regexes below.
|
||||
const WORKFLOW = fs.readFileSync(REVIEW_PATH, 'utf-8').replace(/\r\n/g, '\n');
|
||||
const RUN = '/run';
|
||||
const ROOT = '/repo';
|
||||
|
||||
// This suite executes the extracted blocks with a real bash and a POSIX CLI
|
||||
// stub on PATH. Gate to non-Windows, mirroring the opencode reconstruction
|
||||
// suite's win32 skip — the guard logic is platform-independent and is asserted
|
||||
// in full on every macOS/Linux CI leg.
|
||||
const skipReason = process.platform === 'win32'
|
||||
? 'extracted block is POSIX shell; guard logic is platform-independent and asserted on macOS/Linux'
|
||||
: false;
|
||||
/** Lanes whose empty-output policy is the shared stub (antigravity owns its own diagnostics). */
|
||||
const STUB_LANES = REVIEWER_LANES.filter(
|
||||
(l) => l.transport === 'spawn' && l.emptyOutput === 'stub-with-stderr',
|
||||
);
|
||||
|
||||
/**
|
||||
* Extract a reviewer dispatch block verbatim from the workflow. If review.md
|
||||
* changes the block's shape these throw and the test fails loudly — intended
|
||||
* coupling, the same contract the opencode suite pins for its jq programs.
|
||||
*/
|
||||
function extractBlock(headingRe, label) {
|
||||
const re = new RegExp(`${headingRe}\\n\`\`\`bash\\n([\\s\\S]*?)\\n\`\`\``);
|
||||
const m = WORKFLOW.match(re);
|
||||
assert.ok(m, `review.md must define the ${label} reviewer dispatch as a bash block (#2494)`);
|
||||
return m[1];
|
||||
function planFor(slug) {
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === slug);
|
||||
const r = resolveLanePlan({ lane, configGet: () => undefined, runDir: RUN, repoRoot: ROOT });
|
||||
assert.equal(r.ok, true, `${slug} failed to resolve`);
|
||||
return r.plan;
|
||||
}
|
||||
|
||||
const GEMINI_BLOCK = extractBlock('\\*\\*Gemini:\\*\\*', 'Gemini');
|
||||
const CLAUDE_BLOCK = extractBlock('\\*\\*Claude \\(separate session\\):\\*\\*', 'Claude');
|
||||
|
||||
const STUB_STDERR = 'gsd-2494-stub: command not found / not authenticated';
|
||||
|
||||
let sandbox;
|
||||
|
||||
before(() => {
|
||||
sandbox = createTempDir('gsd-2494-');
|
||||
});
|
||||
|
||||
after(() => {
|
||||
cleanup(sandbox);
|
||||
});
|
||||
|
||||
/**
|
||||
* Run one extracted block with `{run_dir}` pointed at a fresh run directory and
|
||||
* `stubBody` installed on PATH under the reviewer's binary name.
|
||||
*
|
||||
* Single exec site for the whole suite so the Windows guard lives in one place:
|
||||
* Git Bash (msys2) ignores Node's chmod exec bit for PATH-executed
|
||||
* extension-less scripts (DEFECT.WINDOWS-TEST-PORTABILITY), and every suite
|
||||
* below is skipped on win32 — this early return keeps the exec unreachable
|
||||
* there rather than relying on the skip alone.
|
||||
*/
|
||||
function runBlockWithStub({ binName, block, stubBody, env = {} }) {
|
||||
if (process.platform === 'win32') return null;
|
||||
|
||||
const caseDir = fs.mkdtempSync(path.join(sandbox, 'run-'));
|
||||
const runDir = path.join(caseDir, 'run');
|
||||
const binDir = path.join(caseDir, 'bin');
|
||||
fs.mkdirSync(runDir);
|
||||
fs.mkdirSync(binDir);
|
||||
|
||||
const stub = path.join(binDir, binName);
|
||||
fs.writeFileSync(stub, stubBody);
|
||||
fs.chmodSync(stub, 0o755);
|
||||
|
||||
fs.writeFileSync(path.join(runDir, 'gsd-review-prompt.md'), '# review prompt\n');
|
||||
|
||||
const script = block.split('{run_dir}').join(runDir);
|
||||
const result = spawnSync('bash', ['-c', script], {
|
||||
encoding: 'utf8',
|
||||
timeout: 30000,
|
||||
killSignal: 'SIGKILL',
|
||||
env: { ...process.env, PATH: `${binDir}${path.delimiter}${process.env.PATH}`, ...env },
|
||||
});
|
||||
|
||||
const reviewPath = path.join(runDir, `gsd-review-${binName}.md`);
|
||||
function deps(spawnResult, files = {}) {
|
||||
return {
|
||||
result,
|
||||
reviewPath,
|
||||
errPath: path.join(runDir, `gsd-review-${binName}.err`),
|
||||
review: fs.existsSync(reviewPath) ? fs.readFileSync(reviewPath, 'utf-8') : null,
|
||||
files,
|
||||
// `kimi-code` declares a `command-capability` probe, so the runner spawns `--help` BEFORE the
|
||||
// review. Answer that separately or the probe fails and the lane never reaches the invocation
|
||||
// this test is about.
|
||||
spawn: (binary, argv) =>
|
||||
argv && argv.length === 1 && argv[0] === '--help'
|
||||
? { status: 0, stdout: '--output-format', stderr: '' }
|
||||
: spawnResult,
|
||||
httpJson: async () => ({ ok: true, status: 200, body: '{}' }),
|
||||
readFile: (p) => { if (!(p in files)) throw new Error(`ENOENT ${p}`); return files[p]; },
|
||||
writeFile: (p, c) => { files[p] = c; },
|
||||
exists: (p) => p in files,
|
||||
hasBinary: () => true,
|
||||
configGet: () => undefined,
|
||||
homeDir: '/home/u',
|
||||
warn: () => {},
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* The issue's own repro harness: a CLI that writes STUB_STDERR to stderr,
|
||||
* nothing to stdout, and exits non-zero.
|
||||
*/
|
||||
function runLeg({ binName, block, env = {} }) {
|
||||
return runBlockWithStub({
|
||||
binName,
|
||||
block,
|
||||
stubBody: `#!/bin/sh\necho "${STUB_STDERR}" >&2\nexit 127\n`,
|
||||
env,
|
||||
describe('#2494 — a failed lane writes a diagnosable stub, not a zero-byte file', () => {
|
||||
for (const lane of STUB_LANES) {
|
||||
test(`${lane.slug}: a lane that exits non-zero with no stdout is stubbed`, async () => {
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps({ status: 127, stdout: '', stderr: 'command not found' });
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
|
||||
assert.equal(r.stubbed, true, 'a failed lane must be reported as stubbed');
|
||||
const review = d.files[p.reviewPath];
|
||||
assert.ok(review !== undefined, 'a review file must exist after a failed lane');
|
||||
assert.notStrictEqual(review.trim(), '', 'the review file must not be empty');
|
||||
assert.ok(
|
||||
review.includes('failed or returned empty output'),
|
||||
'the stub must be distinguishable from a real review',
|
||||
);
|
||||
});
|
||||
|
||||
test(`${lane.slug}: stderr is captured to a .err sidecar, never discarded`, async () => {
|
||||
// The sidecar is the difference between "this lane failed" and "this lane failed BECAUSE…".
|
||||
// Without it every failure mode looks identical to every other.
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps({ status: 1, stdout: '', stderr: 'HTTP 429 rate limited' });
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
|
||||
assert.equal(d.files[p.errPath], 'HTTP 429 rate limited', 'stderr must reach the sidecar');
|
||||
assert.ok(
|
||||
d.files[p.reviewPath].includes('HTTP 429 rate limited'),
|
||||
'and must be surfaced in the stub, where a reader will actually see it',
|
||||
);
|
||||
assert.ok(p.reviewPath.endsWith('.md'), 'review output path unchanged');
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* The guard's contract: the review file exists, is NOT empty, names the leg as
|
||||
* failed-or-empty, and carries the captured stderr. The last assertion is what
|
||||
* separates this fix from a bare `[ ! -s ]` stub — the evidence has to survive.
|
||||
*/
|
||||
function assertDiagnosable({ review, reviewPath, errPath }, legLabel) {
|
||||
assert.ok(review !== null, `${legLabel}: review file must exist after a failed lane (#2494)`);
|
||||
assert.notStrictEqual(review.trim(), '', `${legLabel}: review file must not be empty after a failed lane (#2494)`);
|
||||
assert.match(
|
||||
review,
|
||||
new RegExp(`${legLabel} review failed or returned empty output`, 'i'),
|
||||
`${legLabel}: review file must carry a diagnosable failure line, distinguishable at consensus synthesis from "ran cleanly, nothing to report" (#2494)`,
|
||||
);
|
||||
test('a successful review passes through untouched', async () => {
|
||||
const p = planFor('gemini');
|
||||
const d = deps({ status: 0, stdout: 'Looks good.\n', stderr: '' });
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
|
||||
assert.equal(r.stubbed, false);
|
||||
assert.ok(d.files[p.reviewPath].includes('Looks good.'));
|
||||
assert.ok(
|
||||
review.includes(STUB_STDERR),
|
||||
`${legLabel}: captured stderr must be appended to the review file, not discarded to /dev/null (#2494)`,
|
||||
!d.files[p.reviewPath].includes('failed or returned empty output'),
|
||||
'a real review must never carry the failure header',
|
||||
);
|
||||
assert.ok(fs.existsSync(errPath), `${legLabel}: stderr must be captured to a .err sidecar (#2494)`);
|
||||
assert.ok(reviewPath.endsWith('.md'), `${legLabel}: review output path unchanged`);
|
||||
});
|
||||
|
||||
test('no lane sends stderr to /dev/null — the sidecar is unconditional', async () => {
|
||||
// The original defect in one line, asserted over the whole roster rather than the two legs the
|
||||
// issue named: the policy is uniform now, and a future lane must not be able to opt out.
|
||||
for (const lane of STUB_LANES) {
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps({ status: 0, stdout: 'ok', stderr: 'a warning' });
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(d.files[p.errPath], 'a warning', `${lane.slug} discarded stderr`);
|
||||
}
|
||||
|
||||
describe('#2494 — gemini reviewer leg fails loudly', { skip: skipReason }, () => {
|
||||
test('a failing gemini CLI produces a diagnosable stub, not an empty file (model configured)', () => {
|
||||
const out = runLeg({ binName: 'gemini', block: GEMINI_BLOCK, env: { GEMINI_MODEL: 'gemini-test-model' } });
|
||||
assertDiagnosable(out, 'Gemini');
|
||||
});
|
||||
|
||||
test('a failing gemini CLI produces a diagnosable stub, not an empty file (no model configured)', () => {
|
||||
const out = runLeg({ binName: 'gemini', block: GEMINI_BLOCK, env: { GEMINI_MODEL: '' } });
|
||||
assertDiagnosable(out, 'Gemini');
|
||||
});
|
||||
});
|
||||
|
||||
describe('#2494 — claude reviewer leg fails loudly', { skip: skipReason }, () => {
|
||||
test('a failing claude CLI produces a diagnosable stub, not an empty file (model configured)', () => {
|
||||
const out = runLeg({
|
||||
binName: 'claude',
|
||||
block: CLAUDE_BLOCK,
|
||||
env: { CLAUDE_MODEL: 'claude-test-model', CLAUDE_EFFORT_ARGS: '' },
|
||||
});
|
||||
assertDiagnosable(out, 'Claude');
|
||||
});
|
||||
|
||||
test('a failing claude CLI produces a diagnosable stub, not an empty file (no model configured)', () => {
|
||||
const out = runLeg({
|
||||
binName: 'claude',
|
||||
block: CLAUDE_BLOCK,
|
||||
env: { CLAUDE_MODEL: '', CLAUDE_EFFORT_ARGS: '' },
|
||||
});
|
||||
assertDiagnosable(out, 'Claude');
|
||||
});
|
||||
});
|
||||
|
||||
describe('#2494 — the guard does not swallow a successful review', { skip: skipReason }, () => {
|
||||
test('gemini stdout is preserved verbatim when the CLI succeeds', () => {
|
||||
const out = runBlockWithStub({
|
||||
binName: 'gemini',
|
||||
block: GEMINI_BLOCK,
|
||||
stubBody: '#!/bin/sh\necho "## Real Review"\necho "Looks good."\nexit 0\n',
|
||||
env: { GEMINI_MODEL: '' },
|
||||
});
|
||||
|
||||
const { review } = out;
|
||||
assert.ok(review !== null, 'a successful review must produce a review file (#2494)');
|
||||
assert.ok(review.includes('Looks good.'), 'a successful review must pass through untouched (#2494)');
|
||||
assert.ok(
|
||||
!/failed or returned empty output/i.test(review),
|
||||
'the stub must not fire on a non-empty review (#2494)',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
@@ -1,399 +1,147 @@
|
||||
// allow-test-rule: source-text-is-the-product (see #2605)
|
||||
// The lm_studio and llama_cpp reviewer dispatch blocks in gsd-core/workflows/review.md
|
||||
// ARE the runtime contract — the workflow's text is what the reviewing agent executes.
|
||||
// This suite extracts those two shell blocks verbatim from the workflow and runs them
|
||||
// under a real bash against a stubbed curl, so the shipped guard is what gets exercised
|
||||
// rather than a reimplementation of it. The assertions on the produced review file are
|
||||
// assertions on that guard's documented output contract (the stub line the consensus
|
||||
// step must be able to tell apart from a clean empty review), not incidental string
|
||||
// matching.
|
||||
|
||||
/**
|
||||
* Regression tests for #2605 — the lm_studio and llama_cpp reviewer legs dropped
|
||||
* empty output with no stub file, the same defect class as #2494 (claude/gemini)
|
||||
* but a worse variant.
|
||||
* #2605 — the local OpenAI-compatible lanes (ollama / lm_studio / llama.cpp) dropped silently.
|
||||
*
|
||||
* Before the fix both blocks ran `curl -s ... 2>/dev/null` and, when the model
|
||||
* returned empty content, wrote NOTHING to
|
||||
* `{run_dir}/gsd-review-<leg>.md` — there was no `[ ! -s … ]` stub at all. The
|
||||
* review file therefore never existed, `write_reviews` silently omitted that
|
||||
* reviewer's section, and the outcome was indistinguishable from the reviewer
|
||||
* never having been selected.
|
||||
* The original defects, all of which made a failed lane indistinguishable from a clean empty
|
||||
* review: bare `curl -s` suppressed curl's own error text; the response was piped straight into
|
||||
* `jq` so the BODY — where an OpenAI-compatible server puts its error JSON on an HTTP 4xx/5xx while
|
||||
* curl still exits 0 — was discarded unread; nothing was written when content was empty, so the
|
||||
* file never existed and `write_reviews` omitted the section entirely; and a whitespace-only reply
|
||||
* passed the byte-counting `[ ! -s … ]` guard as a successful review.
|
||||
*
|
||||
* Two distinct diagnostic holes are covered here, because an OpenAI-compatible
|
||||
* server can fail in two ways that leave evidence in different places:
|
||||
* - transport failure (endpoint unreachable): curl writes to STDERR and exits
|
||||
* non-zero. `-s` suppressed that error text entirely, so `-sS` is required.
|
||||
* - application failure (HTTP 4xx/5xx): curl exits 0 and the error JSON is in
|
||||
* the response BODY, so stderr is empty and only the body is diagnosable.
|
||||
* The llama_cpp leg additionally piped curl straight into jq, discarding the
|
||||
* body before anything could inspect it.
|
||||
*
|
||||
* These tests fail against pre-fix review.md: no review file is produced at all,
|
||||
* so the existence assertion trips first.
|
||||
* Phase 5b (#2799) replaced the curl/jq pipeline with an in-process HTTP call and `JSON.parse`, so
|
||||
* this suite drives the runner instead of extracting and executing shell. Two of the original
|
||||
* defects are now structurally impossible rather than merely guarded: there is no pipe to discard
|
||||
* the body, and no `echo` to swallow a `-n`-shaped reply.
|
||||
*/
|
||||
|
||||
'use strict';
|
||||
|
||||
const { describe, test, before, after } = require('node:test');
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const { spawnSync, execFileSync } = require('node:child_process');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
const { createTempDir, cleanup } = require('./helpers.cjs');
|
||||
|
||||
const ROOT = path.join(__dirname, '..');
|
||||
const REVIEW_PATH = path.join(ROOT, 'gsd-core', 'workflows', 'review.md');
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const { runLane, runOpenAiCompatible } = require('../gsd-core/bin/lib/review-lane-runner.cjs');
|
||||
|
||||
// Normalize CRLF: on a Windows git-autocrlf checkout every line carries a
|
||||
// trailing \r, which would leave the extracted block's redirect tokens mangled
|
||||
// and defeat the fence regexes below.
|
||||
const WORKFLOW = fs.readFileSync(REVIEW_PATH, 'utf-8').replace(/\r\n/g, '\n');
|
||||
const RUN = '/run';
|
||||
const ROOT = '/repo';
|
||||
const HTTP_LANES = REVIEWER_LANES.filter((l) => l.transport === 'openai-http');
|
||||
|
||||
// These blocks shell out to a real `jq` (the request body is built with it).
|
||||
// Gate to jq-present non-Windows hosts, mirroring the opencode reconstruction
|
||||
// suite — the guard logic is platform-independent and is asserted in full on
|
||||
// every macOS/Linux CI leg.
|
||||
let jqAvailable = false;
|
||||
try {
|
||||
execFileSync('jq', ['--version'], { stdio: 'ignore', timeout: 10000, killSignal: 'SIGKILL' });
|
||||
jqAvailable = true;
|
||||
} catch { /* no jq on PATH */ }
|
||||
|
||||
const skipReason = process.platform === 'win32'
|
||||
? 'extracted block is POSIX shell; guard logic is platform-independent and asserted on macOS/Linux'
|
||||
: (jqAvailable ? false : 'jq not on PATH');
|
||||
|
||||
/**
|
||||
* Extract a reviewer dispatch block verbatim from the workflow. If review.md
|
||||
* changes the block's shape these throw and the test fails loudly — intended
|
||||
* coupling, the same contract the #2494 suite pins for the claude/gemini legs.
|
||||
*/
|
||||
function extractBlock(headingRe, label) {
|
||||
// Some legs (Ollama, CodeRabbit) put an explanatory paragraph between the
|
||||
// heading and the fence, so allow non-fence content in between. The lazy
|
||||
// quantifiers take the FIRST ```bash fence after the heading, which is that
|
||||
// leg's own block.
|
||||
const re = new RegExp(`${headingRe}\\n[\\s\\S]*?\`\`\`bash\\n([\\s\\S]*?)\\n\`\`\``);
|
||||
const m = WORKFLOW.match(re);
|
||||
assert.ok(m, `review.md must define the ${label} reviewer dispatch as a bash block (#2605)`);
|
||||
return m[1];
|
||||
function planFor(slug, config = {}) {
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === slug);
|
||||
const r = resolveLanePlan({ lane, configGet: (k) => config[k], runDir: RUN, repoRoot: ROOT });
|
||||
assert.equal(r.ok, true);
|
||||
return r.plan;
|
||||
}
|
||||
|
||||
const LEGS = [
|
||||
{
|
||||
key: 'lm_studio',
|
||||
label: 'LM Studio',
|
||||
budgetVar: 'LM_STUDIO_REVIEWER_BUDGET',
|
||||
block: extractBlock('\\*\\*LM Studio \\(local, OpenAI-compatible\\):\\*\\*', 'LM Studio'),
|
||||
},
|
||||
{
|
||||
key: 'llama_cpp',
|
||||
label: 'llama.cpp',
|
||||
budgetVar: 'LLAMA_CPP_REVIEWER_BUDGET',
|
||||
block: extractBlock('\\*\\*llama\\.cpp \\(local, OpenAI-compatible\\):\\*\\*', 'llama.cpp'),
|
||||
},
|
||||
// Ollama was the least diagnosable of the three local-server legs (bare `-s`,
|
||||
// stderr to /dev/null, response piped straight into jq). It emitted a stub, so
|
||||
// it never silently vanished — but it is the same family and is held to the
|
||||
// same contract here.
|
||||
{
|
||||
key: 'ollama',
|
||||
label: 'Ollama',
|
||||
budgetVar: 'OLLAMA_REVIEWER_BUDGET',
|
||||
block: extractBlock('\\*\\*Ollama \\(local, OpenAI-compatible\\):\\*\\*', 'Ollama'),
|
||||
},
|
||||
];
|
||||
|
||||
const STUB_STDERR = 'gsd-2605-stub: curl: (7) Failed to connect to localhost';
|
||||
const STUB_ERROR_BODY = '{"error":{"message":"gsd-2605-stub: model not loaded","code":503}}';
|
||||
|
||||
let sandbox;
|
||||
|
||||
before(() => { sandbox = createTempDir('gsd-2605-'); });
|
||||
after(() => { cleanup(sandbox); });
|
||||
|
||||
/**
|
||||
* Run one extracted block with `{run_dir}` pointed at a fresh run directory and
|
||||
* `curlBody` installed on PATH as `curl`.
|
||||
*
|
||||
* `gsd_run` is deliberately NOT stubbed: every call site is
|
||||
* `$(gsd_run … || echo "<default>")`, so an absent binary takes the documented
|
||||
* default path (no prompt budget, default host, model probed from the server) —
|
||||
* which is the configuration the issue reproduces against.
|
||||
*
|
||||
* Single exec site so the Windows guard lives in one place: Git Bash (msys2)
|
||||
* ignores Node's chmod exec bit for PATH-executed extension-less scripts
|
||||
* (DEFECT.WINDOWS-TEST-PORTABILITY), and every suite below is skipped on win32 —
|
||||
* this early return keeps the exec unreachable there rather than relying on the
|
||||
* skip alone.
|
||||
* These lanes declare an `http-reachable` probe, so the runner performs a GET on /v1/models BEFORE
|
||||
* the chat call. The stub must answer that separately — otherwise the lane is reported unreachable
|
||||
* and never reaches the invocation these tests are actually about.
|
||||
*/
|
||||
function runLeg({ leg, curlBody, preamble = '' }) {
|
||||
if (process.platform === 'win32') return null;
|
||||
if (!jqAvailable) return null;
|
||||
function reachableThen(chatResponse) {
|
||||
return async (url, opts) =>
|
||||
opts.method === 'GET'
|
||||
? { ok: true, status: 200, body: JSON.stringify({ data: [{ id: 'stub-model' }] }) }
|
||||
: (typeof chatResponse === 'function' ? chatResponse(url, opts) : chatResponse);
|
||||
}
|
||||
|
||||
const caseDir = fs.mkdtempSync(path.join(sandbox, 'run-'));
|
||||
const runDir = path.join(caseDir, 'run');
|
||||
const binDir = path.join(caseDir, 'bin');
|
||||
fs.mkdirSync(runDir);
|
||||
fs.mkdirSync(binDir);
|
||||
|
||||
const stub = path.join(binDir, 'curl');
|
||||
fs.writeFileSync(stub, curlBody);
|
||||
fs.chmodSync(stub, 0o755);
|
||||
|
||||
fs.writeFileSync(path.join(runDir, 'gsd-review-prompt.md'), '# review prompt\n');
|
||||
|
||||
const script = preamble + leg.block.split('{run_dir}').join(runDir);
|
||||
const result = spawnSync('bash', ['-c', script], {
|
||||
encoding: 'utf8',
|
||||
timeout: 30000,
|
||||
killSignal: 'SIGKILL',
|
||||
env: { ...process.env, PATH: `${binDir}${path.delimiter}${process.env.PATH}` },
|
||||
});
|
||||
|
||||
const reviewPath = path.join(runDir, `gsd-review-${leg.key}.md`);
|
||||
function deps(httpJson, files = { [`${RUN}/gsd-review-prompt.md`]: 'PLAN' }) {
|
||||
const warnings = [];
|
||||
return {
|
||||
result,
|
||||
reviewPath,
|
||||
errPath: path.join(runDir, `gsd-review-${leg.key}.err`),
|
||||
review: fs.existsSync(reviewPath) ? fs.readFileSync(reviewPath, 'utf-8') : null,
|
||||
files,
|
||||
warnings,
|
||||
spawn: () => ({ status: 0, stdout: '', stderr: '' }),
|
||||
httpJson,
|
||||
readFile: (p) => { if (!(p in files)) throw new Error(`ENOENT ${p}`); return files[p]; },
|
||||
writeFile: (p, c) => { files[p] = c; },
|
||||
exists: (p) => p in files,
|
||||
hasBinary: () => true,
|
||||
configGet: () => undefined,
|
||||
homeDir: '/home/u',
|
||||
warn: (m) => warnings.push(m),
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* A curl stub. The block may call curl twice — once to probe `/v1/models` for a
|
||||
* model id, once to POST `/v1/chat/completions` — so the stub branches on the URL
|
||||
* and only applies the failure mode to the completion call.
|
||||
*/
|
||||
function curlStub({ completionStdout = '', completionStderr = '', exitCode = 0 }) {
|
||||
return `#!/bin/sh
|
||||
for a in "$@"; do
|
||||
case "$a" in
|
||||
*/v1/models) echo '{"data":[{"id":"stub-model"}]}'; exit 0 ;;
|
||||
esac
|
||||
done
|
||||
${completionStderr ? `echo "${completionStderr}" >&2` : ':'}
|
||||
${completionStdout ? `cat <<'GSD_EOF'\n${completionStdout}\nGSD_EOF` : ':'}
|
||||
exit ${exitCode}
|
||||
`;
|
||||
}
|
||||
|
||||
/**
|
||||
* The guard's contract: the review file EXISTS (the #2605 core defect — it did
|
||||
* not), is not empty, and names the leg as failed-or-empty so consensus
|
||||
* synthesis can tell it apart from "ran cleanly, nothing to report".
|
||||
*/
|
||||
function assertDiagnosable(out, legLabel) {
|
||||
assert.ok(out.review !== null, `${legLabel}: review file must exist after a failed lane — its absence is what write_reviews silently omitted (#2605)`);
|
||||
assert.notStrictEqual(out.review.trim(), '', `${legLabel}: review file must not be empty after a failed lane (#2605)`);
|
||||
assert.match(
|
||||
out.review,
|
||||
new RegExp(`${legLabel.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')} review failed or returned empty output`, 'i'),
|
||||
`${legLabel}: review file must carry a diagnosable failure line (#2605)`,
|
||||
);
|
||||
}
|
||||
|
||||
for (const leg of LEGS) {
|
||||
describe(`#2605 — ${leg.label} reviewer leg fails loudly`, { skip: skipReason }, () => {
|
||||
test('an unreachable endpoint produces a diagnosable stub carrying curl stderr', () => {
|
||||
// Transport failure: curl writes to stderr and exits non-zero. `-s` alone
|
||||
// would have suppressed this text, which is why the fix uses `-sS`.
|
||||
const out = runLeg({
|
||||
leg,
|
||||
curlBody: curlStub({ completionStderr: STUB_STDERR, exitCode: 7 }),
|
||||
const okBody = (content) => ({
|
||||
ok: true, status: 200, body: JSON.stringify({ choices: [{ message: { content } }] }),
|
||||
});
|
||||
|
||||
assertDiagnosable(out, leg.label);
|
||||
assert.ok(
|
||||
out.review.includes(STUB_STDERR),
|
||||
`${leg.label}: captured stderr must be appended to the review file, not discarded to /dev/null (#2605)`,
|
||||
);
|
||||
assert.ok(
|
||||
fs.existsSync(out.errPath),
|
||||
`${leg.label}: stderr must be captured to a .err sidecar (#2605)`,
|
||||
);
|
||||
describe('#2605 local OpenAI-compatible lanes produce diagnosable output', () => {
|
||||
for (const lane of HTTP_LANES) {
|
||||
test(`${lane.slug}: an unreachable endpoint produces a stub carrying the transport error`, async () => {
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen({ ok: false, status: 0, body: '', error: 'ECONNREFUSED' }));
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, true);
|
||||
assert.ok(d.files[p.reviewPath].includes('ECONNREFUSED'),
|
||||
'the transport error must be visible — bare `curl -s` used to swallow it');
|
||||
});
|
||||
|
||||
test('an HTTP error body produces a diagnosable stub carrying the raw response', () => {
|
||||
// Application failure: an OpenAI-compatible server returns its error JSON in
|
||||
// the BODY and curl exits 0, so stderr is empty and only the body is
|
||||
// evidence. The llama_cpp leg used to pipe curl straight into jq, throwing
|
||||
// the body away before anything could inspect it.
|
||||
const out = runLeg({
|
||||
leg,
|
||||
curlBody: curlStub({ completionStdout: STUB_ERROR_BODY, exitCode: 0 }),
|
||||
test(`${lane.slug}: an HTTP error body is preserved in the stub`, async () => {
|
||||
// The body is the ONLY evidence on a 4xx/5xx: such a server returns its error JSON there and
|
||||
// curl still exits 0, so stderr is empty. The old pipe into jq discarded it.
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen({ ok: false, status: 404, body: '{"error":"model not found"}' }));
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.ok(d.files[p.reviewPath].includes('Raw response body:'));
|
||||
assert.ok(d.files[p.reviewPath].includes('model not found'));
|
||||
});
|
||||
|
||||
assertDiagnosable(out, leg.label);
|
||||
assert.ok(
|
||||
out.review.includes('gsd-2605-stub: model not loaded'),
|
||||
`${leg.label}: the raw response body must be preserved — it is the only evidence when curl exits 0 (#2605)`,
|
||||
);
|
||||
test(`${lane.slug}: an empty 200 response still produces a file`, async () => {
|
||||
// Previously nothing was written, so the file never existed, write_reviews omitted the
|
||||
// section, and the result was indistinguishable from the reviewer never being selected.
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen(okBody('')));
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, true);
|
||||
assert.ok(d.files[p.reviewPath] !== undefined, 'a file must exist even on an empty reply');
|
||||
assert.ok(d.files[p.reviewPath].includes('failed or returned empty output'));
|
||||
});
|
||||
|
||||
test('an empty 200 response still produces a stub rather than no file', () => {
|
||||
// Boundary: a well-formed response whose content is the empty string. This
|
||||
// is the exact case the issue reproduces — the old code took the `else`
|
||||
// branch and wrote nothing at all.
|
||||
const out = runLeg({
|
||||
leg,
|
||||
curlBody: curlStub({
|
||||
completionStdout: '{"choices":[{"message":{"content":""}}]}',
|
||||
exitCode: 0,
|
||||
}),
|
||||
test(`${lane.slug}: a whitespace-only response is empty, not a successful review`, async () => {
|
||||
// `[ ! -s … ]` counted BYTES, so " " passed as a real review.
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen(okBody(' \n\t ')));
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, true);
|
||||
});
|
||||
|
||||
assertDiagnosable(out, leg.label);
|
||||
test(`${lane.slug}: a reply that is exactly an echo option is NOT misclassified`, async () => {
|
||||
// `echo "$VAR"` would write 0 bytes for `-n`/`-e`/`-E`. Nothing here goes through echo, so
|
||||
// this is structurally impossible now — locked anyway.
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen(okBody('-n')));
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, false);
|
||||
assert.ok(d.files[p.reviewPath].includes('-n'));
|
||||
assert.ok(!d.files[p.reviewPath].includes('failed or returned empty output'));
|
||||
});
|
||||
|
||||
test('a whitespace-only response is treated as empty, not as a successful review', () => {
|
||||
// `[ ! -s … ]` counts BYTES, so a reply of " " was written out and passed
|
||||
// the guard as a "successful" but vacuous review — the same
|
||||
// indistinguishable-from-success outcome the guard exists to prevent.
|
||||
// Command substitution strips trailing newlines but NOT spaces, so this
|
||||
// case is not covered by the empty-string case above.
|
||||
const out = runLeg({
|
||||
leg,
|
||||
curlBody: curlStub({
|
||||
completionStdout: '{"choices":[{"message":{"content":" "}}]}',
|
||||
exitCode: 0,
|
||||
}),
|
||||
});
|
||||
|
||||
assertDiagnosable(out, leg.label);
|
||||
});
|
||||
|
||||
test('a reply that is exactly an echo option is not misclassified as empty', () => {
|
||||
// `echo "$VAR" > file` writes 0 bytes when VAR is exactly `-n`, which would
|
||||
// trip the empty guard and DISCARD a genuine reply. printf is required.
|
||||
const out = runLeg({
|
||||
leg,
|
||||
curlBody: curlStub({
|
||||
completionStdout: '{"choices":[{"message":{"content":"-n"}}]}',
|
||||
exitCode: 0,
|
||||
}),
|
||||
});
|
||||
|
||||
assert.ok(out.review !== null, `${leg.label}: a reply of "-n" must still produce a review file (#2605)`);
|
||||
assert.ok(
|
||||
out.review.includes('-n'),
|
||||
`${leg.label}: a reply of "-n" must be written verbatim, not swallowed by echo's option parsing (#2605)`,
|
||||
);
|
||||
assert.ok(
|
||||
!/failed or returned empty output/i.test(out.review),
|
||||
`${leg.label}: a genuine "-n" reply must not be misclassified as empty (#2605)`,
|
||||
);
|
||||
});
|
||||
|
||||
test('a budget skip leaves a visible stub rather than no file', () => {
|
||||
// The skip path sits one `if` away from the guard and dropped the lane just
|
||||
// as silently: no file, so write_reviews omitted the section entirely and
|
||||
// the only trace was a stderr warning nothing persists.
|
||||
const out = runLeg({
|
||||
leg,
|
||||
curlBody: curlStub({ completionStdout: '{"choices":[{"message":{"content":"unused"}}]}' }),
|
||||
// `gsd_run` supplies the budget, so stubbing the variable directly is
|
||||
// useless — the block's first line overwrites it. Drive the skip through
|
||||
// `gsd_run` itself: `config-get` yields a non-null budget so the trim
|
||||
// branch is entered, and `prompt-budget` returns 2 — the documented
|
||||
// "budget too small for the minimum review set" code. Stubbing only the
|
||||
// helper would not work for the Ollama leg, whose fence also carries the
|
||||
// shared `prepare_trimmed_prompt_for_reviewer` definition and so
|
||||
// overrides any stub of it.
|
||||
preamble: 'gsd_run() { case "$2" in prompt-budget) return 2 ;; esac; echo 1; }\n'
|
||||
+ 'prepare_trimmed_prompt_for_reviewer() { return 2; }\n',
|
||||
});
|
||||
|
||||
assert.ok(out.review !== null, `${leg.label}: a budget-skipped lane must still produce a review file (#2605)`);
|
||||
assert.match(
|
||||
out.review,
|
||||
/review skipped: prompt budget/i,
|
||||
`${leg.label}: a budget-skipped lane must say so in the review file (#2605)`,
|
||||
);
|
||||
});
|
||||
|
||||
test('a successful review passes through untouched', () => {
|
||||
// The guard must not fire on real content, and must not wrap or annotate it.
|
||||
const out = runLeg({
|
||||
leg,
|
||||
curlBody: curlStub({
|
||||
completionStdout: '{"choices":[{"message":{"content":"## Real Review\\nLooks good."}}]}',
|
||||
exitCode: 0,
|
||||
}),
|
||||
});
|
||||
|
||||
assert.ok(out.review !== null, `${leg.label}: a successful review must produce a review file (#2605)`);
|
||||
assert.ok(
|
||||
out.review.includes('Looks good.'),
|
||||
`${leg.label}: a successful review must pass through untouched (#2605)`,
|
||||
);
|
||||
assert.ok(
|
||||
!/failed or returned empty output/i.test(out.review),
|
||||
`${leg.label}: the stub must not fire on a non-empty review (#2605)`,
|
||||
);
|
||||
});
|
||||
test(`${lane.slug}: a successful review passes through untouched`, async () => {
|
||||
const p = planFor(lane.slug);
|
||||
const d = deps(reachableThen(okBody('## Findings\nreal review')));
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, false);
|
||||
assert.ok(d.files[p.reviewPath].includes('## Findings'));
|
||||
assert.ok(!d.files[p.reviewPath].includes('failed or returned empty output'));
|
||||
});
|
||||
}
|
||||
|
||||
describe('#2605 — the CodeRabbit leg fails loudly', { skip: skipReason }, () => {
|
||||
// CodeRabbit was the last CLI leg still shaped like pre-#2494 code:
|
||||
// `2>/dev/null > file` with no stub. A missing or unauthenticated binary left a
|
||||
// zero-byte file that write_reviews rendered as "ran cleanly, nothing to report".
|
||||
const CODERABBIT_BLOCK = extractBlock('\\*\\*CodeRabbit:\\*\\*', 'CodeRabbit');
|
||||
|
||||
test('a failing coderabbit CLI produces a diagnosable stub, not a zero-byte file', () => {
|
||||
if (process.platform === 'win32') return;
|
||||
|
||||
const caseDir = fs.mkdtempSync(path.join(sandbox, 'cr-'));
|
||||
const runDir = path.join(caseDir, 'run');
|
||||
const binDir = path.join(caseDir, 'bin');
|
||||
fs.mkdirSync(runDir);
|
||||
fs.mkdirSync(binDir);
|
||||
|
||||
const stub = path.join(binDir, 'coderabbit');
|
||||
fs.writeFileSync(stub, `#!/bin/sh\necho "${STUB_STDERR}" >&2\nexit 127\n`);
|
||||
fs.chmodSync(stub, 0o755);
|
||||
|
||||
const script = CODERABBIT_BLOCK.split('{run_dir}').join(runDir);
|
||||
spawnSync('bash', ['-c', script], {
|
||||
encoding: 'utf8',
|
||||
timeout: 30000,
|
||||
killSignal: 'SIGKILL',
|
||||
env: { ...process.env, PATH: `${binDir}${path.delimiter}${process.env.PATH}` },
|
||||
test('a served-model mismatch is warned about, not silently accepted', async () => {
|
||||
const p = planFor('lm_studio', { 'review.models.lm_studio': 'asked-for' });
|
||||
const d = deps(async () => ({
|
||||
ok: true, status: 200,
|
||||
body: JSON.stringify({ model: 'actually-served', choices: [{ message: { content: 'R' } }] }),
|
||||
}));
|
||||
const out = await runOpenAiCompatible(p, 'PLAN', d);
|
||||
assert.equal(out.review, 'R');
|
||||
assert.ok(d.warnings.some((w) => w.includes('actually-served') && w.includes('asked-for')));
|
||||
});
|
||||
|
||||
const reviewPath = path.join(runDir, 'gsd-review-coderabbit.md');
|
||||
assert.ok(fs.existsSync(reviewPath), 'CodeRabbit: review file must exist after a failed lane (#2605)');
|
||||
const review = fs.readFileSync(reviewPath, 'utf-8');
|
||||
assert.notStrictEqual(review.trim(), '', 'CodeRabbit: review file must not be zero-byte after a failed lane (#2605)');
|
||||
assert.match(
|
||||
review,
|
||||
/CodeRabbit review failed or returned empty output/i,
|
||||
'CodeRabbit: review file must carry a diagnosable failure line (#2605)',
|
||||
);
|
||||
assert.ok(
|
||||
review.includes(STUB_STDERR),
|
||||
'CodeRabbit: captured stderr must be appended, not discarded to /dev/null (#2605)',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
describe('#2605 — every local-server leg uses -sS so curl errors are not suppressed', { skip: skipReason }, () => {
|
||||
test('the chat/completions call captures stderr to a sidecar instead of /dev/null', () => {
|
||||
// Pins the two changes that make the stub's evidence real rather than empty.
|
||||
// Asserted on the extracted block text because the redirect target is the
|
||||
// contract; the behavioural consequence is covered by the suites above.
|
||||
for (const leg of LEGS) {
|
||||
assert.match(
|
||||
leg.block,
|
||||
new RegExp(`curl -sS[^\\n]*\\n[^\\n]*-d @- 2>[^\\n]*gsd-review-${leg.key}\\.err`),
|
||||
`${leg.label}: the completion call must use -sS and redirect stderr to its .err sidecar (#2605)`,
|
||||
);
|
||||
assert.ok(
|
||||
!new RegExp(`curl -s --max-time 120`).test(leg.block),
|
||||
`${leg.label}: the completion call must not use bare -s, which suppresses curl's error text (#2605)`,
|
||||
);
|
||||
test('neither jq nor curl is required by any of these lanes', () => {
|
||||
// The dependency is gone, not merely satisfied: parsing is JSON.parse and the request is
|
||||
// in-process. `jq` is absent on stock Windows/Git-Bash (#2589), which gated these lanes.
|
||||
for (const lane of HTTP_LANES) {
|
||||
assert.deepStrictEqual([...lane.requiresBinaries], [], `${lane.slug} still declares a binary`);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
@@ -1,166 +1,93 @@
|
||||
// allow-test-rule: source-text-is-the-product (see #2794)
|
||||
// The qwen reviewer dispatch block in gsd-core/workflows/review.md IS the runtime
|
||||
// contract — the workflow's text is what the reviewing agent executes. This suite
|
||||
// extracts that shell block verbatim and runs it under a real bash against a stubbed
|
||||
// `qwen`, so the shipped guard is what gets exercised rather than a reimplementation
|
||||
// of it. Assertions are on the guard's documented output contract (the stub line the
|
||||
// consensus step must be able to tell apart from a clean empty review, and the
|
||||
// presence of the captured stderr), not incidental string matching.
|
||||
|
||||
/**
|
||||
* Regression tests for #2794 — the qwen reviewer leg was the last one still
|
||||
* sending stderr to /dev/null.
|
||||
* #2794 — the qwen reviewer leg was the last one still sending stderr to /dev/null.
|
||||
*
|
||||
* Every other lane captures stderr to a `.err` sidecar and appends it to the
|
||||
* empty-output stub (#2494 for claude/gemini, #2605 for the local servers and
|
||||
* CodeRabbit). qwen alone ran `2>/dev/null` and wrote a bare
|
||||
* "Qwen review failed or returned empty output." with no diagnostic, so a missing
|
||||
* binary, an auth prompt, a rate-limit, and a genuinely empty review were all
|
||||
* indistinguishable — the same cross-cutting defect landing per-leg that ADR-2782
|
||||
* exists to end.
|
||||
* Every other lane captured stderr to a `.err` sidecar and appended it to the stub (#2494/#2605);
|
||||
* qwen wrote a bare "failed or returned empty output." with no diagnostic at all, so a missing
|
||||
* binary, an auth prompt and a rate-limit were indistinguishable from each other AND from a clean
|
||||
* empty review.
|
||||
*
|
||||
* These tests fail against pre-fix review.md: the stub carries no stderr, so the
|
||||
* diagnostic assertion in E2 trips.
|
||||
* Phase 5b (#2799) deleted the per-CLI bash this suite used to extract and run under a real bash.
|
||||
* The invariant is now structural rather than per-leg — the sidecar is written by the shared runner
|
||||
* for every lane — so these tests drive the runner directly. That also means the defect this issue
|
||||
* describes can no longer recur for ONE lane: there is no longer a per-lane place to get it wrong.
|
||||
*/
|
||||
|
||||
'use strict';
|
||||
|
||||
const { describe, test, before, after } = require('node:test');
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const { spawnSync } = require('node:child_process');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
const { createTempDir, cleanup } = require('./helpers.cjs');
|
||||
|
||||
const ROOT = path.join(__dirname, '..');
|
||||
const REVIEW_PATH = path.join(ROOT, 'gsd-core', 'workflows', 'review.md');
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const { runLane } = require('../gsd-core/bin/lib/review-lane-runner.cjs');
|
||||
|
||||
// Normalize CRLF: on a Windows autocrlf checkout every line carries a trailing
|
||||
// \r, which would mangle the extracted block's redirect tokens and defeat the
|
||||
// fence regex below.
|
||||
const WORKFLOW = fs.readFileSync(REVIEW_PATH, 'utf-8').replace(/\r\n/g, '\n');
|
||||
const RUN = '/run';
|
||||
const ROOT = '/repo';
|
||||
|
||||
/**
|
||||
* Extract the qwen dispatch block verbatim. If review.md changes the block's
|
||||
* shape this throws and the test fails loudly — intended coupling, the same
|
||||
* contract the #2494 and #2605 suites pin for their legs.
|
||||
*/
|
||||
function extractQwenBlock() {
|
||||
// `\r?\n` throughout: WORKFLOW is CRLF-normalized above, but the anchors stay
|
||||
// CRLF-tolerant so the regex cannot silently miss on a Windows checkout.
|
||||
const m = WORKFLOW.match(
|
||||
/<!--\s*reviewer-lane:\s*qwen\s*-->\r?\n[\s\S]*?```bash\r?\n([\s\S]*?)\r?\n```/,
|
||||
);
|
||||
assert.ok(m, 'review.md must define the qwen reviewer dispatch as a bash block (#2794)');
|
||||
return m[1];
|
||||
function planFor(slug) {
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === slug);
|
||||
const r = resolveLanePlan({ lane, configGet: () => undefined, runDir: RUN, repoRoot: ROOT });
|
||||
assert.equal(r.ok, true);
|
||||
return r.plan;
|
||||
}
|
||||
|
||||
const QWEN_BLOCK = extractQwenBlock();
|
||||
|
||||
// The extracted block is POSIX shell. Git Bash on Windows ignores Node's chmod
|
||||
// exec bit for PATH-executed extension-less scripts, so the stub would never run
|
||||
// there; the guard logic is platform-independent and is asserted in full on every
|
||||
// macOS/Linux leg.
|
||||
const WIN32_SKIP =
|
||||
'extracted block is POSIX shell; guard logic is platform-independent and asserted on macOS/Linux';
|
||||
|
||||
let sandbox;
|
||||
|
||||
before(() => { sandbox = createTempDir('gsd-2794-'); });
|
||||
after(() => { cleanup(sandbox); });
|
||||
|
||||
/**
|
||||
* Run the extracted block with `{run_dir}` pointed at a fresh run directory and
|
||||
* `qwenBody` installed on PATH as `qwen`. Pass `qwenBody: null` to run with no
|
||||
* qwen on PATH at all.
|
||||
*/
|
||||
function runQwenLeg(qwenBody) {
|
||||
// Single exec site so the Windows guard lives in one place: Git Bash (msys2)
|
||||
// ignores Node's chmod exec bit for PATH-executed extension-less scripts
|
||||
// (DEFECT.WINDOWS-TEST-PORTABILITY). Every test below is skipped on win32 —
|
||||
// this early return keeps the exec unreachable there rather than relying on
|
||||
// the skip alone.
|
||||
if (process.platform === 'win32') return null;
|
||||
|
||||
const caseDir = fs.mkdtempSync(path.join(sandbox, 'run-'));
|
||||
const runDir = path.join(caseDir, 'run');
|
||||
const binDir = path.join(caseDir, 'bin');
|
||||
fs.mkdirSync(runDir);
|
||||
fs.mkdirSync(binDir);
|
||||
|
||||
if (qwenBody !== null) {
|
||||
const stub = path.join(binDir, 'qwen');
|
||||
fs.writeFileSync(stub, qwenBody);
|
||||
fs.chmodSync(stub, 0o755);
|
||||
}
|
||||
|
||||
fs.writeFileSync(path.join(runDir, 'gsd-review-prompt.md'), '# review prompt\n');
|
||||
|
||||
const script = QWEN_BLOCK.split('{run_dir}').join(runDir);
|
||||
const result = spawnSync('bash', ['-c', script], {
|
||||
encoding: 'utf8',
|
||||
timeout: 30000,
|
||||
killSignal: 'SIGKILL',
|
||||
// An empty PATH would also break `cat`; prepend the stub dir instead so the
|
||||
// no-qwen case still resolves the shell builtins the block needs.
|
||||
env: { ...process.env, PATH: `${binDir}${path.delimiter}${process.env.PATH}` },
|
||||
});
|
||||
|
||||
const reviewPath = path.join(runDir, 'gsd-review-qwen.md');
|
||||
function deps(spawnResult, files = {}, hasBinary = () => true) {
|
||||
return {
|
||||
result,
|
||||
reviewPath,
|
||||
errPath: path.join(runDir, 'gsd-review-qwen.err'),
|
||||
review: fs.existsSync(reviewPath) ? fs.readFileSync(reviewPath, 'utf-8') : null,
|
||||
files,
|
||||
spawn: () => spawnResult,
|
||||
httpJson: async () => ({ ok: true, status: 200, body: '{}' }),
|
||||
readFile: (p) => { if (!(p in files)) throw new Error(`ENOENT ${p}`); return files[p]; },
|
||||
writeFile: (p, c) => { files[p] = c; },
|
||||
exists: (p) => p in files,
|
||||
hasBinary,
|
||||
configGet: () => undefined,
|
||||
homeDir: '/home/u',
|
||||
warn: () => {},
|
||||
};
|
||||
}
|
||||
|
||||
const STUB_STDERR = 'gsd-2794-stub: qwen: authentication required';
|
||||
|
||||
describe('qwen reviewer leg — empty-output guard (#2794)', () => {
|
||||
test('writes the review on success', { skip: process.platform === 'win32' ? WIN32_SKIP : false }, () => {
|
||||
const out = runQwenLeg('#!/bin/sh\necho "## Qwen findings"\nexit 0\n');
|
||||
assert.ok(out.review !== null, 'review file must exist');
|
||||
assert.match(out.review, /Qwen findings/);
|
||||
// A successful lane must NOT be decorated with the failure stub.
|
||||
assert.doesNotMatch(out.review, /failed or returned empty output/i);
|
||||
describe('#2794 qwen reviewer stderr capture', () => {
|
||||
test('writes the review on success', async () => {
|
||||
const p = planFor('qwen');
|
||||
const d = deps({ status: 0, stdout: '## Qwen findings\nall good\n', stderr: '' });
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, false);
|
||||
assert.ok(d.files[p.reviewPath].includes('## Qwen findings'));
|
||||
});
|
||||
|
||||
test('a failed lane surfaces its stderr in the review stub', { skip: process.platform === 'win32' ? WIN32_SKIP : false }, () => {
|
||||
// THE regression row. Pre-fix the block ran `2>/dev/null`, so this stderr
|
||||
// was discarded and the stub carried no diagnostic at all.
|
||||
const out = runQwenLeg(`#!/bin/sh\necho "${STUB_STDERR}" >&2\nexit 1\n`);
|
||||
assert.ok(out.review !== null, 'review file must exist after a failed lane');
|
||||
assert.notStrictEqual(out.review.trim(), '', 'review file must not be empty');
|
||||
assert.match(
|
||||
out.review,
|
||||
/Qwen review failed or returned empty output/i,
|
||||
'stub must name the lane as failed-or-empty so consensus can tell it apart',
|
||||
);
|
||||
assert.ok(
|
||||
out.review.includes(STUB_STDERR),
|
||||
`stub must carry the captured stderr; got: ${JSON.stringify(out.review)}`,
|
||||
);
|
||||
assert.ok(fs.existsSync(out.errPath), 'stderr sidecar must be written');
|
||||
test('a failed lane surfaces its stderr in the review stub', async () => {
|
||||
const p = planFor('qwen');
|
||||
const d = deps({ status: 1, stdout: '', stderr: 'auth required: run `qwen login`' });
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.ok(d.files[p.reviewPath].includes('auth required'),
|
||||
'the diagnostic must reach the review, not just the sidecar');
|
||||
assert.equal(d.files[p.errPath], 'auth required: run `qwen login`');
|
||||
});
|
||||
|
||||
test('a silently empty lane still produces a diagnosable stub', { skip: process.platform === 'win32' ? WIN32_SKIP : false }, () => {
|
||||
// Boundary: nothing on either stream. The stub line is the only signal, and
|
||||
// it must still be present so write_reviews does not render the lane as a
|
||||
// reviewer that ran cleanly with nothing to report.
|
||||
const out = runQwenLeg('#!/bin/sh\nexit 0\n');
|
||||
assert.ok(out.review !== null, 'review file must exist');
|
||||
assert.notStrictEqual(out.review.trim(), '', 'review file must not be zero-byte');
|
||||
assert.match(out.review, /Qwen review failed or returned empty output/i);
|
||||
test('a silently empty lane still produces a diagnosable stub', async () => {
|
||||
const p = planFor('qwen');
|
||||
const d = deps({ status: 0, stdout: '', stderr: '' });
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.stubbed, true);
|
||||
assert.ok(d.files[p.reviewPath].includes('failed or returned empty output'));
|
||||
});
|
||||
|
||||
test('a missing qwen binary produces a diagnosable stub', { skip: process.platform === 'win32' ? WIN32_SKIP : false }, () => {
|
||||
const out = runQwenLeg(null);
|
||||
assert.ok(out.review !== null, 'review file must exist when the binary is absent');
|
||||
assert.notStrictEqual(out.review.trim(), '', 'review file must not be zero-byte');
|
||||
assert.match(out.review, /Qwen review failed or returned empty output/i);
|
||||
// The shell's own "command not found" lands in the sidecar and is appended,
|
||||
// which is the whole point of capturing instead of discarding.
|
||||
assert.ok(fs.existsSync(out.errPath), 'stderr sidecar must be written');
|
||||
test('a missing qwen binary reports unavailable rather than an empty review', async () => {
|
||||
// Stronger than the original: the lane is now reported with a TYPED reason before it is ever
|
||||
// spawned, instead of producing a stub that looked the same as every other failure.
|
||||
const p = planFor('qwen');
|
||||
const d = deps({ status: 0, stdout: '', stderr: '' }, {}, () => false);
|
||||
const r = await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(r.ok, false);
|
||||
assert.equal(r.reason, 'missing_binary');
|
||||
assert.equal(d.files[p.reviewPath], undefined, 'no review file for a lane that never ran');
|
||||
});
|
||||
|
||||
test('the sidecar is structural — no lane can opt out of it', async () => {
|
||||
// The #2794 defect was one lane diverging from a convention every other lane followed. There is
|
||||
// no per-lane place to diverge any more; this asserts that directly.
|
||||
for (const lane of REVIEWER_LANES.filter((l) => l.transport === 'spawn')) {
|
||||
const p = planFor(lane.slug);
|
||||
assert.ok(p.errPath.endsWith('.err'), `${lane.slug} must declare a stderr sidecar`);
|
||||
assert.notEqual(p.errPath, '/dev/null');
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
@@ -1,224 +1,125 @@
|
||||
// allow-test-rule: source-text-is-the-product (see #1936)
|
||||
// The OpenCode reviewer reconstructs its review from opencode's --format json
|
||||
// event stream using two embedded jq programs in gsd-core/workflows/review.md.
|
||||
// Those programs ARE the runtime contract; this test extracts them verbatim from
|
||||
// the workflow and exercises the real jq (not a reimplementation) so the shipped
|
||||
// reconstruction logic is what gets property-tested.
|
||||
//
|
||||
// ARCHITECTURE (#2099): the shipped jq program is run over a WHOLE fast-check corpus
|
||||
// in ONE jq process, not once per generated case. Each generated event stream is
|
||||
// written as one compact-JSON array per line to a temp file, then `jq -c <PROGRAM>`
|
||||
// (no `-s`) applies PROGRAM to each array — `.` is that array, exactly what
|
||||
// production's `jq -rs` sees after slurping opencode's one-value-per-line stream —
|
||||
// emitting one compact-JSON result per line. This is empirically identical to the
|
||||
// per-stream `-rs` form (verified across embedded-newline/empty/quote/unicode/
|
||||
// null-drop cases) AND reads from a file like production (`jq -rs '…' <file>`), so
|
||||
// there is no stdin pipe to deadlock on large I/O. The prior design spawned ~600
|
||||
// synchronous jq subprocesses (numRuns × 3 properties); a single one freezing on a
|
||||
// contended CI runner hung the whole unit-test chunk to its 600s kill (macOS CI,
|
||||
// #2099). `node --test`'s --test-force-exit cannot interrupt a synchronous
|
||||
// execFileSync, so the cure is to stop spawning per case — not just to time-bound it.
|
||||
/**
|
||||
* OpenCode review reconstruction — property tests (#1936).
|
||||
*
|
||||
* The OpenCode lane rebuilds its review from the assistant `text` parts of a `--format json` event
|
||||
* stream. `--format json` is the PRIMARY invocation, not a fallback: the default formatter drops
|
||||
* the assistant text when the agent ends its turn with no final message, silently losing the
|
||||
* reviewer.
|
||||
*
|
||||
* This suite used to extract two embedded `jq` programs from `review.md` and exercise the real jq
|
||||
* so the shipped logic was what got property-tested. Phase 5b (#2799) replaced those programs with
|
||||
* a named first-party handler in JavaScript, so the property now runs directly against the shipped
|
||||
* function.
|
||||
*
|
||||
* That deletes the entire #2099 problem this file was architected around. The old design spawned
|
||||
* ~600 synchronous jq subprocesses (numRuns × 3 properties), and one freezing on a contended CI
|
||||
* runner hung the whole unit-test chunk to its 600s kill — `--test-force-exit` cannot interrupt a
|
||||
* synchronous `execFileSync`. The batching workaround (one jq process over a whole corpus) existed
|
||||
* solely to avoid that. There is now no subprocess at all, so the hazard is gone by construction
|
||||
* rather than mitigated.
|
||||
*/
|
||||
|
||||
'use strict';
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const { execFileSync } = require('node:child_process');
|
||||
const fs = require('node:fs');
|
||||
const os = require('node:os');
|
||||
const path = require('node:path');
|
||||
const fc = require('./helpers/fast-check-setup.cjs');
|
||||
const { cleanup } = require('./helpers.cjs');
|
||||
const fc = require('fast-check');
|
||||
|
||||
const reviewPath = path.resolve(__dirname, '..', 'gsd-core', 'workflows', 'review.md');
|
||||
const workflow = fs.readFileSync(reviewPath, 'utf-8');
|
||||
const { handleOpencodeOutput } = require('../gsd-core/bin/lib/review-lane-runner.cjs');
|
||||
|
||||
// Extract the two shipped jq programs verbatim. If review.md changes their shape,
|
||||
// these throw and the test fails loudly (intended coupling — #1936).
|
||||
function extractJqProgram(varName) {
|
||||
const re = new RegExp(`${varName}=\\$\\(jq -rs '([^']*)'`);
|
||||
const m = workflow.match(re);
|
||||
assert.ok(m, `review.md must define ${varName} via jq -rs '<program>' (#1936)`);
|
||||
return m[1];
|
||||
}
|
||||
const TEXT_PROGRAM = extractJqProgram('OPENCODE_REVIEW'); // review reconstruction
|
||||
const DIAG_PROGRAM = extractJqProgram('OPENCODE_DIAG'); // empty-output diagnostic
|
||||
/** Deterministic: pinned seed, bounded runs, replay data printed on failure. */
|
||||
const FC = { seed: 42, numRuns: 200 };
|
||||
|
||||
// This suite shells out to `jq`. On Windows, Node's child_process argument quoting
|
||||
// mangles the jq program (it embeds quotes) — jq then raises a parse error — and
|
||||
// jq isn't guaranteed on the host regardless. The reconstruction logic is
|
||||
// platform-independent (the review workflow runs jq in its Unix-y runtime), so gate
|
||||
// the suite to jq-present non-Windows hosts, mirroring golden-install-parity's win32
|
||||
// skip. The assertions run in full on every macOS/Linux CI leg.
|
||||
let jqAvailable = false;
|
||||
try { execFileSync('jq', ['--version'], { stdio: 'ignore', timeout: 10000, killSignal: 'SIGKILL' }); jqAvailable = true; } catch { /* no jq on PATH */ }
|
||||
const skipReason = process.platform === 'win32'
|
||||
? 'jq invocation is not portable under Node child_process arg-quoting on Windows; logic is platform-independent and asserted on macOS/Linux'
|
||||
: (jqAvailable ? false : 'jq not on PATH');
|
||||
const opts = { skip: skipReason };
|
||||
const NL = String.fromCharCode(10);
|
||||
|
||||
// Bound each jq subprocess so a frozen spawn on a contended runner fails fast +
|
||||
// diagnosably (ETIMEDOUT) rather than hanging the chunk. jq over this corpus
|
||||
// completes in ~10ms, so 30s is an enormous margin that never trips on a healthy run.
|
||||
const JQ_EXEC_OPTS = { encoding: 'utf8', timeout: 30000, killSignal: 'SIGKILL', maxBuffer: 64 * 1024 * 1024 };
|
||||
/** An assistant text event, as opencode emits it. */
|
||||
const textEvent = (text) => ({ type: 'text', part: { text } });
|
||||
/** A terminal step event carrying the stop reason and token counts. */
|
||||
const stepFinish = (reason, output) => ({ type: 'step_finish', part: { reason, tokens: { output } } });
|
||||
|
||||
// Run a shipped jq program over an array of event streams in a SINGLE jq process.
|
||||
// Returns one result per input stream (order preserved), decoded from jq's compact
|
||||
// (`-c`) JSON output back to the raw string production's `-r` would have captured.
|
||||
// Both shipped programs yield exactly one value per stream (join(...) / last|"...");
|
||||
// the length assertion pins that invariant so a future program change that broke it
|
||||
// (0 or >1 outputs) fails loudly instead of silently misaligning results.
|
||||
function runJqBatch(program, streams) {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-jq-batch-'));
|
||||
try {
|
||||
const file = path.join(dir, 'streams.jsonl');
|
||||
fs.writeFileSync(file, streams.map((events) => JSON.stringify(events)).join('\n') + '\n');
|
||||
const out = execFileSync('jq', ['-c', program, file], JQ_EXEC_OPTS);
|
||||
const lines = out.split('\n').filter((line) => line.length > 0);
|
||||
assert.equal(
|
||||
lines.length,
|
||||
streams.length,
|
||||
`jq must emit exactly one result per stream (got ${lines.length} for ${streams.length})`,
|
||||
/** Serialize events the way opencode does: one compact JSON value per line. */
|
||||
const streamOf = (events) => events.map((e) => JSON.stringify(e)).join(NL);
|
||||
|
||||
/** Text that survives a JSON round-trip, including the shapes that broke the jq version. */
|
||||
const hostileText = fc.oneof(
|
||||
fc.string(),
|
||||
fc.constantFrom('', ' ', ' ', NL, 'a' + NL + 'b', '"quoted"', '\\backslash', 'emoji 🎉', 'null', '-n'),
|
||||
fc.string({ unit: 'grapheme' }),
|
||||
);
|
||||
return lines.map((line) => JSON.parse(line));
|
||||
} finally {
|
||||
cleanup(dir);
|
||||
}
|
||||
}
|
||||
|
||||
// The intended reconstruction, computed independently of jq. jq is the unit under
|
||||
// test; this JS is the spec it must match on every generated case.
|
||||
function expectedReview(events) {
|
||||
return events
|
||||
.filter((e) => e.type === 'text' && e.part && typeof e.part.text === 'string')
|
||||
.map((e) => e.part.text)
|
||||
.join('\n');
|
||||
}
|
||||
|
||||
// Text values safe to round-trip through JSON → jq (utf8) → string. Excludes lone
|
||||
// surrogates (which don't survive utf8) but keeps the interesting cases: newlines,
|
||||
// quotes, backslashes, braces, unicode.
|
||||
const safeText = fc
|
||||
.string({ minLength: 0, maxLength: 40 })
|
||||
.filter((s) => Buffer.from(s, 'utf8').toString('utf8') === s);
|
||||
|
||||
// A `text` event whose `.part.text` is a string, or null/absent (dropped by `// empty`).
|
||||
const textEvent = fc.record({
|
||||
type: fc.constant('text'),
|
||||
part: fc.oneof(
|
||||
fc.record({ text: safeText }),
|
||||
fc.record({ text: fc.constant(null) }), // null → jq `// empty` drops it
|
||||
fc.record({}), // absent → jq `// empty` drops it
|
||||
),
|
||||
});
|
||||
const stepFinishEvent = fc.record({
|
||||
type: fc.constant('step_finish'),
|
||||
part: fc.record({
|
||||
reason: fc.constantFrom('stop', 'length', 'tool_calls'),
|
||||
tokens: fc.record({ output: fc.integer({ min: 0, max: 100000 }) }),
|
||||
describe('opencode reconstruction — properties', () => {
|
||||
test('every assistant text part appears, in order, joined by newlines', () => {
|
||||
fc.assert(
|
||||
fc.property(fc.array(hostileText, { minLength: 1, maxLength: 8 }), (texts) => {
|
||||
const stream = streamOf(texts.map(textEvent));
|
||||
return handleOpencodeOutput(stream).review === texts.join(NL);
|
||||
}),
|
||||
});
|
||||
const nonTextEvent = fc.oneof(
|
||||
stepFinishEvent,
|
||||
fc.record({ type: fc.constant('tool_use'), part: fc.record({ tool: safeText }) }),
|
||||
fc.record({ type: fc.constant('step_start'), part: fc.record({}) }),
|
||||
);
|
||||
// Weight text events higher so streams routinely mix real review text with noise,
|
||||
// but also generate text-free streams (the #1936 zero-output case).
|
||||
const eventStream = fc.array(fc.oneof(textEvent, textEvent, nonTextEvent), {
|
||||
minLength: 1,
|
||||
maxLength: 30,
|
||||
});
|
||||
|
||||
// Corpus size per property. Matches the prior fast-check-setup numRuns:200 so
|
||||
// coverage is unchanged; distinct seeds give each property an independent corpus,
|
||||
// and fixed seeds keep the corpus deterministic across CI runs/OS legs.
|
||||
const CORPUS = 200;
|
||||
|
||||
describe('#1936 OpenCode review reconstruction — jq properties', () => {
|
||||
test('review == the newline-join of every assistant text part, over a generated corpus (order preserved)', opts, () => {
|
||||
const streams = fc.sample(eventStream, { numRuns: CORPUS, seed: 42 });
|
||||
const actual = runJqBatch(TEXT_PROGRAM, streams); // one jq process for the whole corpus
|
||||
streams.forEach((events, i) => {
|
||||
assert.equal(
|
||||
actual[i],
|
||||
expectedReview(events),
|
||||
`case ${i}: shipped jq review must equal the spec for events=${JSON.stringify(events)}`,
|
||||
FC,
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
test('a stream with no assistant text part reconstructs to empty (drives the #1936 stub), over a corpus', opts, () => {
|
||||
// The exact failure the bug describes: the agent runs tool calls and ends with
|
||||
// step_finish, emitting no text. Reconstruction must be empty so the content-gate
|
||||
// (`[ -n "$OPENCODE_REVIEW" ]`) falls through to the stub.
|
||||
const streams = fc.sample(fc.array(nonTextEvent, { minLength: 1, maxLength: 20 }), { numRuns: CORPUS, seed: 43 });
|
||||
const actual = runJqBatch(TEXT_PROGRAM, streams);
|
||||
streams.forEach((events, i) => {
|
||||
assert.equal(actual[i], '', `case ${i}: text-free stream must reconstruct to '' for ${JSON.stringify(events)}`);
|
||||
});
|
||||
});
|
||||
|
||||
test('text parts that are null/absent are dropped, never rendered as "null", over a corpus', opts, () => {
|
||||
const nullish = fc.array(
|
||||
fc.oneof(
|
||||
fc.record({ type: fc.constant('text'), part: fc.record({ text: fc.constant(null) }) }),
|
||||
fc.record({ type: fc.constant('text'), part: fc.record({}) }),
|
||||
test('non-text events never contribute to the review', () => {
|
||||
fc.assert(
|
||||
fc.property(
|
||||
fc.array(hostileText, { minLength: 1, maxLength: 5 }),
|
||||
fc.array(fc.record({ reason: fc.string(), output: fc.nat() }), { maxLength: 4 }),
|
||||
(texts, steps) => {
|
||||
const events = [...texts.map(textEvent), ...steps.map((s) => stepFinish(s.reason, s.output))];
|
||||
return handleOpencodeOutput(streamOf(events)).review === texts.join(NL);
|
||||
},
|
||||
),
|
||||
{ minLength: 1, maxLength: 10 },
|
||||
FC,
|
||||
);
|
||||
const streams = fc.sample(nullish, { numRuns: CORPUS, seed: 44 });
|
||||
const actual = runJqBatch(TEXT_PROGRAM, streams);
|
||||
streams.forEach((events, i) => {
|
||||
assert.equal(actual[i], '', `case ${i}: null/absent text must drop to ''`);
|
||||
assert.doesNotMatch(actual[i], /null/, `case ${i}: must never render "null"`);
|
||||
});
|
||||
});
|
||||
|
||||
// Explicit boundary + happy examples (deterministic, not sampled) — all in one
|
||||
// batched jq spawn. Pins the exact contract the corpus only covers probabilistically.
|
||||
test('boundary + happy example streams reconstruct exactly (batched)', opts, () => {
|
||||
const cases = [
|
||||
{ events: [{ type: 'text', part: { text: 'only' } }], expect: 'only' }, // single text part
|
||||
{ events: [{ type: 'text', part: { text: '' } }], expect: '' }, // empty-string text is kept
|
||||
{ events: [{ type: 'text', part: { text: 'a' } }, { type: 'tool_use', part: { tool: 'r' } }, { type: 'text', part: { text: 'b' } }], expect: 'a\nb' }, // text interleaved with noise
|
||||
{ events: [{ type: 'text', part: { text: 'x\ny' } }], expect: 'x\ny' }, // embedded newline preserved
|
||||
{ events: [{ type: 'tool_use', part: { tool: 'read' } }, { type: 'step_finish', part: { reason: 'stop', tokens: { output: 3 } } }], expect: '' }, // no text at all
|
||||
{ events: Array.from({ length: 30 }, (_v, i) => ({ type: 'text', part: { text: `p${i}` } })), expect: Array.from({ length: 30 }, (_v, i) => `p${i}`).join('\n') }, // max-size all-text stream
|
||||
];
|
||||
const actual = runJqBatch(TEXT_PROGRAM, cases.map((c) => c.events));
|
||||
cases.forEach((c, i) => assert.equal(actual[i], c.expect, `example ${i}: ${JSON.stringify(c.events)}`));
|
||||
test('a malformed line is skipped without losing the surrounding review', () => {
|
||||
// Losing an entire review to one unparseable line would be strictly worse than the bug this
|
||||
// handler exists to fix, so a partial stream must still yield its usable text.
|
||||
fc.assert(
|
||||
fc.property(
|
||||
fc.array(hostileText, { minLength: 1, maxLength: 4 }),
|
||||
fc.constantFrom('NOT JSON', '{"truncated":', '}{', '[', 'null', ' '),
|
||||
(texts, garbage) => {
|
||||
const lines = [...texts.map((t) => JSON.stringify(textEvent(t))), garbage];
|
||||
return handleOpencodeOutput(lines.join(NL)).review === texts.join(NL);
|
||||
},
|
||||
),
|
||||
FC,
|
||||
);
|
||||
});
|
||||
|
||||
// Diagnostic path (empty-output stub). The finding calls out `missing .tokens.output`
|
||||
// and no-step_finish as real edges — pin them with examples against the shipped jq,
|
||||
// all in one batched spawn.
|
||||
describe('diagnostic reconstruction (stop reason + output tokens)', () => {
|
||||
test('reports reason/tokens from the LAST step_finish and degrades missing fields to "?"', opts, () => {
|
||||
const cases = [
|
||||
{ events: [
|
||||
{ type: 'step_finish', part: { reason: 'tool_calls', tokens: { output: 5 } } },
|
||||
{ type: 'tool_use', part: {} },
|
||||
{ type: 'step_finish', part: { reason: 'stop', tokens: { output: 0 } } },
|
||||
], expect: 'stop reason=stop, output tokens=0' }, // LAST step_finish wins
|
||||
{ events: [{ type: 'step_finish', part: { reason: 'stop', tokens: {} } }], expect: 'stop reason=stop, output tokens=?' }, // missing .tokens.output → "?"
|
||||
{ events: [{ type: 'tool_use', part: { tool: 'read' } }], expect: 'stop reason=?, output tokens=?' }, // no step_finish at all → both "?"
|
||||
];
|
||||
const actual = runJqBatch(DIAG_PROGRAM, cases.map((c) => c.events));
|
||||
cases.forEach((c, i) => assert.equal(actual[i], c.expect, `diag ${i}: ${JSON.stringify(c.events)}`));
|
||||
});
|
||||
});
|
||||
|
||||
// The primary reconstruction runs before any content gate; on non-JSON stdout
|
||||
// (e.g. an opencode crash that printed a plain-text error) jq must fail rather
|
||||
// than emit that text as a "review" — the workflow's `2>/dev/null` + empty
|
||||
// capture then routes to the diagnostic stub.
|
||||
test('non-JSON stdout does not masquerade as a reconstructed review', opts, () => {
|
||||
let threw = false;
|
||||
test('is total — never throws on arbitrary input', () => {
|
||||
fc.assert(
|
||||
fc.property(fc.anything(), (anything) => {
|
||||
try {
|
||||
execFileSync('jq', ['-rs', TEXT_PROGRAM], { ...JQ_EXEC_OPTS, input: 'auth token expired\n' });
|
||||
const r = handleOpencodeOutput(anything);
|
||||
return typeof r.review === 'string' && typeof r.diagnostic === 'string';
|
||||
} catch {
|
||||
threw = true;
|
||||
return false;
|
||||
}
|
||||
assert.ok(threw, 'jq must reject non-JSON input so it cannot be captured as a review');
|
||||
}),
|
||||
FC,
|
||||
);
|
||||
});
|
||||
|
||||
test('CRLF streams reconstruct identically to LF', () => {
|
||||
fc.assert(
|
||||
fc.property(fc.array(fc.string({ minLength: 1 }), { minLength: 1, maxLength: 5 }), (texts) => {
|
||||
const lf = streamOf(texts.map(textEvent));
|
||||
const crlf = lf.split(NL).join(String.fromCharCode(13) + NL);
|
||||
return handleOpencodeOutput(lf).review === handleOpencodeOutput(crlf).review;
|
||||
}),
|
||||
FC,
|
||||
);
|
||||
});
|
||||
|
||||
test('a stream with no assistant text yields an empty review and a diagnostic', () => {
|
||||
fc.assert(
|
||||
fc.property(fc.string({ minLength: 1 }), fc.nat(), (reason, output) => {
|
||||
const r = handleOpencodeOutput(streamOf([stepFinish(reason, output)]));
|
||||
return r.review === '' && r.diagnostic.includes(String(output));
|
||||
}),
|
||||
FC,
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
@@ -171,71 +171,73 @@ const path = require('node:path');
|
||||
const reviewPath = path.resolve(__dirname, '..', 'gsd-core', 'workflows', 'review.md');
|
||||
const read = () => fs.readFileSync(reviewPath, 'utf-8');
|
||||
|
||||
describe('bug #687 → #2073: agy print mode bounded by --print-timeout PAIRED with an external timeout', () => {
|
||||
// #687 established that agy print mode must be bounded (its native
|
||||
// --print-timeout, default 5m). #2073 superseded the "no external killer"
|
||||
// half of that contract with documentation:
|
||||
// - agy's own print-mode guidance says to PAIR --print-timeout with an
|
||||
// external terminal `timeout` ("Pair with the terminal timeout= so the
|
||||
// outer call doesn't cut the run short"), because --print-timeout cannot
|
||||
// fire before agy creates a session (a pre-session stall otherwise hangs
|
||||
// unbounded). The external cap is set HIGHER than --print-timeout so it
|
||||
// only backstops a stall, never cuts a healthy run.
|
||||
// - agy gained `--model` in ~1.0.3 (issue #3782's "no --model flag" note
|
||||
// was correct at the time, stale now); review.models.agy is passed as
|
||||
// --model so a pinned model that 404s has an escape hatch.
|
||||
// - the prompt is now a file reference: inline `-p "$(cat …)"` overflows
|
||||
// the exec arg list on a large review prompt (Linux MAX_ARG_STRLEN
|
||||
// 128 KB/single-arg → rc 126).
|
||||
describe('bug #687 → #2073: agy print mode is bounded, and its fallback chain fires', () => {
|
||||
// Phase 5b (#2799) moved the agy invocation out of review.md's bash into the declared lane plus
|
||||
// the named `antigravity` handler, so these assertions follow it.
|
||||
//
|
||||
// ONE INVARIANT DELIBERATELY CHANGED. The old contract was "--print-timeout PAIRED with an
|
||||
// external timeout/gtimeout killer, falling back to bare agy on macOS". That fallback WAS the
|
||||
// bug: stock macOS ships neither killer, so the lane ran unbounded there — and --print-timeout
|
||||
// cannot fire before agy creates a session (#2073 mode 3), which is the whole reason a second
|
||||
// bound existed. spawnSync's native timeout is always available, so the outer bound is now
|
||||
// unconditional on every platform. Strictly stronger; only the mechanism changed.
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const { runLane } = require('../gsd-core/bin/lib/review-lane-runner.cjs');
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === 'antigravity');
|
||||
const planFor = () => {
|
||||
const r = resolveLanePlan({ lane, configGet: (k) => ({ 'review.models.agy': 'agy-m' })[k], runDir: '/run', repoRoot: '/repo' });
|
||||
assert.equal(r.ok, true);
|
||||
return r.plan;
|
||||
};
|
||||
|
||||
test('invokes agy with --print-timeout AND a paired external killer when available', () => {
|
||||
const c = read();
|
||||
assert.match(c, /--print-timeout \d+s?/, 'review.md must pass agy its native --print-timeout');
|
||||
// Capability probe for GNU `timeout` / macOS `gtimeout` (stock macOS has neither).
|
||||
assert.match(c, /command -v timeout/, 'review.md must probe for the `timeout` killer');
|
||||
assert.match(c, /command -v gtimeout/, 'review.md must probe for `gtimeout` (macOS Homebrew)');
|
||||
// The external cap (600s) is applied ahead of agy and is >= --print-timeout (540s).
|
||||
assert.match(c, /600 agy --print-timeout 540s/,
|
||||
'review.md must pair an external cap (600s) >= --print-timeout (540s) with agy (agy guidance)');
|
||||
test('keeps its native --print-timeout AND carries an unconditional outer bound', () => {
|
||||
const argv = planFor().argv;
|
||||
const i = argv.indexOf('--print-timeout');
|
||||
assert.notEqual(i, -1, 'the tool-native inner bound must survive');
|
||||
assert.match(argv[i + 1], /^\d+s$/);
|
||||
assert.ok(lane.timeoutFloorMs > 0, 'and an outer wall-clock bound must be declared');
|
||||
});
|
||||
|
||||
test('external cap is >= --print-timeout, and falls back to bare agy on macOS', () => {
|
||||
const c = read();
|
||||
const bound = c.match(/(\d+)\s+agy --print-timeout (\d+)s/);
|
||||
assert.ok(bound, 'review.md must encode the external-cap + --print-timeout pair');
|
||||
assert.ok(
|
||||
Number(bound[1]) >= Number(bound[2]),
|
||||
'external cap (seconds) must be >= --print-timeout (seconds) so it only backstops a stall',
|
||||
);
|
||||
// Graceful fallback when no external killer is available (stock macOS).
|
||||
assert.match(c, /else\n\s*agy --print-timeout/,
|
||||
'review.md must fall back to --print-timeout alone when no external killer is available (macOS)');
|
||||
test('the outer bound is larger than the inner one, so it only backstops', () => {
|
||||
const argv = planFor().argv;
|
||||
const innerSec = parseInt(argv[argv.indexOf('--print-timeout') + 1], 10);
|
||||
assert.ok(lane.timeoutFloorMs / 1000 > innerSec,
|
||||
'the outer cap must not pre-empt a healthy run bounded by --print-timeout');
|
||||
});
|
||||
|
||||
test('uses a file-reference prompt, not inline "$(cat …)" (arg-list overflow, #2073)', () => {
|
||||
const c = read();
|
||||
assert.doesNotMatch(c, /agy[^\n]*-p "\$\(cat/,
|
||||
'review.md must not feed agy the prompt inline via "$(cat …)" — a large review prompt overflows the exec arg list (rc 126)');
|
||||
assert.match(c, /Read the file at \{run_dir\}\/gsd-review-prompt\.md/,
|
||||
'review.md should pass agy a file-reference prompt (mirrors the Cursor block)');
|
||||
assert.equal(lane.invoke.promptChannel, 'argv-file-ref');
|
||||
const last = planFor().argv.slice(-1)[0];
|
||||
assert.ok(last.includes('/run/gsd-review-prompt.md'));
|
||||
assert.ok(last.length < 1000, 'the reference must never carry the prompt body');
|
||||
});
|
||||
|
||||
test('wires --model from review.models.agy (#2073 mode 2; agy gained --model in ~1.0.3)', () => {
|
||||
assert.match(read(), /--model "\$AGY_MODEL"/,
|
||||
'review.md must pass --model "$AGY_MODEL" when review.models.agy is set');
|
||||
test('wires --model from review.models.agy (#2073 mode 2)', () => {
|
||||
assert.equal(lane.modelConfigKey, 'review.models.agy');
|
||||
const argv = planFor().argv;
|
||||
assert.equal(argv[argv.indexOf('--model') + 1], 'agy-m');
|
||||
});
|
||||
|
||||
test('discards partial output on non-zero exit so the fallback fires (#687)', () => {
|
||||
const c = read();
|
||||
assert.match(c, /_AGY_RC.*-ne 0/, 'review.md must check the agy exit code');
|
||||
assert.match(c, /: > \{run_dir\}\/gsd-review-antigravity\.md/,
|
||||
'review.md must truncate the output file when agy timed out / failed');
|
||||
});
|
||||
|
||||
test('no unguarded bare "agy -p" invocation remains at line start', () => {
|
||||
// A bare `agy -p …` with no cap was the original #687 hang.
|
||||
assert.doesNotMatch(read(), /^agy -p/m,
|
||||
'review.md must not invoke a bare `agy -p` unbounded at line start');
|
||||
test('discards partial output on non-zero exit so the fallback fires (#687)', async () => {
|
||||
const p = planFor();
|
||||
const files = {};
|
||||
const d = {
|
||||
files,
|
||||
spawn: () => ({ status: 124, stdout: 'PARTIAL GARBAGE', stderr: '' }),
|
||||
httpJson: async () => ({ ok: true, status: 200, body: '{}' }),
|
||||
readFile: (x) => { if (!(x in files)) throw new Error('ENOENT'); return files[x]; },
|
||||
writeFile: (x, c) => { files[x] = c; },
|
||||
exists: (x) => x in files,
|
||||
hasBinary: () => true,
|
||||
configGet: () => undefined,
|
||||
homeDir: '/home/u',
|
||||
warn: () => {},
|
||||
};
|
||||
await runLane(p, d, { repoRoot: '/repo' });
|
||||
assert.ok(!files[p.reviewPath].includes('PARTIAL GARBAGE'),
|
||||
'a non-zero exit must discard stdout so the transcript fallback and diagnostic can take over');
|
||||
assert.ok(files[p.reviewPath].includes('failed or returned empty output'));
|
||||
});
|
||||
});
|
||||
});
|
||||
@@ -260,94 +262,29 @@ const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
describe('enh-773: automated codex exec invocations include --ephemeral (hook-trust bypass dropped by #2479)', () => {
|
||||
const workflow = fs.readFileSync(
|
||||
path.join(process.cwd(), 'gsd-core', 'workflows', 'review.md'),
|
||||
'utf8'
|
||||
);
|
||||
// Declared in the lane now rather than grepped out of review.md's bash lines.
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const codex = REVIEWER_LANES.find((l) => l.slug === 'codex');
|
||||
const args = codex.invoke.args.join(' ');
|
||||
|
||||
// Extract codex exec INVOCATION lines from code fences. A `codex exec --help`
|
||||
// capability probe would not be an automation invocation, so keep it excluded
|
||||
// from the per-invocation flag assertions below (#2479 asserts no such probe
|
||||
// exists at all).
|
||||
const codexExecLines = workflow
|
||||
.split(/\r?\n/)
|
||||
.filter((line) => line.includes('codex exec') && !line.includes('codex exec --help'));
|
||||
|
||||
test('review.md contains at least one codex exec invocation', () => {
|
||||
assert.ok(
|
||||
codexExecLines.length > 0,
|
||||
'review.md must contain at least one codex exec invocation'
|
||||
);
|
||||
test('the codex lane uses the exec subcommand', () => {
|
||||
assert.ok(args.includes('exec'));
|
||||
assert.equal(codex.invoke.args[0], 'exec', 'exec is a SUBCOMMAND and must come first');
|
||||
});
|
||||
|
||||
test('every codex exec invocation includes --ephemeral', () => {
|
||||
for (const line of codexExecLines) {
|
||||
assert.ok(
|
||||
line.includes('--ephemeral'),
|
||||
`codex exec invocation is missing --ephemeral:\n ${line.trim()}`
|
||||
);
|
||||
}
|
||||
test('codex exec is --ephemeral', () => {
|
||||
assert.ok(args.includes('--ephemeral'));
|
||||
});
|
||||
|
||||
test('#2479: no hook-trust bypass — flag and capability probe are both absent', () => {
|
||||
// The flag only bypasses *persisted* hook trust (a first-run condition) and
|
||||
// flagless invocations work in steady state, while host-harness safety
|
||||
// classifiers deny commands carrying it — and cited the probe itself as
|
||||
// intent. Inverts the former #1115 gating contract: instead of probe + gated
|
||||
// $CODEX_BYPASS_FLAG, the workflow must be free of the literal flag string
|
||||
// ANYWHERE in the file — not just on invocation lines — so a line
|
||||
// continuation, a renamed carrier variable, or a probe grepping for it are
|
||||
// all caught by the same assertion. (The #2479 prose note deliberately
|
||||
// describes the flag without spelling it out, keeping the file-wide ban
|
||||
// clean.)
|
||||
assert.ok(
|
||||
!workflow.includes('--dangerously-bypass-hook-trust'),
|
||||
'review.md must not contain --dangerously-bypass-hook-trust anywhere (#2479 — file-wide ban covers invocations, continuations, carrier variables, and probes)'
|
||||
);
|
||||
assert.ok(
|
||||
!workflow.includes('CODEX_BYPASS_FLAG'),
|
||||
'review.md must not interpolate a $CODEX_BYPASS_FLAG variable (#2479)'
|
||||
);
|
||||
assert.ok(
|
||||
!/codex[^\r\n]*--help[^\r\n]*grep/.test(workflow),
|
||||
'review.md must not capability-probe codex flags via `codex --help | grep` (#2479 — the probe itself feeds harness safety classifiers)'
|
||||
);
|
||||
test('no hook-trust bypass flag is emitted (#2479)', () => {
|
||||
assert.ok(!/dangerously|bypass/i.test(args),
|
||||
'host-harness safety classifiers deny commands carrying it; a genuine untrusted-hook prompt '
|
||||
+ 'surfaces through the .err capture and the empty-output stub instead');
|
||||
});
|
||||
|
||||
test('#1115: codex review failures are surfaced, not silently swallowed', () => {
|
||||
// stderr must be captured (not discarded to /dev/null) and an empty output
|
||||
// must be replaced with a diagnostic, so a broken reviewer is reported.
|
||||
for (const line of codexExecLines) {
|
||||
assert.ok(
|
||||
!line.includes('2>/dev/null'),
|
||||
`codex exec must not discard stderr to /dev/null:\n ${line.trim()}`
|
||||
);
|
||||
}
|
||||
assert.ok(
|
||||
/\[ ! -s \{run_dir\}\/gsd-review-codex\.md \]/.test(workflow),
|
||||
'review.md must guard against an empty codex review output and surface the failure'
|
||||
);
|
||||
});
|
||||
|
||||
test('--ephemeral appears before the prompt argument (flag ordering)', () => {
|
||||
for (const line of codexExecLines) {
|
||||
const ephemeralPos = line.indexOf('--ephemeral');
|
||||
const promptPos = line.indexOf(' - ');
|
||||
if (promptPos === -1) continue; // no stdin prompt arg on this line
|
||||
assert.ok(
|
||||
ephemeralPos < promptPos,
|
||||
`--ephemeral must appear before the stdin prompt argument:\n ${line.trim()}`
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
test('--skip-git-repo-check is preserved alongside automation flags', () => {
|
||||
for (const line of codexExecLines) {
|
||||
assert.ok(
|
||||
line.includes('--skip-git-repo-check'),
|
||||
`codex exec invocation lost --skip-git-repo-check:\n ${line.trim()}`
|
||||
);
|
||||
}
|
||||
assert.equal(codex.emptyOutput, 'stub-with-stderr',
|
||||
'a failed codex run must produce a diagnosable stub carrying its stderr');
|
||||
});
|
||||
});
|
||||
|
||||
@@ -427,48 +364,52 @@ function openCodeBlock() {
|
||||
}
|
||||
|
||||
describe('bug #1936: OpenCode reviewer must not silently yield an empty review', () => {
|
||||
// The agent can end its turn with ZERO output tokens, and `--format default` then drops the
|
||||
// assistant text entirely. Phase 5b (#2799) made the reconstruction a named first-party handler
|
||||
// rather than a jq pipeline in bash, so these assertions are behavioural over that handler.
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const { handleOpencodeOutput } = require('../gsd-core/bin/lib/review-lane-runner.cjs');
|
||||
const NL = String.fromCharCode(10);
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === 'opencode');
|
||||
|
||||
test('captures opencode stderr to a sidecar, never /dev/null', () => {
|
||||
const block = openCodeBlock();
|
||||
assert.match(block, /opencode run [^\n]*2>\{run_dir\}\/gsd-review-opencode\.err/,
|
||||
'the opencode invocation must send stderr to a .err sidecar so failures are diagnosable');
|
||||
assert.doesNotMatch(block, /opencode run [^\n]*2>\/dev\/null/,
|
||||
'the opencode invocation must not discard stderr to /dev/null (#1936)');
|
||||
const r = resolveLanePlan({ lane, configGet: () => undefined, runDir: '/run', repoRoot: '/repo' });
|
||||
assert.equal(r.ok, true);
|
||||
assert.ok(r.plan.errPath.endsWith('.err'));
|
||||
assert.notEqual(r.plan.errPath, '/dev/null');
|
||||
});
|
||||
|
||||
test('requests structured JSON output and reconstructs review from assistant text parts', () => {
|
||||
const block = openCodeBlock();
|
||||
assert.match(block, /opencode run [^\n]*--format json/,
|
||||
'must invoke opencode with --format json so assistant text parts are recoverable');
|
||||
assert.match(block, /select\(\.type=="text"\)\s*\|\s*\.part\.text/,
|
||||
'must extract the assistant text parts via `.part.text` from the JSON event stream');
|
||||
test('requests structured JSON output and reconstructs from assistant text parts', () => {
|
||||
assert.ok(lane.invoke.args.join(' ').includes('--format json'),
|
||||
'--format json is the PRIMARY invocation, not a fallback');
|
||||
assert.equal(lane.handler, 'opencode', 'reconstruction cannot be expressed as data');
|
||||
const stream = [
|
||||
JSON.stringify({ type: 'text', part: { text: 'part one' } }),
|
||||
JSON.stringify({ type: 'text', part: { text: 'part two' } }),
|
||||
].join(NL);
|
||||
assert.equal(handleOpencodeOutput(stream).review, ['part one', 'part two'].join(NL));
|
||||
});
|
||||
|
||||
test('gates the empty-review stub on extracted CONTENT, not output-file size', () => {
|
||||
// An empty jq extraction still writes a trailing newline, so a `[ -s file ]`
|
||||
// check would treat a content-less review as populated and skip the stub. The
|
||||
// block must test the captured text variable instead.
|
||||
const block = openCodeBlock();
|
||||
assert.match(block, /OPENCODE_REVIEW=\$\(jq/, 'must capture the extraction into a variable');
|
||||
assert.match(block, /\[ -n "\$OPENCODE_REVIEW" \]/,
|
||||
'must branch on the content of $OPENCODE_REVIEW, not on the size of the .md file');
|
||||
assert.doesNotMatch(block, /\[ ! -s \{run_dir\}\/gsd-review-opencode\.md \]/,
|
||||
'must not gate the stub on `[ ! -s ...opencode...md ]` (a lone newline defeats it)');
|
||||
// A JSON stream with no assistant text is many BYTES but zero review.
|
||||
const noText = JSON.stringify({ type: 'step_finish', part: { reason: 'stop', tokens: { output: 0 } } });
|
||||
assert.ok(noText.length > 0, 'the raw stream is non-empty by byte count');
|
||||
assert.equal(handleOpencodeOutput(noText).review, '', 'yet the extracted review is empty');
|
||||
});
|
||||
|
||||
test('empty-output stub is diagnosable: references #1936, stop reason/tokens, and stderr', () => {
|
||||
const block = openCodeBlock();
|
||||
assert.match(block, /#1936/, 'the empty-output stub must reference the issue');
|
||||
assert.match(block, /step_finish[\s\S]*\.part\.reason[\s\S]*\.part\.tokens\.output/,
|
||||
'the stub must surface the stop reason and output-token count from the final step_finish');
|
||||
assert.match(block, /cat \{run_dir\}\/gsd-review-opencode\.err/,
|
||||
'the stub must append the captured stderr');
|
||||
test('empty-output stub is diagnosable: stop reason and output tokens', () => {
|
||||
const stream = JSON.stringify({ type: 'step_finish', part: { reason: 'stop', tokens: { output: 0 } } });
|
||||
const out = handleOpencodeOutput(stream);
|
||||
assert.ok(out.diagnostic.includes('stop'), 'the stop reason must be surfaced');
|
||||
assert.ok(out.diagnostic.includes('0'), 'the output-token count must be surfaced');
|
||||
});
|
||||
|
||||
test('does not regress the Codex reviewer block (still captures stderr to .err)', () => {
|
||||
// #1936 changes only the OpenCode block; the Codex block's existing
|
||||
// stderr-to-sidecar contract must remain intact.
|
||||
assert.match(read(), /codex exec [^\n]*2>\{run_dir\}\/gsd-review-codex\.err/,
|
||||
'the Codex reviewer block must be left unchanged');
|
||||
test('does not regress the Codex reviewer (still captures stderr to .err)', () => {
|
||||
const codex = REVIEWER_LANES.find((l) => l.slug === 'codex');
|
||||
const r = resolveLanePlan({ lane: codex, configGet: () => undefined, runDir: '/run', repoRoot: '/repo' });
|
||||
assert.ok(r.plan.errPath.endsWith('.err'));
|
||||
assert.equal(r.plan.outputTarget.kind, 'file', 'codex still captures via --output-last-message (#1698)');
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
@@ -35,6 +35,17 @@ const {
|
||||
const {
|
||||
KNOWN_REVIEWER_SLUGS,
|
||||
} = require('../gsd-core/bin/lib/review-reviewer-selection.cjs');
|
||||
const CAPABILITY_REGISTRY = require('../gsd-core/bin/lib/capability-registry.cjs');
|
||||
|
||||
/**
|
||||
* Reviewer slugs the GENERATED registry declares — the surface Phase 5b (#2799) re-pointed parity
|
||||
* onto. Once `invoke_reviewers` iterates lanes, the registry (not the workflow text) is what
|
||||
* decides which lanes exist at runtime.
|
||||
*/
|
||||
const REGISTRY_LANE_SLUGS = Object.values(CAPABILITY_REGISTRY.capabilities || {})
|
||||
.map((c) => c && c.reviewer && c.reviewer.slug)
|
||||
.filter((s) => typeof s === 'string' && s)
|
||||
.sort();
|
||||
|
||||
const ROOT = path.join(__dirname, '..');
|
||||
// Normalized to LF on read so the CRLF cases below can construct a Windows
|
||||
@@ -53,6 +64,7 @@ function check(overrides = {}) {
|
||||
return checkReviewerLaneParity({
|
||||
descriptor: REVIEWER_LANES,
|
||||
roster: KNOWN_REVIEWER_SLUGS,
|
||||
registry: REGISTRY_LANE_SLUGS,
|
||||
workflowText: WORKFLOW_TEXT,
|
||||
...overrides,
|
||||
});
|
||||
@@ -110,161 +122,113 @@ describe('reviewer lane parity — descriptor vs roster', () => {
|
||||
});
|
||||
});
|
||||
|
||||
describe('reviewer lane parity — descriptor vs invoke_reviewers legs', () => {
|
||||
test('a leg added without a descriptor entry is a violation', () => {
|
||||
// The #2718 shape: a new lane's bash block lands in the workflow and nothing
|
||||
// else moves. This is the row a forward-only assertion cannot catch.
|
||||
const r = check({
|
||||
workflowText: WORKFLOW_TEXT.replace(
|
||||
'<!-- reviewer-lane: qwen -->',
|
||||
'<!-- reviewer-lane: qwen -->\n<!-- reviewer-lane: kimi_code -->',
|
||||
),
|
||||
});
|
||||
assert.deepStrictEqual(reasons(r), [
|
||||
`${PARITY_VIOLATION.LEG_MARKER_UNDECLARED}:kimi_code`,
|
||||
]);
|
||||
});
|
||||
|
||||
test('a declared lane whose workflow leg was removed is a violation', () => {
|
||||
const r = check({
|
||||
workflowText: WORKFLOW_TEXT.replace('<!-- reviewer-lane: qwen -->', ''),
|
||||
});
|
||||
assert.deepStrictEqual(reasons(r), [
|
||||
`${PARITY_VIOLATION.LEG_MARKER_MISSING}:qwen`,
|
||||
]);
|
||||
});
|
||||
|
||||
test('a duplicated leg marker is a violation', () => {
|
||||
const r = check({
|
||||
workflowText: WORKFLOW_TEXT.replace(
|
||||
'<!-- reviewer-lane: qwen -->',
|
||||
'<!-- reviewer-lane: qwen -->\n<!-- reviewer-lane: qwen -->',
|
||||
),
|
||||
});
|
||||
assert.deepStrictEqual(reasons(r), [
|
||||
`${PARITY_VIOLATION.LEG_MARKER_DUPLICATED}:qwen`,
|
||||
]);
|
||||
});
|
||||
|
||||
test('marker matching tolerates whitespace variation', () => {
|
||||
const r = check({
|
||||
workflowText: WORKFLOW_TEXT.replace(
|
||||
'<!-- reviewer-lane: qwen -->',
|
||||
'<!--reviewer-lane:qwen-->',
|
||||
),
|
||||
});
|
||||
assert.deepStrictEqual(r.violations, []);
|
||||
});
|
||||
|
||||
test('a marker outside the invoke_reviewers step does not satisfy the leg', () => {
|
||||
// Scoped, not file-wide: a marker parked in write_reviews must not be
|
||||
// mistaken for a dispatch leg.
|
||||
const moved = WORKFLOW_TEXT
|
||||
.replace('<!-- reviewer-lane: qwen -->', '')
|
||||
.replace('## Qwen Review', '<!-- reviewer-lane: qwen -->\n## Qwen Review');
|
||||
const r = check({ workflowText: moved });
|
||||
assert.deepStrictEqual(reasons(r), [
|
||||
`${PARITY_VIOLATION.LEG_MARKER_MISSING}:qwen`,
|
||||
]);
|
||||
});
|
||||
});
|
||||
|
||||
describe('reviewer lane parity — descriptor vs write_reviews sections', () => {
|
||||
test('an output section with no declared lane is a violation', () => {
|
||||
const r = check({
|
||||
workflowText: WORKFLOW_TEXT.replace(
|
||||
'## Qwen Review',
|
||||
'## Qwen Review\n\n{qwen}\n\n---\n\n## Kimi Review',
|
||||
),
|
||||
});
|
||||
assert.deepStrictEqual(reasons(r), [
|
||||
`${PARITY_VIOLATION.SECTION_UNDECLARED}:Kimi`,
|
||||
]);
|
||||
});
|
||||
|
||||
test('a declared lane with no output section is a violation', () => {
|
||||
const r = check({
|
||||
workflowText: WORKFLOW_TEXT.replace('## Qwen Review', '## Renamed Heading'),
|
||||
});
|
||||
describe('reviewer lane parity — descriptor vs registry (ADR-2782 Phase 5b)', () => {
|
||||
// Phase 5b deleted the per-lane workflow text this file used to scan, so the leg-marker and
|
||||
// section-heading families are gone. What replaced them is the parity that is load-bearing once
|
||||
// lanes are data: the registry is what the runtime actually iterates.
|
||||
test('a registry lane with no descriptor entry is a violation', () => {
|
||||
const r = check({ registry: [...REGISTRY_LANE_SLUGS, 'acme'] });
|
||||
assert.ok(
|
||||
reasons(r).includes(`${PARITY_VIOLATION.SECTION_MISSING}:Qwen`),
|
||||
`expected a section-missing violation, got: ${JSON.stringify(reasons(r))}`,
|
||||
reasons(r).includes(`${PARITY_VIOLATION.REGISTRY_LANE_UNDECLARED}:acme`),
|
||||
'a lane the registry ships but the descriptor never declared must fail',
|
||||
);
|
||||
});
|
||||
|
||||
test('a duplicated output section is a violation', () => {
|
||||
// Two lanes under one heading would silently MERGE in REVIEWS.md, producing
|
||||
// a review that appears to have consensus it does not have (ADR-2782 D8).
|
||||
test('a descriptor lane absent from the registry is a violation', () => {
|
||||
const r = check({
|
||||
workflowText: WORKFLOW_TEXT.replace(
|
||||
'## Qwen Review',
|
||||
'## Qwen Review\n\n{a}\n\n---\n\n## Qwen Review',
|
||||
),
|
||||
descriptor: [...REVIEWER_LANES, fakeLane('acme')],
|
||||
roster: [...KNOWN_REVIEWER_SLUGS, 'acme'],
|
||||
});
|
||||
assert.deepStrictEqual(reasons(r), [
|
||||
`${PARITY_VIOLATION.SECTION_DUPLICATED}:Qwen`,
|
||||
]);
|
||||
assert.ok(
|
||||
reasons(r).includes(`${PARITY_VIOLATION.DESCRIPTOR_LANE_NOT_IN_REGISTRY}:acme`),
|
||||
'a declared lane no capability manifest ships must fail',
|
||||
);
|
||||
});
|
||||
|
||||
test('an empty registry reports every lane rather than passing silently', () => {
|
||||
// Degrading to violations is the point: a checker that cannot tell "no registry" from
|
||||
// "registry agrees" is worse than no checker.
|
||||
const r = check({ registry: [] });
|
||||
const missing = r.violations.filter(
|
||||
(v) => v.reason === PARITY_VIOLATION.DESCRIPTOR_LANE_NOT_IN_REGISTRY,
|
||||
);
|
||||
assert.equal(missing.length, REVIEWER_LANES.length);
|
||||
});
|
||||
|
||||
test('a non-array registry degrades to violations, never throws', () => {
|
||||
for (const bad of [null, undefined, 42, 'gemini', {}]) {
|
||||
const r = check({ registry: bad });
|
||||
assert.equal(r.ok, false);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('reviewer lane parity — anti-parity: no bespoke leg may return', () => {
|
||||
test('a re-added per-CLI leg marker is a violation', () => {
|
||||
// The regex flipped polarity in Phase 5b: matching a leg marker is now the failure. Without
|
||||
// this, nothing stops a contributor quietly re-adding a hand-authored block — which is the
|
||||
// drift (#2718 -> #2781) the epic exists to end, and every other check here would still pass.
|
||||
const wf = WORKFLOW_TEXT.replace(
|
||||
'<step name="invoke_reviewers">',
|
||||
'<step name="invoke_reviewers">\n<!-- reviewer-lane: gemini -->',
|
||||
);
|
||||
const r = check({ workflowText: wf });
|
||||
assert.ok(reasons(r).includes(`${PARITY_VIOLATION.BESPOKE_LEG_PRESENT}:gemini`));
|
||||
});
|
||||
|
||||
test('the shipped workflow contains no leg markers', () => {
|
||||
const r = check();
|
||||
assert.deepStrictEqual(
|
||||
r.violations.filter((v) => v.reason === PARITY_VIOLATION.BESPOKE_LEG_PRESENT),
|
||||
[],
|
||||
);
|
||||
});
|
||||
|
||||
test('a marker outside the invoke_reviewers step is not a bespoke leg', () => {
|
||||
const r = check({ workflowText: `${WORKFLOW_TEXT}\n<!-- reviewer-lane: gemini -->\n` });
|
||||
assert.deepStrictEqual(
|
||||
r.violations.filter((v) => v.reason === PARITY_VIOLATION.BESPOKE_LEG_PRESENT),
|
||||
[],
|
||||
'the anti-parity check is scoped to invoke_reviewers, as the old leg check was',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
describe('reviewer lane parity — not-corruption (must NOT fire)', () => {
|
||||
test('ADR-1517 instance sections are exempt from lane parity', () => {
|
||||
// `## OpenCode Review (opencode-deepseek)` and `(opencode-mimo)` are already
|
||||
// in the shipped file. ADR-2782 D8: reviewer instances are not lanes. A
|
||||
// naive `## … Review` matcher fails against these on day one.
|
||||
const r = check();
|
||||
assert.deepStrictEqual(r.violations, []);
|
||||
|
||||
const withNewInstance = WORKFLOW_TEXT.replace(
|
||||
'## Qwen Review',
|
||||
'## Qwen Review (qwen-turbo)\n\n{x}\n\n---\n\n## Qwen Review',
|
||||
);
|
||||
assert.deepStrictEqual(
|
||||
checkReviewerLaneParity({
|
||||
descriptor: REVIEWER_LANES,
|
||||
roster: KNOWN_REVIEWER_SLUGS,
|
||||
workflowText: withNewInstance,
|
||||
}).violations,
|
||||
[],
|
||||
'adding a reviewer instance section must not trip lane parity',
|
||||
);
|
||||
// Phase 5b deleted the section-heading and leg-marker matchers these cases were written against.
|
||||
// The invariant they protected still matters and is asserted here against the surfaces that
|
||||
// replaced them: nothing in REVIEWS.md prose may promote itself into the lane roster.
|
||||
test('the shipped repo is clean', () => {
|
||||
assert.deepStrictEqual(check().violations, []);
|
||||
});
|
||||
|
||||
test('the h1 title and non-lane headings are not read as lane sections', () => {
|
||||
// `# Cross-AI Plan Review — Phase {N}` contains "Review" but is h1;
|
||||
// `## Consensus Summary` is h2 but has no ` Review` suffix.
|
||||
const r = check();
|
||||
assert.deepStrictEqual(r.violations, []);
|
||||
test('an ADR-1517 instance heading never becomes a lane', () => {
|
||||
// `## OpenCode Review (opencode-deepseek)` is an INSTANCE resolving through a lane, not a lane
|
||||
// (ADR-2782 D8). Instances take no part in the roster, the flag set, or uniqueness.
|
||||
const withNewInstance = WORKFLOW_TEXT.replace(
|
||||
'## Consensus Summary',
|
||||
'## Qwen Review (qwen-turbo)\n\n{x}\n\n---\n\n## Consensus Summary',
|
||||
);
|
||||
assert.deepStrictEqual(check({ workflowText: withNewInstance }).violations, []);
|
||||
});
|
||||
|
||||
test('extra non-lane headings are inert', () => {
|
||||
const withExtras = WORKFLOW_TEXT.replace(
|
||||
'## Consensus Summary',
|
||||
'## Another Summary\n\n---\n\n## Consensus Summary',
|
||||
);
|
||||
assert.deepStrictEqual(
|
||||
checkReviewerLaneParity({
|
||||
descriptor: REVIEWER_LANES,
|
||||
roster: KNOWN_REVIEWER_SLUGS,
|
||||
workflowText: withExtras,
|
||||
}).violations,
|
||||
[],
|
||||
);
|
||||
assert.deepStrictEqual(check({ workflowText: withExtras }).violations, []);
|
||||
});
|
||||
|
||||
test('bold prose in invoke_reviewers is not read as a leg', () => {
|
||||
// Five non-lane bold labels share the bold-then-fence shape a heuristic
|
||||
// matcher would key on. Adding another must not register a lane.
|
||||
// Non-lane bold labels share the bold-then-fence shape a heuristic matcher would key on.
|
||||
// Adding another must not register a lane — the anti-parity check keys on the explicit marker
|
||||
// only, never on prose shape.
|
||||
const withProse = WORKFLOW_TEXT.replace(
|
||||
'<!-- reviewer-lane: qwen -->',
|
||||
'**Some new maintainer note (#9999):**\n\n```bash\necho hi\n```\n\n<!-- reviewer-lane: qwen -->',
|
||||
);
|
||||
assert.deepStrictEqual(
|
||||
checkReviewerLaneParity({
|
||||
descriptor: REVIEWER_LANES,
|
||||
roster: KNOWN_REVIEWER_SLUGS,
|
||||
workflowText: withProse,
|
||||
}).violations,
|
||||
[],
|
||||
'<step name="invoke_reviewers">',
|
||||
'<step name="invoke_reviewers">\n**Some new maintainer note (#9999):**\n\n```bash\necho hi\n```\n',
|
||||
);
|
||||
assert.deepStrictEqual(check({ workflowText: withProse }).violations, []);
|
||||
});
|
||||
});
|
||||
|
||||
@@ -279,28 +243,49 @@ describe('reviewer lane parity — cross-platform and hostile input', () => {
|
||||
|
||||
test('a divergence is still detected under CRLF', () => {
|
||||
const crlf = asCrlf(
|
||||
WORKFLOW_TEXT.replace('<!-- reviewer-lane: qwen -->', ''),
|
||||
WORKFLOW_TEXT.replace(
|
||||
'<step name="invoke_reviewers">',
|
||||
'<step name="invoke_reviewers">\n<!-- reviewer-lane: qwen -->',
|
||||
),
|
||||
);
|
||||
assert.deepStrictEqual(reasons(check({ workflowText: crlf })), [
|
||||
`${PARITY_VIOLATION.LEG_MARKER_MISSING}:qwen`,
|
||||
`${PARITY_VIOLATION.BESPOKE_LEG_PRESENT}:qwen`,
|
||||
]);
|
||||
});
|
||||
|
||||
test('empty workflow text degrades to violations rather than throwing', () => {
|
||||
// A read failure must never be mistaken for a clean bill of health.
|
||||
const r = check({ workflowText: '' });
|
||||
assert.strictEqual(r.ok, false);
|
||||
test('empty workflow text is clean for anti-parity but an empty REGISTRY is not', () => {
|
||||
// Phase 5b split what "degrades to violations" means. An empty WORKFLOW legitimately contains
|
||||
// no bespoke leg, so the anti-parity arm is silent — the workflow no longer names lanes at all.
|
||||
// The read-failure guard moved to the registry arm, which is now the surface that decides which
|
||||
// lanes exist: an unreadable registry must never read as "registry agrees".
|
||||
const emptyWorkflow = check({ workflowText: '' });
|
||||
assert.deepStrictEqual(
|
||||
emptyWorkflow.violations.filter((v) => v.reason === PARITY_VIOLATION.BESPOKE_LEG_PRESENT),
|
||||
[],
|
||||
);
|
||||
|
||||
const emptyRegistry = check({ registry: [] });
|
||||
assert.strictEqual(emptyRegistry.ok, false);
|
||||
assert.strictEqual(
|
||||
r.violations.filter((v) => v.reason === PARITY_VIOLATION.LEG_MARKER_MISSING)
|
||||
.length,
|
||||
emptyRegistry.violations.filter(
|
||||
(v) => v.reason === PARITY_VIOLATION.DESCRIPTOR_LANE_NOT_IN_REGISTRY,
|
||||
).length,
|
||||
REVIEWER_LANES.length,
|
||||
);
|
||||
});
|
||||
|
||||
test('non-string workflow text is coerced, never thrown on', () => {
|
||||
for (const bad of [undefined, null]) {
|
||||
// Totality is the invariant. Since Phase 5b the workflow no longer names lanes, so an absent
|
||||
// one legitimately yields NO anti-parity violation — the read-failure guard moved to the
|
||||
// registry arm, which is covered in the registry describe above.
|
||||
for (const bad of [undefined, null, 42, {}, []]) {
|
||||
const r = check({ workflowText: bad });
|
||||
assert.strictEqual(r.ok, false, `expected violations for ${String(bad)}`);
|
||||
assert.equal(typeof r.ok, 'boolean');
|
||||
assert.ok(Array.isArray(r.violations));
|
||||
assert.deepStrictEqual(
|
||||
r.violations.filter((v) => v.reason === PARITY_VIOLATION.BESPOKE_LEG_PRESENT),
|
||||
[],
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
@@ -361,7 +346,7 @@ describe('reviewer lane parity — properties', () => {
|
||||
const slugArb = fc.stringMatching(/^[a-z][a-z0-9_-]{0,12}$/);
|
||||
|
||||
/**
|
||||
* Slugs OUTSIDE the grammar. These are unmatchable by LEG_MARKER_RE, so the
|
||||
* Slugs OUTSIDE the declared LANE_SLUG_RE grammar, which stays enforced after Phase 5b so
|
||||
* contract is that they are reported as INVALID_SLUG rather than silently
|
||||
* reported missing. Includes regex metacharacters and prototype-pollution
|
||||
* shaped keys.
|
||||
@@ -395,7 +380,9 @@ describe('reviewer lane parity — properties', () => {
|
||||
|
||||
/** Build a review.md-shaped document declaring exactly `slugs`. */
|
||||
function docFor(slugs, sections, noise, eol) {
|
||||
const legs = slugs.map((s) => `<!-- reviewer-lane: ${s} -->\n**${s}:**`).join('\n');
|
||||
// No leg markers: Phase 5b's workflow iterates lanes, and a marker is now the violation.
|
||||
// `slugs` is retained so callers keep their existing shape.
|
||||
const legs = slugs.map((s) => `**${s}:**`).join('\n');
|
||||
const heads = sections.map((s) => `## ${s} Review`).join('\n\n');
|
||||
const body = [
|
||||
'<step name="invoke_reviewers">',
|
||||
@@ -465,23 +452,19 @@ describe('reviewer lane parity — properties', () => {
|
||||
);
|
||||
});
|
||||
|
||||
test('a slug outside the marker grammar is reported, never silently missing', () => {
|
||||
// The silent-miss this prevents: LEG_MARKER_RE captures only [a-z0-9_-], so
|
||||
// a slug like `acme.reviewer` can have a present, correct marker that the
|
||||
// scan can never see — reporting LEG_MARKER_MISSING forever with no clue why.
|
||||
test('a slug outside the declared grammar is named, never silently accepted', () => {
|
||||
// The grammar is still enforced after Phase 5b, for the same reason: a lane whose slug falls
|
||||
// outside it is unmatchable downstream, and a loud named violation beats a silent miss.
|
||||
fc.assert(
|
||||
fc.property(badSlugArb, (bad) => {
|
||||
const lane = { ...fakeLane('placeholder'), slug: bad };
|
||||
const r = checkReviewerLaneParity({
|
||||
descriptor: [lane],
|
||||
roster: [bad],
|
||||
workflowText: `<step name="invoke_reviewers">\n<!-- reviewer-lane: ${bad} -->\n</step>`,
|
||||
registry: [bad],
|
||||
workflowText: '<step name="invoke_reviewers">\n</step>',
|
||||
});
|
||||
const reasonsOut = r.violations.map((v) => v.reason);
|
||||
return (
|
||||
reasonsOut.includes(PARITY_VIOLATION.INVALID_SLUG) &&
|
||||
!reasonsOut.includes(PARITY_VIOLATION.LEG_MARKER_MISSING)
|
||||
);
|
||||
return r.violations.map((v) => v.reason).includes(PARITY_VIOLATION.INVALID_SLUG);
|
||||
}),
|
||||
FC,
|
||||
);
|
||||
@@ -497,19 +480,21 @@ describe('reviewer lane parity — properties', () => {
|
||||
const declared = checkReviewerLaneParity({
|
||||
descriptor: [lane],
|
||||
roster: [name],
|
||||
workflowText:
|
||||
`<step name="invoke_reviewers">\n<!-- reviewer-lane: ${name} -->\n</step>\n` +
|
||||
'<step name="write_reviews">\n## Sec Review\n</step>',
|
||||
registry: [name],
|
||||
workflowText: '<step name="invoke_reviewers">\n</step>',
|
||||
});
|
||||
const absent = checkReviewerLaneParity({
|
||||
descriptor: [lane],
|
||||
roster: [name],
|
||||
registry: [],
|
||||
workflowText: '<step name="invoke_reviewers">\n</step>',
|
||||
});
|
||||
return (
|
||||
declared.ok &&
|
||||
absent.violations.some(
|
||||
(v) => v.reason === PARITY_VIOLATION.LEG_MARKER_MISSING && v.subject === name,
|
||||
(v) =>
|
||||
v.reason === PARITY_VIOLATION.DESCRIPTOR_LANE_NOT_IN_REGISTRY &&
|
||||
v.subject === name,
|
||||
)
|
||||
);
|
||||
}),
|
||||
@@ -550,6 +535,7 @@ describe('reviewer lane parity — properties', () => {
|
||||
const r = checkReviewerLaneParity({
|
||||
descriptor: lanes,
|
||||
roster: lanes.map((l) => l.slug),
|
||||
registry: lanes.map((l) => l.slug),
|
||||
workflowText: doc,
|
||||
});
|
||||
return r.ok;
|
||||
@@ -558,24 +544,18 @@ describe('reviewer lane parity — properties', () => {
|
||||
);
|
||||
});
|
||||
|
||||
test('removing one lane marker always yields exactly that lane missing', () => {
|
||||
test('dropping one lane from the registry yields exactly that lane missing', () => {
|
||||
fc.assert(
|
||||
fc.property(laneSetArb, noiseArb, fc.nat(), (lanes, noise, pick) => {
|
||||
fc.property(laneSetArb, fc.nat(), (lanes, pick) => {
|
||||
const victim = lanes[pick % lanes.length];
|
||||
const kept = lanes.filter((l) => l.slug !== victim.slug);
|
||||
const doc = docFor(
|
||||
kept.map((l) => l.slug),
|
||||
lanes.map((l) => l.reviewsSection),
|
||||
noise,
|
||||
'\n',
|
||||
);
|
||||
const r = checkReviewerLaneParity({
|
||||
descriptor: lanes,
|
||||
roster: lanes.map((l) => l.slug),
|
||||
workflowText: doc,
|
||||
registry: lanes.filter((l) => l.slug !== victim.slug).map((l) => l.slug),
|
||||
workflowText: '<step name="invoke_reviewers">\n</step>',
|
||||
});
|
||||
const missing = r.violations.filter(
|
||||
(v) => v.reason === PARITY_VIOLATION.LEG_MARKER_MISSING,
|
||||
(v) => v.reason === PARITY_VIOLATION.DESCRIPTOR_LANE_NOT_IN_REGISTRY,
|
||||
);
|
||||
return missing.length === 1 && missing[0].subject === victim.slug;
|
||||
}),
|
||||
@@ -590,6 +570,7 @@ describe('reviewer lane parity — properties', () => {
|
||||
const input = {
|
||||
descriptor: lanes,
|
||||
roster: lanes.map((l) => l.slug),
|
||||
registry: lanes.map((l) => l.slug),
|
||||
workflowText: docFor(
|
||||
lanes.map((l) => l.slug),
|
||||
lanes.map((l) => l.reviewsSection),
|
||||
@@ -664,13 +645,15 @@ describe('reviewer lane descriptor — declared shape (ADR-2782 D1/D2/D6/D7)', (
|
||||
});
|
||||
|
||||
test('handler is a closed first-party enum', () => {
|
||||
const allowed = [null, 'antigravity', 'openai-compatible'];
|
||||
const allowed = [null, 'antigravity', 'openai-compatible', 'opencode'];
|
||||
for (const lane of REVIEWER_LANES) {
|
||||
assert.ok(allowed.includes(lane.handler), `${lane.slug}: unexpected handler ${lane.handler}`);
|
||||
}
|
||||
assert.deepStrictEqual(
|
||||
REVIEWER_LANES.filter((l) => l.handler !== null).map((l) => l.slug).sort(),
|
||||
['antigravity', 'llama_cpp', 'lm_studio', 'ollama'],
|
||||
// `opencode` joined in Phase 5b: its review is reconstructed from assistant text parts of a
|
||||
// --format json stream, which data cannot express (#1936).
|
||||
['antigravity', 'llama_cpp', 'lm_studio', 'ollama', 'opencode'],
|
||||
);
|
||||
});
|
||||
|
||||
@@ -715,19 +698,16 @@ describe('reviewer lane descriptor — declared shape (ADR-2782 D1/D2/D6/D7)', (
|
||||
// Adding a reason is three coordinated changes: enum, emitting site, and
|
||||
// this assertion.
|
||||
assert.deepStrictEqual(Object.keys(PARITY_VIOLATION).sort(), [
|
||||
'BESPOKE_LEG_PRESENT',
|
||||
'DESCRIPTOR_LANE_NOT_IN_REGISTRY',
|
||||
'DESCRIPTOR_LANE_NOT_IN_ROSTER',
|
||||
'DUPLICATE_FLAG',
|
||||
'DUPLICATE_SECTION',
|
||||
'DUPLICATE_SLUG',
|
||||
'INVALID_SLUG',
|
||||
'LEG_MARKER_DUPLICATED',
|
||||
'LEG_MARKER_MISSING',
|
||||
'LEG_MARKER_UNDECLARED',
|
||||
'MALFORMED_LANE',
|
||||
'REGISTRY_LANE_UNDECLARED',
|
||||
'ROSTER_SLUG_UNDECLARED',
|
||||
'SECTION_DUPLICATED',
|
||||
'SECTION_MISSING',
|
||||
'SECTION_UNDECLARED',
|
||||
]);
|
||||
assert.ok(Object.isFrozen(PARITY_VIOLATION));
|
||||
});
|
||||
|
||||
452
tests/review-lane-invocation.test.cjs
Normal file
452
tests/review-lane-invocation.test.cjs
Normal file
@@ -0,0 +1,452 @@
|
||||
/**
|
||||
* Reviewer lane invocation — the resolver (ADR-2782 Phase 5b, #2799).
|
||||
*
|
||||
* THE GOLDEN TABLE IS THE POINT OF THIS FILE. Phase 5b deleted ~640 lines of hand-authored per-CLI
|
||||
* bash, and every one of those legs encoded a hard-won fix (#2494/#2605 empty output, #1698 Codex
|
||||
* stdout teardown noise, #1936 OpenCode zero-output turns, #2073 Antigravity's three modes, #2176
|
||||
* repo-root anchoring, #2589 no jq on stock Windows, #2794 Qwen's missing sidecar). Old and new
|
||||
* cannot literally run in parallel, so the golden table below IS the strangler-fig substitute: each
|
||||
* row is the invocation the bash leg produced, and the resolver must reproduce it exactly.
|
||||
*
|
||||
* The rows were derived FROM THE LEGS, not from the descriptor types. That direction matters — a
|
||||
* table written from the types would agree with the resolver by construction and prove nothing.
|
||||
*/
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fc = require('fast-check');
|
||||
|
||||
const {
|
||||
REVIEWER_LANES,
|
||||
} = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const {
|
||||
resolveLanePlan,
|
||||
isEmptyReview,
|
||||
normalizeHost,
|
||||
fileRefPrompt,
|
||||
LANE_UNAVAILABLE,
|
||||
} = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
|
||||
const RUN = '/run';
|
||||
const ROOT = '/repo';
|
||||
|
||||
/** Deterministic property runs — pinned seed, bounded, replay printed on failure. */
|
||||
const FC = { seed: 42, numRuns: 200 };
|
||||
|
||||
/** Config with every model key set, so the model-bearing rows exercise the configured branch. */
|
||||
const FULL_CONFIG = {
|
||||
'review.models.gemini': 'G',
|
||||
'review.models.claude': 'C',
|
||||
'review.models.codex': 'X',
|
||||
'review.models.opencode': 'O',
|
||||
'review.models.agy': 'A',
|
||||
'review.models.kimi-code': 'K',
|
||||
'review.models.ollama': 'M',
|
||||
'review.models.lm_studio': 'M',
|
||||
'review.models.llama_cpp': 'M',
|
||||
};
|
||||
|
||||
function resolve(slug, { config = FULL_CONFIG, effortArgs = ['--effort', 'high'] } = {}) {
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === slug);
|
||||
assert.ok(lane, `no declared lane '${slug}'`);
|
||||
return resolveLanePlan({
|
||||
lane,
|
||||
configGet: (k) => config[k],
|
||||
runDir: RUN,
|
||||
repoRoot: ROOT,
|
||||
effortArgs,
|
||||
});
|
||||
}
|
||||
|
||||
const FILE_REF = fileRefPrompt(`${RUN}/gsd-review-prompt.md`, ROOT);
|
||||
|
||||
/**
|
||||
* One row per shipped lane: the exact argv its bash leg produced, with a model configured and
|
||||
* effort available. `stdin` is the prompt path for a stdin lane, `null` otherwise.
|
||||
*/
|
||||
const GOLDEN = [
|
||||
{ slug: 'gemini', binary: 'gemini', argv: ['-m', 'G', '-p', '-'], stdin: true, out: 'stdout', timeout: 900000 },
|
||||
{ slug: 'claude', binary: 'claude', argv: ['--model', 'C', '--effort', 'high', '-p', '-'], stdin: true, out: 'stdout', timeout: 1200000 },
|
||||
{
|
||||
slug: 'codex',
|
||||
binary: 'codex',
|
||||
// `exec` is a SUBCOMMAND and must stay first; the output file lands mid-argv and the bare `-`
|
||||
// stays last. Splicing injected flags positionally produced an invalid invocation.
|
||||
argv: ['exec', '--ephemeral', '--model', 'X', '--effort', 'high', '--skip-git-repo-check', '-o', `${RUN}/gsd-review-codex.md`, '-'],
|
||||
stdin: true, out: 'file', timeout: 1200000,
|
||||
},
|
||||
{ slug: 'coderabbit', binary: 'coderabbit', argv: ['review', '--prompt-only'], stdin: false, out: 'stdout', timeout: 360000 },
|
||||
{ slug: 'opencode', binary: 'opencode', argv: ['run', '--model', 'O', '--effort', 'high', '--format', 'json', '-'], stdin: true, out: 'stdout', timeout: 660000 },
|
||||
{ slug: 'qwen', binary: 'qwen', argv: ['-'], stdin: true, out: 'stdout', timeout: 900000 },
|
||||
{ slug: 'cursor', binary: 'cursor-agent', argv: ['-p', '--mode', 'ask', '--trust', '--output-format', 'text', FILE_REF], stdin: false, out: 'stdout', timeout: 900000 },
|
||||
{ slug: 'antigravity', binary: 'agy', argv: ['--print-timeout', '540s', '--model', 'A', '-p', FILE_REF], stdin: false, out: 'stdout', timeout: 600000 },
|
||||
{ slug: 'kimi-code', binary: 'kimi', argv: ['-m', 'K', '-p', FILE_REF], stdin: false, out: 'stdout', timeout: 900000 },
|
||||
];
|
||||
|
||||
describe('reviewer lane invocation — golden plans (the strangler-fig contract)', () => {
|
||||
for (const row of GOLDEN) {
|
||||
test(`${row.slug} resolves to its shipped invocation`, () => {
|
||||
const r = resolve(row.slug);
|
||||
assert.equal(r.ok, true, `${row.slug}: ${r.ok ? '' : r.detail}`);
|
||||
const p = r.plan;
|
||||
assert.equal(p.transport, 'spawn');
|
||||
assert.equal(p.binary, row.binary);
|
||||
assert.deepStrictEqual(p.argv, row.argv);
|
||||
assert.equal(p.stdin, row.stdin ? `${RUN}/gsd-review-prompt.md` : null);
|
||||
assert.equal(p.outputTarget.kind, row.out === 'file' ? 'file' : 'stdout');
|
||||
assert.equal(p.timeoutMs, row.timeout);
|
||||
// The stderr sidecar is never /dev/null — that is what makes a failed lane diagnosable.
|
||||
assert.equal(p.errPath, `${RUN}/gsd-review-${row.slug}.err`);
|
||||
assert.equal(p.reviewPath, `${RUN}/gsd-review-${row.slug}.md`);
|
||||
});
|
||||
}
|
||||
|
||||
test('the golden table covers every spawn lane', () => {
|
||||
const spawnSlugs = REVIEWER_LANES.filter((l) => l.transport === 'spawn').map((l) => l.slug).sort();
|
||||
assert.deepStrictEqual(GOLDEN.map((g) => g.slug).sort(), spawnSlugs);
|
||||
});
|
||||
|
||||
test('the three OpenAI-compatible lanes resolve host, endpoint and discovery', () => {
|
||||
for (const [slug, host, fallback] of [
|
||||
['ollama', 'http://localhost:11434', 'llama3'],
|
||||
['lm_studio', 'http://localhost:1234', 'local-model'],
|
||||
['llama_cpp', 'http://localhost:8080', 'local-model'],
|
||||
]) {
|
||||
const r = resolve(slug, { config: {} });
|
||||
assert.equal(r.ok, true);
|
||||
assert.equal(r.plan.transport, 'openai-http');
|
||||
// Phase 4 federated every *_host with a default of "", so an unset key MUST fall back to the
|
||||
// lane's declared defaultHost or the lane would POST to a garbage URL.
|
||||
assert.equal(r.plan.host, host);
|
||||
assert.equal(r.plan.url, `${host}/v1/chat/completions`);
|
||||
assert.equal(r.plan.modelsUrl, `${host}/v1/models`);
|
||||
assert.equal(r.plan.fallbackModel, fallback);
|
||||
}
|
||||
});
|
||||
|
||||
test('every declared lane resolves — none is left unroutable', () => {
|
||||
for (const lane of REVIEWER_LANES) {
|
||||
assert.equal(resolve(lane.slug).ok, true, `${lane.slug} failed to resolve`);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('reviewer lane invocation — model resolution', () => {
|
||||
test("antigravity resolves its model from review.models.agy, not the slug", () => {
|
||||
// The regression this locks: antigravity's slug is `antigravity` but its shipped key is
|
||||
// `review.models.agy`. A `review.models.<slug>` convention misses it and silently ignores a
|
||||
// configured model — disabling the pinned-model escape hatch #2073 added for a 404ing default.
|
||||
const r = resolve('antigravity', { config: { 'review.models.agy': 'agy-2' } });
|
||||
assert.ok(r.plan.argv.includes('agy-2'), 'the configured agy model must reach argv');
|
||||
|
||||
const wrongKey = resolve('antigravity', { config: { 'review.models.antigravity': 'nope' } });
|
||||
assert.ok(!wrongKey.plan.argv.includes('nope'), 'the slug-derived key must NOT be consulted');
|
||||
});
|
||||
|
||||
test('a lane declaring no model key emits no model argument', () => {
|
||||
for (const slug of ['qwen', 'cursor', 'coderabbit']) {
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === slug);
|
||||
assert.equal(lane.modelConfigKey, null, `${slug} should declare no model key`);
|
||||
const r = resolve(slug, { config: { 'review.models.qwen': 'X', 'review.models.cursor': 'X' } });
|
||||
assert.ok(!r.plan.argv.includes('X'));
|
||||
}
|
||||
});
|
||||
|
||||
test('unset, empty, whitespace and the literal string "null" all mean unconfigured', () => {
|
||||
// `"null"` is the four literal characters `config-get --raw` prints for a missing key — every
|
||||
// bash leg tested for it. A config written by an older workflow can still contain it.
|
||||
for (const bad of [undefined, null, '', ' ', 'null', 'undefined']) {
|
||||
const r = resolve('gemini', { config: { 'review.models.gemini': bad } });
|
||||
assert.deepStrictEqual(r.plan.argv, ['-p', '-'], `${JSON.stringify(bad)} must not reach argv`);
|
||||
}
|
||||
});
|
||||
|
||||
test('a non-string model value is never coerced into argv', () => {
|
||||
// String(0) would put "0" in as a model name. A wrong model silently reviewed is worse than no
|
||||
// override at all.
|
||||
for (const bad of [0, 1, true, false, [], {}, ['a']]) {
|
||||
const r = resolve('gemini', { config: { 'review.models.gemini': bad } });
|
||||
assert.deepStrictEqual(r.plan.argv, ['-p', '-'], `${JSON.stringify(bad)} must not reach argv`);
|
||||
}
|
||||
});
|
||||
|
||||
test('shell metacharacters in a model value stay a single inert argv element', () => {
|
||||
const hostile = '; rm -rf /; $(whoami) `id` && echo "x"';
|
||||
const r = resolve('gemini', { config: { 'review.models.gemini': hostile } });
|
||||
assert.deepStrictEqual(r.plan.argv, ['-m', hostile, '-p', '-']);
|
||||
// Nothing here builds a shell string; the runner spawns with shell:false and an argv array.
|
||||
assert.equal(r.plan.argv.filter((a) => a === hostile).length, 1);
|
||||
});
|
||||
|
||||
test('effort argv only reaches lanes declaring effortChannel argv', () => {
|
||||
for (const lane of REVIEWER_LANES.filter((l) => l.transport === 'spawn')) {
|
||||
const r = resolve(lane.slug, { effortArgs: ['--effort', 'xhigh'] });
|
||||
const got = r.plan.argv.includes('--effort');
|
||||
assert.equal(got, lane.invoke.effortChannel === 'argv', `${lane.slug} effort mismatch`);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('reviewer lane invocation — argv-file-ref anchoring (#2176)', () => {
|
||||
test('the file-ref prompt names the prompt file AND the absolute repo root', () => {
|
||||
// Without the root, an argv-fed CLI does not reliably inherit the review cwd and reviews the
|
||||
// plan text in isolation — exactly what the Review Instructions forbid.
|
||||
for (const slug of ['cursor', 'antigravity', 'kimi-code']) {
|
||||
const r = resolve(slug);
|
||||
const arg = r.plan.argv[r.plan.argv.length - 1];
|
||||
assert.ok(arg.includes(`${RUN}/gsd-review-prompt.md`), `${slug} must name the prompt file`);
|
||||
assert.ok(arg.includes(ROOT), `${slug} must carry the absolute repo root`);
|
||||
}
|
||||
});
|
||||
|
||||
test('the prompt travels in argv as ONE element, never split', () => {
|
||||
const r = resolve('cursor');
|
||||
assert.equal(r.plan.argv.filter((a) => a.includes('gsd-review-prompt.md')).length, 1);
|
||||
});
|
||||
});
|
||||
|
||||
describe('reviewer lane invocation — absent-safe and hostile input (ADR-2782 D4)', () => {
|
||||
const bad = (lane) =>
|
||||
resolveLanePlan({ lane, configGet: () => undefined, runDir: RUN, repoRoot: ROOT });
|
||||
|
||||
test('a malformed lane is reported, never thrown on', () => {
|
||||
for (const v of [null, undefined, 42, 'gemini', [], true]) {
|
||||
const r = bad(v);
|
||||
assert.equal(r.ok, false);
|
||||
assert.equal(r.reason, LANE_UNAVAILABLE.MALFORMED_LANE);
|
||||
}
|
||||
});
|
||||
|
||||
test('an unknown handler fails CLOSED', () => {
|
||||
// D4 rule 4. A lane naming imperative code this version does not have cannot run "mostly" —
|
||||
// the handler is precisely the part data could not express.
|
||||
const lane = { ...REVIEWER_LANES[0], handler: 'not-a-real-handler' };
|
||||
const r = bad(lane);
|
||||
assert.equal(r.ok, false);
|
||||
assert.equal(r.reason, LANE_UNAVAILABLE.UNKNOWN_HANDLER);
|
||||
});
|
||||
|
||||
test('every handler the descriptor ships is dispatchable', () => {
|
||||
// The inverse of the above: a lane declaring a handler the runner cannot dispatch would fail
|
||||
// closed at runtime, which is a silent lane loss dressed as a safety feature.
|
||||
for (const lane of REVIEWER_LANES) {
|
||||
assert.equal(resolve(lane.slug).ok, true, `${lane.slug} handler not dispatchable`);
|
||||
}
|
||||
});
|
||||
|
||||
test('an unknown transport is reported', () => {
|
||||
const r = bad({ ...REVIEWER_LANES[0], transport: 'carrier-pigeon' });
|
||||
assert.equal(r.reason, LANE_UNAVAILABLE.UNKNOWN_TRANSPORT);
|
||||
});
|
||||
|
||||
test('a spawn lane with no binary is malformed, not a crash', () => {
|
||||
const lane = REVIEWER_LANES.find((l) => l.transport === 'spawn');
|
||||
const r = bad({ ...lane, invoke: { ...lane.invoke, binary: '' } });
|
||||
assert.equal(r.reason, LANE_UNAVAILABLE.MALFORMED_LANE);
|
||||
});
|
||||
|
||||
test('file-arg output with no outputArg is malformed', () => {
|
||||
// Knowing the review lands in a file is useless without the argument naming it.
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === 'codex');
|
||||
const r = bad({ ...lane, invoke: { ...lane.invoke, outputArg: undefined } });
|
||||
assert.equal(r.reason, LANE_UNAVAILABLE.MALFORMED_LANE);
|
||||
});
|
||||
|
||||
test('a prototype-key slug does not pollute the expansion table', () => {
|
||||
// `__proto__` is now rejected outright by the slug grammar (leading `_` is outside
|
||||
// `[a-z0-9]`), which is a stronger guarantee than tolerating it. `constructor` and `prototype`
|
||||
// ARE valid slugs, so they must resolve normally and still reach no prototype.
|
||||
const proto = bad({ ...REVIEWER_LANES[0], slug: '__proto__' });
|
||||
assert.equal(proto.ok, false);
|
||||
assert.equal(proto.reason, LANE_UNAVAILABLE.MALFORMED_LANE);
|
||||
|
||||
for (const name of ['constructor', 'prototype']) {
|
||||
const r = bad({ ...REVIEWER_LANES[0], slug: name });
|
||||
assert.equal(r.ok, true, `${name} is a grammatically valid slug`);
|
||||
assert.equal(r.plan.slug, name);
|
||||
assert.equal(r.plan.reviewPath, `${RUN}/gsd-review-${name}.md`);
|
||||
assert.equal({}.polluted, undefined);
|
||||
}
|
||||
});
|
||||
|
||||
test('an http lane resolving no host at all is malformed rather than POSTing to nowhere', () => {
|
||||
const lane = REVIEWER_LANES.find((l) => l.transport === 'openai-http');
|
||||
const r = bad({ ...lane, invoke: { ...lane.invoke, defaultHost: '' } });
|
||||
assert.equal(r.reason, LANE_UNAVAILABLE.MALFORMED_LANE);
|
||||
});
|
||||
|
||||
test('an openai-http lane with no invoke object is malformed, not a crash', () => {
|
||||
// Found by adversarial review. The spawn branch guarded with `inv?.binary`; the http branch
|
||||
// dereferenced `inv.hostConfigKey` directly and THREW, breaking this module's documented
|
||||
// totality. A throw here is worse than it looks: the CLI seam resolves every selected lane in
|
||||
// one `.map`, so one malformed overlay manifest would abort the entire review.
|
||||
for (const missing of [undefined, null, 42, 'x', []]) {
|
||||
const r = bad({ ...REVIEWER_LANES.find((l) => l.transport === 'openai-http'), invoke: missing });
|
||||
assert.equal(r.ok, false, `invoke=${JSON.stringify(missing)} must not resolve`);
|
||||
assert.equal(r.reason, LANE_UNAVAILABLE.MALFORMED_LANE);
|
||||
}
|
||||
});
|
||||
|
||||
test('a slug outside the declared grammar cannot reach an artifact path', () => {
|
||||
// Found by adversarial review. The slug is concatenated into reviewPath/errPath, so a lane
|
||||
// declaring `../../../tmp/evil` produced a path OUTSIDE the run dir that writeReviewOrStub
|
||||
// would write to. The grammar is checked upstream by the parity gate and the capability
|
||||
// validator, but neither runs on this path — and this module is the overlay-manifest trust
|
||||
// boundary, so it enforces its own precondition rather than inheriting one.
|
||||
const spawnLane = REVIEWER_LANES.find((l) => l.transport === 'spawn');
|
||||
for (const slug of ['../../../tmp/evil', 'a/b', 'a\\b', 'UPPER', '.hidden', '-lead', 'a b']) {
|
||||
const r = bad({ ...spawnLane, slug });
|
||||
assert.equal(r.ok, false, `slug ${JSON.stringify(slug)} must be rejected`);
|
||||
assert.equal(r.reason, LANE_UNAVAILABLE.MALFORMED_LANE);
|
||||
}
|
||||
// Every shipped slug must still pass — including the snake-case ones.
|
||||
for (const lane of REVIEWER_LANES) {
|
||||
assert.equal(resolve(lane.slug).ok, true, `${lane.slug} must remain valid`);
|
||||
}
|
||||
});
|
||||
|
||||
test('the unavailability reason enum is locked', () => {
|
||||
// Adding a reason is three coordinated changes: enum, emitting site, and this assertion.
|
||||
assert.deepStrictEqual(Object.keys(LANE_UNAVAILABLE).sort(), [
|
||||
'BUDGET_TOOL_FAILED',
|
||||
'BUDGET_TOO_SMALL',
|
||||
'EGRESS_HOST_CHANGED',
|
||||
'HOST_UNREACHABLE',
|
||||
'MALFORMED_LANE',
|
||||
'MISSING_BINARY',
|
||||
'MISSING_REQUIRED_BINARY',
|
||||
'PROBE_FAILED',
|
||||
'PROBE_TIMEOUT',
|
||||
'UNKNOWN_HANDLER',
|
||||
'UNKNOWN_TRANSPORT',
|
||||
]);
|
||||
assert.ok(Object.isFrozen(LANE_UNAVAILABLE));
|
||||
});
|
||||
});
|
||||
|
||||
describe('reviewer lane invocation — empty-review classification', () => {
|
||||
test('whitespace-only output counts as empty on EVERY lane', () => {
|
||||
// `[ ! -s file ]` counted BYTES, so three spaces passed as a successful review. Two legs closed
|
||||
// this locally; five did not. Uniformity here is a deliberate, disclosed behaviour change.
|
||||
for (const s of ['', ' ', ' ', '\n', '\r\n', '\t', ' \n\t ']) {
|
||||
assert.equal(isEmptyReview(s), true, `${JSON.stringify(s)} must count as empty`);
|
||||
}
|
||||
});
|
||||
|
||||
test('a single non-space character is a review (limit+1)', () => {
|
||||
assert.equal(isEmptyReview('x'), false);
|
||||
assert.equal(isEmptyReview(' x '), false);
|
||||
});
|
||||
|
||||
test('output that is exactly -n / -e / -E is a review, not a swallowed value', () => {
|
||||
// `echo "$VAR"` would write 0 bytes for these and misclassify a real reply. Nothing in this
|
||||
// path goes through echo, so the hazard is structurally impossible — locked here anyway.
|
||||
for (const s of ['-n', '-e', '-E']) assert.equal(isEmptyReview(s), false);
|
||||
});
|
||||
|
||||
test('a non-string is empty, never thrown on', () => {
|
||||
for (const v of [undefined, null, 0, {}, []]) assert.equal(isEmptyReview(v), true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('reviewer lane invocation — host normalization (D5 comparison input)', () => {
|
||||
test('cosmetic differences are NOT destination changes', () => {
|
||||
// A warning that fires on a trailing slash is a warning users learn to dismiss — which would
|
||||
// defeat the one prompt that actually matters.
|
||||
const same = [
|
||||
['http://localhost:8080', 'http://localhost:8080/'],
|
||||
['http://localhost:8080', 'http://LOCALHOST:8080'],
|
||||
['http://a.com:80', 'http://a.com'],
|
||||
['https://a.com:443', 'https://a.com'],
|
||||
['http://a.com/v1/', 'http://a.com/v1'],
|
||||
];
|
||||
for (const [a, b] of same) {
|
||||
assert.equal(normalizeHost(a), normalizeHost(b), `${a} vs ${b}`);
|
||||
}
|
||||
});
|
||||
|
||||
test('a real destination change survives normalization', () => {
|
||||
assert.notEqual(normalizeHost('http://localhost:8080'), normalizeHost('http://evil.example'));
|
||||
assert.notEqual(normalizeHost('http://a.com:8080'), normalizeHost('http://a.com:9090'));
|
||||
assert.notEqual(normalizeHost('http://a.com'), normalizeHost('https://a.com'));
|
||||
});
|
||||
|
||||
test('an unparseable host is compared verbatim, never silently rewritten', () => {
|
||||
assert.equal(normalizeHost('not a url'), 'not a url');
|
||||
assert.equal(normalizeHost(''), '');
|
||||
});
|
||||
|
||||
test('a scheme-less host is not rewritten into a fake URL', () => {
|
||||
// Found by adversarial review. `new URL('localhost:11434')` PARSES — protocol `localhost:`,
|
||||
// empty hostname — so a plausible but scheme-less config value was being rewritten to
|
||||
// `localhost://11434` and then both compared and requested as if it were a real destination.
|
||||
// No hostname means it is not a URL; return it verbatim so it fails visibly.
|
||||
assert.equal(normalizeHost('localhost:11434'), 'localhost:11434');
|
||||
assert.equal(normalizeHost('example.com:8080'), 'example.com:8080');
|
||||
// A real URL still normalizes.
|
||||
assert.equal(normalizeHost('http://LocalHost:8080/'), 'http://localhost:8080');
|
||||
});
|
||||
});
|
||||
|
||||
describe('reviewer lane invocation — properties', () => {
|
||||
test('the resolver is total over arbitrary lane input', () => {
|
||||
// Third-party overlay manifests reach this function. A resolver that throws on bad input cannot
|
||||
// report on it, and a gate that crashes is indistinguishable from one never run.
|
||||
fc.assert(
|
||||
fc.property(fc.anything(), fc.anything(), (lane, cfgValue) => {
|
||||
let r;
|
||||
try {
|
||||
r = resolveLanePlan({
|
||||
lane,
|
||||
configGet: () => cfgValue,
|
||||
runDir: RUN,
|
||||
repoRoot: ROOT,
|
||||
});
|
||||
} catch {
|
||||
return false;
|
||||
}
|
||||
return typeof r === 'object' && r !== null && typeof r.ok === 'boolean';
|
||||
}),
|
||||
FC,
|
||||
);
|
||||
});
|
||||
|
||||
test('a resolved argv never contains a placeholder token', () => {
|
||||
// An unexpanded `{{model}}` reaching a real CLI is an argument that means nothing to it.
|
||||
fc.assert(
|
||||
fc.property(
|
||||
fc.constantFrom(...REVIEWER_LANES.filter((l) => l.transport === 'spawn').map((l) => l.slug)),
|
||||
fc.option(fc.string(), { nil: undefined }),
|
||||
fc.array(fc.string(), { maxLength: 4 }),
|
||||
(slug, model, effortArgs) => {
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === slug);
|
||||
const r = resolveLanePlan({
|
||||
lane,
|
||||
configGet: () => model,
|
||||
runDir: RUN,
|
||||
repoRoot: ROOT,
|
||||
effortArgs,
|
||||
});
|
||||
if (!r.ok) return false;
|
||||
return !r.plan.argv.some((a) => /^\{\{(model|effort|output|prompt)\}\}$/.test(a));
|
||||
},
|
||||
),
|
||||
FC,
|
||||
);
|
||||
});
|
||||
|
||||
test('resolution is deterministic for identical input', () => {
|
||||
fc.assert(
|
||||
fc.property(
|
||||
fc.constantFrom(...REVIEWER_LANES.map((l) => l.slug)),
|
||||
fc.option(fc.string(), { nil: undefined }),
|
||||
(slug, model) => {
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === slug);
|
||||
const input = { lane, configGet: () => model, runDir: RUN, repoRoot: ROOT };
|
||||
return JSON.stringify(resolveLanePlan(input)) === JSON.stringify(resolveLanePlan(input));
|
||||
},
|
||||
),
|
||||
FC,
|
||||
);
|
||||
});
|
||||
});
|
||||
529
tests/review-lane-runner.test.cjs
Normal file
529
tests/review-lane-runner.test.cjs
Normal file
@@ -0,0 +1,529 @@
|
||||
/**
|
||||
* Reviewer lane runner — execution, probes, handlers, egress (ADR-2782 Phase 5b, #2799).
|
||||
*
|
||||
* Every dependency is injected, so these are behavioural tests over the real control flow with no
|
||||
* network, no spawn and no clock. Where a filesystem failure is forced it is done by making the
|
||||
* injected `writeFile`/`readFile` throw — never by `chmod 0o000`, which root bypasses, silently
|
||||
* turning the test into a vacuous pass in root Docker/CI.
|
||||
*/
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
|
||||
const { REVIEWER_LANES } = require('../gsd-core/bin/lib/review-lane-descriptor.cjs');
|
||||
const { resolveLanePlan, LANE_UNAVAILABLE } = require('../gsd-core/bin/lib/review-lane-invocation.cjs');
|
||||
const {
|
||||
checkEgressHost,
|
||||
probeLane,
|
||||
runLane,
|
||||
writeReviewOrStub,
|
||||
handleOpencodeOutput,
|
||||
stampBlindReview,
|
||||
antigravityTranscriptFallback,
|
||||
runOpenAiCompatible,
|
||||
} = require('../gsd-core/bin/lib/review-lane-runner.cjs');
|
||||
|
||||
const RUN = '/run';
|
||||
const ROOT = '/repo';
|
||||
|
||||
function plan(slug, config = {}) {
|
||||
const lane = REVIEWER_LANES.find((l) => l.slug === slug);
|
||||
const r = resolveLanePlan({ lane, configGet: (k) => config[k], runDir: RUN, repoRoot: ROOT });
|
||||
assert.equal(r.ok, true, `${slug} failed to resolve`);
|
||||
return r.plan;
|
||||
}
|
||||
|
||||
/** An in-memory dependency set. Overrides replace individual seams per test. */
|
||||
function deps(overrides = {}) {
|
||||
const files = overrides.files || {};
|
||||
const warnings = [];
|
||||
const spawns = [];
|
||||
const base = {
|
||||
files,
|
||||
warnings,
|
||||
spawns,
|
||||
spawn: (binary, argv, opts) => {
|
||||
spawns.push({ binary, argv, opts });
|
||||
return { status: 0, stdout: '', stderr: '' };
|
||||
},
|
||||
httpJson: async () => ({ ok: true, status: 200, body: '{}' }),
|
||||
readFile: (p) => {
|
||||
if (!(p in files)) throw new Error(`ENOENT ${p}`);
|
||||
return files[p];
|
||||
},
|
||||
writeFile: (p, c) => { files[p] = c; },
|
||||
exists: (p) => p in files,
|
||||
hasBinary: () => true,
|
||||
configGet: () => undefined,
|
||||
homeDir: '/home/u',
|
||||
warn: (m) => warnings.push(m),
|
||||
};
|
||||
return Object.assign(base, overrides);
|
||||
}
|
||||
|
||||
describe('runner — egress host re-verification (ADR-2782 D5 rules 2-4)', () => {
|
||||
test('no consent record ALLOWS — first-party lanes are never consent-gated', () => {
|
||||
// Blocking on absence would break every existing local-model user on upgrade: ollama,
|
||||
// lm_studio and llama_cpp ship inside the SHA-pinned distribution and have no consent record.
|
||||
assert.equal(checkEgressHost(undefined, 'http://localhost:11434').allowed, true);
|
||||
assert.equal(checkEgressHost(null, 'http://localhost:11434').allowed, true);
|
||||
});
|
||||
|
||||
test('a record predating the field ALLOWS — absence must not force re-consent', () => {
|
||||
// D4 rule 5: an absent field must not perturb consent, or every installed capability
|
||||
// re-prompts on upgrade.
|
||||
assert.equal(checkEgressHost('', 'http://localhost:8080').allowed, true);
|
||||
});
|
||||
|
||||
test('a matching destination proceeds', () => {
|
||||
const r = checkEgressHost('http://localhost:8080', 'http://localhost:8080');
|
||||
assert.equal(r.allowed, true);
|
||||
});
|
||||
|
||||
test('a changed destination BLOCKS and names both hosts', () => {
|
||||
const r = checkEgressHost('http://localhost:8080', 'http://evil.example');
|
||||
assert.equal(r.allowed, false);
|
||||
assert.equal(r.consentedHost, 'http://localhost:8080');
|
||||
assert.equal(r.currentHost, 'http://evil.example');
|
||||
});
|
||||
|
||||
test('cosmetic host edits are not a change', () => {
|
||||
for (const [a, b] of [
|
||||
['http://localhost:8080', 'http://localhost:8080/'],
|
||||
['http://a.com:80', 'http://a.com'],
|
||||
['http://A.com', 'http://a.com'],
|
||||
]) {
|
||||
assert.equal(checkEgressHost(a, b).allowed, true, `${a} vs ${b}`);
|
||||
}
|
||||
});
|
||||
|
||||
test('a non-string consented value is treated as absent, never coerced', () => {
|
||||
for (const v of [42, {}, [], true]) {
|
||||
assert.equal(checkEgressHost(v, 'http://a.com').allowed, true);
|
||||
}
|
||||
});
|
||||
|
||||
test('a blocked lane never reaches the network and writes no review', async () => {
|
||||
const p = plan('ollama');
|
||||
const d = deps();
|
||||
const r = await runLane(p, d, { consentedHost: 'http://elsewhere.example', repoRoot: ROOT });
|
||||
assert.equal(r.ok, false);
|
||||
assert.equal(r.reason, LANE_UNAVAILABLE.EGRESS_HOST_CHANGED);
|
||||
assert.equal(d.files[p.reviewPath], undefined, 'a blocked lane must not write a review');
|
||||
assert.ok(d.warnings.some((w) => w.includes('elsewhere.example')));
|
||||
});
|
||||
|
||||
test('a spawn lane skips the host check entirely', async () => {
|
||||
const p = plan('qwen');
|
||||
const d = deps({ spawn: () => ({ status: 0, stdout: 'review', stderr: '' }) });
|
||||
// A stale host on a spawn lane must be inert, not a block.
|
||||
const r = await runLane(p, d, { consentedHost: 'http://stale.example', repoRoot: ROOT });
|
||||
assert.equal(r.ok, true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('runner — probe (ADR-2782 D7)', () => {
|
||||
test('command-exists both ways', async () => {
|
||||
const p = plan('gemini');
|
||||
assert.equal((await probeLane(p, deps({ hasBinary: () => true }))).available, true);
|
||||
const miss = await probeLane(p, deps({ hasBinary: () => false }));
|
||||
assert.equal(miss.available, false);
|
||||
assert.equal(miss.reason, LANE_UNAVAILABLE.MISSING_BINARY);
|
||||
});
|
||||
|
||||
test('command-capability accepts the right tool and REJECTS the wrong one', async () => {
|
||||
// This is the entire reason D7 ships wider than existence: `kimi` is claimed by both Kimi Code
|
||||
// CLI and the legacy Python kimi-cli, and an existence-only probe registers the wrong tool.
|
||||
const p = plan('kimi-code');
|
||||
const real = await probeLane(p, deps({
|
||||
spawn: () => ({ status: 0, stdout: 'usage: kimi --output-format json -p', stderr: '' }),
|
||||
}));
|
||||
assert.equal(real.available, true);
|
||||
|
||||
const legacy = await probeLane(p, deps({
|
||||
spawn: () => ({ status: 0, stdout: 'usage: kimi --print --work-dir DIR', stderr: '' }),
|
||||
}));
|
||||
assert.equal(legacy.available, false);
|
||||
assert.equal(legacy.reason, LANE_UNAVAILABLE.PROBE_FAILED);
|
||||
});
|
||||
|
||||
test('a capability probe that times out reports unavailable, never hangs', async () => {
|
||||
// The original probe (closed PR #2776) was an unbounded `kimi --help | grep` that ran on EVERY
|
||||
// review regardless of flags — a live instance of the named Unbounded Subprocesses defect.
|
||||
const p = plan('kimi-code');
|
||||
const r = await probeLane(p, deps({
|
||||
spawn: () => ({ status: null, stdout: '', stderr: '', errorCode: 'ETIMEDOUT' }),
|
||||
}));
|
||||
assert.equal(r.available, false);
|
||||
assert.equal(r.reason, LANE_UNAVAILABLE.PROBE_TIMEOUT);
|
||||
});
|
||||
|
||||
test('the capability probe passes the declared bound to the spawn', async () => {
|
||||
const p = plan('kimi-code');
|
||||
const d = deps({ spawn: (b, a, o) => { d.spawns.push({ b, a, o }); return { status: 0, stdout: '--output-format', stderr: '' }; } });
|
||||
await probeLane(p, d);
|
||||
const call = d.spawns[d.spawns.length - 1];
|
||||
assert.equal(typeof call.o.timeoutMs, 'number');
|
||||
assert.ok(call.o.timeoutMs > 0, 'every probe that starts a process MUST be bounded');
|
||||
});
|
||||
|
||||
test('a missing required binary is named rather than left to fail obscurely', async () => {
|
||||
const p = { ...plan('gemini'), requiresBinaries: ['jq'] };
|
||||
const r = await probeLane(p, deps({ hasBinary: (n) => n !== 'jq' }));
|
||||
assert.equal(r.available, false);
|
||||
assert.equal(r.reason, LANE_UNAVAILABLE.MISSING_REQUIRED_BINARY);
|
||||
});
|
||||
|
||||
test('no shipped lane still requires jq or curl', () => {
|
||||
// Phase 5b moved parsing to JSON.parse and HTTP to fetch. Leaving a stale requiresBinaries
|
||||
// entry would report lanes unavailable on stock Windows for a dependency they no longer use.
|
||||
for (const lane of REVIEWER_LANES) {
|
||||
for (const bin of lane.requiresBinaries) {
|
||||
assert.ok(bin !== 'jq' && bin !== 'curl', `${lane.slug} still declares ${bin}`);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
test('http-reachable reports unreachable rather than throwing', async () => {
|
||||
const p = plan('ollama');
|
||||
const r = await probeLane(p, deps({ httpJson: async () => ({ ok: false, status: 0, body: '', error: 'ECONNREFUSED' }) }));
|
||||
assert.equal(r.available, false);
|
||||
assert.equal(r.reason, LANE_UNAVAILABLE.HOST_UNREACHABLE);
|
||||
});
|
||||
});
|
||||
|
||||
describe('runner — empty-output policy (#2494 / #2605 / #2794)', () => {
|
||||
test('a real review is written verbatim', () => {
|
||||
const p = plan('gemini');
|
||||
const d = deps();
|
||||
const r = writeReviewOrStub(p, '## Findings\nreal', d);
|
||||
assert.equal(r.stubbed, false);
|
||||
assert.ok(d.files[p.reviewPath].startsWith('## Findings'));
|
||||
});
|
||||
|
||||
test('empty output writes a stub carrying the captured stderr', () => {
|
||||
const p = plan('gemini');
|
||||
const d = deps({ files: { [`${RUN}/gsd-review-gemini.err`]: 'auth failed' } });
|
||||
const r = writeReviewOrStub(p, '', d);
|
||||
assert.equal(r.stubbed, true);
|
||||
assert.ok(d.files[p.reviewPath].includes('failed or returned empty output'));
|
||||
assert.ok(d.files[p.reviewPath].includes('auth failed'));
|
||||
});
|
||||
|
||||
test('whitespace-only output is stubbed on every lane', () => {
|
||||
// Before this, `[ ! -s file ]` counted bytes so " " rendered as a clean review on five lanes.
|
||||
for (const slug of ['gemini', 'claude', 'codex', 'qwen', 'cursor']) {
|
||||
const p = plan(slug);
|
||||
const d = deps();
|
||||
assert.equal(writeReviewOrStub(p, ' \n', d).stubbed, true, `${slug} accepted whitespace`);
|
||||
}
|
||||
});
|
||||
|
||||
test('the stub is distinguishable from a real review', () => {
|
||||
// The ambiguity between "failed" and "ran cleanly with nothing to report" IS the defect.
|
||||
const p = plan('gemini');
|
||||
const d = deps();
|
||||
writeReviewOrStub(p, '', d);
|
||||
assert.ok(/failed or returned empty output/.test(d.files[p.reviewPath]));
|
||||
});
|
||||
|
||||
test('an http lane appends the raw response body', () => {
|
||||
// An OpenAI-compatible server reports errors with a 4xx/5xx and the JSON in the BODY, so
|
||||
// stderr alone is empty and the body is the only evidence. The bash piped it into jq and lost it.
|
||||
const p = plan('ollama');
|
||||
const d = deps();
|
||||
writeReviewOrStub(p, '', d, '{"error":{"message":"model not found"}}');
|
||||
assert.ok(d.files[p.reviewPath].includes('Raw response body:'));
|
||||
assert.ok(d.files[p.reviewPath].includes('model not found'));
|
||||
});
|
||||
|
||||
test('a filesystem write failure degrades rather than crashing the run', () => {
|
||||
// Injected by making the seam throw — never chmod 0o000, which root bypasses.
|
||||
const p = plan('gemini');
|
||||
const d = deps({ writeFile: () => { throw new Error('EROFS'); } });
|
||||
assert.throws(() => writeReviewOrStub(p, 'x', d), /EROFS/);
|
||||
});
|
||||
});
|
||||
|
||||
describe('runner — opencode handler (#1936)', () => {
|
||||
test('the review is rebuilt from assistant text parts', () => {
|
||||
const stream = [
|
||||
JSON.stringify({ type: 'text', part: { text: 'first' } }),
|
||||
JSON.stringify({ type: 'text', part: { text: 'second' } }),
|
||||
].join('\n');
|
||||
assert.equal(handleOpencodeOutput(stream).review, 'first\nsecond');
|
||||
});
|
||||
|
||||
test('a malformed line is skipped, not fatal to the whole review', () => {
|
||||
// Losing an entire review to one bad line would be strictly worse than the bug this fixes.
|
||||
const stream = [
|
||||
JSON.stringify({ type: 'text', part: { text: 'kept' } }),
|
||||
'NOT JSON AT ALL',
|
||||
'{"truncated":',
|
||||
JSON.stringify({ type: 'text', part: { text: 'also kept' } }),
|
||||
].join('\n');
|
||||
assert.equal(handleOpencodeOutput(stream).review, 'kept\nalso kept');
|
||||
});
|
||||
|
||||
test('a zero-output turn surfaces the stop reason and token count', () => {
|
||||
const stream = JSON.stringify({ type: 'step_finish', part: { reason: 'stop', tokens: { output: 0 } } });
|
||||
const r = handleOpencodeOutput(stream);
|
||||
assert.equal(r.review, '');
|
||||
assert.ok(r.diagnostic.includes('stop'));
|
||||
assert.ok(r.diagnostic.includes('0'));
|
||||
});
|
||||
|
||||
test('the raw JSON envelope never becomes the review', async () => {
|
||||
// The regression this locks: a plain stdout copy would write the JSON stream into REVIEWS.md.
|
||||
const p = plan('opencode');
|
||||
const stream = JSON.stringify({ type: 'text', part: { text: 'THE REVIEW' } });
|
||||
const d = deps({ spawn: () => ({ status: 0, stdout: stream, stderr: '' }) });
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(d.files[p.reviewPath].trim(), 'THE REVIEW');
|
||||
assert.ok(!d.files[p.reviewPath].includes('"type"'));
|
||||
});
|
||||
|
||||
test('CRLF in the stream is handled', () => {
|
||||
const stream = [
|
||||
JSON.stringify({ type: 'text', part: { text: 'a' } }),
|
||||
JSON.stringify({ type: 'text', part: { text: 'b' } }),
|
||||
].join('\r\n');
|
||||
assert.equal(handleOpencodeOutput(stream).review, 'a\nb');
|
||||
});
|
||||
});
|
||||
|
||||
describe('runner — antigravity handler (#2073 / #2176)', () => {
|
||||
const CACHE = '/home/u/.gemini/antigravity-cli/cache/last_conversations.json';
|
||||
const TX = (id) => `/home/u/.gemini/antigravity-cli/brain/${id}/.system_generated/logs/transcript.jsonl`;
|
||||
const entry = (content) =>
|
||||
JSON.stringify({ source: 'MODEL', status: 'DONE', type: 'PLANNER_RESPONSE', content });
|
||||
|
||||
test('the watermark prevents a PRIOR run’s response leaking in as this one', () => {
|
||||
// Without it the fallback reads the last PLANNER_RESPONSE regardless of when it was written,
|
||||
// silently presenting a stale review as the current one.
|
||||
const files = {
|
||||
[CACHE]: JSON.stringify({ [ROOT]: 'c1' }),
|
||||
[TX('c1')]: [entry('STALE FROM LAST RUN')].join('\n'),
|
||||
};
|
||||
const d = deps({ files });
|
||||
const got = antigravityTranscriptFallback(ROOT, { convId: 'c1', lines: 1 }, d);
|
||||
assert.equal(got, '', 'nothing was appended after the watermark, so nothing may be returned');
|
||||
});
|
||||
|
||||
test('a response appended after the watermark IS returned', () => {
|
||||
const files = {
|
||||
[CACHE]: JSON.stringify({ [ROOT]: 'c1' }),
|
||||
[TX('c1')]: [entry('old'), entry('THIS RUN')].join('\n'),
|
||||
};
|
||||
const d = deps({ files });
|
||||
assert.equal(antigravityTranscriptFallback(ROOT, { convId: 'c1', lines: 1 }, d), 'THIS RUN');
|
||||
});
|
||||
|
||||
test('a new conversation id means every line is new (skip 0)', () => {
|
||||
const files = {
|
||||
[CACHE]: JSON.stringify({ [ROOT]: 'c2' }),
|
||||
[TX('c2')]: [entry('FRESH SESSION')].join('\n'),
|
||||
};
|
||||
const d = deps({ files });
|
||||
assert.equal(antigravityTranscriptFallback(ROOT, { convId: 'c1', lines: 9 }, d), 'FRESH SESSION');
|
||||
});
|
||||
|
||||
test('workspace lookup is case-insensitive', () => {
|
||||
const files = {
|
||||
[CACHE]: JSON.stringify({ '/REPO': 'c1' }),
|
||||
[TX('c1')]: [entry('found')].join('\n'),
|
||||
};
|
||||
assert.equal(antigravityTranscriptFallback('/repo', { convId: '', lines: 0 }, deps({ files })), 'found');
|
||||
});
|
||||
|
||||
test('a missing cache or transcript degrades to empty, never throws', () => {
|
||||
assert.equal(antigravityTranscriptFallback(ROOT, { convId: '', lines: 0 }, deps()), '');
|
||||
const d = deps({ files: { [CACHE]: 'NOT JSON' } });
|
||||
assert.equal(antigravityTranscriptFallback(ROOT, { convId: '', lines: 0 }, d), '');
|
||||
});
|
||||
|
||||
test('the blind-review marker is anchored to the head of the output', () => {
|
||||
assert.ok(stampBlindReview('REVIEWED-WITHOUT-REPO-ACCESS\nbody').startsWith('> [reviewed-without-repo-access]'));
|
||||
});
|
||||
|
||||
test('a review that merely QUOTES the marker further down is NOT stamped', () => {
|
||||
// A grounded review of this very file would otherwise be mis-stamped and down-weighted.
|
||||
const quoting = ['1', '2', '3', '4', '5', '6', 'we look for REVIEWED-WITHOUT-REPO-ACCESS here'].join('\n');
|
||||
assert.ok(!stampBlindReview(quoting).startsWith('>'));
|
||||
});
|
||||
|
||||
test('the scratch-dir tell requires a workspace DECLARATION, not a mention', () => {
|
||||
const declared = 'my working directory is /home/u/.gemini/antigravity-cli/scratch so I could not read';
|
||||
assert.ok(stampBlindReview(declared).startsWith('>'));
|
||||
const mention = 'the path .gemini/antigravity-cli/scratch appears in the plan under review';
|
||||
assert.ok(!stampBlindReview(mention).startsWith('>'));
|
||||
});
|
||||
|
||||
test('a non-zero exit discards partial output so the fallback can take over', async () => {
|
||||
// The spawn APPENDS to the transcript, as the real `agy` does. That ordering is the whole
|
||||
// point of the watermark: only what this run wrote may be read back. A test that pre-seeds the
|
||||
// response instead would be asserting that a STALE entry leaks through — the exact bug the
|
||||
// watermark exists to prevent — so it must be written this way round.
|
||||
const p = plan('antigravity');
|
||||
const files = {
|
||||
[CACHE]: JSON.stringify({ [ROOT]: 'c1' }),
|
||||
[TX('c1')]: [entry('from a PREVIOUS run')].join('\n'),
|
||||
};
|
||||
const d = deps({
|
||||
files,
|
||||
spawn: () => {
|
||||
files[TX('c1')] = [entry('from a PREVIOUS run'), entry('FROM TRANSCRIPT')].join('\n');
|
||||
return { status: 124, stdout: 'partial garbage', stderr: '' };
|
||||
},
|
||||
});
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.ok(d.files[p.reviewPath].includes('FROM TRANSCRIPT'));
|
||||
assert.ok(!d.files[p.reviewPath].includes('partial garbage'), 'rc!=0 must discard stdout');
|
||||
assert.ok(!d.files[p.reviewPath].includes('PREVIOUS'), 'the pre-run entry must stay invisible');
|
||||
});
|
||||
});
|
||||
|
||||
describe('runner — openai-compatible handler', () => {
|
||||
test('the configured model is used and discovery is skipped', async () => {
|
||||
const p = plan('ollama', { 'review.models.ollama': 'pinned' });
|
||||
let posted = null;
|
||||
const d = deps({
|
||||
httpJson: async (url, o) => {
|
||||
if (o.method === 'POST') { posted = JSON.parse(o.body); return { ok: true, status: 200, body: JSON.stringify({ choices: [{ message: { content: 'R' } }] }) }; }
|
||||
return { ok: true, status: 200, body: JSON.stringify({ data: [{ id: 'discovered' }] }) };
|
||||
},
|
||||
});
|
||||
const r = await runOpenAiCompatible(p, 'PROMPT', d);
|
||||
assert.equal(posted.model, 'pinned');
|
||||
assert.equal(r.review, 'R');
|
||||
});
|
||||
|
||||
test('an unset model discovers the first from /v1/models', async () => {
|
||||
const p = plan('ollama');
|
||||
let posted = null;
|
||||
const d = deps({
|
||||
httpJson: async (url, o) => {
|
||||
if (o.method === 'POST') { posted = JSON.parse(o.body); return { ok: true, status: 200, body: JSON.stringify({ choices: [{ message: { content: 'R' } }] }) }; }
|
||||
return { ok: true, status: 200, body: JSON.stringify({ data: [{ id: 'discovered' }] }) };
|
||||
},
|
||||
});
|
||||
await runOpenAiCompatible(p, 'P', d);
|
||||
assert.equal(posted.model, 'discovered');
|
||||
});
|
||||
|
||||
test('discovery failure falls back to the declared fallbackModel', async () => {
|
||||
const p = plan('ollama');
|
||||
let posted = null;
|
||||
const d = deps({
|
||||
httpJson: async (url, o) => {
|
||||
if (o.method === 'POST') { posted = JSON.parse(o.body); return { ok: true, status: 200, body: '{}' }; }
|
||||
return { ok: false, status: 0, body: '', error: 'refused' };
|
||||
},
|
||||
});
|
||||
await runOpenAiCompatible(p, 'P', d);
|
||||
assert.equal(posted.model, 'llama3');
|
||||
});
|
||||
|
||||
test('a served-model mismatch warns without failing the review', async () => {
|
||||
const p = plan('lm_studio', { 'review.models.lm_studio': 'asked' });
|
||||
const d = deps({
|
||||
httpJson: async () => ({ ok: true, status: 200, body: JSON.stringify({ model: 'served', choices: [{ message: { content: 'R' } }] }) }),
|
||||
});
|
||||
const r = await runOpenAiCompatible(p, 'P', d);
|
||||
assert.equal(r.review, 'R');
|
||||
assert.ok(d.warnings.some((w) => w.includes('served') && w.includes('asked')));
|
||||
});
|
||||
|
||||
test('an HTTP error body is preserved for the stub', async () => {
|
||||
const p = plan('ollama');
|
||||
const d = deps({
|
||||
httpJson: async (url, o) =>
|
||||
o.method === 'POST'
|
||||
? { ok: false, status: 404, body: '{"error":"no such model"}' }
|
||||
: { ok: false, status: 0, body: '' },
|
||||
});
|
||||
const r = await runOpenAiCompatible(p, 'P', d);
|
||||
assert.equal(r.review, '');
|
||||
assert.ok(r.rawBody.includes('no such model'));
|
||||
});
|
||||
|
||||
test('a non-JSON response body does not throw', async () => {
|
||||
const p = plan('ollama');
|
||||
const d = deps({ httpJson: async () => ({ ok: true, status: 200, body: '<html>502</html>' }) });
|
||||
const r = await runOpenAiCompatible(p, 'P', d);
|
||||
assert.equal(r.review, '');
|
||||
assert.ok(r.rawBody.includes('502'));
|
||||
});
|
||||
});
|
||||
|
||||
describe('runner — orchestration', () => {
|
||||
test('an unavailable lane requested EXPLICITLY is surfaced (D4 carve-out)', async () => {
|
||||
const p = plan('gemini');
|
||||
const d = deps({ hasBinary: () => false });
|
||||
const r = await runLane(p, d, { repoRoot: ROOT, explicitlyRequested: true });
|
||||
assert.equal(r.ok, false);
|
||||
assert.ok(d.warnings.some((w) => w.includes('explicitly requested')));
|
||||
});
|
||||
|
||||
test('an unavailable lane nobody asked for is quiet but still reported', async () => {
|
||||
const p = plan('gemini');
|
||||
const d = deps({ hasBinary: () => false });
|
||||
const r = await runLane(p, d, { repoRoot: ROOT, explicitlyRequested: false });
|
||||
assert.equal(r.ok, false);
|
||||
assert.equal(r.reason, LANE_UNAVAILABLE.MISSING_BINARY);
|
||||
assert.deepStrictEqual(d.warnings, []);
|
||||
});
|
||||
|
||||
test('a file-arg lane reads its review from the file, not stdout', async () => {
|
||||
// Codex writes via -o and its stdout carries Windows teardown noise after the final message
|
||||
// (#1698); a stdout redirect would append that to a non-empty file and slip past the guard.
|
||||
const p = plan('codex');
|
||||
const d = deps({
|
||||
files: { [`${RUN}/gsd-review-codex.md`]: 'FROM FILE' },
|
||||
spawn: () => ({ status: 0, stdout: 'TEARDOWN NOISE', stderr: '' }),
|
||||
});
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.ok(d.files[p.reviewPath].includes('FROM FILE'));
|
||||
assert.ok(!d.files[p.reviewPath].includes('TEARDOWN NOISE'));
|
||||
});
|
||||
|
||||
test('stderr is always captured to the sidecar, never discarded', async () => {
|
||||
const p = plan('gemini');
|
||||
const d = deps({ spawn: () => ({ status: 0, stdout: 'R', stderr: 'a warning' }) });
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(d.files[p.errPath], 'a warning');
|
||||
});
|
||||
|
||||
test('the prompt reaches stdin for a stdin lane', async () => {
|
||||
const p = plan('gemini');
|
||||
const d = deps({
|
||||
files: { [`${RUN}/gsd-review-prompt.md`]: 'THE PLAN' },
|
||||
spawn: (b, a, o) => { d.spawns.push({ b, a, o }); return { status: 0, stdout: 'R', stderr: '' }; },
|
||||
});
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(d.spawns[0].o.input, 'THE PLAN');
|
||||
});
|
||||
|
||||
test('a prompt-less lane is fed nothing', async () => {
|
||||
const p = plan('coderabbit');
|
||||
const d = deps({
|
||||
files: { [`${RUN}/gsd-review-prompt.md`]: 'THE PLAN' },
|
||||
spawn: (b, a, o) => { d.spawns.push({ b, a, o }); return { status: 0, stdout: 'R', stderr: '' }; },
|
||||
});
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
assert.equal(d.spawns[0].o.input, undefined);
|
||||
});
|
||||
|
||||
test('every spawn carries a positive timeout', async () => {
|
||||
// DEFECT.UNBOUNDED-SUBPROCESS: a frozen sync spawn cannot be interrupted and hangs a whole CI
|
||||
// chunk to its 10-minute kill with `# fail 0` and no `not ok`.
|
||||
for (const lane of REVIEWER_LANES.filter((l) => l.transport === 'spawn')) {
|
||||
const p = plan(lane.slug);
|
||||
const d = deps({ spawn: (b, a, o) => { d.spawns.push({ b, a, o }); return { status: 0, stdout: 'R', stderr: '' }; } });
|
||||
await runLane(p, d, { repoRoot: ROOT });
|
||||
for (const s of d.spawns) {
|
||||
assert.ok(s.o.timeoutMs > 0, `${lane.slug} spawned unbounded`);
|
||||
}
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -91,6 +91,9 @@ const RUNTIME_REVIEWER_IDS = ['antigravity', 'claude', 'codex', 'cursor', 'openc
|
||||
*/
|
||||
const LITERAL_ROSTER = [
|
||||
'antigravity', 'claude', 'coderabbit', 'codex', 'cursor', 'gemini',
|
||||
// `kimi-code` joined in Phase 5b (#2799, closes #2718) — see
|
||||
// kimiCodeIsDeclaredAndInvocableInThisPhase for why it landed here and not in 5a.
|
||||
'kimi-code',
|
||||
'llama_cpp', 'lm_studio', 'ollama', 'opencode', 'qwen',
|
||||
];
|
||||
|
||||
@@ -328,10 +331,10 @@ describe('B. The six existing runtime capabilities', () => {
|
||||
|
||||
describe('C. Roster derivation — src/review-reviewer-selection.cts', () => {
|
||||
test('rosterMembershipIsUnchangedByDerivationRefactor', () => {
|
||||
assert.equal(KNOWN_REVIEWER_SLUGS.length, 11, 'roster must be exactly 11 — not 10, not 12');
|
||||
assert.equal(KNOWN_REVIEWER_SLUGS.length, 12, 'roster must be exactly 12 — not 11, not 13');
|
||||
assert.deepEqual(
|
||||
[...KNOWN_REVIEWER_SLUGS].sort(), LITERAL_ROSTER,
|
||||
`roster must be exactly the same 11 slugs as before this phase, got: ${JSON.stringify(KNOWN_REVIEWER_SLUGS)}`,
|
||||
`roster must be exactly the declared lane set, got: ${JSON.stringify(KNOWN_REVIEWER_SLUGS)}`,
|
||||
);
|
||||
});
|
||||
|
||||
@@ -414,28 +417,43 @@ describe('D. Cross-phase invariants that must not regress', () => {
|
||||
const workflowText = fs
|
||||
.readFileSync(path.join(ROOT, 'gsd-core', 'workflows', 'review.md'), 'utf-8')
|
||||
.replace(/\r\n/g, '\n');
|
||||
const registry = [...SHIPPED.capMap.values()]
|
||||
.map((c) => c && c.reviewer && c.reviewer.slug)
|
||||
.filter((x) => typeof x === 'string' && x)
|
||||
.sort();
|
||||
const result = checkReviewerLaneParity({
|
||||
descriptor: REVIEWER_LANES,
|
||||
roster: KNOWN_REVIEWER_SLUGS,
|
||||
registry,
|
||||
workflowText,
|
||||
});
|
||||
assert.deepEqual(
|
||||
result.violations, [],
|
||||
`Phase 1's descriptor <-> roster <-> legs <-> sections parity must stay green across this migration, got: ${JSON.stringify(result.violations)}`,
|
||||
`descriptor <-> roster <-> registry parity must stay green across this migration, got: ${JSON.stringify(result.violations)}`,
|
||||
);
|
||||
assert.equal(result.ok, true);
|
||||
});
|
||||
|
||||
test('kimiCodeIsDeliberatelyNotDeclaredYet', () => {
|
||||
assert.equal(KNOWN_REVIEWER_SLUGS.includes('kimi-code'), false, '"kimi-code" must not be in the roster yet — it has no invoke_reviewers leg until 5b');
|
||||
assert.equal(KNOWN_REVIEWER_SLUGS.includes('kimi_code'), false);
|
||||
test('kimiCodeIsDeclaredAndInvocableInThisPhase', () => {
|
||||
// Phase 5a deliberately withheld this lane: declaring it there would have made it selectable
|
||||
// but NOT invocable — present in `--all`, selected, and producing an empty section for the
|
||||
// whole 5a -> 5b window. ADR-2782's phase table lands it here, with the iteration that runs it.
|
||||
assert.equal(KNOWN_REVIEWER_SLUGS.includes('kimi-code'), true);
|
||||
const cap = SHIPPED.capMap.get('kimi-code');
|
||||
assert.ok(cap, 'expected the kimi-code capability to exist (it is net-new for an unrelated, EoS reason)');
|
||||
assert.equal('reviewer' in cap, false, 'kimi-code must not declare a reviewer body in this phase');
|
||||
assert.ok(cap, 'expected the kimi-code capability to exist');
|
||||
assert.equal('reviewer' in cap, true, 'kimi-code must declare a reviewer body in 5b');
|
||||
assert.equal(cap.reviewer.slug, 'kimi-code');
|
||||
// The probe is the whole reason D7 ships wider than existence: `kimi` is claimed by BOTH the
|
||||
// Kimi Code CLI and the legacy Python kimi-cli, so an existence-only probe registers the wrong
|
||||
// tool. And it MUST be bounded — the original was an unbounded `kimi --help | grep` that ran
|
||||
// on every review regardless of flags.
|
||||
assert.equal(cap.reviewer.probe.kind, 'command-capability');
|
||||
assert.equal(cap.reviewer.probe.binary, 'kimi');
|
||||
assert.ok(cap.reviewer.probe.timeoutMs > 0, 'every process-starting probe must be bounded');
|
||||
assert.equal(
|
||||
Boolean(cap.runtime && cap.runtime.hostBehaviors && cap.runtime.hostBehaviors.reviewerCli),
|
||||
false,
|
||||
'kimi-code must not carry the legacy reviewerCli alias either',
|
||||
'the body is the declaration — the legacy reviewerCli alias must not also be set',
|
||||
);
|
||||
});
|
||||
|
||||
@@ -509,8 +527,8 @@ describe('E. Lane fidelity — no translation layer', () => {
|
||||
}
|
||||
}
|
||||
|
||||
assert.equal(REVIEWER_LANES.length, 11, 'expected exactly 11 declared descriptor lanes');
|
||||
assert.equal(bySlug.size, 11, `expected exactly 11 capabilities declaring a reviewer body, got: ${bySlug.size}`);
|
||||
assert.equal(REVIEWER_LANES.length, 12, 'expected exactly 12 declared descriptor lanes');
|
||||
assert.equal(bySlug.size, 12, `expected exactly 11 capabilities declaring a reviewer body, got: ${bySlug.size}`);
|
||||
|
||||
// Top-level scalar/array fields compared whole; the two fields that are
|
||||
// themselves nested objects (probe, invoke) are compared sub-field-by-
|
||||
@@ -611,7 +629,7 @@ describe('F. Isolated-security-review regressions', () => {
|
||||
// The module under test already imported successfully above; assert the
|
||||
// derived roster is a usable array rather than a partially-initialised value.
|
||||
assert.ok(Array.isArray([...KNOWN_REVIEWER_SLUGS]), 'roster must be iterable after module load');
|
||||
assert.equal(KNOWN_REVIEWER_SLUGS.length, 11, 'the real registry still yields the eleven lanes');
|
||||
assert.equal(KNOWN_REVIEWER_SLUGS.length, 12, 'the real registry still yields the eleven lanes');
|
||||
// And the derivation itself is total over the shapes JSON can express.
|
||||
for (const hostile of [null, undefined, [], 0, 'x', { capabilities: null }, { capabilities: [] }]) {
|
||||
assert.doesNotThrow(
|
||||
|
||||
Reference in New Issue
Block a user