`review-lane plan` resolved every cross-AI reviewer lane's reasoning effort by spawning `query resolve-execution gsd-plan-checker --host <slug>`. The agent id was a hardcoded literal, so `--host` chose only the argv RENDERING while the LEVEL always came from the installed plan-checker's frontmatter — `low` under every shipped model profile. Every prompt-fed lane therefore ran at a fast structural verifier's effort, and because the rendered argument is a CLI config override it silently beat the effort the operator had configured for that CLI. At `low` a large source-grounded prompt makes a model end its turn with no final message, so the lane came back empty and its stub read as a crash. Effort is a property of the review, so the lane declares it. Two new fields on ReviewerLane — `effortConfigKey` (`review.effort.<slug>`) and `defaultEffort` — carried through each capability manifest and the generated registry, set on the three lanes with an argv effort channel and null on the other nine. A new pure `resolveLaneEffort()` resolves config key -> lane default -> nothing, where "nothing" emits no effort argument at all and the reviewer CLI's own configuration decides; `inherit` selects that path explicitly and an unrecognized level falls back to the lane default rather than being forwarded to a CLI that would reject it. The host's negotiated effortSurface still gates the rendering, so ADR-1239/#2481's trust boundary holds on this path too. Resolving in-process also removes up to twelve subprocess spawns per review. The empty-output stub now names the effort the lane ran at and distinguishes a clean exit from a timeout kill, a non-zero exit, and a process that never ran — `status` is null for both a timeout and a signal, so those were indistinguishable before. The hint is hedged: a clean empty exit is most often a model stopping short, but it is also consistent with a CLI writing its output elsewhere. Also: the capability validator now knows both fields, rejects a malformed key or an out-of-vocabulary default, and rejects a default declared without a config key (a level the operator could never override). An existing end-to-end row in tests/effort-surface-axis.test.cjs asserted the old coupling; it now configures the lane's own key and pins the decoupling in the same real spawn, with the agent execution tier set to a level that must not appear. Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort; leaving the new key undocumented there is the same invisibility that made the plan-checker coupling survive this long. Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort, so leaving the new key undocumented there is the same invisibility that let the plan-checker coupling survive. Claude-Session: https://claude.ai/code/session_01CRMEuzNMWn3gs5uUW2ghcF Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
67 lines
2.7 KiB
JSON
67 lines
2.7 KiB
JSON
{
|
|
"id": "llama-cpp",
|
|
"role": "reviewer",
|
|
"version": "1.12.0",
|
|
"title": "llama.cpp",
|
|
"description": "llama.cpp server — cross-AI /gsd:review reviewer lane only; not a GSD install target (no runtime body, no artifacts). OpenAI-compatible HTTP transport against a user-configured `review.llama_cpp_host` (POST /v1/chat/completions); model discovered via GET /v1/models piped through jq. Capability id/folder are kebab (`llama-cpp`, required by KEBAB_RE); `reviewer.slug` stays snake (`llama_cpp`) to match the shipped roster and the `review.llama_cpp_host` config key (ADR-2782's three-namespace trap).",
|
|
"tier": "full",
|
|
"requires": [],
|
|
"engines": {
|
|
"gsd": ">=1.8.0"
|
|
},
|
|
"reviewer": {
|
|
"slug": "llama_cpp",
|
|
"flags": [
|
|
"--llama-cpp"
|
|
],
|
|
"transport": "openai-http",
|
|
"probe": {
|
|
"kind": "http-reachable",
|
|
"hostConfigKey": "review.llama_cpp_host",
|
|
"path": "/v1/models",
|
|
"timeoutMs": 2000
|
|
},
|
|
"invoke": {
|
|
"hostConfigKey": "review.llama_cpp_host",
|
|
"defaultHost": "http://localhost:8080",
|
|
"path": "/v1/chat/completions",
|
|
"modelDiscovery": "first-from-models-endpoint",
|
|
"fallbackModel": "local-model",
|
|
"effortChannel": "none"
|
|
},
|
|
"timeoutFloorMs": 120000,
|
|
"timeoutConfigKey": "review.timeouts.llama_cpp",
|
|
"emptyOutput": "stub-with-stderr",
|
|
"reviewsSection": "llama.cpp",
|
|
"evidenceClass": "source-grounded",
|
|
"requiresBinaries": [],
|
|
"promptBudgetKey": "review.max_prompt_tokens_per_reviewer.llama_cpp",
|
|
"modelConfigKey": "review.models.llama_cpp",
|
|
"effortConfigKey": null,
|
|
"defaultEffort": null,
|
|
"handler": "openai-compatible"
|
|
},
|
|
"config": {
|
|
"review.models.llama_cpp": {
|
|
"type": "string",
|
|
"default": "",
|
|
"description": "Model requested from the llama.cpp reviewer lane; empty discovers the first model from /v1/models."
|
|
},
|
|
"review.llama_cpp_host": {
|
|
"type": "string",
|
|
"default": "",
|
|
"description": "Base URL of the llama.cpp OpenAI-compatible server."
|
|
},
|
|
"review.max_prompt_tokens_per_reviewer.llama_cpp": {
|
|
"type": "number",
|
|
"default": -1,
|
|
"description": "Prompt-token budget for the llama.cpp reviewer lane. Unset is -1, a sentinel: 0 is a legitimate value meaning \"do not trim this lane\", so it cannot double as \"not configured\"."
|
|
},
|
|
"review.timeouts.llama_cpp": {
|
|
"type": "number",
|
|
"default": -1,
|
|
"description": "Outer wall-clock timeout override (seconds) for the llama.cpp reviewer lane. Unset is -1, a sentinel: 0 or a negative number is also treated as unset (a timeout has no legitimate zero/negative value), so no second sentinel is needed. Falls back to the lane's built-in timeoutFloorMs when unset."
|
|
}
|
|
}
|
|
}
|