Files
msd-core/capabilities/llama-cpp/capability.json
Tom Boucher 62b0d939b6 feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey (#4083)
* feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey

Add an optional `timeoutConfigKey` field to the reviewer lane descriptor,
resolved in `resolveLanePlan` at invocation time and falling back to the
frozen `timeoutFloorMs` when unset or invalid, in the same spirit as the
existing `promptBudgetKey`/`modelConfigKey` fields. All 12 shipped lanes
declare `review.timeouts.<slug>` on both surfaces (the descriptor and their
capability.json manifest), validated by capability-validator.cjs.

For the antigravity lane, the native `agy --print-timeout` flag — previously
a second hardcoded literal (`540s`) independent of the outer cap — is now
derived from the same resolved outer timeout in `antigravityArgv`, preserving
the existing 60-second buffer relationship (ADR-2782 D6: the outer bound is
declared data, the inner one is handler-owned).

The antigravity default timeoutFloorMs stays at 600s per the maintainer's
disposition; users raise it through the new config key instead.

* docs(#3274): document review.timeouts.* and extract resolveTimeoutMs helper

Address code-review findings on the timeoutConfigKey change: extract the
inline timeout-resolution logic into a named, exported, directly-tested
resolveTimeoutMs helper (matching the file's existing configString/
normalizeHost convention); document the new review.timeouts.* federated
config keys in docs/CONFIGURATION.md, docs/reference/capability-manifest.md,
and docs/how-to/ship-a-reviewer-lane.md; add the changeset fragment.

* fix(#3274): resolve native antigravity timeout in resolveLanePlan, not the runner

gsd-test caught two design mistakes in the prior commits:

1. SpawnPlan.argv is documented and tested as fully resolved by
   resolveLanePlan (model/effort/output/prompt already folded in) — leaving
   the antigravity '{{nativeTimeout}}' marker unresolved until the runner's
   antigravityArgv violated that contract and broke tests that read
   plan.argv directly (tests/antigravity-reviewer.test.cjs,
   tests/review-default-reviewers-workflow.test.cjs). Fix: '{{nativeTimeout}}'
   is now a fifth ARGV_PLACEHOLDER member, resolved by resolveLanePlan itself
   via the new nativeTimeoutToken() helper, exactly like the other four.
   antigravityArgv reverts to its pre-#3274 four-argument form. Also missed
   updating capabilities/antigravity/capability.json's invoke.args to match
   the descriptor, which broke the manifest/descriptor parity test.

2. tests/reviewer-config-federation.test.cjs enforces a deliberate, narrow
   invariant (#3691 narrows #2797): qwen, cursor, and coderabbit — the three
   lanes with neither a model flag nor a host — may own no config key beyond
   their own prompt-budget key. Adding review.timeouts.<slug> to all 12 lanes
   violated it. Fix: those three keep timeoutConfigKey: null and own no
   review.timeouts.* key, matching their existing modelConfigKey: null. The
   other 9 lanes are unaffected.

* chore(#3274): backfill changeset PR number (pr:0 -> 4083)

---------

Co-authored-by: sim <sim@local>
2026-08-30 14:26:02 -04:00

65 lines
2.7 KiB
JSON

{
"id": "llama-cpp",
"role": "reviewer",
"version": "1.12.0",
"title": "llama.cpp",
"description": "llama.cpp server — cross-AI /gsd:review reviewer lane only; not a GSD install target (no runtime body, no artifacts). OpenAI-compatible HTTP transport against a user-configured `review.llama_cpp_host` (POST /v1/chat/completions); model discovered via GET /v1/models piped through jq. Capability id/folder are kebab (`llama-cpp`, required by KEBAB_RE); `reviewer.slug` stays snake (`llama_cpp`) to match the shipped roster and the `review.llama_cpp_host` config key (ADR-2782's three-namespace trap).",
"tier": "full",
"requires": [],
"engines": {
"gsd": ">=1.8.0"
},
"reviewer": {
"slug": "llama_cpp",
"flags": [
"--llama-cpp"
],
"transport": "openai-http",
"probe": {
"kind": "http-reachable",
"hostConfigKey": "review.llama_cpp_host",
"path": "/v1/models",
"timeoutMs": 2000
},
"invoke": {
"hostConfigKey": "review.llama_cpp_host",
"defaultHost": "http://localhost:8080",
"path": "/v1/chat/completions",
"modelDiscovery": "first-from-models-endpoint",
"fallbackModel": "local-model",
"effortChannel": "none"
},
"timeoutFloorMs": 120000,
"timeoutConfigKey": "review.timeouts.llama_cpp",
"emptyOutput": "stub-with-stderr",
"reviewsSection": "llama.cpp",
"evidenceClass": "source-grounded",
"requiresBinaries": [],
"promptBudgetKey": "review.max_prompt_tokens_per_reviewer.llama_cpp",
"modelConfigKey": "review.models.llama_cpp",
"handler": "openai-compatible"
},
"config": {
"review.models.llama_cpp": {
"type": "string",
"default": "",
"description": "Model requested from the llama.cpp reviewer lane; empty discovers the first model from /v1/models."
},
"review.llama_cpp_host": {
"type": "string",
"default": "",
"description": "Base URL of the llama.cpp OpenAI-compatible server."
},
"review.max_prompt_tokens_per_reviewer.llama_cpp": {
"type": "number",
"default": -1,
"description": "Prompt-token budget for the llama.cpp reviewer lane. Unset is -1, a sentinel: 0 is a legitimate value meaning \"do not trim this lane\", so it cannot double as \"not configured\"."
},
"review.timeouts.llama_cpp": {
"type": "number",
"default": -1,
"description": "Outer wall-clock timeout override (seconds) for the llama.cpp reviewer lane. Unset is -1, a sentinel: 0 or a negative number is also treated as unset (a timeout has no legitimate zero/negative value), so no second sentinel is needed. Falls back to the lane's built-in timeoutFloorMs when unset."
}
}
}