* test(#3691): failing-first coverage for the reviewer prompt budget No prompt cap can reach any CLI reviewer lane, by any configuration. Two independent defects compound: all nine `transport: spawn` lanes declare `promptBudgetKey: null`, so `budgetFor` returns on its first line; and the documented global `review.max_prompt_tokens` is advertised in the schema manifest but declared nowhere, so the resolver never materializes it and `budgetFor`'s fallback is dead code. Adds to tests/reviewer-config-federation.test.cjs, which already owns the per-reviewer budget config-set/config-get idiom: - a CLI lane inherits the global cap (RED: reports null) - an http lane with the -1 sentinel inherits the global cap (RED: reports null) - the resolved review surface carries max_prompt_tokens at all (RED: absent) - per-lane overrides the global on a CLI lane - the sentinel boundary: -1 inherits, 0 means do-not-trim and must NOT read as unset, 1 is the smallest real budget — the regression budgetFor's own comment warns about - anti-tightening pins that must stay green: an empty config leaves every lane null, the three existing budgeted lanes are unchanged, and config-set still rejects a per-reviewer key naming something that is not a declared lane - a fast-check property over the resolution contract itself, with -1, 0 and non-finite inputs generated explicitly rather than left to chance Every row was reproduced by hand against the real CLI before being written, so the RED/GREEN split is observed rather than predicted. Refs #3691 * fix(#3691): let every reviewer lane take a prompt cap, and make the global resolve No prompt cap could reach any CLI reviewer lane, by any configuration. Two independent defects compounded. The nine spawn-transport lanes — claude, coderabbit, antigravity, cursor, gemini, codex, kimi-code, opencode, qwen — declared `promptBudgetKey: null`, so `budgetFor` returned on its first line and `review-lane plan` reported `promptBudget: null` no matter what was configured. Each now declares `review.max_prompt_tokens_per_reviewer.<slug>` with the same `-1`-is-unset sentinel the three local-server lanes already use. Separately, the central `review.max_prompt_tokens` was listed in the schema manifest's validKeys and documented as a supported setting, but declared nowhere — the resolved surface is built from capability declarations plus the defaults manifest, and neither carried it. `configGet` returned undefined and `budgetFor`'s documented fallback was dead code. It is now declared with a `null` default, exactly as docs/CONFIGURATION.md already specified, so the default behavior is unchanged: nothing configured means nothing trims. Two things the diagnosis had not predicted, found and fixed while implementing: - `REVIEWER_LANES` in src/review-lane-descriptor.cts is a second, hardcoded registration site that `mergeReviewerLanes` prefers over the capability registry on a slug collision. Editing only the capability files left every CLI lane still null. Both sites now agree. - The generated `gsd-core/bin/lib/capability-registry.cjs` was stale and masked the capability edits; regenerated with `npm run gen:capability-registry` rather than hand-edited. docs/CONFIGURATION.md said "Only lanes that declare a budget key accept one — today ollama, lm_studio and llama_cpp". That is false as of this change and is corrected rather than left to rot. The trim-versus-refuse question the issue raises is deliberately not taken up here: the refusal path already exists for the case that matters — a reviewer whose minimum set exceeds its budget is skipped rather than sent a misleading prompt — and trimming above that floor is the documented, shipped design of the feature. Changing it would alter behavior for the three lanes that already work, which is not what the issue asks for. Fixes #3691 * fix(#3691): document the new global and narrow an invariant this change obsoleted The full suite surfaced two consequences of giving every CLI lane a budget key. `review.max_prompt_tokens` entered CONFIG_DEFAULTS without a matching entry in the planning-config reference, which config-field-docs guards. Documented, including the sentinel semantics a reader needs: a per-lane value overrides the global, `-1` means unset and inherits it, and `0` means "do not trim that lane" and is not unset. The #2797 federation guard asserted that "a lane with no model flag and no host owns no config keys". That held only because budget keys existed solely on the three local-server lanes, all of which have hosts. A lane can now legitimately own a config key for a third reason, so qwen tripped it. The assertion is narrowed rather than weakened: such a lane must still own no model key and no host key, and may own at most its own `review.max_prompt_tokens_per_reviewer.<slug>` — never another lane's. That is strictly more specific in the dimensions that still matter. Proven to still bite: hypothetically giving qwen a `review.models.qwen` key fails it with `model/host: review.models.qwen`. The name and comment cite #3691 for why the premise changed, so a reader sees a deliberate narrowing, not erosion. Checked the sibling assertions in that describe block; the other three do not rest on the obsolete premise and are untouched. Refs #3691 * fix(#3685): port the write-flag content-change contract to its three sibling sites #3685 fixed `phase complete`'s `roadmap_updated` / `state_updated`, which reported `fs.existsSync(path)` rather than whether the transaction wrote anything. Three sibling sites carried the identical defect and are ported here. - `cmdPhaseRemove` reported `roadmap_updated: true`, hardcoded. `updateRoadmapAfterPhaseRemoval` now returns whether the content changed and the flag reports it. #2640/#2974 already fixed `state_updated` at this same call site and left this one behind, so the correct shape was adjacent. - `cmdMilestoneComplete` reported `state_updated: fs.existsSync(statePath)` — byte-identical to #3685's bug in a different command. - `cmdMilestoneComplete` reported `milestones_updated: true`, hardcoded, never consulting the MILESTONES.md write. `gsd-core/workflows/remove-phase.md:100` extracts `roadmap_updated` for display and never branches on it, so the flip from always-true to content-based changes no workflow behavior. Verified by reading the step, not assumed. One trap found while implementing: the obvious in-memory `finalContent !== originalStateContent` comparison — copying `cmdPhaseComplete`'s shipped shape verbatim — gives a FALSE POSITIVE for milestone completion. `platformWriteSync` normalizes Markdown at write time, and the milestone-closure transform regenerates `## Current Position` fresh on every call, so its pre-normalize output always differs from the already-normalized file on disk even when the persisted bytes are identical. The comparison is therefore made against the post-write on-disk content. `cmdPhaseComplete`'s own comparisons are left untouched — their repeat-no-op tests pass, so they are not exposed to this artifact. `milestones_updated` has no reachable no-op: the MILESTONES.md write unconditionally appends an entry every call. Only the true direction is pinned, documented inline rather than faked with a passing test. Refs #3685 * fix(#3685): compare write-flag content through the writer's own normalizer An independent reviewer disproved a claim made while porting #3685's contract to its sibling sites: that `cmdPhaseComplete`'s comparisons were not exposed to the Markdown-normalization artifact already diagnosed in `cmdMilestoneComplete`. `platformWriteSync` normalizes on write — CRLF stripped, blank-line runs collapsed, a blank line inserted after a heading, a single trailing newline enforced. Every flag that compares the PRE-normalization in-memory string against the on-disk pre-image can therefore report a change when the persisted bytes are identical. `cmdMilestoneComplete` had been worked around by re-reading the file after the write; the other sites compared raw strings. All of them now go through one exported seam, `contentChangedAfterNormalize(filePath, before, after)`, which normalizes both sides exactly as the writer does. That removes the extra disk read the milestone workaround needed, and makes the sites agree by construction rather than by four independent implementations of one rule — the divergence the repo names as an anti-pattern. Reachability, stated precisely rather than uniformly: the seam is load-bearing at `cmdPhaseComplete`'s `roadmapUpdated`, `requirementsUpdated` and `stateUpdated`, where section-rewrite logic genuinely regenerates content into a different-but-normalization-equivalent shape. At `updateRoadmapAfterPhaseRemoval` it is defense-in-depth: the no-match branch never reassigns `content`, so the raw comparison was already correct there. The first analysis claimed the reverse; this is the corrected finding. Also fixes an unsound test premise the remote suite caught. The byte-identity precondition in `roadmap_updated is false when ROADMAP.md comes out byte-identical` asserted against a hand-authored, un-normalized fixture — so the very first write reformatted it and the file could not come back identical. The fixture is now written already-normalized, so the assertion compares a normalized pre-image against a normalized post-image and still fails if the flag regresses to a hardcoded `true`. Not platform-specific; it reproduces on macOS too, and the earlier local check simply never exercised it. The sibling true-direction and milestone tests were checked for the same premise and do not share it — they assert `notEqual`, or compare two post-write states produced through the same normalizing seam. Refs #3685 * chore(changeset): backfill PR number for #3691 fragment --------- Co-authored-by: sim <sim@local>
168 lines
4.3 KiB
JSON
168 lines
4.3 KiB
JSON
{
|
|
"id": "claude",
|
|
"role": "runtime",
|
|
"version": "1.11.0",
|
|
"title": "Claude Code",
|
|
"description": "Anthropic Claude Code — primary development runtime; tier-1 support with full hook surface and skills-based global install.",
|
|
"tier": "core",
|
|
"requires": [],
|
|
"engines": {
|
|
"gsd": ">=1.6.0"
|
|
},
|
|
"runtime": {
|
|
"configHome": {
|
|
"kind": "dot-home",
|
|
"name": ".claude",
|
|
"env": [
|
|
"CLAUDE_CONFIG_DIR"
|
|
]
|
|
},
|
|
"localConfigDir": ".claude",
|
|
"configFormat": "settings-json",
|
|
"artifactLayout": {
|
|
"global": [
|
|
{
|
|
"kind": "skills",
|
|
"destSubpath": "skills",
|
|
"prefix": "gsd-",
|
|
"nesting": "flat",
|
|
"recursive": false,
|
|
"converter": "convertClaudeCommandToClaudeSkill"
|
|
},
|
|
{
|
|
"kind": "agents",
|
|
"destSubpath": "agents",
|
|
"prefix": "gsd-",
|
|
"nesting": "flat",
|
|
"recursive": false,
|
|
"converter": null
|
|
}
|
|
],
|
|
"local": [
|
|
{
|
|
"kind": "commands",
|
|
"destSubpath": "commands",
|
|
"prefix": "gsd-",
|
|
"nesting": "flat",
|
|
"recursive": false,
|
|
"converter": null
|
|
},
|
|
{
|
|
"kind": "agents",
|
|
"destSubpath": "agents",
|
|
"prefix": "gsd-",
|
|
"nesting": "flat",
|
|
"recursive": false,
|
|
"converter": null
|
|
}
|
|
]
|
|
},
|
|
"triggerPrecedence": [
|
|
"skills",
|
|
"commands"
|
|
],
|
|
"commandStyle": "slash-hyphen",
|
|
"hooksSurface": "settings-json",
|
|
"hookEvents": "claude",
|
|
"sandboxTier": "none",
|
|
"supportTier": 1,
|
|
"installSurface": "settings-json",
|
|
"writesSharedSettings": true,
|
|
"permissionWriter": null,
|
|
"extendedHookEvents": [
|
|
"SubagentStop",
|
|
"Stop",
|
|
"PreCompact",
|
|
"FileChanged"
|
|
],
|
|
"hostIntegration": {
|
|
"embeddingMode": "imperative",
|
|
"commandSurface": "slash-file",
|
|
"dispatch": {
|
|
"namedDispatch": true,
|
|
"nested": true,
|
|
"maxDepth": 5,
|
|
"background": true,
|
|
"subagentToolkit": "full",
|
|
"backgroundDispatch": false,
|
|
"isolation": "harness-worktree"
|
|
},
|
|
"modelMode": "passive",
|
|
"hookBus": "host",
|
|
"stateIO": "filesystem",
|
|
"transport": "mcp",
|
|
"runtime": "node",
|
|
"effortSurface": "argv"
|
|
},
|
|
"harnessIsolationFlag": "isolation=\"worktree\"",
|
|
"hostBehaviors": {
|
|
"attributionSource": "settings-json-commit",
|
|
"authorsCanonicalWorkflow": true,
|
|
"localInstallStyle": "legacy-flat",
|
|
"permissionsSchema": "claude",
|
|
"settingsFileByScope": {
|
|
"local": "settings.local.json",
|
|
"global": "settings.json"
|
|
},
|
|
"sourceMarkerFile": ".gsd-source",
|
|
"agentFrontmatterExtensions": [
|
|
"effort"
|
|
],
|
|
"ownsClaudePaths": true,
|
|
"nativeModelAliases": true,
|
|
"skillsGlobalOnboarding": true,
|
|
"legacyCommandsGsdInstallMigration": true,
|
|
"legacyCommandsGsdUninstall": "global",
|
|
"hyphenNameAgentBody": true
|
|
}
|
|
},
|
|
"reviewer": {
|
|
"slug": "claude",
|
|
"flags": [
|
|
"--claude"
|
|
],
|
|
"transport": "spawn",
|
|
"probe": {
|
|
"kind": "command-exists",
|
|
"binary": "claude"
|
|
},
|
|
"invoke": {
|
|
"binary": "claude",
|
|
"args": [
|
|
"{{model}}",
|
|
"{{effort}}",
|
|
"-p",
|
|
"-"
|
|
],
|
|
"promptChannel": "stdin",
|
|
"outputChannel": "stdout",
|
|
"modelArg": "--model",
|
|
"effortChannel": "argv",
|
|
"env": {
|
|
"CLAUDE_CODE_DISABLE_CLAUDE_MDS": "1",
|
|
"CLAUDE_CODE_DISABLE_AUTO_MEMORY": "1"
|
|
}
|
|
},
|
|
"timeoutFloorMs": 1200000,
|
|
"emptyOutput": "stub-with-stderr",
|
|
"reviewsSection": "Claude",
|
|
"evidenceClass": "source-grounded",
|
|
"requiresBinaries": [],
|
|
"promptBudgetKey": "review.max_prompt_tokens_per_reviewer.claude",
|
|
"modelConfigKey": "review.models.claude",
|
|
"handler": null
|
|
},
|
|
"config": {
|
|
"review.models.claude": {
|
|
"type": "string",
|
|
"default": "",
|
|
"description": "Model passed to the Claude reviewer lane."
|
|
},
|
|
"review.max_prompt_tokens_per_reviewer.claude": {
|
|
"type": "number",
|
|
"default": -1,
|
|
"description": "Prompt-token budget for the Claude reviewer lane. Unset is -1, a sentinel: 0 is a legitimate value meaning \"do not trim this lane\", so it cannot double as \"not configured\"."
|
|
}
|
|
}
|
|
}
|