Commit Graph

141 Commits

Author SHA1 Message Date
JusticeWay
7b204ad2ac enhance(#2530): extend UAT checkpoint frame language pack (9 more languages) (#2564)
* feat: extend UAT checkpoint frame language pack (9 more languages)

response_language is a free-form config value, but CHECKPOINT_FRAMES only
covered 9 languages — any other configured language silently fell back to
the English frame. Add Dutch, Polish, Russian, Ukrainian, Turkish, Hindi,
Arabic, Vietnamese, and Indonesian frames plus their aliases, with a
regression test asserting each resolves instead of falling back.

Follow-up to #2402 (PR #2457).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: add changeset for #2527

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(#2530): list UAT checkpoint frame languages in CONFIGURATION.md

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2530): point changeset fragment at PR #2557

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2530): address Unicode language-pack review

* fix: address checkpoint language review

* fix: count spacing combining marks in checkpoint width

* test: verify checkpoint aliases structurally

* fix: isolate RTL checkpoint frames

* fix: isolate RTL checkpoint frames correctly

* test(#2530): assert checkpoint aliases neither collide nor go unreachable

Review Minor #1. A duplicate alias key was invisible to the existing
catalog tests: the runtime object is well-formed after JS collapses the
literal, the self-alias assertion still holds, and the losing language
just stops resolving. tsc catches the byte-equal case (TS1117), but not
the two that survive compilation — an alias whose NFC-lowercase form
already belongs to another language, and an alias not in lookup form at
all, which resolveCheckpointFrame() can never produce.

The check reads the source literal rather than the object, since the
object no longer records what was written. Both assertions are
independently load-bearing: an NFD twin of an existing alias trips the
collision check, an uppercase alias trips the unreachability check.

Review Minor #2: changeset retyped Changed -> Added. Nine wholly new
supported response_language values are an addition under Keep a
Changelog, not a modification of existing behavior.

* test(#2530): check alias collisions on the catalog, not its source

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Rezolv <dave@sienkowski.com>
2026-07-31 21:01:50 -04:00
Tom Boucher
3f6b063fbb chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans

Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers
iterate declared lanes instead of hand-authored per-CLI bash.

Five additive descriptor amendments, each forced by a lane that ships today:
- LaneHandler gains 'opencode' — the lane rebuilds its review from assistant
  text parts of a --format json stream; a plain stdout copy re-breaks #1936.
- modelConfigKey — antigravity's key is review.models.agy, not .antigravity,
  so resolving by slug silently dropped a configured model.
- defaultHost/fallbackModel — Phase 4 federated every *_host with a default of
  empty string; the real fallback only existed in the bash.
- args becomes an argv template with a closed four-placeholder vocabulary.
  Positional splicing produced 'codex --model M -o F exec --ephemeral', which
  is not a valid invocation: codex injects in the middle, twice.
- kimi-code lane, with the bounded command-capability probe (needle
  --output-format) that tells Kimi Code from the legacy python kimi-cli.

Parity gate re-pointed: the workflow-text families it scanned are the text this
phase deletes, so they are replaced by descriptor-to-registry parity plus an
anti-parity check that no bespoke leg returns.

jq, curl and external timeout/gtimeout all drop out of the review path.

Refs #2782

* chore(#2799): add review-lane query surface and widen the manifest vocabulary

Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow
loops over, projects all twelve lanes into their capability manifests, and
widens capability-validator for the amendments.

opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's
own admission rule: one lane, justified by a documented upstream defect data
cannot express (#1936 — the agent can end its turn with zero output tokens and
--format default then drops the assistant text entirely).

Two bugs caught by an end-to-end stub run and fixed here:
- loadConfigResolved returns a provenance wrapper, not the config; using it
  directly resolved every key to undefined, which reads as 'nothing
  configured' and silently dropped every model override.
- hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced
  with a PATH scan that spawns nothing at all.

Refs #2782

* chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews

Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved
lanes, and renders REVIEWS.md sections from each lane's declared
reviewsSection instead of thirteen hardcoded headings. review.md drops from
1104 lines to 507 (61KB to 28.7KB).

Parity gate re-pointed, as agreed: the leg-marker and section-heading families
scanned exactly the text this phase deletes, so they are replaced by
descriptor-to-registry parity in both directions, plus an anti-parity check
that fires if a bespoke leg is ever re-added. Enum, emitting sites and the
Object.keys lock moved together.

The budget-trim helper is hoisted out of the Ollama leg: it was always
lane-agnostic, and any lane may now declare a promptBudgetKey.

Refs #2782

* feat(#2799): bind the consented egress host and re-verify it at invocation

Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3
but was not implemented: ConsentRecord had no host field and nothing in the
tree bound one, so this phase's rule-4 comparison had no baseline.

ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design:
isValidConsentRecord does not require it, so every record already on disk
stays valid and no re-consent storm fires (D4 rule 5). It is deliberately
excluded from disclosureSignature — the loader has no config resolver, so
folding a config-derived value in would make loader and lifecycle compute
different signatures for the same manifest and re-prompt forever.

Install resolves hostConfigKey (falling back to the lane's declared
defaultHost, which is what the invocation path uses) and records it.
Invocation re-resolves and blocks on mismatch rather than silently
redirecting. Absence allows: no record, or a record predating the field,
means nothing to compare — denying there would break every existing
local-model user on upgrade.

Refs #2782

* test(#2799): cover the resolver, runner and handlers; retarget the parity suites

Adds the golden invocation-plan table (one row per shipped lane, derived from
the bash legs rather than the descriptor types) plus runner coverage for the
probe, empty-output policy, the three handlers and the egress check.

Retargets the existing suites onto the new contract: descriptor-to-registry
parity, the anti-parity check, the opencode handler, and the twelfth lane.

Two corrections found by running them:
- modelConfigKey was required; that breaks D4 rule 2, since a reviewer
  manifest authored before this phase would fail validation on upgrade. It is
  optional, read as null when absent.
- the antigravity non-zero-exit test pre-seeded the transcript, which asserted
  that a STALE entry leaks through — the exact bug the watermark prevents. The
  spawn now appends, as the real tool does.

Refs #2782

* fix(#2799): restore agy --add-dir and the self-report prompt in the handler

Retargeting the three legacy reviewer suites off the deleted bash surfaced two
real regressions in the port, both #2176:

- --add-dir was dropped. Without it agy's permission context never receives the
  cwd repo, so the agent anchors on its own scratch dir and reviews the plan
  text in isolation — the exact failure the Review Instructions forbid. It is
  capability-probed, because an older agy rejects the unknown flag outright and
  a lane that fails to start is worse than one running on the prompt anchor.
- the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS
  self-report, which is what makes a blind review distinguishable from a
  grounded one. antigravity now builds its own prompt variant.

Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned
model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own
log is the only evidence that anything failed.

The three suites now assert against the plan and the handler instead of
matching fence text, so they no longer need allow-test-rule exemptions.

Refs #2782

* docs(#2799): document the declared lanes, the new flag, and dropped prerequisites

COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph,
which is now false: no lane requires jq, curl or an external timeout. Adds the
changed-egress-destination behavior, since a blocked lane is something a user
can hit.

CONFIGURATION.md records that the model config key is declared per lane rather
than derived from the flag — antigravity's is review.models.agy — and adds
review.models.kimi-code.

reviewer-instances.md now routes an instance through its lane's single
invocation seam instead of a copied per-adapter bash block, which is what lets
a cross-cutting fix reach instances for free. That required implementing the
--model/--agent/--as flags it documents; --model re-resolves through the lane's
argv template rather than splicing, so the flag lands where the lane declares
it rather than ahead of a subcommand.

CONTEXT.md glossary gains both new modules.

Refs #2782

* chore(#2799): drop the stale emitted-drift acknowledgment

The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md.
That file now shrinks by ~32KB and every emitted hash that moved is
attributable to this diff, so the ack no longer explains anything. Removing
the last entry means removing the file: its presence is the alarm, and an
empty one signals nothing.

Verified by deleting it and re-running the attribution and provenance gates
plus lint:ci — all green without it.

Refs #2782

* docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782

Five additive amendments, each forced by a lane that ships today, plus two
corrections the phase had to make rather than work around:

- D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so
  this phase's rule-4 comparison had no baseline. Recorded because an ADR
  asserting a rule was delivered is exactly what stops a later phase checking.
- The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families
  scanned the text this phase deletes.

Also records that D7's 'skip the probe where no bounding mechanism exists'
carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded
on every stock macOS host, which ships neither timeout nor gtimeout.

Refs #2782

* fix(#2799): close four defects found by adversarial review

Two confirmed bugs, both reproduced before fixing:

- resolveLanePlan was not total. An openai-http lane with a missing or
  non-object invoke dereferenced inv.hostConfigKey and threw, contradicting
  the module's own documented contract; the spawn branch guarded correctly and
  the http branch did not. The CLI seam resolves every selected lane in one
  map, so one malformed overlay manifest would have aborted the whole review
  rather than dropping its own lane. Guarded, plus a per-lane try/catch at the
  seam so a throw can never take down siblings.
- A reviewer-instance model was silently dropped for any lane declaring
  modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates
  that cli is a known slug but never that the slug accepts a model, so a user
  could configure one, get a clean run, and never learn a different model
  reviewed their plan. Now warns explicitly.

Two hardening fixes:

- The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in
  the resolver rather than inherited from a validator that does not run on this
  path — the module documents itself as the overlay-manifest trust boundary, so
  it should not depend on someone else having checked.
- normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses
  with an empty hostname, so it became 'localhost://11434' and was compared and
  requested as if real. An empty hostname now means not-a-URL.

Also documents the one gap that cannot be closed here: the antigravity
watermark is keyed by workspace, so two concurrent reviews of the same repo
share a transcript. agy exposes no per-invocation id to filter on, so the
handler now states which half of its never-stale guarantee actually holds.

Refs #2782

* test(#2799): retarget the remaining eight review.md-asserting suites

The remote runner found 37 failures the local sweep missed (it hit the shell's
two-minute cap before reaching these). All eight extract per-CLI bash from
review.md that this phase deletes; each protects a real invariant, so each is
retargeted onto the plan, the runner or the handler rather than removed.

Three real defects surfaced by doing so:

- effort args never reached ANY lane. model-resolver.cjs exports no
  resolveExecution, so effortFor silently returned [] every time. Restored by
  calling the same bounded resolve-execution query the bash legs used — and
  NOT with --raw, which prints the resolved effort rather than the picked
  field, so claude got 'low' instead of '--effort low'.
- the timeout guidance lost 'a silent empty output is a timeout kill, not a
  crash' — the operator note that exists because of the Codex 0xc0000142
  misdiagnosis. Restored.
- the opencode handler dropped EMPTY assistant text parts. The shipped jq was
  , and  only substitutes for false/null — an empty
  string is truthy in jq and contributed a blank line. Found by a property
  test shrinking to ['', ''].

The opencode property suite no longer spawns jq at all, which deletes the
#2099 hang mechanism it was architected around rather than mitigating it.

Refs #2782

* fix(#2799): register the two new generated modules, and untrack them

The remote runner caught build output committed to git. Both new modules
compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated
that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own
review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each
bin/lib/*.cjs is linted xor ignored according to migration state" failed.

Registered both in .gitignore and eslint.config.mjs alongside the Phase 1
module, and dropped them from the index. Nothing about the shipped behaviour
changes; the artifacts are rebuilt by build:lib.

This is the new-.cts-module registration ripple, and it is the one part of it I
had not completed - the CONTEXT.md glossary and the inventory manifest were
already done.

Refs #2782

* chore(#2799): backfill changeset pr number to 2861

* chore(#2799): backfill changeset pr number to 2861

---------

Co-authored-by: Test <test@example.com>
2026-07-30 12:48:06 -04:00
Tom Boucher
0408276791 chore(#2797): federate reviewer config keys off the central schema (#2841)
* chore(#2797): federate reviewer config keys off the central schema

Phase 4 of epic #2782 (ADR-2782 D9, config half). Runs AFTER 5a per the
ADR's swap amendment: a federated config slice lives inside a
capabilities/<id>/capability.json, and three of the five key families had
no capability directory until 5a created them.

Four key families move to the lanes that use them; the central-schema
removal and the federated addition land in this one commit because the
exclusivity invariant fails the build on a key present in both.
review.max_prompt_tokens, review.default_reviewers and
review.reviewer_instances describe policy ACROSS lanes and stay central.

Two things the issue did not name, both found while building it:

1. THE EXCLUSIVITY GATE WAS BLIND TO PATTERNS. It compared federated keys
   against manifest.validKeys only, and two of the four families
   (review.models.<slug>, review.max_prompt_tokens_per_reviewer.<slug>)
   were pattern-backed. That is not cosmetic: isCentralConfigKey consults
   those patterns and mergeFederatedConfig skips every key for which it
   returns true, so declaring a slice while the pattern survived would
   have shipped an INERT slice behind a green gate — the exact
   half-migrated shape the invariant exists to prevent. The gate now
   loads the patterns from the same manifest the runtime reads.

2. AN UNSET PER-LANE BUDGET NOW RESOLVES TO 0, NOT NOT-FOUND, because a
   federated key always resolves to its declared default. The three
   fallback guards in review.md checked only empty-or-"null", so a user
   who set the GLOBAL review.max_prompt_tokens would have silently lost
   trimming on the HTTP lanes. The guards now treat 0 as unset.

D9 says review.models.<slug> is owned by "the lane whose slug it names".
That is false for one lane: the shipped key is review.models.agy while
the slug is antigravity. Ownership follows the lane; the key name is
preserved, because renaming would break every config that sets it.

Existing tests updated rather than left asserting the old world:
config-get on a cleared federated key yields empty instead of
not-found (what #2046 actually protects — never persisting the literal
"null" — is unchanged and still asserted); the config-schema dynamic
pattern representative moves to reviewer_instances; the
prototype-pollution guard case moves to a surviving dynamic prefix so
alert #26 keeps its coverage, with a new case asserting the old key is
now rejected earlier; and Phase 2's harvest-widening inertness assertion
becomes an ownership assertion, since Phase 4 is what consumes it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2797): use a -1 sentinel so an explicit per-lane budget of 0 survives

A federated config key always resolves to its declared default, so an
unset per-lane prompt budget needed a value the workflow could treat as
'not configured'. The first cut used 0 — which is wrong: 0 is already a
LEGITIMATE per-lane budget meaning 'do not trim this lane' (the
early-return guard in prepare_trimmed_prompt_for_reviewer). Treating it
as unset would have silently switched a user who deliberately disabled
trimming for one lane onto the global budget.

The sentinel is now -1, which is not a valid token budget, so all three
states stay distinguishable: unset falls back to global, an explicit 0
disables trimming for that lane, and an explicit N is used. Locked by
three CLI round-trip tests.

Surfaced by the isolated security reviewer before it crashed mid-run;
verified independently against the shipped trim guard rather than taken
on trust.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2797): update central-registration assertions and stay under the review.md cap

The remote runner caught both; my local sweep missed the files.

1. tests/plan-review-convergence.test.cjs asserted the three local-server
   host keys are in VALID_CONFIG_KEYS. They are federated to their lane
   capabilities now, and the exclusivity invariant forbids a key living
   in both places. What #2306-local actually protects is that config-set
   ACCEPTS them, so that is what is asserted — via isValidConfigKey, the
   predicate config-set itself uses, which spans central and federated.
   A second assertion pins federated ownership, so a silent reversion
   back to the central schema fails too.

2. review.md exceeded the LARGE tier hard cap (62583 > 61440). That cap
   is a red line, not a budget to raise. The three per-lane budget guard
   comments were near-identical; condensed to one terse line each. 61371
   bytes, 69 to spare. Real extraction to workflows/review/modes/ is
   Phase 5b/6 work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2797): fail closed on a broken config-schema manifest; reconcile stale docs

Isolated security review findings.

MAJOR — loadCentralConfigPatterns failed OPEN. It swallowed a JSON parse
error and returned [], while its sibling loadCentralConfigKeys, reading
the SAME file, writes to stderr and throws ExitError(1) on that identical
failure class. Fail-open here defeats the gate this function exists to
feed: with zero patterns, validateCrossCapability's pattern-collision
check silently passes and an inert federated slice ships green. It was
masked in the one production call site only because loadCentralConfigKeys
runs first against the same path — a coincidence of ordering, not a
guarantee, and this function is exported and called standalone. The two
now share a contract: ENOENT is the legitimate absent case, anything else
throws loudly. A single unparseable PATTERN is still skipped, which
degrades to "checked less" rather than blocking every build. The branch
had zero coverage; it now has two tests (malformed JSON, EISDIR).

MINOR — docs/CONFIGURATION.md still listed review.models.qwen and
review.models.cursor as settable, ~770 lines below this PR's own new
Ownership section. Those lanes take no model flag, so they declare no
model key and config-set now rejects them. Rows removed; the missing
review.models.agy row added; the per-reviewer budget row corrected to
name only the lanes that own a budget key, and to document that a
per-lane 0 disables trimming for that lane.

Also fixes a shadowed "raw" binding introduced by the fail-closed change,
which made the generator unrequirable — caught immediately by its own
--check.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2797): backfill changeset pr number to 2841

* fix(#2452): make the base-ref mutation test hermetic against leaked GIT_* env

tests/mutation-workflow-base-ref.test.cjs fails on PR branches while next
stays green, and it is currently blocking at least three unrelated PRs
(#2841, #2832, #2827) with:

  error: invalid object 100644 <sha> for 'base-N.txt'
  error: Error building trees

The existing loop comment attributes this to `git add .` rehashing O(n^2)
blobs "before the object write had landed" and works around it by staging
one path per iteration. That is not the cause: sequential execFileSync
calls cannot race each other's object writes, and the failure persisted
after that change — it simply moved to a lower commit index.

The cause is that the git() helper inherited the runner's environment. A
leaked GIT_INDEX_FILE makes `git add` write into a DIFFERENT repository's
index; GIT_OBJECT_DIRECTORY / GIT_ALTERNATE_OBJECT_DIRECTORIES send the
blob to another object store; GIT_DIR / GIT_WORK_TREE redirect the whole
operation. In every case `git commit` then cannot resolve a blob it just
staged, which is precisely the error above.

Verified by negative control: with GIT_DIR exported, this test fails on
the unfixed helper (the git commands operate on the wrong repository
entirely); with the helper stripping GIT_* it passes. The single-path
staging is kept — it is genuinely less work — but it is no longer load
bearing.

Found while shipping #2797. Fixed in place rather than deferred: it is a
defect surfaced during the work, and it is blocking other contributors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2452): build the base-advance commits empty, removing the lost-object class

The base-ref guard has been failing in CI with:

  error: invalid object 100644 <sha> for 'base-N.txt'
  error: Error building trees

It is currently red on at least three unrelated PRs (#2841, #2832,
#2827) while next stays green.

Two theories have now been tried and neither held. #1881 blamed `git
add .` rehashing O(n^2) blobs and switched to staging one path per
iteration; the failure moved from commit 32 to commit 25 and carried on.
The preceding commit here made the git helper hermetic against leaked
GIT_* environment — that IS a real vulnerability (with GIT_DIR exported
the helper operates on the wrong repository entirely, proven by negative
control) but it produces a different error than CI reports, so it is not
demonstrably the cause either.

Neither trigger reproduces off-CI, so this stops guessing at the trigger
and removes the failure CLASS instead. The loop needs base-branch DEPTH
and nothing else: no assertion reads these commits' contents, and
base-side files cannot appear in `origin/base...HEAD` regardless.
`--allow-empty` writes no blob and no tree, so there is no object for the
index to reference and lose. It is also far less work than 60
write+hash+index cycles.

The guard still proves its mechanism: the test asserts that a --depth=1
base fetch FAILS and a full fetch resolves, so a broken topology would
surface immediately rather than passing vacuously.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Test <test@example.com>
2026-07-30 08:52:56 -04:00
Tom Boucher
8b44a0da43 chore(#2794): single-source the reviewer invocation contract + parity assertion (#2820)
* chore(#2794): single-source the reviewer invocation contract

Phase 1 of epic #2782 (ADR-2782). Introduces one core descriptor table as
the declared contract for all 11 cross-AI reviewer lanes, and the
DEFECT.GENERATIVE-FIX parity assertion the roster has never had.

The lane contract lived in three unrelated surfaces — the roster, ~640
lines of hand-authored per-CLI bash in invoke_reviewers, and the
write_reviews section headings — so cross-cutting fixes landed per-leg
(#2494 and #2605 were the same empty-output defect filed twice).

- src/review-lane-descriptor.cts: frozen table declaring per lane the
  slug, flags, probe, invoke shape, timeout floor, empty-output policy,
  REVIEWS.md section, evidence class, required binaries, prompt-budget
  key and handler. Field names track ADR-2782 D1/D2/D6/D7 verbatim so
  Phase 2 harvests the shape with no translation layer. It declares;
  it does not execute — invoke_reviewers iterates in Phase 5b.
- checkReviewerLaneParity: bidirectional parity across descriptor,
  roster, invoke_reviewers legs and write_reviews sections. Forward-only
  would miss the failure it exists to catch (#2718 added a leg, #2781
  was the drift). ADR-1517 instance headings are exempt per D8.
- Legs carry an explicit <!-- reviewer-lane: slug --> marker; five
  non-lane bold labels share the bold-then-fence shape a heuristic
  matcher would key on.
- ADR-2782 D4: an explicitly-flagged reviewer that cannot run is now an
  error in both the core module and the workflow prose that mirrors it.
  A code-only change would be unobservable — the module has no
  production caller; the workflow narrates the policy. Discovery paths
  (--all, review.default_reviewers) stay lenient.
- Fixes the qwen leg, the last one discarding stderr to /dev/null.

Two ADR-2782 D2 vocabulary widenings were forced by surveying the
shipped legs: promptChannel 'none' (CodeRabbit is fed no prompt) and
outputChannel 'file-arg' (Codex writes via -o and discards stdout,
#1698). Both are additive and closed; Phase 2 owns the validator.

Closes #2690

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2794): make the parity checker total and pin the lane slug grammar

Findings from the orthogonal review passes.

Spec axis — the module claimed its vocabulary tracked ADR-2782 D1/D2
"verbatim" while diverging in three undisclosed ways, which is the
translation layer Phase 2 was supposed to be spared:
- `transport` moves from `invoke.transport` to the LANE level, a sibling
  of `probe`/`invoke`, exactly as D1's manifest example places it. The
  nested form read better as a TS discriminated union; the union is now
  discriminated at the lane level instead, which costs nothing.
- The header and the CONTEXT.md glossary now enumerate all FOUR
  widenings (adding `outputArg` and `flags[]`), not two.

Standards axis — CLAUDE.md requires a fast-check property test for a
parser, and `checkReviewerLaneParity` parses markdown for markers and
headings. Adding one found two real defects that the hand-written
matrix missed:
- NOT TOTAL: a malformed descriptor entry threw on `lane.flags`
  iteration, contradicting the module's own "never throws" claim. Every
  field is now narrowed from `unknown` at the trust boundary and
  reported as MALFORMED_LANE / INVALID_SLUG. This matters because
  Phase 2 feeds this function third-party overlay data, and a parity
  gate that crashes is indistinguishable from one never run.
- SILENT GRAMMAR MISMATCH: LEG_MARKER_RE captures only [a-z0-9_-], so a
  slug outside that class was unmatchable — its marker could be present
  and correct and the scan would still report LEG_MARKER_MISSING
  forever. LANE_SLUG_RE now pins the grammar and a violating slug is
  reported INVALID_SLUG. A loud named violation beats a silent miss.

Generators are document-shaped, not writer-seeded (CONTRIBUTING #2371):
seeding from the module's own matchers could only produce documents
those matchers already recognize.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2794): register the new bin/lib module in the ESLint ignore list

The remote runner caught this; lint:ci did not, because the invariant
lives in the test suite rather than the lint chain:

  tests/repo-invariants.test.cjs
  "each bin/lib/*.cjs is linted xor ignored according to migration state"
  -> tsc-generated bin/lib modules not yet added to ESLint ignore list:
     review-lane-descriptor.cjs

Adding a src/*.cts module ripples to six surfaces (.gitignore, the
ESLint ignore list, docs/INVENTORY-MANIFEST.json, the CONTEXT.md
glossary, the capability/inventory manifests, and any size baseline).
The other five were covered; this was the miss.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2794): amend ADR-2782 D1/D2/D8 with the vocabulary Phase 1 surfaced

Building the Phase 1 descriptor table against all eleven shipped legs is
the first time every lane's contract was written in one place, and it
surfaced four cases the ADR's original survey did not cover. Amending
the design lock rather than diverging from it, so Phase 2 (#2795)
implements the manifest validator against the amended vocabulary instead
of rediscovering the gaps.

All four are additive widenings of closed enums; no decision reverses:

- D2 promptChannel gains `none` — coderabbit is fed no prompt at all, it
  reviews the working-tree diff.
- D2 outputChannel gains `file-arg` — the ADR called a file-writing lane
  a shape a real CLI *could* take; codex already is one, writing via
  -o/--output-last-message and discarding stdout (#1698).
- D2 gains `outputArg`, required iff file-arg — knowing the review lands
  in a file is useless without the argument naming it.
- D1 `flag` becomes `flags[]` and D8's uniqueness flattens across lanes —
  antigravity is selected by both --antigravity and --agy, which a
  single-valued field cannot express.

This is the same evidence path that produced the openai-http transport:
the vocabulary widens on a lane that exists, under review, never on
speculation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2794): backfill changeset pr number to 2820

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 07:31:32 -04:00
Rezolv
16e59d0db5 fix(#2691): repair seven dangling references in the ADR corpus and contributor docs (#2692)
* fix(#2691): repair five dangling references in the ADR corpus and contributor docs

Found by the 2026-07-24 ADR corpus audit; each mechanism re-reproduced live
against next @ 3eb1cede before filing.

1. docs/adr/1239-gsd-embeddable-orchestration-engine.md linked the
   host-integration capability matrix as `reference/...` from inside
   docs/adr/, which resolves to the nonexistent docs/adr/reference/.
   Three occurrences (the #2584 amendment added two after the audit).
   All now `../reference/...`.

2. src/plan-drift-guard.cts cited docs/adr/0022-source-grounding-drift-guard.md,
   a path that has never existed (git log --all --diff-filter=A returns
   nothing). Corrected to docs/adr/22-plan-drift-guard.md. Comments survive
   tsc and ADR-457 builds at publish, so the bad citation shipped to users --
   verified by grepping the compiled gsd-core/bin/lib/plan-drift-guard.cjs.

3. CONTRIBUTING.md and docs/contributor-standards.md illustrated the ADR
   naming convention with issue #3485 -- a pre-rename number from the
   predecessor repo (get-shit-done-redux) that does not resolve in
   open-gsd/gsd-core. It is dangling, not invented: ADR filenames 3524 and
   3660 show pre-rename numbering reached the 3000s. The worked example now
   uses #2264, which resolves; the one genuinely historical mention is
   annotated rather than rewritten.

4. docs/adr/857-capability-system.md's H1 carried a stale [Proposed] bracket
   contradicting its "Accepted -- ratified 2026-07-17" Status field.
   gen-adr-index.cjs:213 strips the bracket rather than comparing it, so the
   contradiction was invisible to the gate; it is also the only in-repo
   consumer, so removal leaves the rendered index byte-identical.

5. gen-adr-index.cjs's back-link comment still described ADR-857 as Proposed
   and its claim over ADR-0011/ADR-58 as a supersession. Both were restated
   at ratification (the claim became Subsumes; the reciprocals were added).

Regression coverage folds into tests/adr-index-gate.test.cjs rather than a
new bug-* file (lint-regression-test-names): four cases covering link
resolution, H1-bracket-vs-Status agreement, the plan-drift-guard citation,
and the naming worked example. All four fail at the pre-fix tree.

No behavior change. lint:ci exit 0; adr-index-gate 35/35; index regenerates
unchanged (67 ADRs).

* chore(#2691): add the required pr: field to the changeset fragment

docs-lint rejected the fragment with fail_malformed_fragment / missing_pr: the
frontmatter needs both `type:` and `pr:`. The fragment was hand-authored before
the PR existed, so it carried only `type:`.

Note for future work: neither lint:docs nor lint:changeset is part of the
lint:ci chain, so a green lint:ci does not cover these two CI checks. Both were
run directly before this push:
  ok docs-lint: ok_no_triggering_fragments
  ok changeset-lint: ok_fragment_present

* fix(#2691): drop the false branch clause and repair two more matrix links

Review round 2 on #2692, both blocking findings.

F1: the worked example at CONTRIBUTING.md:101 asserted
'on branch docs/2264-golden-parity-redesign'. That branch never existed --
the ADR file carries the epic number (#2264) while the branch and commit
carry the Phase-0 sub-issue number (#2265, PR #2270, branch
docs/2265-golden-parity-adr). #2264 is therefore the one ADR in the corpus
where filename and branch numbers deliberately disagree, making it the worst
available illustration of 'the issue number becomes your prefix and your
branch'. The branch clause is dropped; the surviving claim is verified
(#2264 is open-and-approved with approved-enhancement, and the file is
2264-golden-parity-redesign.md).

F3: docs/how-to/install-on-your-runtime.md:460 and :476 linked the
host-integration capability matrix as a bare filename from inside
docs/how-to/, which resolves to the nonexistent
docs/how-to/host-integration-capability-matrix.md. Same defect class as
repair #1 in this PR. Both now use ../reference/..., matching the form
already used by the sibling add-or-update-a-host-integration.md:167.

* fix(#2691): take the review minors -- anchors, bracket, ADR path, token order

Review round 2 on #2692, non-blocking findings.

F4: ADR-1239:198,206 read '[capability matrix §codex](...matrix.md)' -- the
link text promised a section, the target carried no fragment. '## codex'
exists at docs/reference/host-integration-capability-matrix.md:85, and the
repo already uses that form (#zcode, #pi in the how-to).

F5: docs/adr/857-capability-system.md:1 had its stale [Proposed] bracket
dropped rather than corrected, making 857 the only ratified ADR with no H1
bracket while four other Accepted ADRs carry one. Restored as [Accepted],
which satisfies the H1-vs-Status invariant and preserves consistency. No
recorded rule mandates the bracket, so this is style, not contract.

F7: docs/CONFIGURATION.md:808 cited adr/1244-runtime-capability-registry-
overlay.md; the actual file is adr/1244-capability-ecosystem.md. Same defect
class as repair #2, one directory over.

F8: test 33 resolved the Status token with STATUS_TOKENS.find(), which
matches by array order rather than by position in the line. Since
STATUS_TOKENS[0] === 'Accepted', a future '- **Status:** Superseded by ADR-X
(was Accepted ...)' paired with an H1 [Superseded] would have reported a
false mismatch. Now resolved by earliest index in the line. No ADR has that
shape today, so this is latent; ADR-857's two-token Status line resolves to
'Accepted' under both the old and new rule.

Changeset updated: five repairs -> seven, adding the two how-to matrix links
and the CONFIGURATION.md ADR path, plus the (#2691) issue backlink that 46
of the other 47 fragments carry.

Not addressed here: F2 (widening the guard to resolve anchors and walk docs/
recursively) is filed separately as #2704, approved-enhancement.

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
2026-07-28 18:02:11 -04:00
Tom Boucher
3eb1cede26 fix(#1880): distinguish a corrupt config from an absent one (epic #1879 Phase 1) (#2688)
* test(#1880): prove corrupt config is indistinguishable from absent

Failing-first. Encodes the issue's runtime repro: a trailing comma in
.planning/config.json currently yields source:builtin-defaults with
degraded:false - byte-identical to the file not existing - and the user's
entire configuration is silently discarded.

Asserts on the typed surface (CONFIG_REASON, _warnedUnusableConfig) rather
than diagnostic prose, per the ADR-1411 amendment's test-methodology clause
and CONTRIBUTING.md's raw-text-matching rule. IO failure is injected by
monkeypatching fs.readFileSync and restoring in t.after(), never chmod 0o000
(root bypasses mode bits).

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1880): distinguish a corrupt config from an absent one

loadConfigResolved wrapped the read, the JSON.parse and the entire config
build in one try with one catch, so ENOENT, EACCES and SyntaxError all fell
through to the same defaults and the branches returned degraded:false -
actively asserting health over discarded configuration. A single trailing
comma in .planning/config.json silently replaced the user's whole config,
reporting source:builtin-defaults degraded:false, byte-identical to having
no config file at all.

ConfigResolution now carries a machine-readable reason. Genuine absence keeps
degraded:false / not_configured; a file that exists but cannot be used sets
degraded:true with config_unparseable or config_unreadable. The same split
applies to the root config and to ~/.gsd/defaults.json.

Control flow is deliberately unchanged. preflight_check reports cyclomatic
141 / cognitive 196 and 93 dependents on this function, with the guidance
that small edits beat one big one, so faults are CAPTURED at the existing
read sites and stamped onto the returns rather than the try/catch being
restructured.

Also carries the ADR-1411 amendment's wiring clause: loadConfig returns
.config alone to ~51 call sites and would never see the new field, so an
unusable file emits a deduplicated stderr diagnostic keyed on resolved path
plus errno. Without it the reason would be an unreachable field and the user
whose config was discarded would still get no signal - the actual defect.

Registers the config-loader seam in lint-resolution-provenance, which until
now guarded only agent-skills.

Caller audit: ConfigResolution.degraded has exactly one consumer outside this
module, cmdAgentSkills (src/init.cts:2259), which destructures
{config, source, degraded} - adding a field does not break it. Its --json IR
now reports degraded:true for a corrupt config, which is the intended fix and
the one observable behavior change.

Closes #1880

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1880): degrade when any config on the path is unusable, not just the last

Two defects found by isolated adversarial review of the first cut.

BLOCKER: the success-path return did not consult configFault. A corrupt ROOT
config whose workstream override happened to parse returned degraded:false /
reason:resolved - the root's settings silently dropped, which is the exact
failure this issue closes, reappearing for any project using workstreams. The
stderr diagnostic fired, so the out-of-band half worked while the in-band half
reported a clean resolve; a --json consumer saw health.

MAJOR: reason was derived from Object.keys(parsed) - the root+workstream MERGE
- so an empty workstream file inheriting a non-empty root reported resolved
despite carrying no settings. Emptiness is now judged on the file actually
read, snapshotted before normalizeLegacyKeys mutates it.

Also: corrects the ConfigResolution JSDoc, which still described the pre-#1880
degraded contract; adds a fast-check property asserting a PRESENT file is
never reported not_configured whatever its bytes (CONTRIBUTING.md parser
rule); and asserts the literal enum values so the provenance lint's
configured_empty/not_configured markers check real assertions rather than
incidental prose in test titles.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1880): reject valid JSON that is not a config object at the read seam

The fast-check property added in the previous commit failed on both node
lanes: a config.json containing 0, "str", [], null or true is valid JSON, so
it parsed "ok", then threw downstream in normalizeLegacyKeys, and the outer
catch reported not_configured - a PRESENT file reported as absent, which is
precisely the collapse this issue exists to close. The property asserts a
present file is never not_configured, and it caught it.

_readConfigFile now validates shape, not just parseability (ADR-227: check
the semantic shape at a trust boundary, not merely the type). A non-object
JSON document is an unusable config, reported config_unparseable.

Adds named regression cases for each non-object form alongside the property,
so the class is documented and not only randomly sampled.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#1880): backfill changeset pr number (pr:0 -> 2688)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2452): record a fetch-time shallow failure instead of crashing

This guard failed CI on ubuntu-24 while passing on ubuntu-22 and
windows-24 for the same commit, and passed on other PRs. Not a flake and not
caused by the change under test - a real fragility in the test.

runnerDiff ran the base fetch OUTSIDE its try and only guarded the diff, so it
assumed the failure mode is always 'fetch succeeds, diff reports no merge
base'. At a shallow boundary that lands short of the merge base, git can
instead fail during the FETCH ('unable to parse commit' - the boundary
commit's parent is not available). Which stage git fails at is version and
transport dependent, so on some runners the error escaped runnerDiff and
crashed the test rather than being recorded as the ok:false the assertions
expect. Both stages mean the same thing for what this guard protects: a
shallow base ref cannot resolve the three-dot diff.

Also drops two assert.match calls against git's stderr prose. 'no merge base'
and 'unable to parse commit' are the same condition reported at different
stages, and CONTRIBUTING prohibits raw text matching on subprocess output.
The typed outcome (ok === false) is the contract; the tests now assert that
plus the presence of a cause.

Found while investigating the red lane on #2688; fixed here per the no-defer
rule rather than filed.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 00:09:13 -04:00
Tom Boucher
46ba02acde feat(#2630): phase-estimation module, smart-zone config key, and cli verbs (#2661)
* feat(#2630): add phase-estimation module, smart-zone config key, and cli verbs

* fix(#2630): document smart_zone_tokens, refresh golden fixtures, fix null-proto property assertions

* fix(#2630): align smart_zone_tokens write/read validation and harden estimation tests

* chore(#2630): backfill changeset pr to 2661
2026-07-26 01:42:47 -04:00
Tom Boucher
6ad30f74b6 feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler.

harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run.

Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard.

Closes #2627

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:50:20 -04:00
Rezolv
155c08facf docs(#2197): drop --validate docs for /gsd-plan-phase and /gsd-execute-phase (#2574)
* docs(#2197): drop --validate docs for /gsd-plan-phase and /gsd-execute-phase

These two commands never parse --validate (silent no-op); the flag is
real only for /gsd-quick. Remove the false flag-table rows and CLI
examples across COMMANDS.md and the how-to guides (en + ja-JP/zh-CN/
ko-KR/pt-BR mirrors), and correct the manager.flags.execute example
from --validate to --cross-ai (a flag execute-phase actually parses).
/gsd-quick's real --validate docs are left untouched.

Ref #2197

* docs(#2197): add changeset for --validate docs removal

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
2026-07-24 12:47:53 -04:00
Tom Boucher
09b535ac00 feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort
ADR-1239 gains a ninth negotiated axis, effortSurface (argv | none), declaring how
a host accepts reasoning effort. ADR-443 is amended in the same change because its
recorded deferral is what the axis resolves: its Unblock condition offered paths
(a) and (b) and stated the choice was 'a maintainer call this file records but does
not make'. Path (a) is selected and satisfied here.

Before this, effort reached a runtime only through install-time channels
(EFFORT_RENDERING's frontmatter/api), so reviewer CLIs spawned as subprocesses
silently inherited whatever effort sat in the user's own global CLI config. The
review lane now resolves one universal effort through the ADR-443 cascade and
renders it per host through the negotiated descriptor.

Every per-host value is documentation-sourced, never inferred:
- claude   argv  -- verified via 'claude --help' (--effort <level>)
- opencode argv  -- verified via 'opencode run --help' (--variant)
- codex    argv  -- codex-rs/exec/src/cli.rs: model_reasoning_effort is NOT a CLI
                    flag (config.toml key only), so the global -c override is the
                    only argv route
- 15 hosts undocumented -- their docs state no reasoning setting; the sentinel
                    fails closed rather than inheriting a profile baseline

No config-file vocabulary member: the only host that ever had one (Gemini CLI's
thinkingConfig) was removed as a sunset runtime by 8f2ebbe9b (#1928, PR #1996),
and neither Antigravity CLI nor ZCode documents a reasoning setting.

Closes #2481

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 19:10:22 -04:00
Tom Boucher
455ad49ae3 feat(#2296): config-gated provider escalation on quota-exceeded (#2458)
* test(#2296): failing-first coverage for provider escalation on quota-exceeded

Covers the provider-escalation ladder layered onto EXEC.CLASSIFY: back-compat
(no escalation block without --failure-class), cap boundaries at
min(max_escalations, list length) at limit-1/limit/limit+1, opt-in gating,
malformed/hostile provider_escalation config, the --failure-class CLI negative
matrix, config-key registration, and a fast-check budget-limit property.

Red until the resolver, CLI flag, and manifest key land.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#2296): config-gated provider escalation on quota-exceeded

The dynamic_routing tier ladder escalates within one provider, which does not
help when that provider is what ran out of quota. Add an opt-in provider ladder
layered on the existing EXEC.CLASSIFY seam.

- model-resolver: resolveProviderEscalation walks dynamic_routing.provider_escalation
  capped at min(max_escalations, list length), reporting from/to/attempted/exhausted.
  Invalid entries are dropped (ADR 227 shape validation). Stays a leaf module —
  the quota-class policy decision is the caller's, per the CONTEXT.md contract.
- agent-command-router: export a frozen AGENT_FAILURE_CLASSES so the new CLI
  validator cannot drift from the classifier that produces the values.
- resolve-execution: --failure-class flag; emits an escalation block ONLY when
  passed, so the existing JSON contract is byte-identical for every caller.
- config-schema.manifest: register dynamic_routing.provider_escalation.
- execute-phase step 7.1: auto-escalate, honor Retry-After, fail loudly naming
  every model tried once the ladder is spent.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2296): extract quota recovery to a reference fragment; regen goldens

The step 7.1a addition pushed gsd-core/workflows/execute-phase.md from 93390 to
95111 LF bytes, past the frozen ADR-857 Phase 6 ceiling (hard <93600, margin
<=93400) asserted by tests/fix-2285-claude-orchestration-wiring.test.cjs. The
base sat 10 bytes under the margin, so no inline wording would have fit.

That gate's own rationale is that optional-feature detail belongs in a fragment,
not the host loop. Moved BOTH the new provider-escalation branch and the
pre-existing manual recovery prompt into
gsd-core/references/execute-phase-quota-recovery.md, leaving step 7.1 as a
one-line pointer. execute-phase.md is now 92880 bytes — 510 SMALLER than base.

Also regenerates the fixtures that legitimately moved because three shipped
files changed (gsd-tools.cjs, config-schema.manifest.json, execute-phase.md):
golden-install-parity + install-tree for all 16 runtimes, INVENTORY.md +
INVENTORY-MANIFEST.json for the new reference, and the workflow size baseline.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2351): make the C1 orphan-reaping test load-independent

tests/run-with-timeout.test.cjs C1 asserted the child heartbeat file exists
after a 1s group-kill window, but the child only wrote it on the first 100ms
setInterval tick. Nothing synchronized the two: on a loaded container the group
is SIGKILLed before that tick lands, the file never appears, and the assertion
fails for a reason unrelated to reaping. Observed failing on both linux-node22
and linux-node24.

The behavior actually under test is the FREEZE assertion (heartbeat stops
advancing => descendant was reaped, not orphaned). That is unaffected by
sampling once more at t=0.

Child now writes its first heartbeat synchronously at startup before arming the
interval, and the kill window widens 1s -> 3s to cover child boot under load.
Both remove the timing dependency; neither weakens what the test proves.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2296): backfill pr:2458 in .changeset/rapid-jays-bark.md

* chore(#2296): regenerate fixtures after rebase onto #2402

The rebase conflicted on the generated golden-install-parity fixtures and
workflow-size-baseline.json because #2402 (b6e6a22fc) regenerated the same
artifacts. Conflict resolution picked a side to unblock the rebase; a true
regeneration on the combined tree then produced further drift, confirming the
resolved content was stale and would have dropped #2402's fixture changes.

Regenerated goldens, install-tree, size baseline, and INVENTORY-MANIFEST from
the merged tree. docs/INVENTORY.md keeps BOTH new reference rows.

execute-phase.md is 92782 LF bytes with both #2402's and this PR's extractions
applied — under the frozen ceiling (hard <93600, margin <=93400).

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 14:59:30 -04:00
Tom Boucher
12e4d93b19 fix(#2393): add GSD_ALLOW_SYMLINKED_DEST opt-in for intentional user-owned symlink layouts (#2445)
* fix(#2393): add GSD_ALLOW_SYMLINKED_DEST opt-in for intentional user-owned symlink layouts

Bug: v1.7.0's destSubpath write-confinement (ADR-1239 Phase B) refused
install/update whenever CLAUDE_CONFIG_DIR (or an artifact-kind child like
skills/, hooks/) was a pre-existing symlink, with no opt-out. Three
legitimate user-owned layouts were blocked:

  - (lars-hh) CLAUDE_CONFIG_DIR=~/.claude-personal with skills/hooks
              symlinked to a user-owned external dir
  - (Mamiki)  ~/.claude/skills is a Windows Junction to a shared skills dir
  - (Azd325)  ~/.claude itself is a symlink to a dotfiles repo (the early
              root-is-symlink return refused before the component loop ran)

Fix: add GSD_ALLOW_SYMLINKED_DEST env var (accepts '1' or 'true'). When
set, hasExistingSymlinkBetween follows symlinks instead of refusing them.
Cross-platform: fs.lstatSync().isSymbolicLink() returns true for both
POSIX symlinks and NTFS junctions (Node ≥ 16), so Mamiki's Junction case
is handled by the same code path.

Threat model preserved (these still refuse EVEN WITH opt-in):
  (a) path-traversal in the destSubpath string itself ('../../etc'-style)
      — ADR-1239 Phase B threat (a), untrusted destSubpath protection
  (b) a symlink whose resolved real path equals the install root itself
      — would let _removeGsdEntries wipe the root; #1704 threat (b)
  (c) broken symlinks (realpathSync throws) — fail-closed

What opt-in RELAXES specifically: the 'pre-existing symlink pointing
outside configHome' refusal — #1704 threat (c). The user has explicitly
asserted they own and trust the symlink target.

Error messages at all 4 call sites (installRuntimeArtifacts, _copyStaged,
migrateLegacyDevPreferencesToSkill, installOpencodeFamilySkills) updated
to (1) name the env var opt-in, (2) be accurate when the root itself is
a symlink (Azd325's complaint that the old message accused destDir of
'containing' a symlink when the root was the actual symlink).

Docs: docs/CONFIGURATION.md Environment Variables table updated.

Regression tests in tests/install-write-confinement.test.cjs cover:
  - child-symlink layout (lars-hh / Mamiki): default refuses, opt-in allows
  - root-is-symlink layout (Azd325): default refuses, opt-in follows
  - path-traversal '../../etc' refused EVEN WITH opt-in (threat a preserved)
  - resolved-target-equals-install-root refused EVEN WITH opt-in (threat b)
  - broken symlink refused EVEN WITH opt-in (fail-closed)

* test(#2393): import beforeEach/afterEach in install-write-confinement suite

The original file imported only { describe, test } from node:test. The new
#2393 opt-in describe block uses beforeEach/afterEach to manage the
GSD_ALLOW_SYMLINKED_DEST env var lifecycle — add them to the import.

* test(#2393): correct broken-symlink test — existsSync follows link → loop terminates early

Initial test expected broken symlinks to be refused even with opt-in. That
was wrong: fs.existsSync follows symlinks, so a broken symlink returns
false from existsSync and the component loop terminates before the symlink
check fires. Both default and opt-in paths share this behavior; the fix
preserves it.

Updates the test to pin the actual current behavior so a future refactor
(e.g. switching to lstatSync for existence) is a deliberate behavior change.

* fix(#2393): realpath the install root — guard against macOS /var ↔ /private/var

Code review (security subagent) flagged a HIGH-severity hole in the threat-(b)
preservation: realTarget (from fs.realpathSync) is fully symlink-resolved,
resolvedRoot (from path.resolve) is lexical-only. On macOS /var is a symlink
to /private/var, so resolvedRoot='/var/foo/.claude' but realConfigHome is
'/private/var/foo/.claude'. A symlink whose realtarget matches the install
root by real path would compare unequal to the lexical resolvedRoot —
defeating the wipe-protection guard exactly in the reporter's case (Azd325,
nix-darwin: ~/.claude is itself a symlink).

Fix: compute realRoot once via fs.realpathSync(resolvedRoot) at function entry
(with fail-closed fallback to lexical form on realpath failure — broken/missing
root, permission denied, exotic FS). Threat (a) path-traversal check above
still confines regardless. Compare against BOTH lexical and real forms in
both the root-symlink and component-symlink branches.

Also adds the reviewer's transitivity-trust clarification comment: once a
symlink is followed under opt-in, the walk continues from the resolved real
path WITHOUT re-checking further segments stay inside a confining boundary.
This is documented opt-in semantics — one opt-in trusts the whole reachable
tree — and the comment makes the design choice explicit so a future
maintainer doesn't add a 'follow one symlink only' expectation.

Regression test added for the macOS /var normalization case (spelled configHome
via os.tmpdir() lexically while pointing the test symlink through its realpath).
Test skips on non-darwin platforms and when os.tmpdir() has no symlink component.

* fix(#2393): root-symlink branch — do not apply threat-(b) check to root itself

Initial fix applied the wipe-threat-(b) check to the root-symlink branch
unconditionally. That was wrong: when root itself is a symlink (Azd325's
nix-darwin case), its realpath IS realRoot by construction — so the check
always fires, defeating the opt-in for exactly the case it was meant to
enable.

The wipe threat (b) does NOT apply to root being a symlink: destDir is a
CHILD of root, and resolving root just gives root's target. There is no
circular back-reference to root from a path that descends from a resolved
root. So the root-symlink branch should just follow the symlink under opt-in
and continue the walk, no threat-(b) check.

Threat (b) only fires in the COMPONENT loop, where a child symlink can
resolve back to the install root. That branch keeps the (b) check using BOTH
lexical and real forms of root (the macOS /var ↔ /private/var fix from the
prior commit).

* fix(#2393): apply opt-in at the 5 bin/install.js call sites + add env-var/transitive tests

Code review (correctness subagent) flagged a Critical coverage gap: the
initial fix updated only the 4 src/install-engine.cts call sites. Five
more call sites in bin/install.js still used the 2-arg signature, so the
opt-in env var was silently ignored on:

  - installCodexConfig (config.toml + agents/ dir + per-agent .toml paths)
    — Codex only
  - copyWithPathReplacement (the generic emit path: workflows, commands,
    staging) — ALL runtimes
  - resolveInstallRelativePath (path resolver used in various places)

Result: a user setting GSD_ALLOW_SYMLINKED_DEST=1 would see SOME refusals
disappear (engine path) and OTHERS remain (bin/install.js paths) — a
partially-applied install and a confusing UX, directly contradicting the
PR's headline claim.

Fix:
- Export isSymlinkedDestOptIn from src/install-engine.cts alongside
  hasExistingSymlinkBetween
- Import it in bin/install.js
- Update all 5 bin/install.js call sites to pass { allowOptInFollow }
- Update all 3 bin/install.js error messages to name the env var, matching
  the engine's phrasing

Also addresses reviewer's Medium test-adequacy findings:
- isSymlinkedDestOptIn env-var parsing now tested directly (accepts only
  documented '1' / 'true'; rejects 'TRUE', 'yes', 'on', '0', 'false',
  empty, unset)
- transitive symlink chain (configHome/outer → outside1 → outside2) test
  pins the documented 'transitive and unbounded' opt-in semantics so a
  future contributor can't accidentally narrow it

* chore(changeset): backfill pr:2445 in .changeset/eager-wasps-swim.md
2026-07-20 07:56:36 -04:00
Tom Boucher
517bae8d6d fix(#2372): widen decision-coverage-plan to all planner-canonical tags, drop misleading "(or body)" (#2443)
* fix(#2372): widen decision-coverage scan to planner-canonical tags, fix message

Bug: check.decision-coverage-plan's remediation message told the user to
cite decisions "(or body)" but extractPlanDesignatedSections only scanned
<objective>/<tasks>/<task>/<action>. A decision cited in <read_first>,
<behavior>, <verify>, <acceptance_criteria>, or <done> was invisible to
the gate — false BLOCKING coverage gap, plus the message's own fix-hint
sent the user to "the body" where re-citing still failed.

Two-part fix (must change together — that drift was the bug):

1. Widen XML_DECISION_TAGS_RE in src/check-command-router.cts to also
   match <read_first>, <behavior>, <verify>, <acceptance_criteria>,
   <done>. These are all planner-canonical tags the planner is told to
   use (plan-phase.md:830-862, plan-phase.md:772). The body negative-
   lookahead mirrors the opening-tag set so each tag's body is captured
   independently.

2. Correct buildPlanMessage to name ONLY the surfaces the extractor
   actually scans (front-matter must_haves/truths/objective,
   designated markdown headings, and the nine planner-canonical tag
   bodies). The misleading "(or body)" clause is gone.

Also updates the planner's documented contract (agents/gsd-planner.md:69)
and user-facing docs (docs/CONFIGURATION.md, docs/USER-GUIDE.md) to
reflect the wider scan.

Regression tests in tests/decisions.test.cjs cover each newly-scanned
tag body, a control (no citation still uncovered), and a message/extractor
parity assertion that names every scanned surface — so the two cannot
drift apart again.

Out of scope (per triage): cmdDecisionCoverageVerify/buildVerifyMessage
is a separate command (decision-coverage-verify) checking shipped
artifacts, not plan citations — untouched.

* chore(#2372): regenerate agent-size-baseline + golden-install-parity fixtures

gsd-planner.md grew 49172 → 49294 (+122 chars) from the widened decision-
coverage contract (5 new scanned tag names + heading clarification).
Growth is justified: the contract surface is itself the fix — the prior
text under-described what the gate scans, which was the bug.

Updates:
- tests/agent-size-baseline.json (gsd-planner.md: 49172 → 49294)
- 17 tests/fixtures/golden-install-parity/*.json (one hash per runtime)
- tests/fixtures/install-tree/*.json (regenerated by gen:golden)

* fix(#2372): per-tag matching — outer-tag citations survive inner-tag nesting

Code review (subagent) flagged a Medium edge-case regression from the
single-alternation regex: when a newly-scanned tag nests inside another
scanned tag, the alternation's negative lookahead halts the outer tag's
body at the inner tag — losing any D-NN citation in the outer tag's
prefix prose. Concretely:

  <action>per D-05 <verify>npm test</verify></action>

  → 3-tag alternation (old):  captured 'per D-05 <verify>npm test</verify>' as <action> body → D-05 caught
  → 9-tag alternation (bug):  captured 'npm test' only (from <verify>); D-05 in <action> prefix LOST

Switches extractXmlTagBodies to per-tag matching: each tag gets its own
regex whose negative-lookahead tempers only against the SAME tag's
reopening. So <verify> inside <action> is absorbed into <action>'s body
(D-05 caught) AND <verify> is matched separately on its own pass.

Per-tag preserves both:
- the reporter's case (sibling tags inside <read_first>)
- nested-tag citations in outer-tag prefix prose
- ReDoS safety (each per-tag regex keeps the #2128 body tempering)

Also adds the reviewer's other requested edge-case tests:
- non-scanned tag (<name>) bearing D-NN must NOT count
- self-closing form <read_first /> safely ignored
- attribute form <verify type="...">D-NN</verify> (canonical planner shape)
- CRLF newlines in tag body do not break capture

* chore(changeset): backfill pr:2443 in .changeset/noble-elks-chatter.md
2026-07-19 23:07:13 -04:00
Cody Anderson
20ff405cb3 feat(#2162): opt-in compact GSD-state format for the statusline (#2175)
* feat(#2162): opt-in compact GSD-state format for the statusline

New statusline.state_format config, enum full|compact (default full —
existing rendering untouched). "compact" renders the state segment as
"<version> · P<phase>/<total> · <status>", e.g. "v1.12 · P7/12 ·
executing" — dropping the milestone name and progress bar (the two
biggest width costs) and collapsing narrative statuses to a single
keyword. Per the #2162 approval conditions, the keyword set is the
canonical vocabulary from normalizeStateStatus() in state-document.cjs
(discussing/planning/executing/verifying/completed/paused) — no
parallel hand-rolled list, so the vocabularies can't drift — and the
canonical stuck state "paused" renders uppercase as PAUSED (no new
"blocked" lifecycle state). Statuses the normalizer passes through
unrecognized fall back to their first word capped at 16 chars.
Lifecycle scenes preserved: active_phase wins over the body phase
number, milestone completion renders "complete", idle-with-next-action
renders "next <action> <phases>".

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* docs(#2162): changeset fragment for PR #2175

* fix(#2162): review fixes — ENUM_KEYS coverage, cap boundary tests, changeset format

- register statusline.state_format in the fix-1628 coercion-bypass matrix
- 15/16/17-char boundary tests for the shortGsdStatus fallback cap
- changeset body ends with the (#2162) citation per house convention

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2162): round-2 review fixes — scene exclusivity, direct config-set coverage

- compact renderer gates the milestone-complete scene behind the absence of
  an in-flight phase id, mirroring formatGsdState's if/else precedence
  (Scene 1 beats Scene 3); regression test covers the non-atomic
  active_phase + percent=100 STATE.md shape
- direct config-set accept/reject test for statusline.state_format plain
  strings (ENUM_KEYS matrix covers only the JSON coercion shapes)

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test: regenerate golden-install-parity fixtures for the statusline hook change

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2162): complete-scene gate matches formatGsdState exactly (+property tests)

Re-review Major: gating done on !phaseId held completion back for the
legacy phaseNum shape — formatGsdState reaches Scene 3 on percent=100
regardless of phaseNum, so compact must too. Gate is now !s.activePhase.
The phaseNum-only test now expects 'complete' and cross-checks the full
renderer; a parity test feeds identical inputs to both renderers.
Re-review Minor: shortGsdStatus gets fast-check property coverage
(totality, canonical fixed points, separator safety, fallback shape).
Golden fixtures regenerated for the hook byte change.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
2026-07-14 21:09:24 -04:00
Cody Anderson
d6672ff926 feat(#2163): opt-in git branch/status segment in the statusline
New statusline.show_git config (default false). When enabled, a git
segment renders after the directory: current branch plus compact
work-state markers (+staged ~unstaged ?untracked ↑ahead ↓behind, or ✓
when clean and in sync), e.g. " │ main+2~1?3".

One git status --porcelain=v2 --branch spawn per render via execFileSync
with a fixed argument array (no shell), a 1.5s timeout, and the
workspace dir passed with -C. Fails silently — segment absent outside a
repo, without git, or on timeout. Default output is unchanged when the
flag is absent.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
2026-07-13 11:33:32 -06:00
Cody Anderson
ad7111e50b feat(#2161): opt-in absolute token count on the statusline context meter (#2174)
* feat(#2161): opt-in absolute token count on the statusline context meter

New statusline.show_context_tokens config (default false). When enabled,
the context meter shows the absolute token total after the percentage,
e.g. "████░░░░░░ 46% (156k)" — summing input, cache-creation, cache-read,
and output tokens from context_window.current_usage (matching /context).

Default output is byte-for-byte unchanged when the flag is absent or
false. The .planning config is now read once per render and shared with
the last-command/position block instead of being re-read.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* docs(#2161): changeset fragment for PR #2174

* fix(#2161): review fixes — k-to-M threshold, boundary tests, changeset format

- formatTokens promotes to the M branch when k-rounding reaches 1000
  (999,500-999,999 rendered "1000k" instead of "1.0M")
- boundary tests at 999499/999500/999999/1000000/1000001
- Number() guards on the four usage fields (silent string-concat gap)
- changeset body ends with the (#2161) citation per house convention

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2161): round-2 review fixes — config-set coverage, precision claim, exports style

- config-set accept/reject tests for statusline.show_context_tokens
  (mirrors the post-planning-gaps precedent the issue scope names)
- changeset + docs no longer claim parity with /context: the suffix sums
  four fields while the meter %% derives from used_percentage (three), so
  the figures can diverge slightly
- module.exports one entry per line

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test: regenerate golden-install-parity fixtures for the statusline hook change

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-13 13:25:47 -04:00
Tom Boucher
79d7657eff feat(#2102): make pi a first-class installable runtime + fix its dispatch (ADR-1239)
Net-new EoS/pi installable runtime — purely additive (no prior runtime==='pi'
branches). pi is a bun-runtime programmatic-CLI whose /gsd command is registered
by a native ExtensionAPI extension and dispatches through the embedded engine.

Stage 1 (install plumbing):
- capabilities/pi/capability.json: full hostIntegration descriptor (imperative /
  slash-programmatic / active-model / native-extension / bun) + hostBehaviors
  {nativePlugin, pluginOnlyInstall}.
- --pi flag + interactive-menu renumber (All 17->18); pi added to RUNTIME_FLAG_IDS,
  RUNTIME_LABELS, RUNTIME_META, allRuntimes/runtimeMap, model-catalog defaults.
- Install mirrors OpenCode: pi installs the gsd.cjs extension + the shared engine
  payload (gsd-core + scripts + config markers) + the shared hooks bundle (spawned
  by the extension at lifecycle events, like OpenCode's plugin). pluginOnlyInstall
  EXCLUDES declarative command/agent/skill markdown, which pi has no host-read
  surface for (its /gsd is programmatic). _installNativePluginIfDeclared (extracted
  from the opencode-family path) copies pi/gsd.cjs -> ~/.pi/agent/extensions/gsd.cjs
  (global) / .pi/extensions/ (local). pi added to package.json files.
- Golden: new pi.json (320 files: extension + engine + 27-file hooks bundle, no
  markdown); the 16 other fixtures + claude-local change only by the shared
  model-catalog hash line.

Stage 2 (real dispatch + upgrades):
- Shared dispatchGsdCommand() (shell-command-projection): bounded, no-throw
  subprocess-shim to gsd-tools.cjs (the only full-surface dispatch path; no
  in-process full-hub factory exists). Fixes pi/gsd.cjs's createHub()-no-args bug
  (every dispatch was UnknownCommand) AND the identical bug in mcp-server.cts's
  gsd_invoke_command, which a vacuous unknown-family-only test had masked (now has
  a real dispatch regression test).
- pi/gsd.cjs: /gsd handler now (args, ctx) - tokenizes (quote-aware, via the
  shipped hooks/lib/git-cmd.js) + dispatches real family/subcommand (not hardcoded
  query/help); gsd_invoke gets a TypeBox (JSON-schema-fallback) parameters schema +
  consumes params; getArgumentCompletions; before_provider_request active-model
  steering (fail-open on null resolution); functional session_start /
  before_agent_start / session_before_compact hook bridges (spawn the shipped GSD
  hook scripts).
- EXTENSION_EVENT_SURFACES.pi expanded from ['tool_call'] to the full 30-event
  vocabulary.

Docs (host-integration matrix + how-to) + changeset (Added).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 21:07:38 -04:00
Tom Boucher
f014ec83bd feat(#2093): migrate Kilo onto EoS imperative adapter + hook-bus/model/MCP/dispatch upgrades (ADR-1239)
Fold remaining isKilo logic branches into descriptor-driven reads:
finishPermissionWriter (uninstall cleanup), skipSharedHooksInstall (hooks
copy), and a skills converter-name registry (the artifactLayout.converter
field is now load-bearing, not decorative). frontmatterDialect stays the
documented dispatch key for frontmatter (no descriptor field for it). Dead
isKilo destructure bindings removed. Byte-identical golden parity for all 16
runtimes (opencode, which shares kilo's combined-family path, verified clean).

UPGRADE 1 (hook bus): install .kilo/plugins/gsd-core.js native plugin +
extensionEvents:"kilo" + EXTENSION_EVENT_SURFACES.kilo (OpenCode-fork bus).
UPGRADE 2 (active model): populate runtimeTierDefaults.kilo + thread
modelOverride through convertClaudeToKiloFrontmatter — model no longer stripped
from agents. UPGRADE 3 (MCP): document the gsd-core MCP companion under kilo's
mcp config key. UPGRADE 4 (named dispatch): agents/*.md mode:subagent roster is
the Task-tool dispatch surface (tested); subagentToolkit stays 'undocumented'
per AC so dispatch degrades to 'degraded' by design.

Model-catalog single-source edit ripples the shared model-catalog.json hash
into all 16 golden fixtures (expected). Inline defect fixes (no-defer): stale-
bake-guard resolveAgentDir 'agent'->'agents' (was a silent no-op for opencode/
codex), hardcoded 'Removed OpenCode plugin' uninstall log -> generic, and the
connect-gsd-mcp-server.md OpenCode mcpServers->mcp doc error.

Tests: kilo-imperative-reference (adapter/axes/fail-closed/degradation/
hostBehaviors + widened isKilo source-grep across 4 modules) + kilo-upgrades
(plugin parity+load, model-override converter, agents dispatch surface, MCP
doc). Matrix + how-to + config docs updated; changeset added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 19:55:15 -04:00
Tom Boucher
185abe2d66 feat(codex): advance Codex/OpenAI model defaults to GPT-5.6 (Sol/Terra/Luna)
Update runtimeTierDefaults.codex and providerPresets.openai in
model-catalog.json to the GPT-5.6 family (gpt-5.6-sol/terra/luna),
advancing from the superseded GPT-5.4/5.5 generation.

Model IDs verified against OpenAI developer API docs:
- gpt-5.6-sol: flagship, /, reasoning xhigh
- gpt-5.6-terra: balanced, .50/, reasoning medium
- gpt-5.6-luna: fast/cheap, /, reasoning medium

Tier mapping is 1:1 (Sol↔flagship, Terra↔balanced, Luna↔fast),
so profile semantics are unchanged — only the underlying IDs advance.

Updates: catalog JSON, test assertions (catalog defaults), docs
(CONFIGURATION.md + zh-CN/pt-BR translations, workflow settings),
and changeset.

Closes #2122
2026-07-10 12:11:55 -04:00
Tom Boucher
5aa7771dd3 Merge branch 'next' into fix/2009-load-failed-capability-injects-blocking- 2026-07-07 23:12:09 -04:00
Tom Boucher
1ee00320c7 Merge branch 'next' into fix/2072-thread-model-into-assumptions-and-review-spawns 2026-07-07 22:03:47 -04:00
Tom Boucher
4483300253 fix(#2072): thread resolved model into routed-agent spawns (assumptions-analyzer, code-reviewer, code-fixer)
model_overrides / models.<phaseType> were silently inert for gsd-assumptions-analyzer,
gsd-code-reviewer, and gsd-code-fixer on Claude Code: resolveModelInternal honors them,
but the workflows spawned these agents with no model= param, so the resolved value
never reached the Agent tool and the agents inherited the session model — no warning.

Fix — thread each agent's resolved model at every spawn site (the established
plan-phase pattern; the architecture-consistent Claude mechanism, since 13 other
agents already thread their model):
- discuss-phase-assumptions.md: `resolve-model gsd-assumptions-analyzer --raw`
  → ANALYZER_MODEL, threaded.
- code-review.md + code-review-fix.md (re-review): `resolve-model gsd-code-reviewer --raw`
  → REVIEWER_MODEL, threaded.
- code-review-fix.md (both fixer spawns): `resolve-model gsd-code-fixer --raw`
  → FIXER_MODEL, threaded (same silently-inert bug, same file — folded in per review).
- quick.md review step: was reusing `{executor_model}` for gsd-code-reviewer (so the
  reviewer's own override was ignored); init.quick now resolves `reviewer_model`
  (gsd-code-reviewer) and the spawn threads it.

resolve-model --raw returns the bare model string (resolve-execution --raw would
return effort — wrong). The resolver maps these agents to phaseType discuss /
verification / execution, so models.<phaseType> apply too.

Scope: the three agents reachable from the two issue-named workflows + quick.md. The
wider systemic class (other agents in UNTOUCHED workflows with the same pattern) stays
documented on the issue for a maintainer-scoped structural decision (thread-at-source
vs embed-at-install like #2256), not widened here.

Docs: the stale "discuss — reserved, no subagent today" model-profile tables now list
gsd-assumptions-analyzer and the verification row includes gsd-code-reviewer, across
the English docs, the shipped gsd-core/references/model-profiles.md reference, and the
ja-JP / zh-CN / ko-KR / pt-BR locale mirrors.

Tests:
- tests/model-resolver.test.cjs: #2072 acceptance — model_overrides and
  models.discuss/verification/execution resolve for all three agents.
- tests/model-routing-spawn-threading.test.cjs: every spawn of the three agents threads
  a resolved model (fails pre-fix); a header-precise parity guard fails the suite if a
  new un-threaded spawn of any of them regresses.
All 16 golden-install-parity fixtures + the workflow size baseline regenerated for the
changed shipped files (4 workflows + the reference doc); bin/lib is excluded from parity.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 21:22:48 -04:00
Tom Boucher
46f8d3814c fix(#2009): load-failed capability gates fail open with a loud warning
Previously a capability that failed to LOAD (e.g. incompatible engines.gsd) but
declared a gate-kind loop hook caused the loop resolver to inject a BLOCKING
synthetic gate (blocking:true, onError:halt) at every declared point, halting
every ship:pre / verify:post project-wide over an unrelated load error, with no
remediation surfaced.

Per maintainer decision (#2009) it now fails OPEN: no gate is injected (the loop
proceeds; --active-cap correctly reports the failed cap inactive) and a loud
warning is emitted — to stderr (the channel host workflows/agents actually see)
and in the envelope 'warnings' array — naming the load reason and the exact
'gsd capability remove <id>' remediation. The loader still records blockedGates;
only the consequence changes from block to warn.

Security (review): capId and reason originate from a third-party manifest /
directory name. capId is validated against the canonical kebab-case id shape
before it is placed in the runnable remediation command (withheld otherwise);
reason is stripped of control chars and backticks. This closes an argument/
prompt-injection vector in the surfaced message.

Docs updated to the fail-open-warning posture (ARCHITECTURE, INVENTORY,
CONFIGURATION, README, capability-overlay-model). Also removes a dead 'before'
import surfaced by lint in the issue-2045 test.
2026-07-07 19:38:01 -04:00
Tom Boucher
6addeccd19 feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates
an external API/SDK/service can no longer seal without a decided coverage matrix.

- src/api-coverage.cts: deterministic detector (compound verb+noun signal +
  <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix
  parse/validate/render with field-length caps.
- check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as
  a token under .planning/phases/ only (traversal-neutralized); validates
  COVERAGE.md or blocks iff a strong integration signal is detected and no
  matrix exists; fail-closed when phases tree exists but phase unresolvable.
- capabilities/ai-integration: workflow.api_coverage_gate config key (default
  true), plan:pre contribution, blocking verify:pre gate. Data-driven.
- gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch.
- Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e.
  Code+security review findings fixed (stopword FP, scope containment, pipe/cap
  rejection, prompt-injection message hygiene).
- Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs.

Closes #1562
2026-07-07 15:11:12 -04:00
Tom Boucher
603593d41d fix(#1857): test gates normalize to one-shot + bounded timeout (no watch-mode hang)
A GSD verification gate resolves a project's test command and runs it. vitest
defaults to WATCH mode in an interactive TTY — exactly where a user runs
`gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never
exited and the orchestrator waited indefinitely. Recovery needed the user to
manually prompt "something blocking?".

Fix — one shared helper + a bounded, surfacing timeout on the test-command gates:
- New pure module src/normalize-test-command.cts + `gsd-tools query
  normalize-test-command` verb: rewrites a resolved command to a best-effort
  one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`;
  a package-manager `test` script whose package.json runner is watch-vitest →
  `CI=true` prefix; handles `--dir`; already-one-shot commands unchanged —
  never double-flagged). Named `normalize-test-command` (not `test-*`) so the
  file does not match node --test's default `test-*` discovery glob.
- The three gates that HUNG or silently-continued route through that ONE helper
  and bound execution with `timeout $(config-get workflow.test_gate_timeout)`
  (new config key, default 600s): the regression gate (extracted to
  execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen —
  it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the
  audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124.
- verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so
  it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint
  on 124, staying under its frozen 40960-byte tier cap.

Security hardening (review): the normalizer only rewrites a runner named as a
standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never
mangled), is length-capped and uses only linear-time split-based scanning (no
super-linear backtracking on an adversarial `workflow.test_command`), and reads
package.json only when it is a regular file (never blocks on a FIFO via `--dir`).

Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the
schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors
workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores,
inventory manifest/index. All 16 golden-install-parity fixtures + workflow size
baseline regenerated for the changed shipped files; bin/lib is excluded from the
parity manifest.

Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security
hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route
through the shared helper + configured timeout + exit-124 watch-mode hint;
verify-phase asserted as normalize-only/already-bounded).
tests/execute-phase-active-flags.test.cjs repointed at the extracted step;
tests/planner-language-regression.test.cjs allowlist comment updated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 12:14:43 -04:00
Tom Boucher
ed79902509 feat(#2007): implement mempalace memory_mode kg_backend and replace routing (#2010)
Wire the two forward-declared mempalace.memory_mode modes so they actually
route recall/capture instead of silently behaving as `augment`:

- kg_backend: the palace temporal KG is the primary knowledge-graph source;
  native .planning/graphs/ is the fallback. Non-KG drawer recall stays additive.
- replace: recall resolves through the palace as the source of truth; native
  artifacts are the fallback.

Every mode stays onError:skip and default-resilient — an unreachable palace
degrades to native memory and GSD keeps writing .planning/graphs/, so no memory
is lost. Cross-mode .planning/graphs/ migration remains a documented open
question (PRD/ADR §17), out of scope here.

Surfaces updated (instruction-only contract): recall/capture commands (+ generated
skills), discuss/wave fragments, curator agent, capability.json schema. Docs:
how-to Step 3, CONFIGURATION, FEATURES, CONTEXT glossary. Regenerated
capability-registry, golden install-parity fixtures (mempalace hashes only),
agent-size-baseline. Added a routing-contract + cross-surface parity test.

Incidental (folded per no-defer rule): removed pre-existing unused imports
(spawnSync in capability-registry.test.cjs; fs in issue-498-package-identity.test.cjs)
that eslint flagged in/alongside the touched files.

Closes #2007

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 15:01:22 -04:00
Tom Boucher
bd77b40107 feat(#1825): configurable graphify graph location (graphify.graph_path) (#2013)
* feat(#1825): configurable graphify graph location (graphify.graph_path)

Add a graphify.graph_path config key (.planning/config.json) that overrides
where /gsd-graphify query|status|diff read the knowledge graph, so one curated
umbrella-level cross-repo graph can serve multiple sibling projects without N
drifting ~5 MB mirror copies. Previously the graph location was hardcoded to
<cwd>/.planning/graphs/.

- src/graphify.cts: resolveGraphLocation(cwd, planningDir) honors the key
  (resolved relative to project root; absolute paths honored via path.resolve);
  falls back to the historical .planning/graphs/graph.json when unset/blank/
  non-string (byte-identical). Wired into graphifyQuery, graphifyStatus,
  graphifyDiff (snapshot travels with the configured graph via dirname), and
  writeSnapshot. Configured-but-missing -> actionable error naming the path.
  Build stays project-scoped (skill hardcodes the cp dest); umbrella graph is
  built in the umbrella project, sub-projects only READ it.
- config-schema.manifest.json: register graphify.graph_path in validKeys.
- tests/graphify-graph-path.test.cjs: boundary matrix (unset byte-identical,
  set+present reads configured graph not default, set+missing actionable error,
  relative resolved vs project root, blank treated as unset, snapshot alongside
  configured graph, diff from configured dir, build project-scoped) +
  VALID_CONFIG_KEYS registration.
- docs: CONFIGURATION.md row, FEATURES.md REQ-GRAPH-06, CONTEXT.md module note,
  .changeset (Added).

Closes #1825

* docs(#1825): backfill changeset pr number 2013
2026-07-05 14:50:01 -04:00
Tom Boucher
8f2ebbe9bf feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity

Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap).

--gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): backfill changeset PR number (#1996)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): drop Gemini CLI from issue templates (review nit)

Removes the sunset Gemini CLI runtime from the two GitHub issue-template
runtime lists that the removal PR missed, per @davesienkowski's review nit:
- feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise
  request a feature for a runtime GSD no longer supports)
- bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json
  retrieval-help line

Leaves the post-removal templates fully consistent with the Antigravity redirect.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 13:32:51 -04:00
Jeremy McSpadden
e5ef323b15 feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command

Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.

* feat: add /gsd-start smart-entry command

State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.

- src/smart-entry.cts: deterministic situation classifier (no-project,
  paused, blocked, verify-failed, needs-first-phase, planning, executing,
  verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
  ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire  case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
  dispatcher presenting an AskUserQuestion menu (with --text fallback for
  non-Claude runtimes) and dispatching to existing commands. Falls back
  to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
  situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
  (markdown-layer invariants + every emitted command resolves to a real
  slash command).

Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.

* refactor: rename smart-entry command to /gsd:next

Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.

All affected tests (188) pass; lint:ci clean.

* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)

Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.

- detectSignals now reads total_phases + percent from nested progress{}
  first, then scalar fm, then body; current_phase falls back to the
  body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
  body Phase field) covering verify-pending + executing situations.

Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.

* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)

Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.

Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
   layer is down — read .planning/STATE.md directly with the Read tool
   and synthesize a minimal situation + actions menu so /gsd:next stays
   useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
   /gsd:progress as before.

Matches the direct-read resilience the live agent already did by hand.

* docs: add gsd-next skill surface

* chore: trigger no-mistakes validation

* no-mistakes(review): Fix smart-entry phase ordering

* no-mistakes(review): Fix decimal smart-entry phase ordering

* no-mistakes(test): Fix smart-entry next test contracts

* no-mistakes(document): Docs synced for smart entry

* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: shorten next.md description and update golden install parity fixtures

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: trigger no-mistakes validation

* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files

Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.

* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: regenerate golden install parity fixtures for /gsd:next

Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.

* Fix smart-entry verify-failed phase scoping and empty resolve shim step

Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.

* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: recapture all 16 golden fixtures with updated smart-entry.md hash

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: regenerate fixtures + inventory manifest after rebase onto next

Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next

Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).

Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)

The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs

Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo

Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
  recommended action, 1-4 unique-id /gsd:* actions (previously the
  one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout

Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:

- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
  zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
  kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
  consolidation file doing dozens of real installs — 41s even on a fast Mac
  (much worse on the slow Windows I/O path), plus an install-heavy cluster.

Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.

Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 12:18:25 -04:00
Tom Boucher
3c13903dcd feat(#1866): agent-side self-load of configured agent_skills
Each of the 22 consumer agents now self-loads its configured agent_skills
in its mandatory init step, so .planning/config.json agent_skills.<type>
reaches the agent on every runtime — including Cursor and /gsd-autonomous,
where Skill()-delegated workflow bash init did not reliably execute.

- gsd-core/references/agent-skills-bootstrap.md: shared contract
  (query + Read + dedup guard that skips when <agent_skills> is already
  in the prompt, so Claude's orchestrator-side injection never doubles)
- 22 agents/gsd-*.md: one self-load line naming the agent's own type
- gsd-core/workflows/autonomous.md: note that delegated agents self-load
- tests/agent-skills-bootstrap.test.cjs: regression + parity (CONSUMER_AGENTS
  bijection + fast-check property) — Generative-Fix-Divergence guard
- docs: ADR-1866, CONFIGURATION dual-injection How It Works, INVENTORY
  row, Changed changeset

Closes #1866
2026-07-01 20:09:01 -04:00
Tom Boucher
da37986cd0 fix(#1847): resolve standard tier to claude sonnet 5
Point the sonnet/standard tier at Claude Sonnet 5 (`claude-sonnet-5`,
GA 2026-06-30) across the Anthropic-backed runtimes and provider presets,
replacing the superseded `claude-sonnet-4-6`. Mirrors the change into the
CONFIGURATION.md and settings-advanced.md runtime-defaults tables (the
#3229 catalog↔docs parity gate) plus the pt-BR/zh-CN translations, and
updates the tests that pin the old ID. Regenerates the workflow size
baseline for the (smaller) settings-advanced.md.

Scope is Sonnet only — opus/haiku IDs are untouched. Prepared as a 1.6.1
hotfix off the v1.6.0 tag.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 33260555b4)
2026-06-30 22:21:18 -04:00
Tom Boucher
fd576528a7 fix(#1747): register four search-provider keys in the config schema (#1814)
* fix(#1747): register four search-provider keys in the config schema

buildNewProjectConfig emits seven search-provider availability flags and
research-provider.cts providerAvailability() consumes all seven, but only
three were registered in VALID_CONFIG_KEYS (config-schema.manifest.json).
config-loader.cts then printed an 'unknown config key(s)' warning for the
four unregistered keys (tavily_search, ref_search, perplexity, jina) on
every freshly generated .planning/config.json.

Register the four missing keys in the schema manifest and document them
alongside brave/exa/firecrawl in CONFIGURATION.md. Add a regression test
plus a structural drift guard that requires every config-driven
research-provider flag to be in VALID_CONFIG_KEYS, so a future provider
addition cannot silently reintroduce the drift.

* fix(#1747): move regression into owning test file + add changeset

lint-regression-test-names rejects new bug-NNNN-*.test.cjs files; move the
#1747 regression (four provider keys in VALID_CONFIG_KEYS + provider-flag
drift guard) into tests/bug-2530-valid-config-keys.test.cjs, the canonical
home for VALID_CONFIG_KEYS regressions, and delete the standalone file.

Add the missing .changeset fragment — config-schema.manifest.json lives
under gsd-core/ (user-facing), so changeset-lint requires a fragment.

* test(#1747): regenerate golden-install-parity fixtures for schema change

Adding four provider keys to config-schema.manifest.json shifts its shipped
content hash (65dea848 -> 7d398e94); recapture all 16 runtime fixtures via
UPDATE_GOLDEN=1. Each fixture changes exactly one line — the manifest hash.
2026-06-28 22:42:36 -04:00
Tom Boucher
4ced0a64cc feat(#1561): assumption-delta advisory checkpoint (#1767)
* feat(#1561): assumption-delta advisory checkpoint

* chore(#1561): backfill changeset PR number (#1767)

---------

Co-authored-by: review-bot <review-bot@gsd>
2026-06-27 08:43:47 -04:00
Tom Boucher
b0d5ca3379 feat(#1517): support custom reviewer instances for /gsd:review (#1766)
* feat(#1517): support custom reviewer instances for /gsd:review

Add a bounded review.reviewer_instances config surface so one model-capable
adapter (e.g. opencode) can run as several independent reviewer identities in a
single /gsd:review pass. Instances participate only via review.default_reviewers,
expand before built-in slugs, are available iff their cli is detected, and a
non-matching entry is a hard error (typo must be loud). >=2 same-cli instances
emit a shared-adapter caveat in REVIEWS.md. Default path with no instances is
byte-for-byte unchanged.

Single-source instance->cli resolution lives in resolveReviewerSelection /
normalizeReviewerInstances (parity-locked in
tests/review-reviewer-instances.test.cjs). cli validated against
KNOWN_REVIEWER_SLUGS only (never arbitrary shell); model/agent opaque, never
shell-interpolated.

Closes #1517

* chore(#1517): backfill changeset pr:1766

---------

Co-authored-by: review-bot <review-bot@gsd>
2026-06-26 23:35:04 -04:00
Tom Boucher
2b215b4163 feat(#1688): warn on stale model bake for static-frontmatter runtimes (#1692)
* docs(#1650): fix stale opencode install-path claim in core settings

* feat(#1688): warn on stale model bake for static-frontmatter runtimes

* chore(#1688): backfill changeset pr field with real PR number

* test(#1688): make resolveAgentDir assertions use path.join for windows

* docs(#1688): codify windows path-literal-in-assert anti-pattern + align test
2026-06-25 09:12:40 -04:00
Alex V.
a63684c222 enhance(#1577): WebFetch/WebSearch injection isolation + opt-in blocking (#1585)
* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking

Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the
largest untrusted channel) in gsd-read-injection-scanner; shared
untrusted-input-boundary reference @-included by the 8 ingest agents
(randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring);
opt-in security.injection_blocking (default advisory — non-breaking).

arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth).

* fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized

- A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a
  circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already
  in the transcript. The prompt-level data/instruction boundary is the primary control.
- A2: registered security.injection_blocking in the config schema + defaults manifests (default
  false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads.
- A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention).
- A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale).
- A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input.
- Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) +
  drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as
  the noted pre-existing follow-up.

* fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate

The new reference quotes injection phrases ('ignore previous instructions',
'you are now…') as examples agents must NOT comply with, tripping the repo's
own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red
on HEAD). Allowlist it alongside the other security docs (security-model.md,
TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS
scanner test doesn't scan references/, so only the shell gate needed it.

Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15.

* fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer

trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so
no named web-ingress agent is uncovered, keeping the two justified additions
(gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10.
 - gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset.
 - gsd-assumptions-analyzer reads 5-15 codebase source files (external/source-
   document ingress per the boundary), though it has no web tools.
INGEST_AGENTS in the isolation test now asserts all 10; size baselines
regenerated (+60 bytes each, both well under the DEFAULT cap); changeset
reworded 8 -> 10.

Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39.

* docs(#1577): document security.injection_blocking + boundary seam

trek-e Major 2 + Minor:
 - docs/CONFIGURATION.md: add the top-level security.injection_blocking key to
   the Full Schema and a Security Settings subsection, distinguishing it from
   the workflow.security_* namespace; honest circuit-breaker-not-redactor
   framing matching ADR-1577 / security-model.
 - CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry.

Verified: lint:docs ok; config-field-docs + contributor-standards green.

* test(#1577): make read-injection property test git-text, not binary

trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as
degenerate-edge inputs. The NUL is what actually made git classify it binary
(git binary = NUL in first 8K). Replace both with text-safe escapes that keep
the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File
now diffs/blames line-by-line.

Verified: property test 2/2; no NUL/raw-noncharacter bytes remain.

* docs(#1577): align untrusted boundary docs

Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner.

* docs(#1577): align ADR ingest agent count

Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 17:07:23 -04:00
Tom Boucher
35478b615e refactor(#1646): route capability routers through Command Routing Hub per ADR-959 (#1647)
* refactor(#1646): route capability routers through Command Routing Hub per ADR-959

Phase 2 of parent #1641. Converts graphify, intel, and audit command
routers from hand-rolled if/else dispatch to routeHubCommandFamily,
implementing the ADR-959 §III(B) line 75 mandate. The three routers
now share the uniform dispatch shape with the 14 host routers.

src/cjs-command-router-adapter.cts
  * Imported ERROR_REASON from io.cjs.
  * UnknownCommand translation now passes ERROR_REASON.SDK_UNKNOWN_COMMAND
    as the second arg to error() — additive for host routers (their
    existing one-arg error callbacks ignore the second arg), required
    for capability routers whose tests assert reason === 'sdk_unknown_command'
    on the JSON-error envelope.

src/graphify-command-router.cts
  * Replaced 4-branch if/else with routeHubCommandFamily + handlers map.
  * Validation handlers (missing term, missing/invalid --budget) now
    return makeInvalidArgs(arg, reason, ERROR_REASON.USAGE) Results
    instead of calling error() directly (Q2=C, Q4=ii from grilling).
  * Success handlers keep direct output() calls.
  * Subcommands array is alphabetical for byte-identical 'Available:'
    text in the unknown-subcommand message.
  * The unknown-subcommand path is now owned by the Hub's manifest
    check (the adapter passes SDK_UNKNOWN_COMMAND).

src/intel-command-router.cts
  * Replaced 9-branch if/else with routeHubCommandFamily + handlers map.
  * Validation handlers (missing term, missing filePath for patch-meta
    and extract-exports) return makeInvalidArgs Results.
  * Preserved the timeAgo mutation in the non-raw status handler.
  * Preserved the lazy require('./intel.cjs') inside the route function.

src/audit-command-router.cts
  * routeAuditUat: routes through the Hub with a synthetic 'run'
    defaultSubcommand (no real subcommands). Gives uniform observability.
  * routeAuditOpen: captures --json in a closure, strips it from args
    before Hub dispatch (so it isn't mistaken for a subcommand by the
    manifest check), then branches on wantJson inside the handler to
    preserve the formatAuditReport success-path quirk.

docs/CONFIGURATION.md
  * Observability section: noted capability commands (graphify, intel,
    audit-uat, audit-open) now emit DispatchEvent records since #1646.

.changeset/capability-routers-via-hub.md
  * Changed fragment describing the user-visible audit-trail expansion.
    pr:0 placeholder will be backfilled after gh pr create returns the
    real PR number (DEFECT.CHANGESET-PR-FIELD-DRIFT).

Verification
  * graphify cutover tests: 119/119 pass (all unit, dispatch, behavior,
    error path, JSON-errors, and registry assertions)
  * intel cutover tests: 39/39 pass
  * audit cutover tests: 24/24 pass
  * bug-974-graphify-budget-missing-value regression test: pass
  * npm run test:unit (full suite): 2384 tests, 0 fail
  * gsd-test-summary on docker: outcome=passed, 0 failures
    (RULESET.PR-FLOW.docker-before-push)

JSON-error envelope parity verified byte-identical: reason values
('usage', 'sdk_unknown_command') and message texts are preserved
across all three routers' error paths.

* chore(#1646): backfill changeset pr: 1647 (DEFECT.CHANGESET-PR-FIELD-DRIFT)
2026-06-23 23:21:06 -04:00
Tom Boucher
207d8f1697 fix(#1626): make the security gate severity-aware via per-threat severity (#1635)
workflow.security_block_on was documented as the minimum threat severity
that blocks advancement, but threats carried no severity and the auditor's
threats_open count (the SECURITY.md gate field) counted every open threat
regardless of severity — so the threshold had no effect, and the auditor's
block_on vocabulary (open/unregistered/none) did not even match the config
enum (critical/high/medium/low/none).

- planner: add a Severity column to the STRIDE threat register; assign
  severity per threat.
- auditor: read severity; reconcile the <config> block_on domain to the
  severity enum; redefine threats_open as the count of OPEN threats whose
  severity is at or above block_on (none => 0). Below-threshold opens are
  reported as non-blocking and excluded from threats_open.
- SECURITY.md template + planning-config.md reconciled.

No gate-check site changed: threats_open == 0 stays the gate everywhere;
only its computation is now severity-filtered.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 19:04:13 -04:00
Dave
d1f7ba82f2 feat(#1592): add plan:pre codebase-drift pre-check before planner runs
Add a non-blocking, warn-only codebase-drift gate at plan:pre so a stale
STRUCTURE.md is surfaced before /gsd:plan-phase spawns the planner, instead
of being discovered mid-execution by the existing execute:wave:post gate.
Gated on a dedicated workflow.plan_drift_precheck toggle (default true),
independent of schema_drift_gate. Never blocks planning, never spawns the
mapper agent at plan time.

Review feedback (#1595):
- Use a documented conventional-commit type (feat, not enhance) per
  CONTRIBUTING.md / gsd-validate-commit.sh.
- Normalize the plan_drift_precheck command references to the colon prose
  form (/gsd:plan-phase, /gsd:map-codebase) to match plan-phase.md §5.65;
  registry regenerated from capability.json.
- Make the test temp dirs hermetic: drain mkdtemp dirs in an after() hook
  via the helpers.cleanup() budget (local/no-raw-rmsync-in-tests-compliant).

Closes #1592

Claude-Session: https://claude.ai/code/session_016JBiXEAofvB3prJim29XMS
2026-06-22 20:30:59 -04:00
Tom Boucher
2c718bf972 fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs (#1537)
* fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs

Generalizes the Codex-only #1515/#1519 fix to every non-Claude runtime, and
wires it into the real install path (where it was previously dead-on-arrival).

Root causes:
1. The runtime-default stamping lived only in `_applyRuntimeRewrites`, but the
   installer emits `gsd-core/workflows/*.md` via `copyWithPathReplacement`, which
   never calls it — so a real `--codex`/`--cursor`/etc. install emitted
   `--default claude` and worktrees-on. RUNTIME mis-resolved to claude and the
   workflow ran executors unisolated against the main checkout. (#1515/#1519 were
   also dead-on-arrival in real installs; this repairs them.)
2. Only `case 'codex'` was stamped; every other non-Claude runtime kept the
   Claude default.

Fix:
- New `_stampNonClaudeRuntimeDefaults(content, runtime)` (single shared helper)
  stamps `--default <runtime>` + `use_worktrees=false` for every `runtime !=
  claude`; called from both `_applyRuntimeRewrites` and, crucially,
  `copyWithPathReplacement` in bin/install.js (the real workflow emit path).
- Generalize the fail-closed worktree guard `= codex` -> `!= claude` in
  execute-phase/quick/diagnose-issues (worktree isolation is Claude-Code-only).
- Flip manager/autonomous inline-vs-background gating to `codex -> background,
  everything-else -> inline` (research: only Codex can background-nest the
  pipeline's subagents; all others run inline, which they support).

Worktree-capability determination is research-backed (official docs for all 14
non-Claude runtimes: none honor GSD's isolation="worktree" mechanism, only Codex
background-nests). New end-to-end real-install test asserts the EMITTED workflow
is stamped — the regression guard that would have caught the dead-on-arrival bug.

Closes #1521

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

* chore(#1521): backfill changeset PR number (#1537)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 13:48:47 -04:00
Tom Boucher
2436b76980 fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees (#1519)
* fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees

A Codex install with a runtime-neutral .planning/config.json resolved
RUNTIME=claude and enabled git worktree isolation, which Codex's
spawn_agent cannot honor. Two root causes:

1. Workflows read `config-get runtime` / `config-get workflow.use_worktrees`
   without `--raw`, so config-get's JSON-quoted output ("codex") was captured
   verbatim into the bash var and broke every `[ "$RUNTIME" = ... ]` check —
   the Codex fail-closed guard was dead even when runtime:codex was explicit,
   and Claude's own worktree degrade-check was dead too. Add `--raw` to those
   reads across execute-phase, autonomous, manager, diagnose-issues, quick.

2. The conversion engine emitted `--default claude` for every runtime. Stamp
   the codex-emitted workflows to `--default codex` (runtime) and
   `--default false` (use_worktrees) in _applyRuntimeRewrites case 'codex', so
   a neutral config on a Codex install resolves runtime=codex / worktrees off.

Also extend the Codex fail-closed worktree guard to quick.md and
diagnose-issues.md (they spawned isolation="worktree" with no runtime guard).

Regression test asserts source<->engine parity across all five workflows
(DEFECT.GENERATIVE-FIX) plus fast-check property coverage of the stamping.

Closes #1515

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

* chore(#1515): backfill changeset PR number (#1519)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 11:56:59 -04:00
Tom Boucher
a77f3c4b3a docs(#1452): document workflow.context_guard_mode in CONFIGURATION.md
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 18:12:03 -04:00
Tom Boucher
c330f70f65 feat(#1494): add workflow.mvp_mode to VALID_CONFIG_KEYS; document code_review_command and plan_chunked in planning-config.md (#1500)
* feat(#1494): add workflow.mvp_mode to VALID_CONFIG_KEYS; document code_review_command, plan_chunked, mvp_mode in planning-config.md

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: backfill PR number 1500 in changeset

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 15:32:07 -04:00
Tom Boucher
0d56f544d2 feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation (#1458)
* feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation

ADR-1244 Phase 6. Adds the capability matrix reference, generated FROM the committed registry so it
can never drift from the actual capability set:
- scripts/gen-capability-matrix.cjs (--write / --check); --check is a CI drift guard.
- tests/capability-matrix-sync.test.cjs (4 tests): drift guard, buildMatrix==committed, every cap
  present, no placeholders.
- docs/reference/capability-matrix.md regenerated from the registry (release-stable: shows engines.gsd,
  omits the lockstep per-cap version that would churn the file every release).
- Consolidated the duplicate trust-model doc: deleted docs/explanation/the-capability-trust-model.md,
  merged its content into capability-trust-model.md, redirected ~10 references; no stale links remain.
- Diataxis verification (now that gsd capability is a real command): corrected the matrix's third-party
  section — the matrix is the first-party catalogue; the overlay-aware view of installed third-party
  capabilities is 'gsd capability list', not this generated file.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1435): Added changeset for the capability matrix reference

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1435): address code-review — non-vacuous matrix test + generator polish

- capability-matrix-sync.test.cjs: assert the 'security registers a ship:pre gate' precondition
  unconditionally so the extension-point check can never degrade to a vacuous pass on registry drift.
- gen-capability-matrix.cjs: warn (stderr) on an unknown loop point at generation time; rename
  enginesOf -> fmtEngines for consistency with the other fmt* helpers (output unchanged).
- capability-trust-model.md: point the two how-to links at the real files
  (import-a-capability-from-a-url.md, version-a-capability.md) instead of the bare directory.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1435): backfill changeset PR number → #1458

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 12:45:26 -04:00
Tom Boucher
9219af3360 feat(#1433): capability trust gate + upgrade/compat (ADR-1244 Phase 4) (#1449)
ADR-1244 Phase 4 (D5 trust + D6 upgrade/compat). capability-trust.cjs (disclosure/consent, strict_known_registries, engines+compatVersions, reserved namespace) + capability-lifecycle.cjs (install/upgrade/remove/reconcile; ledger-as-commit-point _pending intent; atomic stage-then-swap; surgical marker-isolated shared-edit strip; owner-token lock) + capability-source promote/skipEnginesGate seams + loader pending-skip + config keys. No sandbox re-derived (consent+integrity+reversibility). 6 Codex adversarial rounds + /security-review (no HIGH) + /code-review; gsd-test green both platforms; CI green. Phase 5 (#1434) wires the CLI dispatch.

Closes #1433.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 21:40:37 -04:00
Tom Boucher
353f63d170 feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) (#1440)
* feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2)

Promote the registry from a frozen data file to loadRegistry({includeInstalled}),
composing the first-party registry with a validated installed overlay (ADR-1244 D2):

- Extract the conformance validator to a shared runtime-callable module
  (gsd-core/bin/lib/capability-validator.cjs); the generator re-exports it
  verbatim, guarded by a generative-parity test (no build-time/runtime drift).
- capability-loader.cts: loadRegistry({includeInstalled}) composes first-party
  ∪ validated overlay from $GSD_HOME/.gsd/capabilities (global) and
  <root>/.gsd/capabilities (project) via the canonical buildRegistry. First-party
  always wins (id/skill/agent/config/command-family + reserved gsd-/anthropic-
  prefixes); full merged-set cross-capability validation; engines.gsd load-time
  re-gate (skip-with-warning); gate-kind capabilities FAIL CLOSED; fragment-path
  escapes rejected.
- semverSatisfies (hand-written, no dep) for the engines.gsd gate, fail-closed.
- Wire surface/state + loop to the overlay; loop injects a blocking gate for each
  skipped gate-kind overlay (fail-closed).
- cwd-aware overlay config-key federation: config-loader _federatedConfigSchema(cwd)
  + config-schema isValidConfigKey(key, cwd) compose the overlay per loadConfig/
  config-set call (never eager at module load, never wrong-cwd); first-party path
  unchanged with no cwd.
- run-tests.cjs sandboxes GSD_HOME (idempotent — nested spawns reuse it) for test
  hermeticity; capability-loader.cjs git+eslint-ignored (tsc artifact);
  capability-validator.cjs stays linted (#551 migration coverage).

Closes #1431

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1431): add changeset for runtime capability registry overlay

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1431): kill config-schema cwd-aware federation mutants (Stryker ≥52)

The cwd-aware overlay config-key federation added to config-schema.cts
(_capabilityConfigSchema(cwd) + isCapabilityConfigKey/isValidConfigKey cwd
threading) introduced mutable surface uncovered by config-schema's mutation
test set, dropping its score to 39.58% (below the 52 break threshold). Add a
real-overlay-fixture describe block exercising every branch (cwd guard, overlay
loadRegistry, found-branch, first-party fallback, cwd threading); local Stryker
score 39.58% -> 77.08%.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 14:41:20 -04:00
Tom Boucher
137760a655 fix(#1296): align config docs/prompts/schema with consumers (#1299)
* fix(#1296): align config docs/prompts/schema with consumers

The user-facing config surface disagreed with what the consumers actually do
(subset of the #1216 audit). No runtime consumption behavior changes.

- workflow.subagent_timeout: settings-advanced.md prompt + docs/CONFIGURATION.md
  said "seconds (default 600)" but the consumer (map-codebase.md) uses
  milliseconds (default 300000). Relabeled all four spots in settings-advanced.md
  (prompt, parse-default list, example, confirmation table) + the CONFIGURATION.md
  row.
- review.models.<cli>: settings-integrations.md, docs/CONFIGURATION.md (Integration
  Settings), and docs/CLI-TOOLS.md documented a shell command, but review.md injects
  the value into a --model/-m flag. Relabeled to a bare model id and reconciled the
  contradictory CONFIGURATION.md sections.
- workflow.test_command + workflow.build_command: consumed via config-get
  (test_command in verify-phase/execute-phase/audit-fix/post-merge-gate;
  build_command in post-merge-gate) and documented, but absent from validKeys so
  `config set` rejected them. Registered both in config-schema.manifest.json and
  documented them in references/planning-config.md (overview + complete reference).

Regression tests: behavioral config-set tests (tests/config.test.cjs) + doc-parity
content guards (tests/config-field-docs.test.cjs).

Deferred to other #1216 clusters: security-gate wiring, search_gitignored wiring,
mvp_mode, source_grounding_authority labeling, and config-set enum enforcement.

Closes #1296
Refs #1216

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(changeset): Fixed fragment for #1296 config-surface alignment

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 19:57:11 -04:00
Tom Boucher
cf68841220 enh(#1243): consume Claude plugin-provided skills in agent_skills (epic #1258 Phase B) (#1261)
* feat(#1243): consume Claude plugin-provided skills via native Skill-tool directive + grant Skill to agent_skills-consumer agents

- Relax global skill name validation to accept namespaced form `^[A-Za-z0-9_-]+(:[A-Za-z0-9_-]+)*$`
- Namespaced names (containing colon) on claude runtime emit a Skill-tool load directive instead of a @-include line
- Namespaced names on non-claude runtimes are skipped with a warning
- Bare unresolved names retain existing warn-and-skip behavior (no promotion to directive)
- Grant `Skill` tool to all 22 agent_skills consumer agents; 5 generated agents updated via research-profiles.cjs + regen, 17 hand-authored agents edited directly
- Add 16 TDD tests in describe('bug #1243') covering happy/mixed/precedence/negative/cross-runtime/regression/grant cases

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#1243): document plugin-provided skills in agent_skills

Update the Agent Skills Injection reference in CONFIGURATION.md with
the three entry forms (project-relative, global:<name>,
global:<plugin>:<skill>), the Claude-only runtime behaviour of the
namespaced form and the warn-skip on other runtimes, the plugin
pre-install prerequisite, and the consumer-agent Skill tool grant.

Add docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md with a
step-by-step guide for installing the plugin, locating the namespaced
skill name, wiring it into agent_skills, and verifying injection.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1243): align agent_skills docs with emitted block format + mixed-block regression test (code-review)

- Replace two-section mixed-block example (bogus "Load these plugin-provided skills using the Skill tool:" header) with the actual single-section inline format in CONFIGURATION.md and docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md
- Fix quoted warning text in how-to doc to exactly match the emitted string: [agent-skills] WARNING: Plugin-namespaced skill "global:<name>" requires a Skill-tool-capable runtime (claude) — skipping on runtime "<runtime>"
- Replace phantom agent slugs (gsd-checker, gsd-researcher, gsd-advisor, gsd-synthesizer) in CONFIGURATION.md Supported Agent Types with real agents/gsd-*.md examples (gsd-plan-checker, gsd-phase-researcher, gsd-code-reviewer, gsd-ui-auditor, gsd-research-synthesizer)
- Add byte-identical mixed-block regression test: one path-resolvable global skill + one plugin-namespaced skill on claude runtime → asserts r.ir.block === single-section interleaved block, no secondary header

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1243): regenerate agent-size baseline for the Skill-tool grant

The 22 agent_skills-consumer agents each grew +7 bytes from adding `Skill`
to their tools list; refresh the committed per-agent size baseline (#1074 guard).

* chore(#1243): add Added changeset fragment

* fix(#1243): traceable allow-test-rule ref + separator-agnostic byte-identical tests (CI)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-15 00:49:12 -04:00
Tom Boucher
a375c4b354 feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) (#1201)
* feat(#956): add MemPalace memory capability (ADR-857 feature plug-in)

Adds an opt-in, default-resilient ADR-857 feature capability that wires
MemPalace (local-first memory: MCP server + CLI) into the GSD loop:
deliberate recall before discuss/plan and verbatim + temporal-KG capture
at phase boundaries. Three memory modes (augment default; kg_backend and
replace forward-declared). Master gate mempalace.enabled (default off);
every hook onError:skip, zero gates; absent/disabled MemPalace => loop
unchanged. Transport is rendered-markdown only — MemPalace runs
out-of-process, no third-party code in gsd-core (ADR-857 §7).

Capability: capabilities/mempalace/ (manifest + 2 fragments), skills
commands/gsd/mempalace-{recall,capture}.md, agent
agents/gsd-mempalace-curator.md. Registration: ns-context router,
utility cluster, KNOWN_SKILLS, help full.md, model-catalog, copilot
install list, size baselines; regenerated capability-registry +
inventory manifest. ship:post wired into ship.md (wire-on-demand).

HELD on #1196: this capability also declares hooks at discuss:pre and
discuss:post, which are structurally un-wireable until the host-loop
conformance model covers the discuss phase (discuss-phase.md is not in
HOST_LOOP_FILES). The phase6-capstone-conformance gate therefore fails
on exactly those two orphaned points by design — see #1196. Once #1196
lands, rebase onto next and the gate goes green with no further change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#956): backfill changeset PR number (#1201)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 02:09:10 -04:00
Tom Boucher
44024aa535 fix(#1133): honor model_policy on the claude runtime (forward-port to next) (#1144)
Forward-port of the #1133 hotfix (commit 22f237a5, 1.4.5 hotfix line) onto
next. The hotfix was authored against src/core.cts (v1.4.4); on next the
resolver logic lives in src/model-resolver.cts (ADR-457 extraction, #888),
so the patch is re-applied there rather than cherry-picked.

resolveModelInternal step 2.5 now honors model_policy on the claude runtime:
the policy-resolved full model ID is mapped back to a Claude Code agent alias
via CLAUDE_POLICY_ID_TO_ALIAS (reverse of MODEL_ALIAS_MAP + claude-fable-5 ->
fable). Bare aliases (opus/sonnet/haiku/fable) pass through; an ID with no
Claude alias warns once to stderr (deduped by agentType::policyModel::tier)
and falls back to the configured tier alias. Non-claude runtimes return full
IDs verbatim (unchanged). resolveModelForTier is intentionally unchanged.

The warn-dedupe cache lives in model-resolver.cts; core.cts composes the
exported _resetRuntimeWarningCacheForTests to clear both that cache and the
config-loader warning cache (config-loader cannot import model-resolver --
circular dependency).

Ports the 6 #1133 tests (rewriting the old claude-no-op test that asserted
the bug) plus one added test covering the MODEL_ALIAS_MAP reverse-map path.

Forward-port of #1133

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 19:56:54 -04:00