Commit Graph

101 Commits

Author SHA1 Message Date
0xdhx
0396d9cab1 enhance(#2483): stop the claude reviewer lane from inheriting CLAUDE.md + auto-memory (#2493)
* enhance(#2483): env-guard the claude reviewer leg against CLAUDE.md injection

The claude reviewer in workflows/review.md was a bare headless `claude -p`
spawn run from the project cwd, so it inherited the invoking user's global
CLAUDE.md, the project CLAUDE.md, and Claude Code auto-memory.

That made it the only reviewer leg seeing anything beyond the prompt file.
gather_context assembles PROJECT.md, the roadmap section, every PLAN file,
CONTEXT.md, RESEARCH.md and REQUIREMENTS.md into the prompt before any
reviewer runs; the gemini leg receives only that prompt and the codex leg
runs --ephemeral. Beyond the measured ~4k tokens/spawn, the asymmetry cuts
at the workflow's own premise: "independent review" meant something
different for the claude leg than for the other two.

Guard both dispatch lines with a per-invocation
`env CLAUDE_CODE_DISABLE_CLAUDE_MDS=1`. `env`, never `export` — the flag
must not leak into the orchestrating session (which may itself be Claude
Code on the SELF_CLI="auto" path) or into any later spawn.

review.md is the only claude -p call site in the installed tree, so this is
two lines on one surface. The self-skip logic is untouched.

* enhance(#2483): fix CRLF-fragile split and regenerate workflow baselines

Two CI failures from the first push, both mine:

1. lint-tests: the new regression test split readFileSync content on a
   literal "\n". On a Windows git-autocrlf checkout that leaves a trailing
   "\r" on every line (local/no-crlf-fragile-split). Use .split(/\r?\n/).

2. golden-install-parity / workflow-size-budget / workflow-compat: editing
   gsd-core/workflows/review.md changes its content hash and byte size, and
   both are pinned in committed baselines. Regenerated via the repo's own
   generators (npm run size:baseline, npm run gen:golden).

The regenerated diffs are review.md-only: exactly one hash line per
golden-install-parity fixture and one size entry in workflow-size-baseline
— no unrelated drift swept in.

Full suite now green locally: 2113 pass, 0 fail, 3 skipped (run with HOME
and CLAUDE_CONFIG_DIR overridden to throwaway dirs; live profile verified
untouched afterward).

* enhance(#2483): adapt guard-test matcher to the effort-args dispatch reshape

The effortSurface wiring (#2481) reshaped the bare-model dispatch to
`claude $CLAUDE_EFFORT_ARGS -p -`; the invocation matcher's dash-first
form could no longer see it, and the count assertion failed exactly as
designed. The matcher now tolerates variable expansions between `claude`
and its first literal flag. Negative-controlled both ways: a stripped
guard and a deleted dispatch line each still fail.

* enhance(#2483): also guard the claude leg against auto-memory injection

CLAUDE_CODE_DISABLE_CLAUDE_MDS suppresses CLAUDE.md file loading;
auto-memory is an independently-toggled mechanism with its own flag.
Add CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 to both dispatch lines, correct
the docs/COMMANDS.md and changeset claims that credited the first flag
with covering auto-memory, and extend the regression test to require
both flags on every claude invocation (negative-controlled: 2/4
assertions fail with the new flag removed).

* enhance(#2483): match the claude binary in command position, not argument position

The line-oriented invocation matcher counted any line where the token
`claude` was followed by a flag. #2589 (landed on next as 920a5f3f)
reshaped the effort-args lookup from

  --host claude 2>/dev/null | jq -r '.effort_argv_string // ""'

to

  --host claude --pick effort_argv_string

which put a flag immediately after `claude` and made the config query
read as a third claude dispatch, failing the count assertion.

The defect class is a binary name in *argument* position being read as a
command. Fixed at the class rather than the instance: tokenise the line
and skip any `claude` whose preceding token is a flag. That also covers
the latent sibling one line away in review.md (`command -v claude`),
which escaped today only because its next token is a redirect.

Negative-controlled four ways: stripping CLAUDE_CODE_DISABLE_AUTO_MEMORY=1
fails, stripping the whole env guard fails, adding a genuine third
unguarded dispatch (`timeout 900 claude --output-format text -p -`) still
fails — so the narrowing did not blind the matcher to reshapes, which is
the property the count assertion exists for — and the pre-#2589 jq form of
the lookup still passes, so the matcher is not pinned to today's base.

* enhance(#2483): carry the claude reviewer's memory guard as declared lane data

ADR-2782 Phase 5b replaced the hand-authored per-CLI dispatch legs in
review.md with the declared lane table, so the two `env`-prefixed shell
lines this PR previously added no longer have a surface to live on. The
guard is reimplemented where the lane contract now lives.

`SpawnInvoke` gains an optional `env`, the claude lane declares the pair,
the resolver folds own string-valued entries into `SpawnPlan.env` (absent
or empty resolves to null, so the runner has one shape to test), and the
runner passes it to spawn. Production merges it OVER `process.env` into a
fresh object for that one child, so nothing reaches the orchestrating
session or any other lane in the run.

Declared data rather than a handler (D6): the pairs are static per lane,
which is precisely what the manifest vocabulary is for. The capability
manifest carries the same field, because the lane-fidelity test compares
manifest and descriptor over the union of `invoke`'s keys.

The regression test is rewritten against the resolver and runner rather
than review.md's text. It gains the property the source-text assertions
could only approximate: that `process.env` is never mutated.

Scope boundary, asserted rather than left in prose: `env` is not part of
the trust-disclosure surface, which is safe only while no manifest body
reaches the resolver — the registry's reviewer bodies contribute slugs to
the parity check and execution resolves from `REVIEWER_LANES`. The new
test fails first if that ever changes.

* enhance(#2483): restate the guard's mechanism in the docs and changeset

Both described the fix as two `env`-prefixed dispatch lines, which is the
surface ADR-2782 Phase 5b removed. The user-visible behaviour is
unchanged; the carrier is not, and a changeset that ships a description
of a mechanism the tree does not have is a CHANGELOG entry nobody can
verify against the code.

* enhance(#2483): cover the production spawn wiring end to end

The unit tests stop at the runner's `deps.spawn` seam — every one injects
a spy. Production supplies that seam in `gsd-core/bin/gsd-tools.cjs` as a
hand-written object no test constructs, so the chain could be correct all
the way to `SpawnPlan.env` and the merge could still be wrong or absent
with the suite green. Deleting those four lines was the one mutation that
left every other control silent.

This runs the real `spawnSync` through `gsd-tools review-lane invoke`,
with a `claude` shim on PATH that records the environment it was handed.
It asserts both halves in one test: the pair arrives, and an unrelated
inherited variable survives — a wiring that REPLACED the environment
rather than merging over it would satisfy the first and break every
lane's PATH and HOME.

POSIX-only; mediating a Windows `.cmd` shim is a separate concern the
repo already tests on its own.

Noted rather than fixed: `timeout`, `killSignal`, `maxBuffer` and
`shell: false` on that same object are equally uncovered. That is the
epic's gap, not this change's, and closing it is not in scope here.

* enhance(#2483): validate the invoke.env shape and register it as spawn-only

`env` was the one spawn-invoke field with no shape enforcement: every sibling in
`validateSpawnInvoke` is checked, and a manifest declaring `env` as an array, a
string, a number, or an object with non-string values passed validation in
silence. That matters more than an ordinary schema gap here, because
`resolveLanePlan` DROPS a non-string value rather than coercing it — so an
unvalidated manifest declares a pair that never reaches the spawn, which is the
failure a memory guard can least afford.

Two registrations, not one. `env` was also absent from
`SPAWN_ONLY_INVOKE_FIELDS`, which is the list the openai-http arm rejects
against — so `invoke.env` was accepted on a transport that issues an HTTP POST
and has no child environment at all. It was the only spawn-shaped field accepted
there; the other six each produce two errors. Self-found while sweeping the
class, not raised in review.

Keys are held to the portable POSIX environment-name grammar. That is a policy,
not a claim about what an environment can hold: measured, only NUL is actually
rejected by `spawnSync`, while `=`, a leading digit, a dash and a space are all
carried through to the child (an `A=B` key arrives as the raw entry `A=B=value`).
They are refused because a name outside the grammar is not portably addressable
by the program meant to read it.

`__proto__` is refused for a different and concrete reason. It passes that
grammar and is a real own key once a manifest is JSON-parsed, but assigning it
onto a plain accumulator goes through the inherited `__proto__` setter rather
than creating an own property — and for the string values this field permits the
setter is a no-op that does not even change the prototype. The pair would
validate and then simply vanish before the spawn. (An environment CAN carry a
literal `__proto__` entry; this is about the resolver's accumulator, and the
error message says so.)

Deliberately narrower than the sibling reserved-name guards in this file, which
also reject `constructor`/`prototype`: those guard bracket lookups that resolve
prototype members, whereas this reads via `Object.keys` plus an own-value read,
where `constructor` assigns as an ordinary key the spawn could carry.

`effortChannel` is deliberately left in neither field list: ADR-2782 D2 defines
it for both transports, so it is shared rather than spawn-only.

Reversion-controlled, three mutations, all three fire a named test: dropping
`env` from the discriminator fails `httpTransportRejectsEnv`; removing the
`__proto__` arm fails `envRejectsProtoKeyThatWouldSilentlyVanish`; disabling
the block fails four.

(#2483)

* enhance(#2483): amend ADR-2782 D2 for the invoke.env vocabulary widening

D2 records the spawn `invoke` shape as a closed vocabulary, and its Amendments
section carries a dated entry for every prior widening (Phase 1 #2794, Phase 2
corrections #2795, Phase 5b #2799). This change extended that vocabulary in code
without touching the ADR governing it, so the ADR contradicted the
implementation — and the repo's own convention, recorded in CONTEXT.md, is that
the ADR is amended in the same PR precisely because the prior widenings did it
correctly.

Adds the `invoke.env` row to the D2 table and a dated Amendments entry.

The entry also corrects the authority this change cited. The source comment
pointed at D6, which governs the closed `handler` enum — imperative behavior
admitted first-party — and says nothing about the `invoke` field vocabulary.
That is D2's territory, so the citation never covered the gap.

Two claims are corrected rather than restated, both about the trust boundary
that justifies leaving `env` out of the D5 disclosure signature:

- The regression test does not enforce that boundary. On one forged lane it
  shows the resolver folds whatever it is handed, so a future path feeding it
  manifest lanes would not make any assertion in that test fail. Its comment
  claimed it "will fail first"; that was wrong, and both the comment and the
  ADR now say the boundary is a property of the production call chain instead.
- The ADR is internally inconsistent on whether third-party manifest lanes
  execute at all: Consequences says adding a reviewer needs "no core patch",
  while `gsd-tools.cjs` rejects every slug absent from the first-party
  REVIEWER_LANES map. CONTEXT.md, `workflows/review.md` and the resolver's own
  header take the first view. #2483 did not create that inconsistency and does
  not resolve it; the entry records it rather than settling it in its own favour.

(#2483)

* enhance(#2483): document invoke.env in the capability-manifest reference

ADR-2782 points capability and plugin authors at
`docs/reference/capability-manifest.md` as where the lane vocabulary must be
visible, and its `invoke` row enumerates the spawn sub-shape field by field.
`env` was absent from that table while being part of the real shape, so the one
document a third-party capability author would actually consult to learn the
field exists did not mention it.

Squarely Diataxis reference material — a field-by-field schema description — so
it goes here rather than in the user-facing prose, which was already updated.
States the constraints a manifest author can actually trip, and is explicit that
the name grammar is a portability policy rather than an OS limit, so a reader
does not take it for a claim about what an environment can hold.

(#2483)

* enhance(#2483): disclose and sign the reviewer lane's env and residual invoke fields

`invoke.env` was undisclosed at install time. That was defensible while manifest
lanes could not execute — the premise this PR's own ADR amendment recorded — and
#2927/#3062 retired it: `routeReviewLane` now merges installed overlay `reviewer`
bodies into its lane map via `mergeReviewerLanes`, which is a field-identical merge
by ADR-2782 D1 and deliberately does not deep-validate. An overlay's whole `invoke`
therefore reaches `resolveLanePlan`, and `env` reaches the spawned child. A consented
third-party capability could set `NODE_OPTIONS=--require ./evil.js` on a reviewer lane
with no install-time disclosure and no re-consent.

The same file already decided what `env` means in a manifest: MCP servers fold it into
the disclosure signature and render each key and value in the consent prompt, with an
inline rationale naming this exact shape. Reviewer lanes get the identical treatment.

`env` was the ninth unsigned invoke field, not the first. `defaultHost` (the manifest's
OWN fallback egress host, used whenever the config key resolves to nothing),
`path`, `outputChannel`/`outputArg`, `modelArg`, `effortChannel` and `modelDiscovery`
all reach `resolveLanePlan` and none was bound. Enumerating a ninth name leaves the
tenth open, so the lane signature carries a RESIDUAL of every other declared `invoke`
key — the completeness backstop `rawConfig` already gives the MCP line (#1459 finding 5),
and the "sign the whole object" remedy the recorded decision on this class prefers.

`defaultHost` is also rendered: `resolvedHost` comes from user config, so a lane whose
key is unset displayed "(unresolved …)" — which reads as "no destination" — while the
runtime egresses the plan and review text to the address the manifest picked.

D4.5 is preserved one level down: the extra element is appended ONLY when the lane
declares something beyond the eight already-bound fields, so an env-free lane's
signature stays byte-identical and no already-consented capability is re-prompted for
a field it does not use. A lane that does declare one re-consents, which is the point.

Execution-primitive env names are FLAGGED in the prompt, not refused in the validator.
A denylist cannot be the boundary here: `PATH` alone is a complete execution primitive
for a spawn lane and can never be refused, the child is an arbitrary third-party binary
so the true set spans every interpreter's injection vars, and the MCP `env` this mirrors
refuses nothing and discloses everything. Missing a name costs a quieter line, never a
boundary.

Refs #2483.

* enhance(#2483): exercise the real overlay merge path in the guard test

The test named for the manifest/first-party boundary did not test it. It built a
forged lane locally, handed it straight to `resolveLanePlan`, and asserted that
`REVIEWER_LANES` did not contain it — so no assertion in it depended on the claim its
name made, and a code path that fed manifest lanes to the resolver would not have made
it fail. Its own comment said as much, and named the production chain as the real
carrier of the guarantee: "gsd-tools.cjs builds its lane map solely from REVIEWER_LANES".

That sentence is now false. #3062 merged overlay reviewer bodies into that map, so the
test's premise and its subject both moved.

The replacement routes through `mergeReviewerLanes` — the real helper the production
path calls — and asserts the overlay lane is admitted, resolves, and carries its `env`
into `SpawnPlan.env`. That makes the security property falsifiable instead of narrated.
It then asserts what now backs it: the env is disclosed on the surface, rendered key
and value in the consent prompt, flagged when the name is an execution primitive, and
bound to the signature so a value change, an addition, or a removal each force
re-consent.

Three further cases, because the finding's generative half is what stops it recurring:
the residual backstop is asserted against five fields including one that does not exist
(`aFieldThatDoesNotExistYet`), so a future vocabulary widening cannot silently re-open
this; a fully-enumerated lane is pinned to its original 8-tuple, which is what keeps the
fix from re-prompting every consented capability; and an http lane's manifest-declared
`defaultHost` is asserted to reach both the prompt and the signature.

Reversion-controlled, seven mutations, all seven fail a named test: env dropped from the
surface, the prompt's env line removed, the execution-primitive warning removed, the
signature's extra element never appended, the residual emptied, the defaultHost line
removed, and the declares-something test un-widened. The last of those was SILENT on its
first run and its test was written in response, then the control re-run.

Refs #2483.

* enhance(#2483): correct the ADR amendment's manifest-lane premise

The amendment argued `env` needed no D5 disclosure because a manifest's `invoke`
fields never reach `resolveLanePlan`. That was true when written and #3062 retired it
22 hours after this branch's last commit: `routeReviewLane` now builds its lane map
from `mergeReviewerLanes(REVIEWER_LANES, loadRegistry({includeInstalled: true}))`, and
D1's no-translation-layer rule makes that a field-identical merge, so an overlay's
whole `invoke` reaches the resolver and executes.

The entry had named this exact trigger — "were manifest lanes ever made executable,
`env` must join the disclosed surface in that change, and nothing here will trip if it
does not." Nothing tripped. The premise is rewritten to current truth rather than
annotated, because an ADR is read in fragments and a superseded paragraph left standing
reads as live reasoning to the next author; a one-line dated tombstone points at git for
the withdrawn text.

The rewritten entry records four things the first draft could not: that the enumeration
itself was the defect (`env` was the ninth unbound `invoke` field, and `defaultHost` and
`path` are egress-relevant on their own), that the residual is what closes the class,
that D4.5's byte-identical-signature property is preserved by appending the residual only
when a lane declares something beyond the eight bound fields, and that consent — not
shape validation — is the boundary, since no honest env denylist can exclude `PATH`.

It also closes the internal inconsistency the previous entry could only record. This ADR,
`CONTEXT.md`, `gsd-core/workflows/review.md` and `resolveLanePlan`'s own header all said
overlay lanes reach the resolver while the runtime said otherwise; #3062 resolved that in
the documents' favour, which is what makes the disclosure mandatory rather than defensive.

Refs #2483.

* enhance(#2483): record in the manifest reference that invoke fields are consent-bound

`docs/reference/capability-manifest.md` is the field table ADR-2782 points capability
authors at, and it described `invoke` purely as a schema. A third-party author reading it
could not learn that everything they declare there is shown to the user at install and
bound to the consent signature — which is exactly what they need to know now that an
overlay reviewer lane executes (#2927/#3062).

States the two things the schema alone cannot: that `env` and `defaultHost` are named in
the consent prompt and the rest is covered by a residual, so any change to a declared
`invoke` field forces re-consent; and that `env`'s validation is a portability policy
rather than a safety boundary, since `PATH` is a complete execution primitive and cannot
be refused. Names that are execution primitives are highlighted in the prompt instead.

Refs #2483.

* enhance(#2483): add a Security changeset for the reviewer-lane disclosure

The existing fragment describes the enhancement this PR was opened for and stays as it
is. The disclosure fix is a separate user-visible change of a different type: a
capability declaring `invoke.env` or `invoke.defaultHost` will ask for consent once
more, and users are entitled to read why in the changelog rather than discover it as an
unexplained prompt.

Type is `Security` rather than `Changed` because the entry describes a closed
code-execution disclosure gap, not a behaviour adjustment.

Refs #2483.

* enhance(#2483): correct this round's own claim about who gets re-prompted

Self-found while auditing the round's claims before publishing them. The changeset and
the ADR entry both stated that a capability declaring `invoke.env` or `defaultHost`
"will ask for consent once more". That is wrong, and it overstated the cost of the fix
in the one direction a maintainer would have had to take on trust.

A code change to `disclosureSignature` re-prompts nobody. `hasProjectConsent` matches on
the recomputed bundle `contentHash` — the signature has not been the security binding
since #1459 CB-1/CB-2 — and the upgrade path's `executableSetChanged(old, new)` compares
two disclosures both computed by the CURRENT code, so widening the signature moves both
sides of that comparison equally. First-party capabilities never reach the path at all:
the install flow blocks a first-party id before trust evaluation.

What the widening actually buys is forward-looking, and is the real argument for it: an
upgrade whose manifest edits a declared `invoke` field now registers as an
executable-surface change and re-consents, where before it could change what the lane
runs in silence.

Also measured and recorded, because the D4.5 property was stated more strongly than it
deserved: of the twelve first-party reviewer capabilities, ZERO are in the
byte-identical-signature class — every real lane declares at least `effortChannel`. The
property is a guarantee about minimal lanes, not a description of the fleet, and the ADR
now says so.

Refs #2483.

* enhance(#2483): sign and disclose the probe binary and the lane's outer fields

Found by this round's own adversarial review, and it is the same defect one level out:
the `invoke` residual cannot reach the lane body's OUTER fields, and `probeLane` SPAWNS
`probe.binary` with `--help` before dispatch (`review-lane-runner.cts`, the
`command-exists`/`command-capability` arms). An overlay naming an arbitrary probe binary
therefore executes it — unsigned and undisclosed, exactly as `invoke.env` was, and
reachable on the same #3062 path.

The lane element now carries a second residual over the outer fields, and the probe
binary is shown in the consent prompt when it differs from the dispatch binary — it is a
program that runs, and the user is entitled to see it.

TWO fields stay excluded, and that is a decision rather than an omission:
`reviewsSection` and `timeoutFloorMs` are ADR-2782's cosmetic carve-outs (matrix
A10/A13), where re-consenting would present a prompt carrying no security information.
A test pins that they remain excluded, so a later widening cannot quietly reverse D4.5
while claiming to complete this fix.

Also corrects a miscount introduced by the previous commit: the source comment said the
enumeration had fallen behind by "seven fields" and omitted `fallbackModel`, while
asserting `env` was the ninth. `resolveLanePlan` reads twelve `inv.*` fields and four
were bound, so the number is eight. The comment now states the derivation rather than
just the total.

Reversion-controlled: emptying the outer residual fails "repointing the probe binary must
force re-consent"; removing the render line fails its own named assertion.

Refs #2483.

* enhance(#2483): refuse execution-primitive env names as defence in depth

Adopts the review's B5 after this round's own adversarial pass refuted my reason for
declining it. I had argued a denylist was worthless because `PATH` can never be refused.
That was wrong on the facts: no shipped reviewer manifest declares `PATH`, so it can be
refused, and it is the most complete primitive in the set — repoint it at a directory
holding a fake binary and the declared `invoke.binary` is irrelevant. A list that cannot
be exhaustive can still close the highest-confidence, lowest-legitimacy routes.

So the validator now rejects `PATH`, `NODE_OPTIONS`, `LD_PRELOAD`, `DYLD_INSERT_LIBRARIES`,
`BASH_ENV`, `PYTHONPATH`, `PERL5OPT`, `RUBYOPT`, `GIT_SSH_COMMAND`, `JAVA_TOOL_OPTIONS`
and their siblings on a reviewer lane. A lane needing a specific executable declares an
absolute `invoke.binary` instead of reshaping the child's environment.

The comment states plainly that this is defence in depth and NOT the boundary — the
boundary is install-time consent, which discloses every declared pair and binds it to the
signature, so an unlisted name is still SEEN before it runs. That framing is load-bearing:
a future reader who mistakes the denylist for the control will under-invest in the one
that is, which is the failure mode I was trying to avoid by declining it outright.

Two tests: the rejection itself across ten names, and a guard asserting no shipped
reviewer capability declares a denied key — so if the list ever outgrows its evidence,
that surfaces as a decision rather than a silent removal.

Refs #2483.

* enhance(#2483): fix two stale D5 enumerations elsewhere in the ADR

The previous commit rewrote the amendment's premise but swept only the amendment. Two
normative passages earlier in the same ADR still enumerated the old closed field list and
now contradicted it: the `executableSetChanged` trigger list, and the split-binding note
asserting the seven manifest-derived fields were "everything that is SHA-pinned".

That is the failure the rewrite-don't-annotate rule exists to prevent, one section over —
an ADR is read in fragments, and a fragment carries no supersession marker, so a reader
landing on either passage would have taken the superseded enumeration as current.

Both now name the residual as the mechanism rather than restating a list, which is also
what stops them going stale the next time the vocabulary widens.

Found by this round's adversarial review, which grepped the whole document rather than
the section under edit.

Refs #2483.

* enhance(#2483): stop the probe disclosure claiming a spawn that does not happen

The probe line added one commit ago rendered "probes by running: <binary> --help" for
every lane. That is false for `kind: "command-exists"`, which only calls `hasBinary` — a
PATH/filesystem scan that starts no process. Only `command-capability` spawns.

A false statement in a consent prompt is worse than a missing one: the prompt is the
surface a user is asked to trust, and this one overstated what a lane does. Worse, the
test I wrote to prove the fix used `command-exists` — the kind that does NOT spawn — so
it pinned the wrong claim and would have kept the error green forever.

The surface now carries `probeKind` and the two kinds render differently: a spawn is
described as a spawn, a presence check as a presence check. The test exercises both, and
asserts the `command-exists` path never emits the spawn wording.

Also corrects the field-count parenthetical to state its derivation unambiguously —
`resolveLanePlan` reads thirteen `inv.*` fields including `env` (twelve before this PR),
four were bound, so eight were unbound before `env` and nine including it. The bare
"twelve" was true only of the pre-PR tree and read as a claim about the current one.

And retires two comments that argued AGAINST the validator denylist this round then
shipped. Leaving them would have handed the next reader the reasoning for removing it.

Reversion-controlled: conflating the two probe kinds fails a named test.

Refs #2483.

* enhance(#2483): match the reviewer-lane env denylist case-insensitively

The denylist added one commit ago compared exact case, so `Path`, `path`, `node_options`
and `Node_Options` all passed it. Windows environment lookup is case-insensitive, so
those reach the child as `PATH` and `NODE_OPTIONS` — the exact inputs the list names.

An exactly-cased denylist is worse than none: it reads as a control while admitting the
input it was written to refuse, and the next reader has no reason to doubt it. Members
are stored uppercase and the key is folded before lookup; the name grammar already
constrains keys to ASCII, so a plain fold is sufficient.

Reversion-controlled: restoring the exact-case compare fails `envDenylistIsCaseInsensitive`
on `Path`.

Refs #2483.

* enhance(#2483): correct the docs that still described the denylist as absent

Both the ADR and the manifest reference still said `env` carries no denylist and that
`PATH` "can never be refused" — written when that was this round's position, and left
standing after the round reversed it. A reader landing on either passage would have taken
the superseded argument as current, which is precisely the failure the rewrite-don't-
annotate rule exists to prevent.

Both now describe the denylist, name `PATH`'s inclusion and the case-insensitive match,
and keep the limit explicit: the list cannot be complete against an arbitrary child and
disclosure runs before validation, so consent remains the boundary.

The ADR's byte-identical-signature claim is also corrected rather than softened. With the
outer residual in place, a lane producing no residual is one the validator rejects — it
declares no `flags`, `probe`, `emptyOutput`, `evidenceClass`, `requiresBinaries` or
`promptBudgetKey`. So the property is about the ENCODING, not a claim that any real
signature is unchanged, and it is not the argument for the change being safe. That
argument is that consent binds to the bundle contentHash and no existing consent is
invalidated at all.

Refs #2483.

* test(#2483): cover the three new lane disclosure fields in the injection-safety parity guard

The PARITY test in section N exists to catch a renderer field that skips
`renderValueForPrompt` (#3248). Its payload manifest is hand-maintained, so it
covers the fields that existed when it was written — slug, binary, args,
hostConfigKey, handler — and none of the fields this PR adds.

This PR renders three further manifest-supplied values into consent-prompt
lines: `invoke.env` (keys and values), `invoke.defaultHost` and `probe.binary`.
The gap was silent rather than theoretical: with the lane env line reverted to
the pre-#3248 raw form, the whole 948-test lane/capability/trust-disclosure
suite stayed green.

Two manifests, because the shapes render disjoint lines — `defaultHost` only on
the openai-http branch, `env`/`probe` only where declared, and the probe line
only when the probe binary differs from the dispatch binary.

Non-vacuity is asserted on the typed disclosure object and on structural line
counts, not by substring-matching rendered prose: CONTRIBUTING.md forbids raw
text matching on test output, and this section's own header promises structural
assertions only, so a prose match here would have made that promise false.

Negative-controlled three ways against the merged tree, each producing exactly
one named failure: env rendered raw, defaultHost rendered raw, probe binary
rendered raw.

* docs(#2483): extend the #3248 render-site comment to the fields this PR adds

The comment enumerates every manifest-supplied value that must pass through
`renderValueForPrompt`, and it stopped at `handler` — the reviewer-lane fields
that existed when #3248 landed. This PR renders three more (`defaultHost`, the
probe binary, and the env keys and values), so the list understated its own
contract in the one place a future author would check before adding a fourth.

A comment enumerating a closed set is a set that can silently fall behind the
code it describes; the parity test added alongside is what makes the omission
fail loudly rather than read as deliberate.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-11 17:42:41 -04:00
Tom Boucher
cf6de5e1c0 feat(#2871): resolve triggers and host precedence, not just placement (#3291)
* test(#2871): failing-first suite for trigger-surface resolution

23 tests over the 50-test-matrix rows. RED by construction:
resolveTriggerSurface and DEFAULT_TRIGGER_PRECEDENCE do not exist yet,
and the validator silently ignores triggerPrecedence today.

Written in the per-runtime describe idiom the other four
runtime-artifact-layout suites use, not a table.

The rows that carry the weight: windsurf must NOT report a shadow it
does not have, since its global scope emits only agents and agents are
not trigger-bearing; agents and kimi-agents must be absent from the
output for every runtime; and reordering a runtime's triggerPrecedence
must flip the winner, which is the only assertion that proves the axis
is read rather than decorative.

Stems are injected, never scanned, so the surface is assertable with no
filesystem.

* feat(#2871): resolve triggers and host precedence, not just placement

resolveTriggerSurface(runtime, scopes) returns every /gsd-<name> trigger
a runtime emits, with the scope and kind that produced it, whether the
host registers it directly or only through a router, and which artifact
shadows it. resolveRuntimeArtifactLayout is untouched -- its 7 callers
need placement only and the issue requires them unchanged.

AGENTS ARE NOT TRIGGER-BEARING, and ADR-2866 said they were. The
host-integration matrix models command and dispatch as separate interface
points: an agent is invoked through the Agent tool's subagent_type, not
by typing a slash trigger, and _copyStaged never applies the kind prefix
to an agents entry. So agents and kimi-agents are excluded from the
surface entirely, and this commit amends ADR-2866 with a dated
correction. #2218's conclusion is unchanged -- the collision is strictly
commands-vs-skills, and claude's local /gsd-* trigger surface is still
fully shadowed -- but the ADR implied the local agents surface was lost
too, and it is not.

That correction is what makes windsurf come out right. Its global scope
emits only agents, so it has no global trigger and its local commands
are unshadowed. Model agents as trigger-bearing and windsurf falsely
reports a full shadow.

The triggerPrecedence axis lands on all 19 descriptors as an ordered
kind list, one value with one owner, rather than a numeric rank spread
across N kind entries with nothing keeping them consistent. Validation
uses a required-with-default shape that has no precedent in this
validator -- every existing axis is hard-required -- so a third-party
capability.json omitting the field still validates, which is what
ADR-894's additive-only contract promises.

Winner resolution reads Phase 1's scope rank first, then the kind
ordering. A test reorders the axis and asserts the winner flips, since
an axis that is added, validated and never consulted would pass every
other assertion.

shadowedBy ships unread. Phase 4 (#2873) is its first consumer, per this
issue's out-of-scope note.

Verified via the remote runner.

* fix(#2871): single-source namespacedByDir and close two test gaps

Four findings from the isolated adversarial review.

The namespacedByDir rule had reached three copies -- install-engine,
surface, and the new trigger resolver -- one of which carried a
hand-written keep-in-sync comment and no assertion. That is this repo's
generative-fix-divergence class. Extracted to one exported predicate all
three now call. Verified by diverging one copy deliberately: the existing
#816 parity test failed, and passes again on revert.

The omission test was vacuous. Row 16 asserted that a descriptor without
triggerPrecedence still validates, but built its fixture from claude's
shipped descriptor -- which this PR had just added the axis to. It now
clones and deletes the key, following the shippedDescriptorWithout
pattern, and asserts both that validation passes and that the resolver
still picks the right winner from the default. The second half is what
makes it prove anything.

resolveTriggerSurface silently dropped an unrecognized scope while every
sibling in this epic throws. Two phases of one epic should not disagree
about whether an invalid scope is an error, so it now rejects through the
same shared validator; an empty scope list still returns empty rather
than throwing.

The ADR amendment had been spliced into the middle of the References
list, orphaning its last bullet. Moved to the top, after the header
block, which is where ADR-3660 and ADR-1016 both put dated amendments.
No lint checks markdown structure, so this was green while malformed.

* fix(#2871): single-source the command filename composition too

The earlier fix shared the namespacedByDir boolean but left the
filename composition around it written twice -- once in _copyStaged as
what actually gets written, once in resolveTriggerSurface as what gets
predicted. The predictor could go stale silently.

One exported helper now composes it for both. The entry.name asymmetry
that looked like it would block extraction does not: entry.name is
filtered to end in .md and stem is entry.name minus those three
characters, so the two branches are the same string by construction.

Divergence proven to fail: injecting a marker into the helper broke the
trigger-surface suite; reverting restored 25/25. The four sibling layout
suites hold at 227 unchanged.

* docs(#2871): correct the ADR timing notes that this phase makes stale

The Amended by back-links on ADR-3660 and ADR-1016 were written in
Phase 0, when the widenings they describe had not shipped. Each carried
a forward-looking clause -- "the module changes at Phase 2, not before,
until then this module resolves placement only" -- which becomes false
the moment this PR merges. ADR-2866's own Amends header and its
reciprocal-notes section carried the same tense.

All four now describe what shipped. This is a tense and status
correction on Accepted ADRs, not a change to any decision.

Worth stating because it is the failure mode this epic keeps meeting:
gen-adr-index.cjs tracks only Supersedes and Subsumes, so nothing in CI
would have caught either the missing back-link in Phase 0 or these stale
clauses now. They stay correct only because someone checks.

* chore(#2871): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-09 22:25:42 -04:00
Tom Boucher
b901d1e06f feat(#1953): complexity-triggered refactor extension point (execute:post) (#3261)
* test(#1953): failing-first suite for the complexity-triggered refactor hook

60 behavioral cases against src/complexity-trigger.cts, which does not exist yet:
decision-point counting, the comment/literal stripping leak surface, threshold and
jump-delta boundaries at limit-1/limit/limit+1, stable-anchor baseline semantics,
and fs fault injection via mock.method. Two fast-check properties assert that
stripping never manufactures a decision point and that comments and string
literals are score-neutral.

Also registers the refactor-trigger capability manifest (inert until
refactor.trigger_enabled) and regenerates the capability registry and matrix.

Verified RED on the remote runner before any implementation exists.

* feat(#1953): complexity-triggered refactor extension point

Adds the opt-in refactor-trigger capability. After a phase executes, an
execute:post step measures per-function complexity for the files the phase
touched and writes a scoped refactor proposal when a function crosses the
configured threshold or drifts past its recorded anchor.

Design notes worth carrying:

- The signal is computed in-core (decision-point counting over comment- and
  literal-stripped source, Node builtins only) rather than via Memtrace or a
  shelled-out analyzer. The hook fires as a deterministic CLI, not an agent
  with MCP tools, and core takes no external dependencies — this is the only
  option a behavioral test can bind to. The metric sits behind a seam.
- The baseline is a stable anchor, not a rolling value: set on first
  observation, moved only on disposition. A rolling baseline makes the delta
  the single-phase change, so a function creeping +2 per phase never trips a
  delta of 5 and the jump check adds nothing over the absolute threshold.
- Strict mode records an open deviation window in the broken-windows ledger
  rather than declaring its own ship:pre gate. ship.md has no generic ship:pre
  gate dispatch — only two hardcoded branches — so a third gate of any kind
  would be declared and never evaluated.
- The gate clears on the proposal being dispositioned, never on the score
  improving. A blocking complexity number is one an executor can satisfy by
  splitting a coherent function in two.

execute-phase.md gains a generic execute:post step-dispatch contract; it
previously matched only ref.skill == "code-review", so any other step
registered there was declared and never run. The code-review branch is
unchanged.

Full rationale in ADR-1953.

Closes #1953

* fix(#1953): close git option injection and symlink escape in the refactor hook

Three findings from the isolated security review, all fixed inline.

HIGH — changedFilesSince interpolated the --since value into a revision
token placed before the -- separator. A -- only stops PATHSPEC parsing of
arguments after it; git still option-parses what comes before. So
--since '--output=/tmp/x' became --output=/tmp/x..HEAD, which git accepts
as --output=<file> and uses to redirect diff output — an arbitrary write.
Fixed with --end-of-options before the revision range plus a conservative
ref validator. The validator deliberately permits ~ ^ @ { } because those
are legitimate git REVISION syntax (HEAD~1, main@{yesterday}) as distinct
from ref-NAME syntax; --end-of-options is the actual barrier. The doc
comment asserting the trailing -- was sufficient was wrong and is corrected.

MEDIUM — resolveConfinedPath confined by string prefix only, so a symlink
committed inside the repo passed the check (its own path is under cwd) and
readFileSync then followed it outside the root. Now lstat-checks for a
regular file and skips anything else with REFACTOR_FILE_UNREADABLE, so one
bad path skips one file and the run continues.

LOW — the new execute:post dispatch contract showed the gsd_run example
before the rule requiring ref.command be validated first. That prose is
executed by an agent, so textual order is execution order. Reordered.

Refs #1953

* fix(#1953): make the analyzer able to see TypeScript at all

Found by running the shipped analyzer over its own source: it reported
functions=1 for a 940-line module with 24 function forms. A return-type
annotation or a generic parameter list made a function invisible —
`function f(a): number {}` and `function f<T>(a: T): T {}` both detected as
zero. Since gsd-core is written in .cts and the capability declares
.ts/.cts/.mts analyzable, the feature silently found nothing in this repo's
own primary language while reporting success. A safety net that reports
"all clear" because it cannot see is worse than no safety net.

All 98 tests passed over this, because every fixture was plain JS — the
exact failure the test matrix's own "assert against the shape production
uses" warning describes. Adds a TypeScript-shapes suite covering return
types (including unions, generics, object literals and type predicates),
generic parameter lists (constrained and defaulted), export/async/generator
combinations, annotated arrows, class-method modifiers, and optional/
default/rest params — plus the two traps: an overload signature has no body
and must not count, and `a < b && c > d` is a comparison, not a generic.
Detection now reports 24/37/21 functions for the three source files, which
matches a hand count exactly.

Also from review:

- The strict-mode ledger dedup identified entries by parsing a prose
  description string. That is banned by CONTRIBUTING's raw-text-matching
  rule and was a real bug: the "exactly one window per untriaged proposal"
  guarantee rested on prose matching, so rewording a description or editing
  WINDOWS.md by hand silently produced duplicates. Now matches structurally
  on kind + phase + file + line.
- A property test asserted on the stripper's output text. Reframed to
  assert the same invariant through analyzeSource's score.
- nextBaseline's `candidates` parameter has been dead since the anchor
  change; removed from the signature and all call sites.
- Extracted the duplicated require-or-degrade and capability-check
  boilerplate.
- ADR-1953's Implementation bullet still named a `refactor.ship-gate` in
  check-command-router.cts — a leftover from the design cut D6 rejects.
  That file is untouched and no such gate exists. Removed.

Refs #1953

* fix(#1953): keep execute-phase.md under its byte ceiling; un-vacuum the large-file test

Five of the seven remote-runner failures were one cause: the execute:post
dispatch contract, written out inline, grew execute-phase.md 1876 bytes
(93,400 -> 95,276) against a frozen PRE_PHASE6 ceiling of 93,600. A drift-ack
does not clear that — tests/phase6-capstone-conformance.test.cjs and
tests/fix-2285-claude-orchestration-wiring.test.cjs assert the file is
literally under the cap.

The contract now lives in gsd-core/references/loop-hook-dispatch.md, which
already claimed to be the point-agnostic dispatch reference and already
documented ref.skill and ref.agent. It gains the ref.command shape, its
in-context validation rule, the advisory-by-construction statement, and a
note that a point whose workflow hand-rolls one kind is not implementing
this contract. execute-phase.md now defers to it in one line: 145 bytes of
growth, 55 B of headroom under the cap. Better placement than the first cut
— the reference was overstating its coverage, and this makes the claim true
rather than duplicating prose next to it.

Acknowledged by appending to tests/emitted-drift-acks/2930-*.json rather
than a new 1953-*.json: two ack sources may never name the same path, and
that fragment is already the accumulating ack for this file.

Sixth and seventh failures: analyzesLargeFileWithinBounds tripped its own
vacuity guard — the fixture generated ~480 KB against a `> 500000` assert,
so the guard fired and the three assertions after it never ran. The test
has been vacuous since it was written. The matrix row specifies ~1 MB, so
N goes 8000 -> 20000 (1.17 MB, 17% margin) and the guard to > 1_000_000.
Verified by reproducing the exact body against the compiled module: 1168888
bytes, 118 ms, all four assertions hold.

Refs #1953

* fix(#1953): fold the execute:post step deferral into the existing resolve line

The remaining two failures were one test: execute-phase.md carries a SECOND,
tighter assertion than the 93,600 ceiling — `<=93400`, which is exactly its
current size. The file cannot grow by a single byte. My previous fix got it
under 93,600 but not under 93,400, so it still failed. ("H." in the report is
just the parent describe of that same test, not a separate defect.)

Rather than add a paragraph, the deferral now REPLACES the existing hook
resolution line. It read:

  Resolve active step hooks from `EXECUTE_POST_HOOKS_JSON` where
  `kind == "step"` and `ref.skill == "code-review"`.

which is the bug itself written down — only code-review was ever dispatched.
It now reads:

  Dispatch each `kind == "step"` hook per
  @gsd-core/references/loop-hook-dispatch.md. For `code-review`:

The following prose already begins "If no active code-review step hook
exists", so it reads correctly and the code-review handling is untouched.
Net effect on the file is -11 bytes: 93,400 -> 93,389, under the margin
assertion rather than merely under the ceiling.

That also removes the need for a drift-ack: the file shrank, so there is no
growth to acknowledge, and the append to the shared 2930-*.json fragment is
reverted. Leaving it would have shipped a claim of "145 bytes of growth"
that is no longer true, on a file six other issues share.

The test's own comment states the principle this ended up honoring: "the host
loop must stay small — optional-feature detail belongs in the capability
fragment, not the host workflow." Putting the dispatch contract in the
reference rather than inline is that rule, applied.

Refs #1953

* fix(#1953): keep the code-review hook literal the workflow test requires

tests/code-review.test.cjs extracts the <step name="code_review_gate"> block
and asserts it contains `ref.skill == "code-review"` verbatim. The previous
commit replaced the line carrying that literal, so the token vanished and the
test went red — a fair assertion: code-review IS the bespoke branch there and
the workflow should still name it.

Restored inside the same one-line deferral, which now reads:

  Dispatch `kind == "step"` hooks per @gsd-core/references/loop-hook-dispatch.md.
  `ref.skill == "code-review"`:

93,396 bytes — still under the `<=93400` margin assertion and 4 bytes below
the base, so the file continues to shrink rather than grow.

Because three consecutive runs were each reddened by a different assertion on
this one file, this change was verified by sweeping ALL of them at once rather
than one run at a time: every test under tests/ that reads execute-phase.md or
references/loop-hook-dispatch.md was located by resolving its path constants,
and each content/size assertion was evaluated directly against the working
tree — 22 assertions, plus two real executions (gen-section-manifest --check,
and emitted-attribution's full real-tree differential). All pass.

That sweep also confirms the earlier judgement call: the net change to
execute-phase.md is a SHRINK, and the size ratchet only gates growth, so
reverting the append to the shared 2930-*.json ack fragment was correct — an
ack would have been both unnecessary and factually wrong.

Refs #1953

* chore(#1953): backfill changeset pr number to 3261

* docs(#1953): add the missing how-to for acting on a refactor proposal

Reference and explanation shipped (COMMANDS.md, CONFIGURATION.md,
FEATURES.md 159, ADR-1953) but the Diataxis how-to quadrant did not, and
that is the one a user reaches for. CONTRIBUTING's required-docs table is
'new command -> COMMANDS.md + FEATURES.md', so CI was green on a gap.

Enabling this feature is genuinely multi-step and no single page walked it:
turn it on, tune the threshold, understand advisory vs strict, discover
that strict needs a SECOND toggle on a DIFFERENT capability, and know what
to do when a proposal appears. The two-toggle subtlety in particular was a
footnote in a config table; here it is a section with both commands.

Follows the shape of its closest siblings, resolve-edge-coverage-findings
and resolve-prohibition-findings — both 'the loop surfaced a finding, here
is what to do with it'. Includes a reason-code table for the silent cases,
since the analyzer is deliberately quiet in six situations and a user who
expected a proposal needs to tell 'nothing to report' from 'could not look'.

Indexed from docs/README.md beside the other loop how-tos.

Docs-only: exempt from the push gate, no re-verification, pass marker on
2af188b4 untouched.

Refs #1953

* feat(#1953): warn when strict mode is on but nothing will actually block

Closes acceptance criterion 5, which I had wrongly marked satisfied.

refactor.trigger_strict records an untriaged proposal as an open deviation
window, but a ship only STOPS if workflow.windows_enforce is also on — a
toggle owned by the broken-windows capability that this feature neither sets
nor requires. So a user could enable strict, believe ship was gated, and find
out otherwise at ship time.

The split itself stays: requires:["broken-windows"] would force-install the
ledger on advisory users who never enable strict, and a ship:pre gate of our
own would never fire because ship.md has no generic ship:pre gate dispatch.
What was missing was discoverability, so that is what this fixes.

`refactor evaluate` now emits a typed REFACTOR_STRICT_NOT_ENFORCING warning,
naming the exact remediation command, whenever strict is on and either
workflow.windows_enforce is off or broken-windows is unavailable. It fires
only on a run that produced a candidate — with nothing to block on there is
nothing to warn about, and warning every run would be noise.

Reads workflow.windows_enforce through the same resolveConfigKey walk the
router already uses for its own keys rather than a second config reader.
Four tests cover the matrix: strict+enforce-off warns, strict+enforce-on does
not, strict+ledger-absent warns, strict-off never warns.

Also corrects a user-facing message in this same file that told the user to
run `gsd-tools config-set` — the wrong form. docs/CONFIGURATION.md and the
broken-windows capability both use `gsd config-set`, and gsd-tools is invoked
as `node gsd-tools.cjs`, so the bare form may not resolve. The two adjacent
messages in this file now agree.

Refs #1953

---------

Co-authored-by: sim <sim@local>
2026-08-09 19:52:47 -04:00
Tom Boucher
653f95e39f chore(#2801): remove the hostBehaviors.reviewerCli deprecated alias (#3272)
* test(#2801): failing-first suite for the hostBehaviors.reviewerCli alias removal

Inverts the Phase 5a rows that assert the derived legacy alias still
contributes a reviewer slug, and adds the removal-warning coverage the
alias's exit needs (ADR-2782 D9).

RED against unmodified production code, by design: the six shipped
manifests still declare the key and collectReviewerWarnings emits nothing
for hostBehaviors.

Refs #2801

* chore(#2801): remove the hostBehaviors.reviewerCli deprecated alias

ADR-2782 D9, Phase 7 — the final phase of epic #2782.

The derived legacy alias survived one release (Phase 5a shipped in 1.9.0;
1.9.1 and 1.10.0 have since gone out), so it goes. A declared reviewer
body is now the only route onto the reviewer roster.

- deriveReviewerSlugs no longer reads runtime.hostBehaviors.reviewerCli
- the key is stripped from the six manifests that carried it; each already
  declares a reviewer body whose slug equals its capability id, so the
  derived roster is unchanged at the same twelve slugs
- collectReviewerWarnings emits a presence-based, non-fatal removal notice
  for any manifest still declaring the key, reaching both the build-time
  registry generation and the third-party overlay load path. The check runs
  before the reviewer-body early-return, because the manifest it exists for
  is the alias-only one that has no body.
- hostBehaviors stays an open, unvalidated bag for its other 59 keys; this
  adds one keyed removal notice, not general validation

Refs #2801

* refactor(#2801): give the reviewer-warning channel a typed IR

Review finding: the new tests asserted with String#includes() on the
warning prose, which CONTRIBUTING.md's 'Prohibited: Raw Text Matching on
Test Outputs' bans in favor of a typed intermediate representation.

Adds the IR beside the renderer rather than replacing it, which is the
shape that section prescribes and bin/verify-reapply-patches.cjs already
models:

- REVIEWER_WARNING, a frozen code enum
- REMOVED_REVIEWER_CLI_FIELD, so the emitting site and its test share one
  symbol instead of duplicating a literal
- collectReviewerWarningRecords(cap), returning typed records

collectReviewerWarnings(cap) keeps its exact string[] contract as a thin
map over the records, so both production consumers are untouched. Every
section-K row now asserts on record.code/field/capId and none on the
rendered message. Locks the code surface, asserts the renderer stays
one-to-one with the records, and migrates the pre-existing Phase 2 test
on the same channel off prose matching.

Refs #2801

* test(#2801): invert the section F alias fall-through regression row

Caught by the remote runner: 2 unique failures on both Node lanes out of
31,692. tests/reviewer-lane-declarations.test.cjs section F — Phase 5a's
isolated-security-review regressions — asserted that a blank reviewer.slug
falls through to the hostBehaviors.reviewerCli alias rather than dropping
the lane. That is the direct inverse of this phase's contract.

The original rationale held only while the alias existed. With it gone
there is nothing to fall through to: a blank body is not a declaration,
and a declaration is the only route onto the roster.

Inverted rather than deleted — the row carries the adversarial-review
provenance for the slug trim, and removing a security regression guard to
make a change pass is backwards. The duplicate row added earlier in
section C is dropped instead; section F is its canonical home.

Also corrects two count strings Phase 5b left at eleven while asserting
twelve, which would misreport on failure.

Refs #2801

* docs(#2801): give the removed reviewerCli flag a migration path

The Reference edit alone satisfied CI — a file under docs/ moved, so
lint-docs-required.cjs was green — while the task-oriented quadrant said
nothing about the removal. A maintainer whose lane had just gone silent
would have found the field documented as removed and no page telling them
what to do about it.

Adds a migration section to the how-to: the symptom, the verbatim warning
they will see, the before/after manifest, and the note to keep the
reviewer slug equal to the capability id so existing
review.default_reviewers entries and --<slug> flags survive.

Refs #2801

* chore(#2801): backfill changeset pr number to 3272

* feat(#2801): close the runtime.hostBehaviors vocabulary

ADR-1016 closes twelve descriptor axes and rejects an open escape hatch
in the descriptor. It never mentioned runtime.hostBehaviors, and that
silence was read as permission: 59 keys across 18 manifests, 39 of them
set by a single capability, validated by nothing. The reference docs went
further and attributed the open seam to ADR-1016, which does not mention
the field at all.

KNOWN_HOST_BEHAVIORS enumerates the vocabulary. An undeclared key yields a
non-fatal UNKNOWN_HOST_BEHAVIOR record on the same D4.3 channel as the
alias removal notice, reaching both build-time generation and overlay
install.

Warning, never error, for the reason this phase exists: an error would
hard-break an out-of-tree descriptor carrying a bespoke key with no
deprecation window, which is what reviewerCli was given a release to
avoid. Escalation is a separate decision.

reviewerCli is excluded from the unknown-key sweep so it keeps its own
notice with the migration pointer rather than drawing two records.

A parity test binds the vocabulary to the shipped manifests in both
directions, and a second asserts no shipped capability draws a notice, so
the closure is provably inert in-tree.

Records the decision and the miscitation as an ADR-1016 amendment.

Refs #2801

* fix(#2801): bound and sanitize the unknown-key diagnostics

Two findings from an isolated adversarial review of the closure commit,
both proven by execution rather than asserted.

MAJOR, introduced by the closure: the new Object.keys(hostBehaviors) sweep
had no ceiling. An installed third-party manifest is bounded only by
MANIFEST_MAX_BYTES, and an 8.69MB manifest with 800,000 keys produced
800,000 records and ~139MB of message text, retained for the registry's
lifetime in OverlayMeta.diagnostics. Now capped at ten records plus a
summary carrying omittedCount, mirroring capability-loader's existing
slice(0,3) idiom. The same manifest now yields 11 records and 1748 chars.

MINOR, newly reachable: manifest-supplied key names were interpolated raw.
Unlike cap.id, which validateCapability gates on KEBAB_RE before these
diagnostics run, hostBehaviors keys have no grammar check anywhere, so
ANSI escapes and CRLF reached stderr and OverlayMeta.warnings intact. New
describeKey replaces C0/C1 controls and clips at 80 chars. The file
already had describeValue for this and applied it only to values.

Both fixes land on the pre-existing reviewer.* sweep too — it carried the
identical pair, and fixing only the new copy would leave the same defect
one screen from its own fix.

Refs #2801

---------

Co-authored-by: sim <sim@local>
2026-08-09 19:18:56 -04:00
github-actions[bot]
4f2a4932ce chore: sync next package version to 1.10.0 2026-08-08 05:07:32 +00:00
Tom Boucher
27aa40f65e fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ directory (#3175)
* test(#3023): failing-first guard — pi must not stage hooks in its reserved dir

pi reserves <configDir>/hooks as its deprecated extension location and warns
on every startup when it exists. Assert a pi install stages the shared hook
bundle under gsd-hooks/ instead, manifests it there, and never creates hooks/.

Also adds pi to the local-scope dir table in install-shared.cjs: pi was in
RUNTIME_META but not LOCAL_DIR_NAME, so scope:'local' resolved
path.join(root, undefined) and no local pi install could be exercised.

Fails before the fix. Verified via the remote runner.

* fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ dir

pi reserves <configDir>/hooks as its now-deprecated extension location and
warns on every startup when that directory merely exists — checkDeprecatedExtensionDirs()
guards the warning with a bare existsSync(), unlike its tools/ sibling. GSD staged
its shared hook bundle exactly there, and pi's advised remediation (move it to
extensions/) would break the adapter's paths and expose GSD's .js helpers to pi's
extension auto-discovery.

The bundle directory name is now runtime-descriptor-driven: hostBehaviors
.sharedHooksDirName, defaulting to 'hooks' so all 18 other runtimes are
byte-identical. pi sets 'gsd-hooks'. The name is validated as a single path
segment — separators, dot-only segments, trailing dots, absolute paths, NUL,
and Windows reserved device names all fall back to the default, because the
value is joined onto a user's config root and written to.

Renamed in place rather than relocated: hook scripts resolve siblings via
__dirname/.., so a depth change would silently break them.

- install / uninstall / manifest sites all read the resolved name
- pi/gsd.cjs probes gsd-hooks then hooks, so dev checkouts and half-upgraded
  trees still resolve; the never-throws contract is preserved
- new migration 009 retires the legacy pi hooks/ dir on upgrade, using a new
  non-recursive remove-empty-dir engine primitive (rmdirSync only,
  symlink-refusing, containment-guarded); ADR-0008 amended accordingly
- fixes two latent name-dependencies the rename exposed: the stale-hook scan
  and the injection scanner's self-exclusion both hardcoded 'hooks'

Verified on the remote runner.

Closes #3023

* fix(#3023): close review findings and align emitted provenance with the rename

Adversarial review found two defects, and the remote runner found four
failure clusters. All fixed here.

Review BLOCKER — detect-custom-files was blind to the renamed bundle.
GSD_PREFIX_MANAGED_DIRS in gsd-tools.cjs hardcoded 'hooks', so for pi the
whole gsd-hooks/ tree was invisible to the custom-file scan and user-added
files there were never backed up before the next update's clean-install wipe.
The dir set now resolves via the .gsd-runtime marker plus the shipped
capability registry (never bin/install.js, which is not shipped into installed
trees), and falls back to scanning every known candidate when the runtime
cannot be determined — over-scanning is safe, under-scanning is the data loss.

Review MAJOR — the pi adapter bound to an empty bundle. resolveSharedHooksDir
accepted any directory, so an interrupted install left gsd-hooks/ winning over
a fully-staged legacy hooks/ and every hook silently no-opped. A candidate now
qualifies only if it is non-empty.

Remote-runner clusters:
- emitted-provenance had no rule for the gsd-hooks/ family; added two pi-scoped
  rules pointing at the same sources the existing hooks/ rules use. The table is
  total, so an unattributed family is a hard failure by design.
- pi tests in install-minimal-hooks and the install integration suite asserted
  the old layout; updated to derive the dir name from the descriptor rather than
  hardcoding either name.
- 19 unrelated-looking failures on node22 only were a leaked fs mock: t.after()
  runs in registration order, cleanup was registered before mock.restoreAll(),
  and node22's JS rimraf calls the public fs.rmdirSync while node24's native
  path does not — so the EACCES stub leaked process-wide on one lane. Restore
  now runs first.

Verified on the remote runner.

* fix(#3023): honor PI_CODING_AGENT_DIR, ack the rename ripple, fix expandTilde

pi resolves its agent dir as PI_CODING_AGENT_DIR ?? ~/<CONFIG_DIR_NAME>/agent
(packages/coding-agent/src/config.ts). GSD's pi descriptor declared an empty
configHome.env, so a user with that variable set had GSD installed where pi
never looks. Added the env name; the dot-home-nested resolver already handled
the override, so no resolver logic changed.

Also fixes expandTilde in the shared runtime-homes resolver, found while adding
that: it hardcoded os.homedir() and ignored the opts.home every caller threads,
so EVERY runtime's tilde-valued env override (claude, antigravity, windsurf, pi)
silently resolved against the real home. That is a correctness bug and a
test-escape hazard — a sandboxed test asserting on a tilde override reached the
developer's actual home directory. Now threaded through every branch; behavior
with no injected home is unchanged.

Adds the emitted-drift ack fragment for the 58 pi paths whose emitted location
moved with the rename. The provenance rules satisfy the totality gate; the
differential gate needs the ack because the hook sources are byte-unchanged —
only the installer's target directory moved. The two hook files this branch
genuinely edits stay attributed and are not double-acked.

Note on piConfig.configDir: it is read from pi's OWN installed package.json
(getPackageDir walks up from pi's __dirname), alongside piConfig.name — a
white-label setting for a redistributed pi fork, not a per-project user setting.
Documented accordingly rather than treated as an unsupported override.

Verified on the remote runner.

* fix(#3023): reject blank env overrides, pin adapter/descriptor parity

Three review findings, all fixed.

A whitespace-only config-dir override was accepted verbatim: the guard was
`if (val)`, falsy only for the empty string, so PI_CODING_AGENT_DIR='   '
resolved to a literal three-space directory name instead of falling back to the
descriptor default. Fixed across every env-consuming branch — dot-home,
dot-home-nested, all three xdg steps, and generic-agents-root — not just pi's.
Non-blank values are still never trimmed, so '~/My Agent Dir' keeps working.

pi/gsd.cjs's probe list and the descriptor were two independent sources of truth
for the bundle directory name; a future rename would have desynced them silently
and left every pi hook quiet with no error. The probe list stays deliberate — it
must resolve in a dev checkout and a half-upgraded tree, where the registry's
answer would be wrong — so this adds the parity assertion the repo's
generative-fix-divergence rule calls for: the descriptor value must be the FIRST
candidate, and the default must remain present.

Changeset body rewritten to cover the two later user-facing fixes it had not
caught up with.

Verified on the remote runner.

* chore(#3023): backfill changeset PR number

* fix(#3023): anchor injection-scan patterns and fix a macOS detection hole

CI's security job flagged CONTEXT.md:124 — pre-existing prose reading 'not the
same fact as a genuinely empty or absent one'. The match was the 'act as a'
INSIDE 'f-act as a': the pattern had no left word boundary, so any word ending
in act tripped it (fact, impact, contract, artifact, interact, redact,
abstract). My four-line CONTEXT.md edit dragged the latent false positive into
this PR because the scan is diff-scoped by file but reads whole files. Anchored
with (^|[^[:alnum:]]) rather than rewording maintainer-owned prose, which would
have left the class alive for the next PR touching any file saying 'fact as a'.

Auditing the rest of the list for the same class surfaced a real detection hole:
the eval/exec/Function patterns matched a quote via \x27, a GNU-grep-only hex
escape. BSD/macOS grep reads it as four literal characters, so single-quoted
eval('...')/exec('...') payloads were NEVER detected there while passing on
GNU-grep CI. Replaced with a literal apostrophe class.

Boundaries were added only where a real word-suffix collision exists; exec,
jailbreak, developer mode and the role-manipulation family were audited and
deliberately left unanchored. 22 new cases cover both directions — the false
positives now scan clean, and every real payload still fires, including the
quote/punctuation/start-of-line boundary forms.

Also builds this branch's injection test fixture at runtime instead of carrying
the literal phrase, so the payload keeps its teeth without tripping the scan.

Verified on the remote runner.

---------

Co-authored-by: sim <sim@local>
2026-08-07 13:41:21 -04:00
Dennis Kim
178ec00040 fix(#2787): clarify broken-windows ship blocking enforcement (#2814)
* fix(#2787): clarify broken-windows ship blocking enforcement

* fix(#2787): update renderLedger header to clarify opt-in enforcement

* fix(#2787): address maintainer scope and wording review

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-03 12:39:16 -04:00
𝚌𝚕𝚎𝚣𝚌𝚘𝚍𝚒𝚗𝚐
88f6d9bd1b fix(#2644): deduplicate Cursor slash menu (#2812)
* fix(#2644): deduplicate Cursor slash menu

* fix: preserve installer executable mode

* chore: add changeset for PR #2812

* test(#2644): acknowledge Cursor emission changes

* test(#2644): drop spent emitted drift acknowledgments

* fix(#2644): remove retired Cursor command converter

---------

Co-authored-by: clezcoding <clezcoding@users.noreply.github.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-03 12:05:45 -04:00
Tom Boucher
51f32d2d40 fix(#2658): detect trae runtime and resolve its instruction file to a concrete rules file (#3006)
* test(#2658): add failing-first regression for trae runtime detection and instruction path

Covers all three collided defects reported in #2658 plus a fourth
instance of defect 1 (ingest-docs.md) found while diagnosing it:
missing trae detection in workflow runtime-detection blocks, the
CLAUDE.md path-mutilation bug in both the js/cjs and md install-time
converters, and the missing projectInstructionFile capability
declaration. Fails against current source; the next commit fixes it.

* fix(#2658): detect trae runtime and resolve its instruction file to a concrete rules file

Three defects collided to produce the reported ".claude/.trae/rules/"
path:

1. new-project.md and ingest-docs.md's runtime-detection blocks only
   recognized codex/gemini/opencode and fell through to RUNTIME=claude
   for trae. Both now recognize the /.trae/ execution-context path and
   the TRAE_CONFIG_DIR env var before the claude fallback.
2. RUNTIME_CONTENT_DISPATCH.trae.js (bin/install.js) replaced bare
   "CLAUDE.md" before the ".claude/" prefix was handled, mutilating
   ".claude/CLAUDE.md" into ".claude/.trae/rules/". Now replaces the
   full ".claude/CLAUDE.md" path first, and targets a concrete file.
3. convertClaudeToTraeMarkdown (mirrored in bin/install.js and
   src/runtime-artifact-conversion.cts per the #2094 output-parity
   test) had the same class of bug with a different wrong output
   (".trae/.trae/rules/", from its generic ".claude/" rewrite firing
   after the bare CLAUDE.md rewrite). Both mirrors now match full-path
   forms before the bare/generic patterns, converging on the same
   concrete file as the js/cjs converter.
4. capabilities/trae/capability.json didn't declare
   hostBehaviors.projectInstructionFile, so getProjectInstructionFile
   fell through to the generic AGENTS.md default even when RUNTIME=trae
   was resolved correctly. Now declares ".trae/rules/rules.md",
   regenerated into gsd-core/bin/lib/capability-registry.cjs via
   npm run gen:capability-registry.

Closes #2658

* chore(#2658): add changeset

* fix(#2658): preserve arbitrary runtime-dir prefixes in the trae path rewrite

Found by the end-to-end --trae install regression test (not by static
trace) across two verification runs:

1. copyWithPathReplacement runs a generic ~/.claude/, $HOME/.claude/,
   and ./.claude/ -> runtime-dir rewrite on every .md file BEFORE
   calling convertClaudeToTraeMarkdown. The prior fix's
   .claude/CLAUDE.md-specific patterns never fire on that
   already-rewritten text, and the bare fallback still doubled
   whatever prefix the generic pass substituted. A first attempt
   handled only the fixed "./.trae/" shape and missed the
   $HOME/.claude/ and ~/.claude/ forms gsd-core/workflows/profile-user.md
   actually uses, which post-rewrite become an arbitrary absolute
   local-install-root path, not the fixed relative shape. Fixed with a
   prefix-preserving pattern that captures whatever precedes a
   ".trae/" tail and fixes only the filename suffix, instead of
   assuming one fixed shape.

2. The fix's own explanatory comments literally spelled out the
   malformed strings and the instruction filename as contiguous text.
   Since these two files ship verbatim into local --trae installs,
   where they are themselves run through the same find/replace, the
   comments got "fixed" right along with the real code, leaking the
   malformed string into the installed tree. Rewrote every comment in
   both mirror copies to never spell either the instruction filename
   or a malformed shape as one contiguous token.

Adds an emitted-drift-ack fragment: the corrected replacement target
for every CLAUDE.md mention (bare directory -> concrete file) changes
trae-emitted output for every repo file that mentions CLAUDE.md, not
only the ones that hit the originally reported bug.

* test(#2658): extend parity test with arbitrary-prefix .trae/ inputs

The bin/install.js vs runtime-artifact-conversion.cjs parity assertion
for convertClaudeToTraeMarkdown only fed the pre-existing bare
.claude/CLAUDE.md input through both implementations. Feed the
prefix-preserving cases (relative, nested-absolute, tilde, $HOME,
backtick-wrapped) plus a property-based check through both, so a
future edit to only one copy of the .trae/-tail regex fails this
test instead of silently diverging.

* chore(#2658): backfill changeset PR number to 3006

---------

Co-authored-by: sim <sim@local>
2026-08-02 18:29:47 -04:00
kyle-the-dev
ce38d44811 fix(#2777): remove stale codex local home metadata (#2831)
* fix(#2777): remove stale codex local home metadata

* chore(#2777): add changeset for codex local layout metadata

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-02 01:00:52 -04:00
github-actions[bot]
854c93533c chore: sync next package version to 1.9.1 2026-07-31 13:12:09 +00:00
github-actions[bot]
4232a79396 chore: sync next package version to 1.9.0 2026-07-31 03:14:04 +00:00
Tom Boucher
3f6b063fbb chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans

Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers
iterate declared lanes instead of hand-authored per-CLI bash.

Five additive descriptor amendments, each forced by a lane that ships today:
- LaneHandler gains 'opencode' — the lane rebuilds its review from assistant
  text parts of a --format json stream; a plain stdout copy re-breaks #1936.
- modelConfigKey — antigravity's key is review.models.agy, not .antigravity,
  so resolving by slug silently dropped a configured model.
- defaultHost/fallbackModel — Phase 4 federated every *_host with a default of
  empty string; the real fallback only existed in the bash.
- args becomes an argv template with a closed four-placeholder vocabulary.
  Positional splicing produced 'codex --model M -o F exec --ephemeral', which
  is not a valid invocation: codex injects in the middle, twice.
- kimi-code lane, with the bounded command-capability probe (needle
  --output-format) that tells Kimi Code from the legacy python kimi-cli.

Parity gate re-pointed: the workflow-text families it scanned are the text this
phase deletes, so they are replaced by descriptor-to-registry parity plus an
anti-parity check that no bespoke leg returns.

jq, curl and external timeout/gtimeout all drop out of the review path.

Refs #2782

* chore(#2799): add review-lane query surface and widen the manifest vocabulary

Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow
loops over, projects all twelve lanes into their capability manifests, and
widens capability-validator for the amendments.

opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's
own admission rule: one lane, justified by a documented upstream defect data
cannot express (#1936 — the agent can end its turn with zero output tokens and
--format default then drops the assistant text entirely).

Two bugs caught by an end-to-end stub run and fixed here:
- loadConfigResolved returns a provenance wrapper, not the config; using it
  directly resolved every key to undefined, which reads as 'nothing
  configured' and silently dropped every model override.
- hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced
  with a PATH scan that spawns nothing at all.

Refs #2782

* chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews

Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved
lanes, and renders REVIEWS.md sections from each lane's declared
reviewsSection instead of thirteen hardcoded headings. review.md drops from
1104 lines to 507 (61KB to 28.7KB).

Parity gate re-pointed, as agreed: the leg-marker and section-heading families
scanned exactly the text this phase deletes, so they are replaced by
descriptor-to-registry parity in both directions, plus an anti-parity check
that fires if a bespoke leg is ever re-added. Enum, emitting sites and the
Object.keys lock moved together.

The budget-trim helper is hoisted out of the Ollama leg: it was always
lane-agnostic, and any lane may now declare a promptBudgetKey.

Refs #2782

* feat(#2799): bind the consented egress host and re-verify it at invocation

Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3
but was not implemented: ConsentRecord had no host field and nothing in the
tree bound one, so this phase's rule-4 comparison had no baseline.

ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design:
isValidConsentRecord does not require it, so every record already on disk
stays valid and no re-consent storm fires (D4 rule 5). It is deliberately
excluded from disclosureSignature — the loader has no config resolver, so
folding a config-derived value in would make loader and lifecycle compute
different signatures for the same manifest and re-prompt forever.

Install resolves hostConfigKey (falling back to the lane's declared
defaultHost, which is what the invocation path uses) and records it.
Invocation re-resolves and blocks on mismatch rather than silently
redirecting. Absence allows: no record, or a record predating the field,
means nothing to compare — denying there would break every existing
local-model user on upgrade.

Refs #2782

* test(#2799): cover the resolver, runner and handlers; retarget the parity suites

Adds the golden invocation-plan table (one row per shipped lane, derived from
the bash legs rather than the descriptor types) plus runner coverage for the
probe, empty-output policy, the three handlers and the egress check.

Retargets the existing suites onto the new contract: descriptor-to-registry
parity, the anti-parity check, the opencode handler, and the twelfth lane.

Two corrections found by running them:
- modelConfigKey was required; that breaks D4 rule 2, since a reviewer
  manifest authored before this phase would fail validation on upgrade. It is
  optional, read as null when absent.
- the antigravity non-zero-exit test pre-seeded the transcript, which asserted
  that a STALE entry leaks through — the exact bug the watermark prevents. The
  spawn now appends, as the real tool does.

Refs #2782

* fix(#2799): restore agy --add-dir and the self-report prompt in the handler

Retargeting the three legacy reviewer suites off the deleted bash surfaced two
real regressions in the port, both #2176:

- --add-dir was dropped. Without it agy's permission context never receives the
  cwd repo, so the agent anchors on its own scratch dir and reviews the plan
  text in isolation — the exact failure the Review Instructions forbid. It is
  capability-probed, because an older agy rejects the unknown flag outright and
  a lane that fails to start is worse than one running on the prompt anchor.
- the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS
  self-report, which is what makes a blind review distinguishable from a
  grounded one. antigravity now builds its own prompt variant.

Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned
model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own
log is the only evidence that anything failed.

The three suites now assert against the plan and the handler instead of
matching fence text, so they no longer need allow-test-rule exemptions.

Refs #2782

* docs(#2799): document the declared lanes, the new flag, and dropped prerequisites

COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph,
which is now false: no lane requires jq, curl or an external timeout. Adds the
changed-egress-destination behavior, since a blocked lane is something a user
can hit.

CONFIGURATION.md records that the model config key is declared per lane rather
than derived from the flag — antigravity's is review.models.agy — and adds
review.models.kimi-code.

reviewer-instances.md now routes an instance through its lane's single
invocation seam instead of a copied per-adapter bash block, which is what lets
a cross-cutting fix reach instances for free. That required implementing the
--model/--agent/--as flags it documents; --model re-resolves through the lane's
argv template rather than splicing, so the flag lands where the lane declares
it rather than ahead of a subcommand.

CONTEXT.md glossary gains both new modules.

Refs #2782

* chore(#2799): drop the stale emitted-drift acknowledgment

The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md.
That file now shrinks by ~32KB and every emitted hash that moved is
attributable to this diff, so the ack no longer explains anything. Removing
the last entry means removing the file: its presence is the alarm, and an
empty one signals nothing.

Verified by deleting it and re-running the attribution and provenance gates
plus lint:ci — all green without it.

Refs #2782

* docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782

Five additive amendments, each forced by a lane that ships today, plus two
corrections the phase had to make rather than work around:

- D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so
  this phase's rule-4 comparison had no baseline. Recorded because an ADR
  asserting a rule was delivered is exactly what stops a later phase checking.
- The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families
  scanned the text this phase deletes.

Also records that D7's 'skip the probe where no bounding mechanism exists'
carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded
on every stock macOS host, which ships neither timeout nor gtimeout.

Refs #2782

* fix(#2799): close four defects found by adversarial review

Two confirmed bugs, both reproduced before fixing:

- resolveLanePlan was not total. An openai-http lane with a missing or
  non-object invoke dereferenced inv.hostConfigKey and threw, contradicting
  the module's own documented contract; the spawn branch guarded correctly and
  the http branch did not. The CLI seam resolves every selected lane in one
  map, so one malformed overlay manifest would have aborted the whole review
  rather than dropping its own lane. Guarded, plus a per-lane try/catch at the
  seam so a throw can never take down siblings.
- A reviewer-instance model was silently dropped for any lane declaring
  modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates
  that cli is a known slug but never that the slug accepts a model, so a user
  could configure one, get a clean run, and never learn a different model
  reviewed their plan. Now warns explicitly.

Two hardening fixes:

- The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in
  the resolver rather than inherited from a validator that does not run on this
  path — the module documents itself as the overlay-manifest trust boundary, so
  it should not depend on someone else having checked.
- normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses
  with an empty hostname, so it became 'localhost://11434' and was compared and
  requested as if real. An empty hostname now means not-a-URL.

Also documents the one gap that cannot be closed here: the antigravity
watermark is keyed by workspace, so two concurrent reviews of the same repo
share a transcript. agy exposes no per-invocation id to filter on, so the
handler now states which half of its never-stale guarantee actually holds.

Refs #2782

* test(#2799): retarget the remaining eight review.md-asserting suites

The remote runner found 37 failures the local sweep missed (it hit the shell's
two-minute cap before reaching these). All eight extract per-CLI bash from
review.md that this phase deletes; each protects a real invariant, so each is
retargeted onto the plan, the runner or the handler rather than removed.

Three real defects surfaced by doing so:

- effort args never reached ANY lane. model-resolver.cjs exports no
  resolveExecution, so effortFor silently returned [] every time. Restored by
  calling the same bounded resolve-execution query the bash legs used — and
  NOT with --raw, which prints the resolved effort rather than the picked
  field, so claude got 'low' instead of '--effort low'.
- the timeout guidance lost 'a silent empty output is a timeout kill, not a
  crash' — the operator note that exists because of the Codex 0xc0000142
  misdiagnosis. Restored.
- the opencode handler dropped EMPTY assistant text parts. The shipped jq was
  , and  only substitutes for false/null — an empty
  string is truthy in jq and contributed a blank line. Found by a property
  test shrinking to ['', ''].

The opencode property suite no longer spawns jq at all, which deletes the
#2099 hang mechanism it was architected around rather than mitigating it.

Refs #2782

* fix(#2799): register the two new generated modules, and untrack them

The remote runner caught build output committed to git. Both new modules
compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated
that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own
review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each
bin/lib/*.cjs is linted xor ignored according to migration state" failed.

Registered both in .gitignore and eslint.config.mjs alongside the Phase 1
module, and dropped them from the index. Nothing about the shipped behaviour
changes; the artifacts are rebuilt by build:lib.

This is the new-.cts-module registration ripple, and it is the one part of it I
had not completed - the CONTEXT.md glossary and the inventory manifest were
already done.

Refs #2782

* chore(#2799): backfill changeset pr number to 2861

* chore(#2799): backfill changeset pr number to 2861

---------

Co-authored-by: Test <test@example.com>
2026-07-30 12:48:06 -04:00
Tom Boucher
0408276791 chore(#2797): federate reviewer config keys off the central schema (#2841)
* chore(#2797): federate reviewer config keys off the central schema

Phase 4 of epic #2782 (ADR-2782 D9, config half). Runs AFTER 5a per the
ADR's swap amendment: a federated config slice lives inside a
capabilities/<id>/capability.json, and three of the five key families had
no capability directory until 5a created them.

Four key families move to the lanes that use them; the central-schema
removal and the federated addition land in this one commit because the
exclusivity invariant fails the build on a key present in both.
review.max_prompt_tokens, review.default_reviewers and
review.reviewer_instances describe policy ACROSS lanes and stay central.

Two things the issue did not name, both found while building it:

1. THE EXCLUSIVITY GATE WAS BLIND TO PATTERNS. It compared federated keys
   against manifest.validKeys only, and two of the four families
   (review.models.<slug>, review.max_prompt_tokens_per_reviewer.<slug>)
   were pattern-backed. That is not cosmetic: isCentralConfigKey consults
   those patterns and mergeFederatedConfig skips every key for which it
   returns true, so declaring a slice while the pattern survived would
   have shipped an INERT slice behind a green gate — the exact
   half-migrated shape the invariant exists to prevent. The gate now
   loads the patterns from the same manifest the runtime reads.

2. AN UNSET PER-LANE BUDGET NOW RESOLVES TO 0, NOT NOT-FOUND, because a
   federated key always resolves to its declared default. The three
   fallback guards in review.md checked only empty-or-"null", so a user
   who set the GLOBAL review.max_prompt_tokens would have silently lost
   trimming on the HTTP lanes. The guards now treat 0 as unset.

D9 says review.models.<slug> is owned by "the lane whose slug it names".
That is false for one lane: the shipped key is review.models.agy while
the slug is antigravity. Ownership follows the lane; the key name is
preserved, because renaming would break every config that sets it.

Existing tests updated rather than left asserting the old world:
config-get on a cleared federated key yields empty instead of
not-found (what #2046 actually protects — never persisting the literal
"null" — is unchanged and still asserted); the config-schema dynamic
pattern representative moves to reviewer_instances; the
prototype-pollution guard case moves to a surviving dynamic prefix so
alert #26 keeps its coverage, with a new case asserting the old key is
now rejected earlier; and Phase 2's harvest-widening inertness assertion
becomes an ownership assertion, since Phase 4 is what consumes it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2797): use a -1 sentinel so an explicit per-lane budget of 0 survives

A federated config key always resolves to its declared default, so an
unset per-lane prompt budget needed a value the workflow could treat as
'not configured'. The first cut used 0 — which is wrong: 0 is already a
LEGITIMATE per-lane budget meaning 'do not trim this lane' (the
early-return guard in prepare_trimmed_prompt_for_reviewer). Treating it
as unset would have silently switched a user who deliberately disabled
trimming for one lane onto the global budget.

The sentinel is now -1, which is not a valid token budget, so all three
states stay distinguishable: unset falls back to global, an explicit 0
disables trimming for that lane, and an explicit N is used. Locked by
three CLI round-trip tests.

Surfaced by the isolated security reviewer before it crashed mid-run;
verified independently against the shipped trim guard rather than taken
on trust.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2797): update central-registration assertions and stay under the review.md cap

The remote runner caught both; my local sweep missed the files.

1. tests/plan-review-convergence.test.cjs asserted the three local-server
   host keys are in VALID_CONFIG_KEYS. They are federated to their lane
   capabilities now, and the exclusivity invariant forbids a key living
   in both places. What #2306-local actually protects is that config-set
   ACCEPTS them, so that is what is asserted — via isValidConfigKey, the
   predicate config-set itself uses, which spans central and federated.
   A second assertion pins federated ownership, so a silent reversion
   back to the central schema fails too.

2. review.md exceeded the LARGE tier hard cap (62583 > 61440). That cap
   is a red line, not a budget to raise. The three per-lane budget guard
   comments were near-identical; condensed to one terse line each. 61371
   bytes, 69 to spare. Real extraction to workflows/review/modes/ is
   Phase 5b/6 work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2797): fail closed on a broken config-schema manifest; reconcile stale docs

Isolated security review findings.

MAJOR — loadCentralConfigPatterns failed OPEN. It swallowed a JSON parse
error and returned [], while its sibling loadCentralConfigKeys, reading
the SAME file, writes to stderr and throws ExitError(1) on that identical
failure class. Fail-open here defeats the gate this function exists to
feed: with zero patterns, validateCrossCapability's pattern-collision
check silently passes and an inert federated slice ships green. It was
masked in the one production call site only because loadCentralConfigKeys
runs first against the same path — a coincidence of ordering, not a
guarantee, and this function is exported and called standalone. The two
now share a contract: ENOENT is the legitimate absent case, anything else
throws loudly. A single unparseable PATTERN is still skipped, which
degrades to "checked less" rather than blocking every build. The branch
had zero coverage; it now has two tests (malformed JSON, EISDIR).

MINOR — docs/CONFIGURATION.md still listed review.models.qwen and
review.models.cursor as settable, ~770 lines below this PR's own new
Ownership section. Those lanes take no model flag, so they declare no
model key and config-set now rejects them. Rows removed; the missing
review.models.agy row added; the per-reviewer budget row corrected to
name only the lanes that own a budget key, and to document that a
per-lane 0 disables trimming for that lane.

Also fixes a shadowed "raw" binding introduced by the fail-closed change,
which made the generator unrequirable — caught immediately by its own
--check.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2797): backfill changeset pr number to 2841

* fix(#2452): make the base-ref mutation test hermetic against leaked GIT_* env

tests/mutation-workflow-base-ref.test.cjs fails on PR branches while next
stays green, and it is currently blocking at least three unrelated PRs
(#2841, #2832, #2827) with:

  error: invalid object 100644 <sha> for 'base-N.txt'
  error: Error building trees

The existing loop comment attributes this to `git add .` rehashing O(n^2)
blobs "before the object write had landed" and works around it by staging
one path per iteration. That is not the cause: sequential execFileSync
calls cannot race each other's object writes, and the failure persisted
after that change — it simply moved to a lower commit index.

The cause is that the git() helper inherited the runner's environment. A
leaked GIT_INDEX_FILE makes `git add` write into a DIFFERENT repository's
index; GIT_OBJECT_DIRECTORY / GIT_ALTERNATE_OBJECT_DIRECTORIES send the
blob to another object store; GIT_DIR / GIT_WORK_TREE redirect the whole
operation. In every case `git commit` then cannot resolve a blob it just
staged, which is precisely the error above.

Verified by negative control: with GIT_DIR exported, this test fails on
the unfixed helper (the git commands operate on the wrong repository
entirely); with the helper stripping GIT_* it passes. The single-path
staging is kept — it is genuinely less work — but it is no longer load
bearing.

Found while shipping #2797. Fixed in place rather than deferred: it is a
defect surfaced during the work, and it is blocking other contributors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2452): build the base-advance commits empty, removing the lost-object class

The base-ref guard has been failing in CI with:

  error: invalid object 100644 <sha> for 'base-N.txt'
  error: Error building trees

It is currently red on at least three unrelated PRs (#2841, #2832,
#2827) while next stays green.

Two theories have now been tried and neither held. #1881 blamed `git
add .` rehashing O(n^2) blobs and switched to staging one path per
iteration; the failure moved from commit 32 to commit 25 and carried on.
The preceding commit here made the git helper hermetic against leaked
GIT_* environment — that IS a real vulnerability (with GIT_DIR exported
the helper operates on the wrong repository entirely, proven by negative
control) but it produces a different error than CI reports, so it is not
demonstrably the cause either.

Neither trigger reproduces off-CI, so this stops guessing at the trigger
and removes the failure CLASS instead. The loop needs base-branch DEPTH
and nothing else: no assertion reads these commits' contents, and
base-side files cannot appear in `origin/base...HEAD` regardless.
`--allow-empty` writes no blob and no tree, so there is no object for the
index to reference and lose. It is also far less work than 60
write+hash+index cycles.

The guard still proves its mechanism: the test asserts that a --depth=1
base fetch FAILS and a full fetch resolves, so a broken topology would
surface immediately rather than passing vacuously.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Test <test@example.com>
2026-07-30 08:52:56 -04:00
Tom Boucher
6a9babda69 chore(#2798): declare the eleven reviewer lanes as manifest data (#2837)
* chore(#2798): declare the eleven reviewer lanes as manifest data

Phase 5a of epic #2782, delivering ADR-2782 D9 (roster half) and D3.

- Five reviewers GSD never installs into become lane-only role:reviewer
  capabilities with no runtime body, no runtimeCompat and no install surface:
  gemini, coderabbit, ollama, lm-studio, llama-cpp. Before this they had no
  descriptor at all and lived as a hardcoded NON_RUNTIME_REVIEWER_SLUGS tail,
  which is now deleted outright.
- The six hosts that are ALSO reviewers gain a reviewer body alongside their
  runtime body. Their runtime bodies are byte-identical to next -- verified per
  capability against the git blob, not asserted -- so no install behaviour moves.
- KNOWN_REVIEWER_SLUGS derives from declared bodies via an exported
  deriveReviewerSlugs(registry). hostBehaviors.reviewerCli survives as a derived
  legacy alias for one release; where a capability carries both, the body wins
  and the slug appears once. Alias removal is Phase 7 (#2801).

THE KEYSTONE: the roster is the SAME ELEVEN SLUGS as before -- antigravity,
claude, coderabbit, codex, cursor, gemini, llama_cpp, lm_studio, ollama,
opencode, qwen. This phase changes HOW the roster is derived, not WHO is in it,
and the test asserts that literal list rather than a count.

kimi-code is deliberately NOT declared here. It is net-new with no
invoke_reviewers leg, so declaring it now would make it selectable but not
invocable -- present in --all, selected, emitting an empty section for the whole
5a-to-5b window -- and would break Phase 1's parity assertion. It lands in 5b
alongside the iteration that can run it. Legacy kimi (the Python CLI) is not a
reviewer at all and gains nothing.

The highest-value test is declaredManifestLanesMatchThePhase1Descriptor: it
deep-compares all eleven declared bodies against REVIEWER_LANES field-by-field,
including probe and invoke sub-fields. All eleven are byte-identical, key order
included. The epic's premise is that the manifest and the core descriptor
describe the same lane with NO translation layer, and Phase 2's review already
caught one divergence that every other test missed.

Two ADR corrections folded in, as Phases 1-3 each did:

1. PHASE ORDER. The ADR runs Phase 4 (federated config) before 5a and #2798
   claims a dependency on 4. That is inverted and makes Phase 4 unsatisfiable:
   D9 assigns review.<host>_host to lane capabilities that do not exist until
   THIS phase creates them, and a federated config slice must live inside
   capabilities/<id>/capability.json. Real graph: Phase 2 -> 5a -> 4.
2. #2798's INVENTORY acceptance item is vacuous. The inventory catalogs
   bin/lib/*.cjs modules, not capability directories -- antigravity, opencode
   and qwen appear zero times in it -- and gen-inventory-manifest --check passes
   with the five new dirs and no edit.

Also corrected a stale line in Phase 2's own ADR amendment: it recorded the slug
pattern as /^[a-z][a-z0-9_-]*$/, but Phase 2's security review widened the
shipped pattern to /^[a-z0-9][a-z0-9_-]*$/ to match Phase 1's exported
LANE_SLUG_RE. The prose had not followed the code.

Closes #2798

* fix(#2798): catalogue reviewer capabilities in the generated matrix

The capability matrix rendered exactly two tables, feature and runtime, via
renderTable(caps, role) filtering on c.role === role. ADR-2782 D3 added a THIRD
role, so every role:"reviewer" capability was silently dropped from the
first-party catalogue.

The drift guard did not catch it, and could not: --check compares generated
output against the committed file, and both omitted the five lanes identically,
so it reported "up to date" while five shipped capabilities were invisible in
the one document that is supposed to list what ships. A guard blind to an entire
role is not guarding.

This phase is what exposed it -- it ships the first role:"reviewer"
capabilities -- so it is fixed here rather than deferred (CLAUDE.md: a defect
found while working is fixed in the current change, which overrides
one-concern-per-PR).

Verified red-before-green: with a lane row deleted from the matrix, --check now
exits 1; restored, it exits 0. Before this fix the lanes were absent entirely, so
there was nothing for the guard to compare.

Phase 6 (#2800) still owns enriching the matrix with lane-specific detail
(slug/flag/transport columns) and the locale parity gate. This is the narrower
fix: the capabilities APPEAR at all.

* fix(#2798): close two hardening gaps and record three limits durably

Isolated security review (5 targets, no blockers) reproduced two gaps in the new
deriveReviewerSlugs. Both are unreachable through the checked-in registry -- it is
generated, JSON-sourced and code-reviewed -- but the function is EXPORTED for
reuse and carries no other validation, so it must not depend on its caller.

- A whitespace-only slug passed the length>0 test verbatim and occupied a roster
  entry it could never match. Slugs are now trimmed before the emptiness test. A
  blank body correctly falls through to the legacy alias rather than DROPPING the
  lane, which would have been worse than the blank slug.
- KNOWN_REVIEWER_SLUGS is computed at require() time, so an uncaught throw there
  breaks import for EVERY consumer rather than degrading selection. It is now
  guarded, yielding an empty roster on a malformed registry. That is a visible
  degradation, not a silent one: under D4 an explicitly requested reviewer that
  is unavailable is an ERROR, so /gsd:review --claude against an empty roster
  fails loudly. This also removes an asymmetry -- the sibling capability-trust
  module documents its collectors as TOTAL and wraps them for exactly this reason.

Also records three findings that previously existed ONLY in squash-merged PR
bodies, which is not a durable record:

- ADR-2782 D5 gains an implementation note explaining why the resolved host is
  deliberately EXCLUDED from the disclosure signature. Rule 1 says consent binds
  the resolved host; the loader has no config resolver, so folding it in would
  make the loader and lifecycle compute different signatures for one manifest and
  re-prompt forever. The binding is split: signature covers the SHA-pinned
  manifest fields, the consent record stores the resolved host, and Phase 5b
  re-resolves at invocation -- which is where rule 4 already puts the check. A
  reader comparing rule 1 to the code would otherwise conclude it is unimplemented.
- CONTEXT.md's capability-trust entry still described THREE executable surfaces.
  Phase 3 added the fourth and made that false; corrected here, since it is drift
  this epic introduced rather than Phase 6's new-glossary-term work.
- stableJson documents the NaN/Infinity/undefined -> null signature collision and
  why it is unreachable (JSON grammar has no such literal, so JSON.parse throws
  first). Reachability rests entirely on the ingest path staying JSON.parse-only,
  so the note lives where someone would break it.

* chore(#2798): backfill changeset pr number to 2837
2026-07-29 16:58:02 -04:00
Tom Boucher
c07734216f fix(#2603): document kimi-code in the host-integration capability matrix (#2687)
* fix(#2603): document kimi-code in the host-integration matrix; correct 3 inherited axes

The matrix — ADR-1239's deployment source-of-truth — had a section for 18 of 19
installed runtimes but none for `kimi-code`, so its `hostIntegration` axes shipped
with no citation and no evidence quote.

Sourcing every axis independently against Kimi Code CLI's own docs (the issue's
explicit requirement — `kimi` and `kimi-code` are distinct products) showed three
values had been inherited from the Python `kimi` descriptor rather than sourced:

- `embeddingMode` imperative -> declarative. Kimi Code plugins are a
  `kimi.plugin.json` manifest plus markdown Skills with no in-process programmatic
  API (docs/en/customization/plugins.md) — the same shape as `codex`.
- `dispatch.nested` false -> true. The `coder` built-in "can dispatch its own
  nested sub-agents when a task decomposes naturally" (docs/en/customization/agents.md).
  The Python `kimi` CLI genuinely prohibits nesting; Kimi Code does not.
- `dispatch.maxDepth` 1 -> "undocumented". Nesting is documented but no depth
  bound is published, so the fail-closed sentinel applies over a guessed integer.

`dispatch.namedDispatch` deliberately stays `false`: GSD's kimi-code artifact
layout installs Agent Skills only (no `agents` kind), so no named GSD subagent is
registered with the host and `resolveDispatchType` maps every role onto
coder/explore/plan. Flipping it would reintroduce the dispatch failure recorded in
docs/migration/kimi-to-kimi-code.md. The matrix records the host-capability nuance
under Documentation gaps instead.

Behaviourally inert: `namedDispatch:false` already caps nested/maxDepth/background/
backgroundDispatch to false/0 in the effective axes (host-integration.cts:493-499),
and the install adapter is not selected by `embeddingMode` (install.js:543 always
uses the imperative adapter). The one visible effect is the curated profile pin,
which moves programmatic-cli -> declarative-cli.

Also fixes the axes legend, which omitted the `built-in-only` subagentToolkit
member that has been in the closed vocabulary since kimi-code shipped.

Same defect class and countermeasure as #2598: pin the corrected values and require
the matrix to agree with the descriptor, because a descriptor/matrix disagreement is
how the gap survived.

Closes #2603

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* fix(#2603): report the maxDepth `undocumented` sentinel as a sentinel, not as malformed

Surfaced by the orthogonal review of this change. `negotiateHostCapabilities`
emits a sentinel-specific warning for every dispatch sub-axis carrying the
documented `undocumented` value — namedDispatch, nested, background,
subagentToolkit, backgroundDispatch, isolation — except `maxDepth`, which fell
through to the numeric guard and reported `host dispatch.maxDepth is missing or
not a number — treating as 0`.

That message is indistinguishable from a genuinely malformed descriptor, so a
correctly fail-closed descriptor reads as broken. Six shipped runtimes carry the
sentinel here (antigravity, augment, opencode, trae, windsurf, zcode) and this
PR's kimi-code correction adds a seventh, which is why it is fixed here rather
than left in place.

The numeric guard keeps firing for genuinely malformed values; both paths still
degrade `effective.dispatch.maxDepth` closed to 0. Covered by three tests,
including the boundary case that the sentinel carve-out must not swallow a real
malformed value.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* chore(#2603): backfill changeset PR number (#2687)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 22:52:57 -04:00
Tom Boucher
a3853472de fix(#2598): declare OpenCode subagent dispatch synchronous, not background (#2682)
* fix(#2598): declare OpenCode subagent dispatch synchronous, not background

capabilities/opencode/capability.json advertised dispatch.background: true and
dispatch.backgroundDispatch: true. negotiateHostCapabilities and every
degradationFor / shouldFlattenDispatch consumer trusts these per-field values, so
declaring a capability the host lacks OVERSTATES it — the opposite of the
fail-closed posture the negotiation is built for.

The issue's own citations needed checking before acting: the host-integration
matrix (ADR-1239's designated deployment source-of-truth) documented `true` with
NEWER evidence than the issue cited, and explicitly marked the issue's
sst/opencode#5887 reference as a stale snapshot superseded by #2087. git log
confirms #2087 deliberately flipped these from false to true, citing OpenCode
v1.15.0/v1.17 as "background subagents enabled by default in all modes". Applying
the issue as filed would, on that evidence, have REGRESSED a deliberate update.

So the claim was verified against current upstream rather than either document.
`packages/opencode/src/effect/runtime-flags.ts` on `dev` today reads:

    experimentalBackgroundSubagents: enabledByExperimental("OPENCODE_EXPERIMENTAL_BACKGROUND_SUBAGENTS")

`enabledByExperimental` falls back to the `experimental` flag and `bool()`
defaults to false — the parameter is hidden from the model unless an operator
opts in by env var. Upstream #29638 is still OPEN and confirms the session loop
`tasks.pop()`s one subtask at a time. #2087's "default-on in all modes" reading
does not hold against current dev.

The issue's CONCLUSION is therefore right even though part of its evidence was
superseded: concurrent dispatch cannot be relied on, so both fields are false.

The matrix rows are corrected with the verified citation rather than reverted to
the old #5887 quote, so the record shows why the value is false TODAY rather than
re-asserting evidence that was legitimately superseded. Neighbouring sub-fields
are untouched and pinned by test: namedDispatch, subagentToolkit, and
isolation:'orchestrator-worktree' (which works via `opencode run --dir` at the OS
process level and is unaffected — #2584 does not depend on this value either way).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* fix(#2598): re-pin the dispatch contract tests to synchronous OpenCode dispatch

gsd-test on the descriptor change came back FAILED (5 unique, both node
versions). The failures were not incidental — they were deliberate contract-pin
tests encoding #2087's decision, one named literally "background UPGRADE":

  tests/host-integration-descriptors.test.cjs
    - EXPECTED_FLATTEN[opencode] === false (background-eligible)
    - the derived background-eligible set pin
  tests/opencode-imperative-reference.test.cjs
    - "descriptor declares background dispatch true/true (v1.15/v1.17 upgrade)"
    - "background UPGRADE changes shouldFlattenDispatch: false now"

So this is a recorded decision being reversed, not drift being corrected, and it
is reversed on evidence: current upstream `dev` gates the capability behind
OPENCODE_EXPERIMENTAL_BACKGROUND_SUBAGENTS (default false) and upstream #29638
(OPEN) confirms the session loop still handles one subtask at a time. The issue
is filed by the maintainer and explicitly directs "update golden-parity /
validator fixtures as needed", which sanctions re-pinning.

Behavioral consequence, verified: shouldFlattenDispatch(opencode) now returns
TRUE, so GSD serializes opencode dispatch instead of trusting concurrency it
cannot get. That is the correct fail-closed direction and is safe today — no
shipped GSD flow drives OpenCode background waves (per the issue), and
isolation:'orchestrator-worktree' is unaffected because it works at the OS
process level via `opencode run --dir`, not via the native subagent.

Each re-pinned test now asserts the retracted contract in the opposite
direction — feeding the #2087 axes back in must still yield "would not flatten" —
so a silent re-flip of either field is caught rather than merely un-asserted.

lint:ci exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* chore(#2598): backfill changeset pr number (#2682)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 21:47:54 -04:00
Tom Boucher
0d08c32048 fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable (#2681)
* fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable

Every emitted script was rejected. Four invalid constructs, the first fatal on
its own, so the Workflow backend could never dispatch a wave:

  1. no `export const meta = {…}` first statement -> whole script rejected
  2. resumeFromRunId("<id>")  -> "resumeFromRunId is not defined". It is a
     Workflow TOOL INPUT parameter, not a script function. The run id still
     reaches the caller via summary.resumeRunId, to pass as that input.
  3. budget(<n>)              -> "budget is not a function". `budget` is a
     read-only object { total, spent(), remaining() } fed by the caller's token
     directive; a script cannot set it. Recorded as intent in a comment.
  4. parallel(agent(…), agent(…)) -> "parallel() expects an array of functions".
     Now parallel([() => agent(…), …]) — passing agent() results directly also
     started every agent eagerly, before parallel() could bound concurrency.

The single-plan stage had its own branch with the same parallel() defect; both
branches are now one array-emitting path. Waves also emit phase() calls whose
titles match meta.phases exactly, so progress groups correctly.

Two secondary defects kept the script from ever being REACHED — which is why
this shipped undetected:

  5. NOTHING resolved the Agent SDK version. The fragment claimed there was "no
     scriptable way" to introspect it and told callers to omit the flag, so
     gate 5 returned agent_sdk_version_unknown on every automated run while
     `capability state` still reported active:true. True for bash, false for
     Node: the router now reads the installed @anthropic-ai/claude-agent-sdk
     version, walking node_modules up the tree and reading package.json
     directly — require.resolve throws ERR_PACKAGE_PATH_NOT_EXPORTED because the
     SDK's exports map does not expose ./package.json. Precedence: explicit flag
     > GSD_AGENT_SDK_VERSION > installed. Fail-closed is preserved; an
     unresolvable version still declines to inline. A too-old SDK now reports
     the truthful agent_sdk_version_below_floor instead of unknown.
  6. The runtime fallback was `--runtime > GSD_RUNTIME > 'unknown'`, diverging
     from the canonical `GSD_RUNTIME > config.runtime > 'claude'`, so any
     invocation without --runtime reported runtime_not_claude on an ordinary
     Claude project. Now delegates to runtime-slash.resolveRuntime.

The fragment's `${AGENT_SDK_VERSION:+--agent-sdk-version "$AGENT_SDK_VERSION"}`
snippet is removed rather than repaired: it was also shell-dependent — zsh does
not word-split unquoted parameter expansions, so it collapsed to a single argv
element, argValue() never matched, and the run failed into the same
agent_sdk_version_unknown, indistinguishable from genuinely unknown. Auto-
resolution removes the need for the construct entirely.

Verified with the issue's own repro: no flags now reaches the version gate; an
SDK above the floor yields backend:"workflow" with a script that parses as a
real ES module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* fix(#2590): sync generated registry, repair sibling tests, reject duplicate wave ids

Findings from the isolated review, all fixed.

HIGH — gsd-core/bin/lib/capability-registry.cjs was stale, and `lint:ci` was
already RED because of it. The registry embeds the fragment text INLINE, so the
shipped/installed copy still taught the exact broken contract this PR fixes:
the old `${AGENT_SDK_VERSION:+…}` bash line and the "OMIT the flag when unknown"
guidance. Regenerated. (I had read `lint:ci` by grepping its output instead of
checking its exit code, so I recorded a red chain as green — checking $? now.)

HIGH — three existing tests asserted the OLD broken shape and would have failed
CI; none was touched by the first commit:
  tests/fix-2285-claude-orchestration-wiring.test.cjs — matched resumeFromRunId("…")
  tests/claude-orchestration.test.cjs                 — .includes('budget(')
  tests/claude-orchestration-command-router.test.cjs  — .includes('budget(')
Each now asserts the corrected contract: the id/pool reaches the caller via
summary, and neither construct is ever CALLED. Two sibling assertions had also
gone vacuous — `.includes('resumeFromRunId')` still passed, but only because the
new explanatory COMMENT contains that substring, not because anything is wired.
Rewritten to assert the real property.

MEDIUM — duplicate wave ids were never rejected. Plan-id uniqueness was checked
within a wave, but nothing checked wave ids across waves. That was harmless
before; it is not now, because each wave emits a `phase("Wave <id>")` call plus a
matching meta.phases entry and the tool matches titles by exact string — two
waves sharing an id would collapse into one progress group and misattribute the
second wave's agents to the first. Rejected at validation, with tests either
side of the boundary.

MEDIUM — the fragment contradicted itself (its "Manifest construction" header
still listed $AGENT_SDK_VERSION as orchestrator-built) and, more seriously, never
told the orchestrator to pass summary.resumeRunId as the Workflow tool's
resumeFromRunId INPUT. Since this PR moves resume from a broken in-script call to
a tool-invocation input, an implementer following only the fragment would have
silently regressed phase-resume to a no-op. Both fixed.

MEDIUM — docs/how-to/enable-claude-orchestration-workflow-backend.md and
docs/explanation/claude-orchestration-capability.md documented
`resumeFromRunId("<id>")` and `budget(<tokens>)` as current correct output —
teaching the bug as the feature. Updated to the real contract, including the
required meta block and the thunk-array parallel() form. (The changeset is
`Fixed`, so the docs gate exempts this; it is corrected because it is wrong,
not because a gate demanded it.)

LOW — the router's top-of-file comment still described the divergent
`--runtime > GSD_RUNTIME > 'unknown'` chain as current, ninety lines above the
fix; and inserting resolveInstalledAgentSdkVersion had orphaned
resolveDetectionArgs' JSDoc above the wrong function. Both repaired.

lint:ci now exits 0 (verified by exit code, not by reading output).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* chore(#2590): backfill changeset pr number (#2681)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:49:38 -04:00
Tom Boucher
6ad30f74b6 feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler.

harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run.

Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard.

Closes #2627

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:50:20 -04:00
Tom Boucher
4a66d62d10 feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver (#2625)
* feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver

Phase 2 of the negotiated executor-isolation feature (ADR-1239 Codex-binding amendment). Two building blocks for `dispatch.isolation: orchestrator-worktree` hosts, both unconsumed — no scheduler wires them yet (that is Phase 3), so no runtime behavior changes.

worktree create verb (planWorktreeCreate / executeWorktreeCreatePlan / cmdWorktreeCreate in worktree-safety.cts, routed via routeWorktree in gsd-tools.cjs): validates the wave base, creates a bounded branch+worktree, records it in the run manifest reusing record-agent 4-field entry shape, returns the executor working directory. Bounded git (10s timeout, degrade-not-throw); all manifest read/parse/validate/dedupe precedes the single git side effect (no unmanifested-orphan on a bad manifest); timeout-only best-effort partial rollback (a clean collision-exit never removes a live peer worktree); fail-closed on bad base, unsafe leading-dash / .. inputs, and malformed/mis-shaped manifest.

resolveOrchestratorExec (host-integration.cts): pure descriptor->argv resolver reading the new runtime.orchestratorExec descriptor field (codex/opencode/kimi/kimi-code), fail-closed on missing/invalid shape. Validator (capability-validator.cjs) + a parity guard asserting every orchestrator-worktree host declares a resolvable orchestratorExec.

Adding the create route edits the installed gsd-core/bin/gsd-tools.cjs, so the golden-install-parity fixtures for all 19 runtimes are regenerated (npm run gen:golden) — the only changed hash is gsd-tools.cjs. CONTEXT.md glossary updated; capability-registry regenerated. Behavioral tests (worktree-safety + host-integration) incl. a fast-check property test and the parity sweep.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: rebuild tracked state-transition.cjs to match #2400 source

The tracked compiled artifact drifted from src/state-transition.cts: #2400 (commit 2bcfaa2e2) added the progress.total_plans frontmatter sync to source but the tracked bin/lib/state-transition.cjs was never rebuilt, so the fix was not shipping to consumers of the compiled artifact. The mandatory build:lib step for Phase 2 surfaced the drift; recompiling makes the already-merged, already-changelogged #2400 fix effective. Artifact-only resync (no source/test change); drift class tracked by #2591.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-24 21:23:06 -04:00
Tom Boucher
ec7978c0b4 feat(#2584): add dispatch.isolation sub-field, descriptors, validator + negotiation (#2604) 2026-07-24 12:51:42 -04:00
Tom Boucher
d579daa3ed docs(#2505): Phase 6 — migration guide + built-in-only subagent-toolkit enum (#2538)
* docs(#2512): Phase 6 — migration guide + built-in-only subagent-toolkit enum

* fix #2512: update CONTRACT-PIN for built-in-only subagentToolkit value

* docs(changeset): backfill PR #2538 for Phase 6 (#2512)
2026-07-22 14:30:26 -04:00
Tom Boucher
c2a305c44d feat(#2505): Phase 2 — kimi-code Agent Skills install layout (#2520)
* feat(#2454): PR 2 — kimi-code Agent Skills converter + install layout

PR 1 registered the kimi-code EoS descriptor with empty artifactLayout
(SKIP_INSTALL_CONTRACT excluded it from the end-to-end install test).
PR 2 fills in the install surface:

- src/runtime-artifact-conversion.cts: new convertClaudeCommandToKimiCodeSkill
  function. Today it delegates to convertClaudeCommandToKimiSkill (Python
  kimi-cli) because Kimi Code uses the same Agent Skills format + /skill:
  invocation per official docs. The distinct function name lets a future
  divergence land cleanly if Kimi Code's skill format evolves independently.
- gsd-core/bin/lib/capability-validator.cjs: add to ALLOWED_SKILLS_CONVERTERS.
- capabilities/kimi-code/capability.json: artifactLayout.global now declares
  the skills kind with converter='convertClaudeCommandToKimiCodeSkill' +
  home='.kimi-code' (auto-discovered at ~/.kimi-code/skills/ per Kimi Code
  docs: merge_all_available_skills = true default).
- tests/installer-migration-install.integration.test.cjs: REMOVE the
  SKIP_INSTALL_CONTRACT exclusion — kimi-code now has a full install surface.
- Regenerated capability-registry + capability-matrix + golden install
  parity + install tree fixtures for kimi-code.

* fix(#2454): wire kimi-code converter into SKILLS_CONVERTER_REGISTRY + count bump

- src/install-engine.cts: add convertClaudeCommandToKimiCodeSkill to
  SKILLS_CONVERTER_REGISTRY so the layout-driven skills install path
  can dispatch off the descriptor's converter string.
- tests/capability-registry.test.cjs: bump VALID_CONVERTER_NAMES count
  26 → 27 (added convertClaudeCommandToKimiCodeSkill).

* fix(#2454): remove home override from kimi-code skills (inherit configDir)

The home:'.kimi-code' override made the install plan resolve skills dest
to ~/.kimi-code/skills instead of <configDir>/skills, causing the test's
temp configDir to miss the install. Removing it lets skills inherit
configDir like most runtimes.

* fix(#2454): kimi-code install contract surface is flat-skills (no agents)

Kimi Code has NO custom named subagents (per official docs: 3 built-in
coder/explore/plan only). The kimi-skills-agents surface expects agents/
gsd.yaml + subagents/*.yaml which kimi-code does not produce. Changed
to flat-skills which only checks for skills/gsd-* dirs.

* docs(changeset): Phase 2 kimi-code install layout Added (#2509)

* docs(changeset): backfill PR #2520 for Phase 2 (#2509)
2026-07-22 00:58:58 -04:00
Tom Boucher
bf8f320083 feat(#2505): Phase 1 — EoS descriptor split (kimi-code capability.json + drift-guard registration) (#2519)
* feat(#2454): add kimi-code as an EoS capability (Node Kimi Code CLI)

PR 1 of N for #2454. Establishes the EoS descriptor foundation for splitting
GSD's kimi support into two distinct products per the user's directive:
- kimi       (existing): Moonshot's Python kimi-cli (~/.kimi, runtime: python)
- kimi-code  (new):      Moonshot's Node Kimi Code CLI (~/.kimi-code,
                         runtime: node, KIMI_CODE_HOME env)

Per ADR-1239 EoS, runtime behavior is driven by capabilities/<id>/capability.json
descriptors, not hardcoded branches in install.js. The new descriptor uses
the existing primitives (dot-home configHome, skills artifactLayout, kimi-hooks-toml
hooksSurface — same TOML [[hooks]] format Kimi Code reads per its docs).

Critical Kimi Code constraint reflected in the descriptor:
  hostIntegration.dispatch.namedDispatch: false
  hostIntegration.dispatch.builtInSubagents: ['coder', 'explore', 'plan']
  hostBehaviors.namedSubagentsSupported: false
Kimi Code's official docs confirm only 3 built-in subagents with NO custom-
subagent registration (the [subagent] table only has timeout_ms). The
kimi-agents YAML layout (used by Python kimi-cli) is therefore NOT in
kimi-code's artifactLayout.

Schema adjustments:
- subagentToolkit set to 'undocumented' (the existing escape hatch); the
  schema enum (full/read-only) lacks a 'limited'/'built-in-only' value.
  A follow-up PR can extend the schema enum to add 'built-in-only' as a
  first-class axis value reflecting Kimi Code's documented model.

Registration:
- capabilities/kimi-code/capability.json (new descriptor, modeled on codex)
- bin/install.js: allRuntimes array + --all list + --kimi-code flag
- gsd-core/bin/shared/runtime-aliases.manifest.json: kimi-code aliases
  (kimi-code, kimicode, kimi_code)
- src/runtime-name-policy.cts: FALLBACK_ALIASES map
- gsd-core/bin/lib/capability-registry.cjs: regenerated via
  scripts/gen-capability-registry.cjs --write

Tests:
- tests/multi-runtime-select.test.cjs updated for the new runtime count (18)
  + new --kimi-code flag test + 'All' shortcut renumbered 18 → 19.

Out of scope for PR 1 (follow-up PRs in the sequence):
- Install-time decision logic (kimi vs kimi-code detection / prompt)
- agent-install-check semantics for kimi-code (verify Agent Skills presence)
- cmdAgentSkills fallback returning subagent prompt content
- Workflow template mapping (named agents → built-in coder/explore/plan)
- Migration guidance for users currently on 'kimi' who are actually on Kimi Code
- Schema enum extension for subagentToolkit: 'built-in-only'

Refs #2454, #2095 (EoS/kimi migration epic), ADR-1239 (EoS).

* fix(#2454): complete drift-guard registrations for kimi-code runtime

The drift guards caught every surface that pins runtime enumeration. Each
update is mechanical, driven by the guard's named failure mode:

- src/runtime-name-policy.cts RUNTIME_LABELS: 'Kimi Code' label for kimi-code
- src/runtime-name-policy.cts RUNTIME_FLAG_IDS: add kimi-code to the
  isKimiCode predicate generator
- bin/install.js runtimeMap: option '11' → 'kimi-code', renumber downstream
  entries (11..17 → 12..18), ALL_RUNTIMES_OPTION 18 → 19
- gsd-core/bin/shared/model-catalog.json runtimeTierDefaults: kimi-code entry
  (null/null/null — same as kimi, no model tier defaults until configured)
- docs/reference/capability-matrix.md: regenerated via
  scripts/gen-capability-matrix.cjs --write (kimi-code row added)
- tests/global-config-home-fragment.test.cjs GOLDEN_FRAGMENT_MAP:
  kimi-code → '.kimi-code'
- tests/fixtures/golden-install-parity/*.json: regenerated via npm run gen:golden
  (the runtime-aliases.manifest.json hash changed; all 17 runtime fixtures updated)

The capability-registry is already regenerated from the prior commit.

* test(#2454): update drift-guard tests for kimi-code runtime registration

Multiple drift guards pin runtime enumeration counts and option numbering.
Each update is mechanical, driven by the guard's named failure mode:

- tests/runtime-flags.test.cjs: EXPECTED_FLAGS gains isKimiCode (16 → 17);
  'all 16 flags' → 'all 17 flags' in test names + messages.
- tests/multi-runtime-select.test.cjs: parseRuntimeInput option renumbering
  cascade — kilo moves 11→12, opencode 12→13, pi 13→14, qwen 14→15,
  trae 15→16, windsurf 16→17, zcode 17→18, All 18→19. New single-choice
  test for kimi-code (option 11). Prompt test updated for new numbering.
- tests/host-integration-descriptors.test.cjs: EXPECTED_PROFILES gains
  kimi-code → 'programmatic-cli' (terminal CLI per Kimi Code docs);
  EXPECTED_FLATTEN gains kimi-code → false (backgroundDispatch:true per
  docs, same as Python kimi/opencode).
- tests/global-config-home-fragment.test.cjs: table-count test renamed
  13 → 14 table runtimes (kimi-code added to GOLDEN_FRAGMENT_MAP earlier).

* fix(#2454): empty artifactLayout for kimi-code (PR 1 scope)

The skills kind requires a converter (existing converters are per-runtime
like convertClaudeCommandToKimiSkill). PR 1 of this multi-PR sequence only
registers the descriptor; the actual Agent Skills converter (and a new
'convertClaudeCommandToKimiCodeSkill' function) lands in PR 2 alongside
the install-time decision logic. Empty artifactLayout.global is valid and
means 'nothing to install yet via the layout seam'.

Also: added kimi-code to RUNTIME_META in tests/helpers/install-shared.cjs
(localDir .kimi-code, globalSuffix .kimi-code), and added Kimi Code as
option 11 in install.js's buildRuntimePromptText (renumbered downstream
options 11..17 → 12..18, All 18 → 19).

* fix(#2454): camelCase runtimeFlags for hyphenated ids (kimi-code → isKimiCode)

The runtimeFlags generator previously produced 'isKimi-code' (hyphen preserved)
for the new kimi-code runtime id. Property names with hyphens are awkward for
consumers (flags['isKimi-code'] instead of flags.isKimiCode). The new
runtimeIdToFlagName helper folds -[a-z] boundaries to uppercase, producing
the conventional PascalCase flag name. The 16 prior single-word runtime ids
are unaffected (the regex finds no hyphens).

* fix(#2454): update remaining drift-guard tests + gen kimi-code fixtures

- tests/runtime-flags.test.cjs drift guard: use proper kebab-case
  conversion (isKimiCode → kimi-code, not 'kimicode') so the registry
  comparison doesn't false-positive on hyphenated runtime ids.
- tests/multi-runtime-select.test.cjs: fix kilo/opencode/pi/qwen/trae
  single-choice tests for the renumbered options (kilo 11→12, opencode
  12→13, pi 13→14, qwen 14→15, trae 15→16).
- tests/install.test.cjs: Kilo integration option 11→12, prompt test
  regex updated.
- tests/fixtures/golden-install-parity/kimi-code.json + install-tree/
  kimi-code.json: generated via UPDATE_GOLDEN=1 + UPDATE_INSTALL_TREE=1.
  The kimi-code install produces the standard GSD install layout (skills,
  contexts, references, etc.) — 436 paths, same shape as other runtimes
  that have no custom converter yet.

* fix(#2454): add kimi-code install contract + global config home fragment

- src/runtime-name-policy.cts GLOBAL_CONFIG_HOME_FRAGMENTS: add kimi-code
  → '.kimi-code' so getGlobalConfigHomeFragment returns the correct path
  instead of falling through to the default '.claude'.
- tests/installer-migration-install.integration.test.cjs
  RUNTIME_INSTALL_CONTRACTS: kimi-code entry (same surface as kimi for
  PR 1; PR 2 will specialize once the Agent Skills converter lands).
- tests/multi-runtime-select.test.cjs: fix space-separated-choices test
  for the renumbered kilo option (11 → 12).
- tests/fixtures/golden-install-parity/kimi-code.json + install-tree/
  kimi-code.json: regenerated after rebasing onto current next (new
  planner-reversibility.md from #2471 etc. now included).

* test(#2454): skip kimi-code install contract until PR 2 ships install layout

The end-to-end install test (tests/installer-migration-install.integration
.test.cjs) asserts every allRuntimes entry installs a runtime-specific
artifact surface. PR 1 of #2454 registers kimi-code in allRuntimes + the
capability descriptor + flags + labels, but the install LAYOUT (Agent
Skills converter + global AGENTS.md at $KIMI_CODE_HOME/AGENTS.md) lands
in PR 2. The SKIP_INSTALL_CONTRACT set marks this exclusion explicit and
self-removing — PR 2 removes the entry alongside adding the install
surface, restoring the contract loop to full coverage.

* fix(#2454): restore compact model-catalog.json format (M1 review)

Per code-review M1: my prior 'fix(#2454): complete drift-guard registrations'
commit used python json.dump(indent=2) which inflated the file from 165→607
lines (every nested entry got expanded) and lost the trailing newline. The
semantic change was just a 3-line kimi-code entry. Restored the original
hybrid format (top-level indent=2 + inner entries' one-line style) and
added kimi-code in matching form.

Regenerated golden install parity + install tree fixtures since the
model-catalog.json hash changed.

* fix(#2454): update CONTEXT.md allRuntimes glossary (17 → 18, add kimi-code)

CI lint-tests job failed on the glossary drift guard
(scripts/check-glossary-refs.cjs --check):
  ✗ CONTEXT.md's allRuntimes enum-count sentence claims 17 values but
    bin/install.js's allRuntimes array has 18.
  ✗ CONTEXT.md's allRuntimes member list has drifted from bin/install.js
    (missing from CONTEXT.md's list: kimi-code).

Missed in the prior commits because gsd-test does not run the glossary
check (it's a CI lint-tests-only check). Updating CONTEXT.md's two claims
to 18 values + kimi-code in the member list.

* chore(#2505): regen capability-registry + stamp kimi-code version 1.8.0 (#2511)

* docs(changeset): Phase 1 kimi-code runtime Added (#2511)

* test(#2511): regen kimi-code golden parity fixture after Phase 0 guard normalization lands

* docs(changeset): backfill PR #2519 for Phase 1 (#2511)
2026-07-22 00:27:35 -04:00
Tom Boucher
d39ca0c004 Merge pull request #2490 from open-gsd/docs/2481-adr-1239-effort-axis
feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort
2026-07-21 20:14:27 -04:00
github-actions[bot]
8cf747724f chore: sync next package version to 1.8.0 2026-07-22 00:06:08 +00:00
Tom Boucher
09b535ac00 feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort
ADR-1239 gains a ninth negotiated axis, effortSurface (argv | none), declaring how
a host accepts reasoning effort. ADR-443 is amended in the same change because its
recorded deferral is what the axis resolves: its Unblock condition offered paths
(a) and (b) and stated the choice was 'a maintainer call this file records but does
not make'. Path (a) is selected and satisfied here.

Before this, effort reached a runtime only through install-time channels
(EFFORT_RENDERING's frontmatter/api), so reviewer CLIs spawned as subprocesses
silently inherited whatever effort sat in the user's own global CLI config. The
review lane now resolves one universal effort through the ADR-443 cascade and
renders it per host through the negotiated descriptor.

Every per-host value is documentation-sourced, never inferred:
- claude   argv  -- verified via 'claude --help' (--effort <level>)
- opencode argv  -- verified via 'opencode run --help' (--variant)
- codex    argv  -- codex-rs/exec/src/cli.rs: model_reasoning_effort is NOT a CLI
                    flag (config.toml key only), so the global -c override is the
                    only argv route
- 15 hosts undocumented -- their docs state no reasoning setting; the sentinel
                    fails closed rather than inheriting a profile baseline

No config-file vocabulary member: the only host that ever had one (Gemini CLI's
thinkingConfig) was removed as a sunset runtime by 8f2ebbe9b (#1928, PR #1996),
and neither Antigravity CLI nor ZCode documents a reasoning setting.

Closes #2481

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 19:10:22 -04:00
Tom Boucher
909a3b180b fix(#2470): install pi's extension as gsd.js so pi actually discovers it (#2478)
* test(#2470): failing-first — pi extension must satisfy pi's auto-discovery filter

pi auto-discovers extensions/ entries through isExtensionFile(), which accepts
only .ts and .js. GSD installs its extension as gsd.cjs, so pi silently skips
it: no /gsd command, no error, no log line.

Encodes pi's discovery PREDICATE rather than a literal filename, so the
contract keeps holding across future renames, and adds the migration-006 test
matrix for retiring the stale gsd.cjs left in pre-fix installs.

Red until the fix lands.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2470): install pi's extension as gsd.js so pi actually discovers it

pi auto-discovers extensions/ entries via isExtensionFile(), which accepts
only .ts and .js and skips everything else silently. capabilities/pi declared
the dest as gsd.cjs, so the extension installed correctly and was then ignored
forever: no /gsd command, no error, no log line.

Install it as gsd.js. The in-repo source stays pi/gsd.cjs — tests require() it
directly and .cjs is unambiguous CommonJS; only the installed name has to
satisfy pi, and pi loads accepted files through jiti, which handles CJS and ESM
alike. (The reporter's premise that ~/.pi/agent/package.json declares
"type":"commonjs" does not hold — pi never writes that file.)

Renaming an installed artifact requires a migration record, so add 006 to
retire the stale gsd.cjs from pre-fix installs; without it the old path drops
out of the manifest and uninstall can never remove it. The migration plans
nothing for an unmanifested gsd.cjs: emitting remove-managed there would have
the executor downgrade it to preserve-user and mark it blocked, failing the
install for anyone who hand-placed their own file.

Also pins body-parser >=2.3.0 (GHSA-v422-hmwv-36x6). The advisory reaches the
production tree transitively via the Claude Agent SDK and fails the
npm-integrity gate, blocking any PR; pinned via the existing overrides idiom.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2470): address orthogonal review findings + register migration checksum

Code review:
- pi/gsd.cjs's install docstring still told readers to copy the file to
  extensions/gsd.cjs — the exact silently-broken state this PR fixes. Anyone
  following it recreated the bug.
- Two stale extensions/gsd.cjs comments in install-minimal-hooks.test.cjs.

Security review:
- _installNativePluginIfDeclared confined nativePlugin.dir but joined
  nativePlugin.file onto the validated directory unchecked, so a descriptor
  whose file carried .., an absolute path, or a NUL byte would have written
  outside configHome. Not reachable in a shipped build (descriptors are
  first-party and compiled into the capability registry), but file is exactly
  the field this PR changes. Confine the full dest path instead; for a
  well-formed descriptor this resolves identically to the previous
  mkdir(dir) + join(dir, file). Covered by four new write-confinement tests.

Also register migration 006 in the #670 EXPECTED_CHECKSUMS baseline — shipped
migration bodies are locked to a committed checksum and a new migration fails
CI until it is listed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2470): never dereference a symlinked managed path when snapshotting

fs.copyFileSync follows symlinks, so a managed path replaced by a link had the
REFERENT's bytes copied into the migration journal's rollback and backup trees
— a gsd.cjs symlinked at a private key would land that key's contents under
gsd-migration-journal/. Deletion was already safe (fs.rmSync unlinks the link,
never the target); the copy was not.

Nothing GSD installs is ever a symlink, so the faithful snapshot of a symlinked
managed path is the link itself. copyPreservingSymlink recreates it, which
keeps rollback fidelity (restore re-creates the same link) while never reading
the referent. Scoped the pre-delete to the symlink branch only, so the
regular-file path keeps copyFileSync's overwrite-in-place and a mid-restore
failure cannot destroy the destination. The restore-side existence check moves
to lstat, since existsSync follows a link whose target is gone and would
silently skip the restore.

This lives in the engine all six migrations share, so 000-005 are hardened too.

Also regenerates the pi golden-parity hash: correcting pi/gsd.cjs's own install
docstring changes the extension's content, which the golden suite caught.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2470): symlink-preserve the in-apply failure-recovery restore too

The previous commit routed three copy sites through copyPreservingSymlink but
missed a fourth: the catch block inside applyInstallerMigrationPlan, which
replays rollback snapshots taken earlier in the SAME apply attempt. Those
snapshots are symlinks precisely because of that commit, so the raw
copyFileSync there dereferenced them and wrote the referent's bytes to the LIVE
install path — worse than the journal-tree leak it was meant to fix, since it
is user-visible and at a predictable location.

Verified by experiment rather than assertion: with the pre-fix line restored,
the managed path comes back as a REGULAR FILE containing the referent's bytes;
with the fix it comes back as a symlink and the bytes appear nowhere.

The accompanying test injects the failure by letting the delete succeed and
then throwing once, modelling a later step failing after the delete. That
ordering is load-bearing — an earlier draft injected before the delete, which
leaves the live path in place, so the pre-fix copyFileSync hit a same-file
collision and threw instead of leaking. That draft passed against the bug it
was written to catch; this one fails against it.

Adds the missing rollback() coverage as well: a restored symlinked managed path
must come back as a link pointing at its original target.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2470): read the backup location from the journal, not the plan

The new backup-content assertion read backupRelPath off result.plan.actions,
where it is always null: the planner reserves the field and apply chooses the
concrete location, recording it in the journal. The assertion therefore failed
on "backup path must be recorded for the user" rather than on anything about
the behavior it was written to check.

Read it from the journal, which is the authoritative record. Verified by
executing all four new test bodies in-process against the built engine — the
backup file exists and holds the locally patched content.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2470): backfill changeset pr number to 2478

* chore(#2470): backfill changeset pr number to 2478

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 08:26:47 -04:00
Tom Boucher
d16a66479a feat(#1950): broken-windows ledger — cross-phase defect register gating ship (#2441)
* feat(#1950): broken-windows ledger — cross-phase defect register gating ship

Adds a new  capability (#1950) that operationalizes GSD's
no-defer discipline as a tracked, enforced artifact:
accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths
across phases, and /gsd-ship blocks while any entry is open.

Implementation:
- src/broken-windows.cts → gsd-core/bin/lib/broken-windows.cjs: typed IR +
  I/O entry points (parseLedger/renderLedger/appendWindow/markWaived/markFixed
  + cmdWindowsStatus/Append/Waive/MarkFixed). Frozen REASON enum for typed
  error assertions. Windows-safe atomic rename with retry on transient
  EPERM/EBUSY/EACCES.
- gsd-tools.cjs: new  subcommand (status | append | waive | fixed),
  wired via routeWindows + HOST_COMMAND_ROUTERS.windows.
- capabilities/broken-windows/capability.json: one ship:pre gate with
  artifact-frontmatter-equals predicate on WINDOWS.md open_count == 0.
  activationKey windows.enabled (default true) + sibling windows.enforce
  (default true, separate so tracking can precede enforcement).
- gsd-core/workflows/ship.md: capId==broken-windows branch in preflight,
  sibling to security — reads gsd_run windows status --raw, fails closed
  on open_count > 0 or unreadable ledger.
- agents/gsd-executor.md: extends the existing ## Known Stubs instruction
  to also append to WINDOWS.md via gsd_run windows append (best-effort,
  never blocks execution).
- agents/gsd-verifier.md: new Step 8b — record unmet truths + human-verify
  items in WINDOWS.md.
- gsd-core/workflows/progress.md: surfaces open + waived counts.
- docs/COMMANDS.md + CONTEXT.md glossary entry + docs/INVENTORY.md:
  document the gate, waiver mechanism, and new module.
- tests/broken-windows.test.cjs: pure + CLI behavioral coverage + fast-check
  roundtrip property; fail-closed on malformed ledger; security boundary on
  path traversal in --file.

Backward-compatible: a project with no .planning/WINDOWS.md reports
open_count: 0 and ships cleanly. Disable enforcement per-project with
gsd config-set windows.enforce false (tracking continues, gate stays open).

* chore(#1950): ratchet size baselines, defer verifier integration

- Workflow size baseline: ship.md 25575→27928, progress.md 31789→32632
  (broken-windows preflight branch + open-windows surface).
- Agent size baseline: gsd-executor.md 46644→47951 (Known Stubs → also
  appends to WINDOWS.md). gsd-verifier.md unchanged.
- LARGE_CAP (49152) preempted the planned verifier integration
  (gsd-verifier.md was at 49140 pre-PR — 12 bytes of headroom, not the
  documented 'real headroom'). Verifier integration deferred to a follow-up
  PR that extracts the VERIFICATION.md template (lines 739-859) to
  gsd-core/references/ — a pre-existing cap-tightness defect this PR
  exposed but does not expand scope to fix. Verifier integration is not in
  the issue's acceptance criteria (executor writes is; unmet-truths
  recording was an enhancement, not a gate).

* fix(#1950): gate default-off, rename to workflow.windows_enforce, regen goldens

Test-failure-driven fixes after first gsd-test run on db8733c8f failed 44
cases (pre-existing structural tests encoded 'ship:pre has 1 gate' / 'all
caps off → empty hooks'):

- capability manifest: rename windows.enabled+windows.enforce (default
  true) → single federated key workflow.windows_enforce (default FALSE,
  opt-in). Matches security's workflow.security_enforce convention and
  makes the adr857 all-caps-off test pass without modification (the test's
  buildAllFalseConfig handles workflow.* out of the box). Default-OFF keeps
  the gate out of the registry's default ship:pre resolution so existing
  loop-hooks-ship-pre-e2e structural assertions (exactly 1 gate, capId
  'security') stay valid; users opt in via
  gsd config-set workflow.windows_enforce true.
- drop activationKey (security doesn't have one either; workflow.* key
  doubles as the activation toggle).
- regenerate docs/reference/capability-matrix.md to include broken-windows
  (capability-matrix-sync test).
- regenerate tests/fixtures/golden-install-parity/*.json (18 runtimes) —
  installer now emits the new capability + lib file.
- update CONTEXT.md, docs/COMMANDS.md, docs/FEATURES.md, ship.md,
  agents/gsd-executor.md to use the new key name and /gsd:colon slash
  syntax (slash-command-namespace test).
- restore accidentally-regressed /gsd:capture in progress.md.

Tracking-only by default; enforcement is opt-in. Acceptance criterion
'/gsd-ship fails while any ledger entry is open' is met when
workflow.windows_enforce=true (test fixture enables it).

* test(#1950): update ship:pre structural invariants for 2-gate registry

- loop-hooks-ship-pre-e2e: the registry now declares 2 gates at ship:pre
  (security + broken-windows), regardless of activation. Activation tests
  above still pin security-only or empty behavior via fixtures; these
  structural tests pin the REGISTRY shape, which has 2 gates as of #1950.
- workflow-size-baseline: ship.md 27928→27945 (workflow.windows_enforce
  rename added 17 bytes).

* fix(#1950): review H1+H2+M1+M2+M3 — fence-injection, EACCES fail-closed, cleanup, strict line, stryker

Adversarial isolated review (Step 6.3) found 2 HIGH findings that block
the PR and 3 mediums. All addressed:

H1 (HIGH): description containing the markdown 3-backtick fence would
terminate the ledger's JSON code block early inside JSON.stringify output
(JSON doesn't escape backticks), corrupting the file and bricking the
next parse. Fix: use a 4-backtick fence (json ... ) which
JSON.stringify cannot produce on its own, AND validate that no entry
text field contains a 4-backtick run (reject at append time with new
WINDOWS_INVALID_TEXT reason code). Locked by a regression test.

H2 (HIGH): readLedgerOrNull swallowed ALL fs errors as 'no ledger',
silently returning open_count:0 on EACCES/EPERM/EIO. The ship gate
would then pass on an unreadable ledger — the precise vector the
workflow doc claims is impossible. Fix: only ENOENT returns null;
every other fs error propagates as WINDOWS_LEDGER_MALFORMED so the
gate blocks and the operator sees a real diagnostic. Locked by a
regression test that chmod 000s a ledger with open_count=1 and
asserts the result is never a false-green 0.

M1: writeLedgerAtomic left an orphaned .tmp file on rename failure.
Wrapped renameWithRetry in try/catch with best-effort unlink.

M2: validateLine silently coerced 'abc' → NaN → null, hiding type
drift. Removed the line === 0 special case (was undocumented) and
made the error message match the strict check. Now any non-positive-
integer line value throws, including strings.

M3: tests/broken-windows.test.cjs (with its fast-check property test)
was not in stryker.config.mjs DEFAULT_TEST_CMD — Stryker would mutate
src/broken-windows.cts but no test would catch the mutations,
producing false surviving-mutant scores. Added to the list.

L1 (dead throw e after error()), L7 (line boundary tests, H1/H2
regression tests, 4-backtick CLI test) also addressed.

* docs(#1950): inline concurrency + busy-wait notes (review L2+L3)

* fix(#1950): regen goldens against latest gsd-tools; correct --line 0 boundary test

gsd-test v4 caught two issues:
- goldens I regenerated earlier (commit 526682084) predated the L1
  routeWindows catch-block cleanup (commit dd844d565). Regenerated
  via 'npm run gen:golden' against current HEAD so the install
  parity hash for gsd-tools.cjs matches.
- 'append --line boundary' test expected --line 0 to succeed with
  null entry.line, but the M2 fix correctly rejects 0 (lines are
  1-indexed; 0 is not a valid source line). Updated the boundary
  test to assert --line 0 fails alongside -1 and 'abc'.

* chore(#1950): regen goldens after rebase onto next

* chore(#1950): quick.md baseline 50699→50993 (correct resolution from next rebase)

* chore(changeset): backfill pr:2441 in .changeset/broken-windows-ledger.md

* fix(#1950): renderTable escapes backslash before pipe (CodeQL incomplete-sanitization)

CodeQL flagged the markdown-table cell escaper:
  String(s ?? '').replace(/\|/g, '\\|')
— it escapes pipe but not backslash first. A description containing '\|'
would render as '\\|' which markdown parses as 'literal backslash' +
'cell separator', splitting the column.

Fix: escape backslash FIRST (each \ → \\), then pipe (each | → \|).
Now a description with '\|' renders as '\\\\|' (literal '\\' + escaped
pipe), which markdown renders as a single '\|' inside the cell. The JSON
code block (the parse source-of-truth) was already correctly escaped via
JSON.stringify; only the display-only table was affected.

Locked by a regression test that:
1. Verifies the JSON block reparses with the description intact.
2. Walks the rendered table row counting unescaped pipes — must be
   exactly 11 (the row separators for 10 cells), proving no in-cell
   pipe added a split.
2026-07-19 20:24:21 -04:00
Behruz Nassre Esfahani
d04e287fa9 fix(#2365): stop api-coverage detector false-positiving non-API phases (#2397)
* fix(#2365): stop the api-coverage detector false-positiving non-API phases

detectApiIntegration fired on any integration verb co-occurring anywhere on a
line with any API noun, treated / as a word boundary (so a first-party Next.js
src/app/api/... route path matched the noun "api"), and read any capitalized
word before API/SDK/REST/GraphQL as a service name behind a fixed stopword
denylist (so threat-model prose like "Resolver-only API" fired). Because the
verify:pre seal gate is BLOCKING, a phase touching no external API could not
reach UAT without fabricating a coverage matrix.

The compound rule now requires the verb and noun to share one clause (sentence
punctuation and table-cell walls end a clause) within a bounded word gap.
Non-prose spans are excluded before matching: fenced code (already), inline
code spans (new stripInlineCode in the markdown-sectionizer seam), and
path-shaped tokens. The <Service> API surface rule requires proper-noun
position — a clause-initial capitalized word is ordinary English and needs
dependency evidence (URL / package reference) on the same line — and rejects
compound modifiers ("Resolver-only", lowercase after the hyphen).

A phase that integrates no external API now has a first-class, reasoned way to
say so: a COVERAGE.md containing "No external API integration: <reason>"
satisfies the gate (declaration + rows is contradictory and blocks). The
true-positive path is pinned by regression tests: every default-vocabulary
positive still fires, including the widest word-gap pairing and the
surface-rule-only shape.

Fixes #2365

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2365): tighten api-coverage detector per Codex review (round 2)

Applies the Codex review findings on the initial #2365 fix:

- S-1: a COVERAGE.md "no external API integration" declaration is the human
  override for a fallible detector, so it must PASS even when detection still
  fires — but the contradiction is now SURFACED in the gate output (overridden
  signal count + terms) instead of passing silently.
- S-2: verb/noun pairing is now a term-group nearest-pair merge walk over
  precomputed word ordinals (computeWordStarts / minWordGap), not a match×match
  cross product — a hostile line repeating one pair thousands of times stays
  linear instead of going quadratic.
- FN-4: package-shaped inline-code spans (`stripe-sdk`, `@stripe/stripe-js`)
  are kept as noun/dependency evidence rather than being fully masked, so a
  genuine dependency reference inside code ticks still corroborates.
- C-1: the <Service> API surface rule now scans every candidate in every
  clause; a rejected first candidate no longer shadows a later genuine service.
- Cross-clause binding: a verb may bind a noun in the immediately following
  clause only when its own clause names a service object, within a tight gap —
  admits "Integrate Stripe, exposing its endpoints …" without re-admitting the
  unrelated-clauses false-positive class.
- Internal-descriptor negative evidence ("internal Payments API",
  "the internal endpoint") never pairs; URL/scheme matching generalized beyond
  http(s).

All 5 acceptance criteria still hold: the three reported false positives are
clean and "integrate the Stripe API" still fires. Built .cjs committed
alongside the .cts. tsc + eslint (incl. no-adhoc-markdown-parsing) +
lint:regression-names clean; affected suites 256/256 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#2365): retune api-coverage detector fail-closed per Codex review (round 3)

Codex's second-round review found the round-2 tightening had over-corrected into
FAIL-OPEN false negatives — realistic external-API prose that the BLOCKING seal
gate silently let through (the catastrophic class, since a missed API surface is
worse than a dismissable false positive). Retuned the detector to be explicitly
fail-closed: lean toward detecting, and let the one-line COVERAGE.md "no external
API integration" declaration dismiss the residual false positives.

Fail-open false negatives fixed (all now detect):
- F1 clause-initial `<Service> API` with a plain follower ("Stripe API for
  payment processing") — dropped the follower-allowlist / corroboration gate on
  clause-initial surfaces; a service that is not a stopword, descriptor, or
  compound modifier is a real name from any clause position.
- F2 scheme-less external host ("api.stripe.com/v1") — a dotted host with an
  alphabetic final label now contributes its API nouns; a first-party route
  path (no dotted host) still does not.
- F3 vendor's first-party SDK ("Integrate Shopify's first-party SDK") — the
  compound path no longer filters nouns on "internal"/"first-party" (Codex: the
  qualifier can describe the vendor's own API, not the consuming project's).
- F4 long single integration clause — removed the word-gap cap entirely: it
  could not separate a 21-word genuine clause from an 18-word internal one, so
  the clause boundary is now the whole relationship test.
- F5 lowercase cross-clause service — cross-clause binding no longer requires a
  capitalized "service object".

New false positives fixed (all now clean):
- F6 a URL token that swallowed a trailing clause comma, merging two clauses —
  trailing clause punctuation is kept literal so the split survives.
- F7 a capitalized internal component authorizing cross-clause binding — the new
  gate requires a dependent elaboration, not a new coordinate clause opened by a
  conjunction ("…, then document…").
- F8 a protocol name read as a service ("REST API", "GraphQL API") — protocol
  and locality descriptors are rejected in the `<Service>` position.

- Finding 9: the inline-code-span scanner was O(n^2) on pathological backtick
  runs; rewritten to linear via a per-length run cursor (2 MB: 4.15 s -> ~6 ms),
  semantics preserved (148 sectionizer tests unchanged).

Net simplification: the fail-closed model removed the round-2 minWordGap /
groupByTerm / follower / corroboration machinery (350 insertions vs 445
deletions across the touched files). Under fail-closed, three round-2 negative
tests now correctly detect (integration verb + "internal"-qualified noun, and
the distant-same-clause case); none were trek-e acceptance FPs.

Verified: 1491/1491 unit tests pass; tsc + eslint (incl. no-adhoc-markdown-
parsing) + lint:regression-names clean; all 8 review findings reproduced as
regression tests, both directions. Built .cjs committed alongside the .cts.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#2365): resolve round-3 Codex review findings (fail-closed, round 4)

Codex's round-3 adversarial review found the fail-closed retune had introduced
new holes in both directions. Resolved:

Fail-open false negatives (now detect):
- External host addressing a PATH ("graph.microsoft.com/v1.0/me") is itself an
  integration surface and contributes an endpoint noun even when the host names
  no vocabulary word. A bare domain link with no path ("https://example.com")
  stays a non-signal, so "Integrate … from example.com, document …" is still
  clean.
- Locality qualification ("internal", "private") no longer leaks across a
  sentence or clause boundary: only plain spaces may separate the descriptor
  from the service, so "The cache is private. Stripe API …" now detects.
- Cross-clause binding: the fragile head-word cap (which could not tell a
  genuine "Connect … to Stripe payments, exposing its endpoints" from an
  unrelated "Integrate … from URL, document …" — both 4 words after the verb)
  is replaced by a participial-continuation rule: a verb binds a noun in the
  next clause only when that clause begins with an "-ing" elaboration. This
  fixes the 4-word-head false negative AND the false positive below at once.

False positives (now clean):
- Cross-clause no longer binds a finite continuation regardless of separator:
  "Wire the settings form. Document endpoint props." / "…; document …" /
  "…, document …" are separate actions, not elaborations.

Perf (quadratic → linear):
- The trailing-punctuation peel is a backward char scan instead of an
  unanchored `[…]+$` regex (16k chars: 156 ms → ~1 ms).
- SERVICE_SURFACE_API_RE bounds the service-name length {1,40} so a hostile
  "A-A-…-x" run cannot drive O(n^2) backtracking (16k: 385 ms → ~3 ms).

Consumer fail-open (blocking gate):
- readPhaseScope now distinguishes "no plans" from a plan that EXISTS but is
  unreadable. On a read error the gate BLOCKS ("could not read the phase
  scope …") instead of silently certifying no-integration from partial scope —
  an unreadable plan could be the one describing the integration.

Documented fail-closed tradeoffs, now pinned with tests so they are not
"fixed" back into a fail-open: a clause-initial capitalized common word before
"API" ("Payment API", "Search API") reads as a service name; a long clause
pairs a verb with a distant noun; and a CommonMark inline code span that wraps
a newline is matched within-line only. Codex judged these acceptable because
the COVERAGE.md declaration is a cheap override.

One documented limitation remains out of scope: "Integrate Stripe, and
authenticate requests with its API" (a coordinate finite clause whose noun
refers back by pronoun) needs coreference resolution, beyond a lexical detector.

Verified: 379/379 affected + command-router tests pass (+14 new regression
tests covering every round-3 finding, both directions); tsc + eslint
(no-adhoc-markdown-parsing) + lint:regression-names clean. Built .cjs committed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#2365): simplify to robust core — remove whack-a-mole heuristics (round 5)

Round-4 review confirmed the detector's two most complex features generate
findings in both directions no matter how they are tuned, because they need a
vendor dictionary + coreference the issue rules out in principle. Per the
operator's "ship the robust core" decision, both are removed and their gaps are
documented rather than chased further:

- Cross-clause binding DELETED (allowsCrossClause / participle rule). It caused
  a fail-open on finite continuations ("Integrate Stripe; use its OAuth
  endpoints" — missed) and a false positive on "-ing"-SPELLED nouns ("…, billing
  endpoint terminology…" — wrongly fired). Detection is now same-clause only.
- URL-path-as-evidence REVERTED. Treating every path-bearing URL as an endpoint
  fired on ordinary asset/link URLs ("…/theme.css", "…?next=/x", a docs/repo
  link) and recreated routine UI-phase false positives. An external URL is
  evidence only when it NAMES an API vocabulary word ("api.stripe.com/v1").

Two fail-open cases are now DOCUMENTED limitations, pinned by tests so a future
maintainer does not re-add the heuristics that caused the false positives above:
a service named only in a clause separate from its API noun, and a bare external
host that names no vocabulary word. Both are cheaply covered by the COVERAGE.md
declaration and rare in real phase prose ("integrate the X API").

Also fixed from the round-4 review:
- Qualification now survives markdown emphasis ("The **internal** Payments API"
  stays clean) while still not crossing a sentence/clause boundary.
- readPhaseScope fail-closes on a REAL read failure (EACCES/EIO) enumerating the
  phase directory or reading the roadmap fallback — not only per-plan-file
  failures; a missing directory/section remains a legitimate no-op. The
  declaration-override path surfaces scope_read_error so an incomplete-scope
  override stays visible.
- SERVICE_SURFACE_API_RE length-bound comment no longer overclaims.

Net: the detector is same-clause verb+noun + `<Service> API` surface, with
path/code/inline masking and a fail-closed posture. All five acceptance criteria
hold. 1573/1573 unit tests pass; tsc + eslint (no-adhoc-markdown-parsing) +
lint:regression-names clean. Built .cjs committed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#2365): close roadmap-fallback fail-open + stale JSDoc (round-5 review)

The round-5 sanity review confirmed the detector simplification is sound (all
acceptance positives fire, all required negatives clean) and flagged one real
blocker plus a nit:

- Blocker: readPhaseScope's roadmap fallback could still silently pass an
  UNREADABLE roadmap. getRoadmapPhaseWithFallback gated on fs.existsSync(), which
  returns false on EACCES/EIO too — so an unreadable ROADMAP.md read as "absent",
  no exception reached isRealReadFailure, and the blocking gate certified empty
  scope. Fixed at the source: read the roadmap directly and honor the function's
  OWN documented contract — null only on ENOENT (genuinely absent), otherwise
  throw. Both existing callers already wrap it in try/catch expecting that throw,
  and readPhaseScope now fail-closes (blocks) via its roadmap catch. Verified by
  a new e2e test (unreadable roadmap fallback → block).

- Nit: the detectApiIntegration JSDoc still described the removed cross-clause
  participial binding and "every external hostname counts" — corrected to the
  actual same-clause-only behavior and the names-a-vocab-word URL rule.

Verified: full unit suite green; tsc + eslint + lint:regression-names clean.
Built .cjs committed (roadmap.cjs is gitignored/rebuilt, per repo convention).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#2365): backfill changeset PR number (#2397)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#2365): sync generated capability-registry + recapture install goldens

CI surfaced two generated-artifact staleness issues (all failing test shards +
lint-tests traced to these, not to a logic defect):

- gsd-core/bin/lib/capability-registry.cjs was stale: the initial fix edited the
  ai-integration `api-coverage-plan-pre.md` fragment (added the "No external API
  integration" declaration section) but did not regenerate the registry, which
  embeds an inline copy of that fragment. Regenerated via
  `gen-capability-registry.cjs --write` — the diff is exactly the fragment text
  sync. Fixes `lint:generated-sync` and the "committed registry is in sync" +
  "registry integration" tests.

- The 18 golden-install-parity fixtures were stale by exactly one hash line each
  — `gsd-core/references/api-coverage.md`, which this PR edits and which is a
  hashed installed artifact. Recaptured with `UPDATE_GOLDEN=1`; the diff is that
  single hash per runtime and nothing else. Fixes the `golden parity — *` tests.

No source or behavior change — generated artifacts only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#2365): flip representative-corpus manifest to assert the fixed behavior

The #2371 representative corpus (merged into next after this branch was cut) is a
known-bug tripwire: it asserts each fixture's currentBuggyOutput so the test
fails loudly the moment #2365 is fixed, at which point — per its own contract in
representative-corpus.test.cjs — the fixer removes currentBuggyOutput so the
assertion checks expectedDetected instead.

This is that moment. Removed currentBuggyOutput from the three detector fixtures
(nextjs-route-path, unrelated-verb-noun, threat-model-prose); the corpus now
asserts detected:false, which the fail-closed same-clause detector satisfies.
Notes updated to describe the fix rather than the bug. The #2366 matrix corpus
is left untouched — that tripwire belongs to its own PR (#2374).

Corpus test: 7/7 pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#2365): skip chmod-000 fail-closed e2e tests on Windows

The three fail-closed gate tests induce an unreadable plan / directory / roadmap
with chmod 000, but Windows does not enforce POSIX mode bits — readFileSync
still succeeds, so the gate never reaches the read-error path and the assertion
fails on the windows-latest CI leg. The fail-closed LOGIC is platform-
independent (readError → block) and is fully exercised on the macOS/Linux legs;
only the method of inducing EACCES is POSIX-specific. Guard the three tests to
skip on win32 as well as root, mirroring golden-install-parity's win32 skip.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#2365): address trek-e review — glossary, clock-seam, IO injection, bounds

Review response to PR #2397 (trek-e, CHANGES_REQUESTED). Fix logic unchanged;
this closes the test/process-hygiene findings.

Major:
- CONTEXT.md "Markdown Sectionizer" glossary now lists the two exports this fix
  relies on, `stripInlineCode` and `scanInlineCodeSpans` (glossary is a PR gate).
- Replaced the banned wall-clock assertion in the "hostile repeated-term line"
  test (Clock Seams rule — no elapsed-time asserts) with a deterministic
  signal-count assertion, which also directly verifies the term-dedup that keeps
  pairing linear (one signal for a 10k-pair line, not thousands).
- Rewrote the three fail-closed read-failure tests: instead of chmod 0o000
  (a no-op under root / on Windows, the pattern the repo's IO-failure convention
  avoids) they now exercise the newly-exported `readPhaseScope` in-process and
  inject the failure by monkeypatching fs.readFileSync/readdirSync to throw,
  restoring in finally. Deterministic and platform-independent (no skip needed),
  and they add the ENOENT-is-absence case that the chmod tests couldn't express.

Minor:
- Added limit / limit+1 boundary tests for SERVICE_SURFACE_API_RE's {1,40}
  service-name bound, QUALIFIER_LOOKBACK's 24-char window, and REASON_MAX_LEN
  (200) on the declaration reason.
- Added a fast-check property that fuzzes the tokenizer / clause splitter /
  masking (scanLineTokens, splitClauses, collectTermMatches) with adversarial
  tokens (slashes, backticks, URLs, clause punctuation) and asserts the detector
  is total (never throws), shape-stable, holds detected <=> signals, and is
  deterministic.

readPhaseScope is exported for the in-process tests. Verified: 125 detector +
19 gate tests pass; tsc + eslint + generated-sync (glossary/registry) +
lint-regression-test-names + lint-test-file-count clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-19 14:59:25 -04:00
Tom Boucher
6b196ef638 fix(#2423): finalize job syncs next package.json after final release (#2437)
* fix(release): finalize job calls sync-next-version.cjs to bump next after final release (#2423)

The release pipeline's 'finalize' job shipped X.Y.0 to npm 'latest' and
merged to main, but never bumped 'next' to match. 'scripts/sync-next-version.cjs'
exists exactly for this — its docstring promises to run 'for every release
type (rc / hotfix / final)' — but it was wired only into the 'rc' job
(release.yml:479), not 'finalize'. As a result, after 1.7.0 shipped on
2026-07-15, 'next' stayed at 1.7.0-rc.6 and every npm script banner on
'next' (and feature branches cut from it) reported the stale rc version.

Regression of #1104 — closed incomplete (covered rc only, not final).

This patch:
  - adds a 'Sync next branch to the published release' step to the
    'finalize' job, mirroring the rc job's pattern at line 479 (uses
    VERSION from inputs.version rather than PRE_VERSION from steps.prerelease,
    since finalize does not run the prerelease step);
  - gates it on !inputs.dry_run + continue-on-error:true (matches rc);
  - adds tests/release-finalize-syncs-next-version.test.cjs — a structural
    YAML assertion that fails against the pre-fix workflow and passes after.
    Failing-first demonstrated during development; the test parses job blocks
    by indentation rather than grep so it stays valid as the file grows.

Root-cause diagnosis: scripts/sync-next-version.cjs:14 docstring admits
'used by release.yml's rc job, which has no next-targeting PR of its own'.
release.yml:506-685 (finalize job) had no sync-next-version step before
this patch. Every prior rc.N release has a matching 'chore: sync next
package version to 1.7.0-rc.N' commit; there is no such commit for 1.7.0.

* chore(release): sync next package version to 1.7.0 (#2423)

Replays the canonical 'chore: sync next package version to <v>' commit
that the release pipeline's rc job auto-produces via scripts/sync-next-version.cjs,
for the 1.7.0 final release that shipped on 2026-07-15 (commit dd4c90f82
'chore: finalize v1.7.0' on main). Without this, 'next' (and every feature
branch cut from it) carried 1.7.0-rc.6 indefinitely and reported it in
every npm script banner (e.g. 'lint:ci').

Bumps 43 synchronized manifests via the npm 'version' lifecycle hook
(scripts/sync-manifest-versions.cjs --stage + scripts/gen-capability-registry.cjs
--write), matching the file set of commit 27f69cc48 ('chore: sync next
package version to 1.7.0-rc.6') and commit dd4c90f82 ('chore: finalize
v1.7.0').

This is the immediate Layer-1 repair for #2423. Layer-2 (prevent recurrence)
is the workflow patch in the previous commit; Layer-3 (regression test)
ships with it. Future X.Y.0 final releases will produce this commit
automatically once the workflow fix lands.

* test(release): tighten #2423 dry-run gate assertion to the sync step

Code review of fix/2423 found that test #3 ('gates sync-next-version on
!inputs.dry_run') asserted too loosely: it scanned the entire finalize
block for any '!inputs.dry_run' line, so it would still pass if the
gate were stripped from the sync-next-version step specifically — the
exact regression the test name promises to catch. The finalize job has
multiple steps with their own !inputs.dry_run gates (e.g. Verify
publish), so the loose version masked the very bug it claimed to detect.

Tighten by extracting the specific YAML step block containing
'scripts/sync-next-version.cjs' and asserting the gate appears within
THAT step's lines, not anywhere in the job. Verified the tightened test:

  - PASSES against the post-fix workflow (sync step has its own gate)
  - FAILS when the sync step's gate is stripped (even when other steps
    in finalize retain their own !inputs.dry_run gates) — the exact
    regression that previously slipped through

Adds extractStepBlockContaining(jobBlock, marker) helper alongside the
existing extractJobBlock(text, jobName). Reuses the same indentation-
based parsing, so it stays valid as the file grows.

* chore(changeset): backfill pr:2437 in .changeset/sturdy-ibex-jump.md

CLAUDE.md changeset convention: 'Use placeholder pr:0 during initial commit.
Backfill immediately after gh api POST /pulls returns the real number.'

PR #2437 created from branch fix/2423-release-finalize-sync-next-version.

* fix(#2423): add see #2423 to allow-test-rule exemption per ADR-456

CI lint-allow-test-rule-refs failed on PR #2437: ADR-456 requires new
allow-test-rule exemptions added after the ADR's acceptance to include a
tracking issue number in the comment, in the form
  // allow-test-rule: <reason> (see #NNN)
The exemption added in commit 976c8b0a2 lacked this ref. Fixed.

Verified locally:
  node scripts/lint-allow-test-rule-refs.cjs
    → ok lint-allow-test-rule-refs: 173 grandfathered exemption(s) tracked, no novel untracked offenders
2026-07-19 14:33:19 -04:00
0xdhx
50efae13ce fix(#2305): stage the shared guard hooks Kilo's native plugin spawns (#2327)
* fix(#2305): stage shared guard hooks for Kilo — drop skipSharedHooksInstall

Kilo's capability descriptor declared BOTH hostBehaviors.nativePlugin (a
plugin that spawns the shared PreToolUse guard scripts as subprocesses)
AND hostBehaviors.skipSharedHooksInstall:true, which suppresses staging
of hooks/*.js into the Kilo config dir. The plugin's runHook treats an
absent hook script as a silent allow, so every guard it spawned
(gsd-prompt-guard, gsd-read-guard, gsd-worktree-path-guard) no-opped on
every Kilo install. OpenCode uses the byte-identical plugin with hook
staging on and is unaffected — it is the reference shape.

The skip flag predates Kilo's plugin surface: it dates to #1821 (hooks
were dead weight for a runtime with no hook consumer), and #2093 added
the hooks-dependent nativePlugin without revisiting it.

- capabilities/kilo/capability.json: remove skipSharedHooksInstall
  (regenerated gsd-core/bin/lib/capability-registry.cjs accordingly)
- bin/install.js: correct the stale #1821 comments claiming Kilo has no
  plugin surface
- tests/kilo-upgrades.test.cjs: install-fixture tests (global + local)
  asserting the guard scripts land where the plugin's walk-up resolves
  them; an end-to-end test driving a disallowed out-of-worktree write
  through the REAL installed Kilo tree and asserting the guard rejects
  it; a cross-runtime descriptor invariant (nativePlugin and
  skipSharedHooksInstall:true must never coexist)
- tests/kilo-imperative-reference.test.cjs: flip the pinned assertion
- golden fixtures regenerated (kilo now stages the 24 hook files, same
  set as OpenCode)

Fixes #2305

* fix(#2305): warn loudly when a guard hook script is missing (runHook)

runHook's absent-file branch returned a silent exit-0 allow — the
mechanism that let #2305 ship undetected: with the hooks bundle never
staged on Kilo, every PreToolUse guard the plugin spawned resolved to
"file not found → allow" with zero signal anywhere.

Keep the adapter's design contract (a missing hook must never break the
tool call — pinned by the existing adapter test) but make the absence
loud: console.error once per hook file, naming the unresolved path and
the remediation. Applied identically to .kilo/ and .opencode/ plugin
copies (byte-parity guard). Golden parity fixtures regenerated (the
installed plugin file's hash changed).

Fixes #2305

* chore(#2305): add changeset fragment

* test(#2305): include gsd-workflow-guard.js in the staged-guards regression list

The native plugin spawns four guards on write-like tool calls — the
regression test's PLUGIN_GUARD_HOOKS list covered three. Staging itself
was already asserted via the golden fixtures (the full bundle), but the
named per-guard assertion should cover every guard the plugin actually
dispatches. Surfaced by cross-AI review of PR #2327.

* test(#2305): update the #1821 tests that encoded Kilo's false no-plugin premise

The #1821 hook-copy test asserted Kilo must receive no staged hooks — the
exact behavior this PR reverses (and the cause of all 8 CI failures). Kilo
moves from the ZCode "no dead hooks" loop to the OpenCode group, with
positive assertions on the new contract: the three guard hooks the plugin
spawns, hooks/lib/git-cmd.js, and plugins/gsd-core.js all staged. The
integration runtime contract flips kilo packageJson to true (the CommonJS
marker ships with the bundle), and the pi contract comment no longer cites
Kilo as a no-plugin runtime.

* chore(#2305): scope the queued #1821 changeset fragment to ZCode only

The fragment still claimed the installer skips hooks for Kilo — rendering
both it and this PR's fragment into the same release would ship two
contradictory statements about Kilo's install behavior. It now claims
ZCode only and notes that #2327 reverses the Kilo half.

* chore(#2305): rename changeset fragment to the generator naming convention

2305-kilo-stage-guard-hooks.md -> loud-guard-hooks.md, matching the
<adjective>-<noun>-<noun> shape npm run changeset generates (review nit).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-18 12:20:32 -04:00
Tom Boucher
ed06b6a4b9 fix(#2329): write opencode slash commands to commands/ (plural), migrate legacy command/ (#2354)
* test(#2329): fail-first tests for opencode commands/ (plural) command dir

Red phase, empirically probed: global/local install lands in command/ (singular)
with 71 gsd-*.md files and no commands/; the manifest records 71 keys under
command/ and zero under commands/; all four declaring sites report 'command'.
Migration coverage is black-box (two sequential install runs against one
configDir) so it holds regardless of how the fix implements cleanup.

The Kilo guard passes today by design — a forward-looking no-collateral check.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2329): write opencode commands to commands/ (plural), migrate legacy command/

OpenCode discovers slash commands from commands/ (plural); the installer wrote
them to command/ (singular), so none of the ~71 /gsd-* commands appeared in the
TUI. Five sites declared the directory and all had to agree:

- capabilities/opencode/capability.json: both artifactLayout destSubpath entries
  (global + local) and hostBehaviors.flatCommandDir
- bin/install.js: the manifest prefix was a SEPARATE hardcoded 'command/' literal,
  so the manifest would have diverged from the descriptor even after a rename. It
  now derives from _hostBehaviors(runtime).flatCommandDir.
- src/install-engine.cts installOpencodeFamilyArtifacts: the actual write target,
  which bypasses resolveRuntimeArtifactLayout via combinedFamilyInstall. This was
  a fifth site the issue did not list — without it the descriptor change alone
  would not have moved a single file.

Migration: an upgrade over a pre-fix install removes only manifest-proven
GSD-managed files from the legacy command/ dir and rmdirs it once empty.
Unmanifested user files are preserved, never deleted.

Kilo shares the opencode family install path and is explicitly unaffected —
pinned by a no-collateral test.

Note on the tests: the migration cases originally built their legacy fixture by
running the installer and relying on it to produce command/ — i.e. they depended
on the bug to set up the fixture, and became unsatisfiable the moment it was
fixed (block 1 requires command/ to be absent after a fresh install). They now
fabricate the legacy layout explicitly, including rewriting the manifest keys to
the command/ prefix — which is load-bearing, since the migration only removes
manifest-proven files and an unrewritten fixture would silently no-op and pass
even against a broken migration.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* test(#2329): regenerate opencode install golden after rebase onto next

The golden conflicted on rebase because #2322 also regenerated it. Resolved by
regenerating from the merged source rather than hand-merging a generated file;
the only delta is the 71 command/gsd-*.md -> commands/gsd-*.md key renames.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* test(#2329): update stale tests that pinned opencode's singular command/ dir

Seven tests encoded the old contract (opencode: command/gsd-help.md exists, the
descriptor's flatCommandDir, the install-integration contract, and the
resolveRuntimeArtifactLayout golden). They passed in the red phase precisely
because they pinned the buggy singular dir; the fix intentionally changes that
contract, so these are stale-test corrections, not regressions.

Kilo shares the opencode family install path and is deliberately NOT changing —
it stays on command/ (singular). The shared opencode/kilo test is now split via
an explicit per-runtime dir map so the two cannot be conflated, and Kilo's own
layout test is untouched. tests/opencode-command-dir-plural.test.cjs
independently pins Kilo unchanged end-to-end.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2329): changeset for opencode commands/ dir fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2329): backfill PR number 2354 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2329): correct the changeset — do not assert opencode ignores command/

The changeset repeated the issue's stated mechanism ("OpenCode discovers them
from commands/ ... a clean install produced no usable commands in the TUI at
all"). OpenCode's source contradicts that: packages/core/src/v1/config/command.ts
globs {command,commands}/**/*.md, so BOTH names resolve, and its own skill doc
still calls .opencode/command/ typical. Shipping that claim as a release note
would document a mechanism that does not exist.

The change is still right, for the stronger reason: OpenCode's config docs list
plural as the convention and singular as backwards compatibility, so GSD was
shipping on the alias the vendor may withdraw. Reworded to describe it as the
alignment it is, decided on OpenCode's source and docs rather than on bug reports
in either repo.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2329): baseline opencode's commands/ surface — closes a data-loss path this PR opened

Not a bookkeeping gap. Moving opencode's command dir to commands/ moved the
install destination to a surface the first-time baseline scan does not cover:
000-first-time-baseline's RUNTIME_SURFACES.opencode lists ['gsd-core','command',
'skills','agents'] — no 'commands'.

installOpencodeFamilyCommands unconditionally unlinks every gsd-*.md under its
destination before writing the fresh set (install-engine.cts:870-873), with zero
manifest or migration involvement. The only thing that protects a pre-existing
file is assertInstallerMigrationsUnblocked, which runs before materialization and
halts when the baseline scan flags an unknown file at a KNOWN surface.

Probed: a pre-existing commands/gsd-plan.md is silently destroyed (install exits
0). The identical file under the legacy, already-baselined command/ surface
correctly halts the install with "installer migration blocked pending user
choice". So this PR would have traded a protected surface for an unprotected one.

Fixed with a NEW fix-forward migration rather than editing 000, per
docs/installer-migrations.md:131-134 — an applied migration never re-runs, so
editing 000 would only protect fresh installs and leave every existing machine
exposed. A new id runs for both populations and drifts no shipped checksum;
adding its entry to EXPECTED_CHECKSUMS is the case that test explicitly sanctions.
All five pre-existing shipped checksums verified byte-identical.

Kilo is excluded by the migration's runtimes filter and keeps command/.

This was previously deferred as a PR-body note claiming "low impact — nothing
else acts on baseline-scan misses". That claim was never probed and was wrong.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2329): drop the parenthetical product description from the changeset

The product-name purity guard (#1777) rejects "Kilo (which still uses
command/)" — fragment prose renders verbatim into CHANGELOG.md, so a product
name must not carry a parenthetical. Reworded to a plain sentence; the meaning
is unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 08:14:59 -04:00
Tom Boucher
ff9cb6069f fix(#2285): wire claude-orchestration Workflow backend into execute-phase (#2314)
The claude-orchestration capability (#1143) shipped registered 'active'
but fully inert: detectWorkflowBackend/emitWorkflowScript had no caller
outside their own CLI router, and execute-phase.md declared an
execute:wave:pre hook point that the workflow body never rendered — so
claude_orchestration.enabled:true had zero effect on real runs.

Approach B (maintainer-chosen):
- execute-phase.md now renders the execute:wave:pre hook
  (gsd_run loop render-hooks execute:wave:pre) at a new step 2.75,
  immediately before each wave's Agent() dispatch — fixing the latent
  dead-hook gap for any pre-wave capability.
- Move the claude-orchestration contribution execute:wave:post ->
  execute:wave:pre (a pre-wave backend selector belongs before dispatch,
  not after); rename fragments/execute-wave-post.md -> execute-wave-pre.md
  with prose instructing the orchestrator to call resolve-wave-dispatch
  before step 3. Unrelated wave:post contributions (ui.safety-gate, drift,
  external-job, mempalace) untouched.
- New .cts seam resolveWaveDispatch(input) composes detectWorkflowBackend
  + emitWorkflowScript into one {backend:'inline'|'workflow', ...} result;
  exposed as gsd-tools claude-orchestration resolve-wave-dispatch. This is
  a real non-CLI-router, non-test caller of both functions.

Fail-closed: any gate miss (disabled, non-Claude runtime, Workflow tool
absent, SDK below floor, execution_backend:inline, malformed input) or an
emit failure resolves to inline with a byte-identical result shape — no
regression to the default-off execute-phase path.

Regression tests (tests/fix-2285-*) cover happy-path activation + SDK-floor
BVA, the fail-closed gate-miss table with detectWorkflowBackend parity, a
fast-check composition property, capability.json contribution assertions,
and a source-contract guard that execute:wave:pre is now actually rendered.
Dependent registry-shape assertions updated in-scope.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 19:18:28 -04:00
Tom Boucher
d8af61be44 fix(#2220): replace invalid mempalace mine --room with detect_room() staging (#2260)
* fix(#2220): replace invalid 'mine --room' with detect_room() staging approach

mempalace mine has no --room flag (only search does) — verified against
MemPalace 3.5.0 official docs (mempalaceofficial.com/reference/cli.html).
The headless capture path used --room, causing every headless/no-MCP run
to fail with 'unrecognized arguments: --room' and silently skip capture.

Fix: replace the flag with a staging-based approach that uses detect_room()'s
documented folder-path match — stage the artifact under a room-named subfolder
with a mempalace.yaml room taxonomy, then run 'mempalace mine <stage> --wing'.

Docs sources cited in-file:
- CLI reference:  https://mempalaceofficial.com/reference/cli.html
- Mining guide:   https://mempalaceofficial.com/guide/mining.html
- Config guide:   https://mempalaceofficial.com/guide/configuration.html

Changes:
- skills/gsd-mempalace-capture/SKILL.md: headless staging instructions
- commands/gsd/mempalace-capture.md: same
- capabilities/mempalace/fragments/capture-problems.md: reference staging
- .gitignore: exclude .planning/.mempalace-stage/
- tests/mempalace-capture-headless-invocation.test.cjs: regression test
- Golden install parity fixtures + workflow-size baseline regenerated

* docs: backfill changeset PR number (#2260)

* fix: regenerate golden fixtures after next merge
2026-07-14 14:52:24 -04:00
github-actions[bot]
27f69cc489 chore: sync next package version to 1.7.0-rc.6 2026-07-12 05:40:45 +00:00
Tom Boucher
a0fafedfa0 feat(#2103): drive VS Code through the Embeddable Orchestration System (ADR-1239)
VS Code is a net-new EoS runtime that — unlike every prior migration — is NOT
CLI-installed (Marketplace/VSIX extension). It has zero runtime==='vscode'
branches in bin/install.js and stays that way (regression-guarded); it is driven
entirely through the negotiated imperative Host-Integration adapter.

Registry + validator (the hard part):
- capabilities/vscode/capability.json (role:runtime): full hostIntegration block
  (imperative / palette / active vscode.lm model / engine hook bus /
  sandboxed-storage / mcp transport / sandboxed-web runtime; dispatch nested,
  maxDepth 5 per VS Code's documented subagent depth).
- capability-validator.cjs extended so a role:runtime capability can legitimately
  declare "extension-distributed, no config directory": new configHome.kind:'none'
  + installSurface:'none' (+ GATE-A pairing + the parity maps), with localConfigDir
  and configHome.name made conditional on kind!=='none'. All 18 runtimes still
  validate; getDirName returns a distinct sentinel (not '.claude') for a no-config
  runtime.
- The add-a-registry-runtime tax: NON_INSTALLABLE_RUNTIMES exemption in the
  runtime-flags drift guard, vscode added to global-config-home SPECIAL_CASED,
  EXPECTED_PROFILES.vscode='ide', and the config-adapter/derivation/pin-count
  guards updated. No golden-install fixture, model-catalog, or CONFIGURATION rows
  (vscode never enters allRuntimes).

Dispatch + extension surface:
- Fixed vscode/extension.js's createHub()-no-args bug (every dispatch was
  UnknownCommand, masked by a vacuous reachability test) — now reuses the shared
  dispatchGsdCommand subprocess-shim (Node/desktop); the reachability test is
  tightened to assert real dispatch.
- Promoted the #1933 host binding to a shipped vscode/host-binding.js; activate()
  now composes the model/hookBus/stateIO seams through it. Corrected the model
  seam to VS Code's real API (vscode.lm.selectChatModels() -> model.sendRequest();
  vscode.lm.sendRequest does not exist) so the binding actually composes on real
  desktop VS Code instead of throwing.
- New vscode/browser.js Web Extension entry with ZERO Node APIs (the engine's
  config/capability loading is Node-bound, so the web entry registers the surface
  and directs full dispatch to the native MCP server — honestly documented).
- UPGRADE 1: GSD skills as native Language Model Tools (contributes.languageModelTools
  + vscode.lm.registerTool), invoke() dispatching through the hub.
- UPGRADE 2: native subagent dispatch wired onto #runSubagent /
  chat.subagents.allowInvocationsFromSubagents (fail-soft on API availability,
  maxDepth 5 enforced).
- vscode/package.json: browser entry, engines.vscode ^1.105, chatParticipants +
  languageModelTools contributions; fixed a stale activationPoints->activationEvents
  manifest key. Added "vscode" to the package files array.

Docs (## vscode matrix section) + changeset (Added).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 23:24:42 -04:00
Tom Boucher
79d7657eff feat(#2102): make pi a first-class installable runtime + fix its dispatch (ADR-1239)
Net-new EoS/pi installable runtime — purely additive (no prior runtime==='pi'
branches). pi is a bun-runtime programmatic-CLI whose /gsd command is registered
by a native ExtensionAPI extension and dispatches through the embedded engine.

Stage 1 (install plumbing):
- capabilities/pi/capability.json: full hostIntegration descriptor (imperative /
  slash-programmatic / active-model / native-extension / bun) + hostBehaviors
  {nativePlugin, pluginOnlyInstall}.
- --pi flag + interactive-menu renumber (All 17->18); pi added to RUNTIME_FLAG_IDS,
  RUNTIME_LABELS, RUNTIME_META, allRuntimes/runtimeMap, model-catalog defaults.
- Install mirrors OpenCode: pi installs the gsd.cjs extension + the shared engine
  payload (gsd-core + scripts + config markers) + the shared hooks bundle (spawned
  by the extension at lifecycle events, like OpenCode's plugin). pluginOnlyInstall
  EXCLUDES declarative command/agent/skill markdown, which pi has no host-read
  surface for (its /gsd is programmatic). _installNativePluginIfDeclared (extracted
  from the opencode-family path) copies pi/gsd.cjs -> ~/.pi/agent/extensions/gsd.cjs
  (global) / .pi/extensions/ (local). pi added to package.json files.
- Golden: new pi.json (320 files: extension + engine + 27-file hooks bundle, no
  markdown); the 16 other fixtures + claude-local change only by the shared
  model-catalog hash line.

Stage 2 (real dispatch + upgrades):
- Shared dispatchGsdCommand() (shell-command-projection): bounded, no-throw
  subprocess-shim to gsd-tools.cjs (the only full-surface dispatch path; no
  in-process full-hub factory exists). Fixes pi/gsd.cjs's createHub()-no-args bug
  (every dispatch was UnknownCommand) AND the identical bug in mcp-server.cts's
  gsd_invoke_command, which a vacuous unknown-family-only test had masked (now has
  a real dispatch regression test).
- pi/gsd.cjs: /gsd handler now (args, ctx) - tokenizes (quote-aware, via the
  shipped hooks/lib/git-cmd.js) + dispatches real family/subcommand (not hardcoded
  query/help); gsd_invoke gets a TypeBox (JSON-schema-fallback) parameters schema +
  consumes params; getArgumentCompletions; before_provider_request active-model
  steering (fail-open on null resolution); functional session_start /
  before_agent_start / session_before_compact hook bridges (spawn the shipped GSD
  hook scripts).
- EXTENSION_EVENT_SURFACES.pi expanded from ['tool_call'] to the full 30-event
  vocabulary.

Docs (host-integration matrix + how-to) + changeset (Added).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 21:07:38 -04:00
Tom Boucher
a2c7b879d0 feat(#2101): dogfood ZCode through the EoS declarative adapter + fold shared-hooks exclusion (ADR-1239)
The issue's "0 conditional branches" premise missed a live one: the `!isZcode`
shared-hooks exclusion (bin/install.js). Fold it onto descriptor-driven
hostBehaviors.skipSharedHooksInstall:true (zcode's golden has zero hook files —
byte-parity verified) and drop the now-unused isZcode destructure. Zero live
runtime==='zcode'/isZcode branches remain (AC2 source-grep guard over
bin/install.js + install-engine.cts + surface.cts + runtime-artifact-conversion.cts).
The 6 CLI-bookkeeping zcode mentions (--zcode flag, menu, roster, help) stay.

Reference test (declarative-reference-zcode.test.cjs): profileOf → declarative-cli;
createDeclarativeAdapter({runtime:'zcode'}).kind → declarative; a real install emits
the invocable nested-skills/commands/agents surface (no hooks); negotiateHostCapabilities
fail-closes (empty/corrupt descriptor; the nested/maxDepth undocumented sub-axes degrade
to most-restrictive); validateCapability clean.

UPGRADES documented as BLOCKED (verified doc gaps — NOT guessed, to avoid a
non-functional false-green): both of ZCode's documented capabilities lack a published
on-disk config format.
- Hook automation: zcode.z.ai/en/docs/plugin documents the Hook component only as
  "automation hooks triggered on specific events" (capability detected from directory
  layout) — no config file format/location/event schema. Cannot faithfully wire.
- MCP registration: zcode.z.ai/en/docs/mcp-services says servers are "stored in the
  .zcode configuration file" (UI-only) with no documented on-disk filename/path/schema —
  exactly the settings-filename gap the issue AC anticipated.
Both are documented (with the search trail) in the capability matrix + how-to, per AC4's
block-documentation clause; hookBus/transport stay declared for when ZCode publishes the
formats. No hook scripts or MCP artifacts added → no golden change, no other-runtime impact.

Golden: byte-identical for all 16 runtimes (the fold is byte-parity; no upgrade artifacts).
Matrix ## zcode EoS note + how-to; changeset (Changed). capability-registry regenerated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 17:14:48 -04:00
Tom Boucher
bd613566cb feat(#2100): drive Windsurf through the EoS descriptor + wire Cascade's blocking hook bus (ADR-1239)
Fold all 10 residual isWindsurf branches in bin/install.js onto descriptor-driven
hostBehaviors (byte-parity — no fold changes any install output):
- 2 dead destructures dropped (uninstall, finishInstall); the dead
  `else if (isWindsurf)` legacy agent-loop arm removed (windsurf ∈
  _DESCRIPTOR_AGENTS_RUNTIMES → unreachable).
- skipSharedHooksInstall:true folds the two `!isWindsurf` shared-hooks exclusions.
- legacyDevinSkillsCleanup:true folds the `.devin`→`.windsurf` one-time cleanup gate.
- installsCommandBodiesForWorkflowDelegation:true folds the #1629 command-body copy
  (workflow-delegation target — load-bearing; local-install verified intact).
- verificationStyle:"windsurf-workflows" folds the workflow-count report.
- Corrected stale _LEGACY_SCAN_SUBDIR_NAMES + hooks-json manifest comments (cursor + windsurf).
Zero live runtime==='windsurf'/isWindsurf branches remain across bin/install.js,
install-engine.cts, surface.cts, runtime-artifact-conversion.cts (AC2 guard scans all four).

UPGRADE (Cascade hook bus): wire GSD's write/command safety guards into Windsurf's
native hook bus. New hooksSurface 'windsurf-hooks-json' (VALID_HOOKS_SURFACES 7→8, GATE A
profile-marker-only allowlist, the HooksSurface union) + writeWindsurfHooksJson
(Cursor-templated, Cascade's flat {hooks:{<event>:[{command}]}} shape) writing
.windsurf/hooks.json with two BLOCKING pre-hooks:
- pre_write_code → gsd-windsurf-pre-write.js: blocks writes to a file outside the
  active git worktree / into .git internals.
- pre_run_command → gsd-windsurf-pre-command.js: conservative destructive-command
  deny-list (rm -rf of root/home incl. sudo/env/path-prefixed forms; fork bombs;
  force-push refspec forms — HEAD:main, +main, --force/-f — to main/master/next).
Both use Cascade's protocol (stdin JSON, exit 2 + stderr to block, exit 0 to allow,
fail-open on error/timeout). Tokenize-based classifier (no catastrophic-backtracking regex;
4096-char cap) with the fail-closed false-positives fixed post-review.

The 4 advisory GSD guards + pre_mcp_tool_use + 5 post_* logging events are deliberately
NOT wired: Cascade has no context-injection channel for advisory hooks and GSD has no MCP
guard — porting them would be non-functional padding (documented; codebuddy #2098 / copilot
#2099 faithful-subset precedent). extendedHookEvents stays [].

Golden: the 2 guard scripts ship in the shared hook bundle (HOOKS_TO_COPY + the shared
managed-hooks-registry), exactly like cursor's 6 gsd-cursor-*.js scripts — so the 8
shared-bundle runtimes' fixtures gain the 2 inert windsurf scripts + the registry hash
(functionally inert for non-windsurf; the established cursor pattern). No install-output
change beyond that (the folds are byte-parity; skip-bundle runtimes untouched). New scripts
registered in managed-hooks-registry + build-hooks + INVENTORY. Tests: declarative-reference-
windsurf (adapter/axes/fail-closed + AC2 guard) + windsurf-hooks-bridge (live exit-2 blocking
+ allow/fail-open + ReDoS-bound + writer/reconcile/remove idempotency); VALID_HOOKS_SURFACES
pin updated to 8. Matrix hookBus delta + changeset (Changed). capability-registry regenerated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 16:04:24 -04:00
Tom Boucher
d1e9491fef feat(#2099): drive GitHub Copilot through the EoS descriptor + multi-event hook bus (ADR-1239)
Fold Copilot's residual runtime-literal branches onto descriptor-driven
hostBehaviors. Several issue premises were inaccurate (verified via research)
and deliberately NOT followed: writesSharedSettings/"legacy exclusion list"
(per-runtime descriptor data, copilot's false is correct); RUNTIME_CONTENT_DISPATCH.copilot
+ installSurface==='copilot-instructions' (already descriptor-driven);
extendedHookEvents (closed Claude/Gemini enum — hooks extended in code instead);
reapply ternary (already folded, kimi #2095).

Real folds (all byte-parity — golden byte-identical for every runtime):
- src/install-engine.cts + src/surface.cts: the two `.agent.md` filename cutovers
  (_copyStaged + _syncGsdDir) unified onto hostBehaviors.agentFileExtension via a
  new exported agentFileExtensionFor() accessor (kills the two-mechanism divergence).
- src/runtime-artifact-conversion.cts: applyAgentPathRewrites' copilot skip →
  hostBehaviors.noPathRewrite:true (antigravity #2096 precedent).
- bin/install.js uninstall: the two isCopilot cleanup branches → installSurface===
  'copilot-instructions' gate (symmetric with install-time).
- bin/install.js: `!isCopilot` in the two skipSharedHooksInstall checks →
  hostBehaviors.skipSharedHooksInstall:true (copilot has no shared gsd-*.js hooks).
- bin/install.js: three dead legacy inline-agent-loop isCopilot refs removed
  (copilot ∈ _DESCRIPTOR_AGENTS_RUNTIMES → unreachable; byte-parity proven by clean
  golden + real reachable-runtime install diffs). isCopilot dropped from 4 destructures.
Zero live `runtime==='copilot'`/`isCopilot` branches remain in bin/install.js,
install-engine.cts, surface.cts, or runtime-artifact-conversion.cts (AC2 guard scans all four).

UPGRADE 1 (multi-event hook bus): buildCopilotHookConfig() now emits preToolUse/
postToolUse/userPromptSubmitted/sessionEnd advisory handlers alongside sessionStart
(static inline bash/powershell — deterministic, golden-trackable). Only copilot.json's
gsd-session.json hash changes.

UPGRADE 2 (background dispatch): surfaced via the negotiated contract only —
dispatch.background:true exceeds the declarative-cli baseline and survives negotiation
with no downgrade warning. NO .agent.md frontmatter field (copilot has none). MCP
companion out of scope (AC4 names only 2 upgrades).

Tests: declarative-reference-copilot (adapter/axes/fail-closed + AC2 4-file source-grep
guard) + copilot-upgrades (live 5-event hook wiring; dispatch.background negotiation).
Matrix EoS note + how-to; changeset (Changed). capability-registry regenerated.

Incidental flaky-test RE-ARCHITECTURE (no-defer, maintainer-directed):
tests/opencode-review-reconstruction.property.test.cjs spawned ~600 synchronous
execFileSync('jq') subprocesses (numRuns:200 × 3 fast-check properties, one jq per
generated stream); a single jq freezing on a contended macos-22 CI runner hung the whole
unit-test chunk to its 600s kill (this PR's CI). --test-force-exit can't interrupt a
synchronous execFileSync, so the cure is to stop spawning per case, not just time-bound
it. Re-architected to run the SHIPPED jq program over the whole fast-check corpus in ONE
jq process: each generated stream is one compact-JSON array per line in a temp file,
`jq -c <PROGRAM>` (no -s) applies PROGRAM to each array (`.` == the array, exactly what
production's `jq -rs <file>` sees after slurping) and emits one result per line —
empirically byte-identical to the per-stream form across embedded-newline/empty/quote/
unicode/null-drop cases, and file-input (like production) so there's no stdin pipe to
deadlock on large I/O. ~600 spawns → 6; coverage unchanged (200-case corpus per property,
deterministic seeds) plus explicit boundary/diagnostic example batches. Still property-
tests the real shipped jq (no JS reimplementation). Per-call jq timeout retained as a
belt-and-suspenders bound.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 13:10:31 -04:00
Tom Boucher
9e1fad1503 feat(#2098): drive CodeBuddy through the EoS descriptor + wire subagent hooks (ADR-1239)
Fold CodeBuddy's residual runtime-literal branches onto descriptor-driven
hostBehaviors (the issue's "zero branches remain" premise was inaccurate — 2
lived):
- Folded the `if (isCodebuddy)` commands/ report duplicate (byte-identical to
  the generic reportCommandsDir block above it) onto
  hostBehaviors.reportCommandsDir:true, matching cursor's #2089 precedent.
- Deleted the dead `else if (isCodebuddy)` agent-conversion arm (codebuddy ∈
  _DESCRIPTOR_AGENTS_RUNTIMES → gated out before the legacy chain, same dead-arm
  pattern removed for augment/trae in #2097).
- Removed isCodebuddy from all 4 destructure sites (now comments only, mirroring
  the #2096 isAntigravity fold). Zero live isCodebuddy reads remain.

UPGRADE 1 (extended hook bus): populate extendedHookEvents with the full
extended set SubagentStop/Stop/PreCompact/SubagentStart (qwen #2092 / kimi
precedent — codebuddy previously had extendedHookEvents:[] so it got NONE of
these). The generic applySettingsJsonHooks loop wires them, so CodeBuddy's
settings.json now gains all four subagent-lifecycle + stop/compact hooks it
previously lacked, matching Qwen/Kimi coverage. No source change (the
HOOK_EVENT_SURFACES SDK catalog is a locked dict; all 4 events are already in
the validator enum; hookEvents already 'claude'). settings.json is golden-
excluded → no golden change.

UPGRADE 2 (background dispatch): surfaced via the negotiated capability contract
only. CodeBuddy's dispatch.background:true legitimately exceeds the declarative-
cli baseline (false) and survives negotiation with no warning. NO agent-file
frontmatter field is emitted: the CodeBuddy CLI (GSD's install target,
~/.codebuddy/agents/) has NO background-dispatch frontmatter field — background
is a caller-side run_in_background invocation param (verified against
codebuddy.ai/docs/cli/sub-agents); the issue's agentMode/enabledAutoRun are
IDE-only (codebuddy.cn, a different product). Emitting them would be a non-
functional false-green, so it is deliberately not done.

Golden: byte-identical for all 16 runtimes (folds are console/failures-report
only; UPGRADE 1 → golden-excluded settings.json; UPGRADE 2 → no artifact
change). Tests: declarative-reference-codebuddy (adapter/axes/fail-closed
negotiation + AC2 source-grep guard) + codebuddy-upgrades (live install asserts
all 4 extended hooks wired; dispatch.background survives negotiation above
baseline). Matrix + install-on-your-runtime updated; changeset (Changed).
capability-registry regenerated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 09:39:34 -04:00
Tom Boucher
17aa35fba7 feat(#2097): migrate Augment onto EoS declarative adapter + settings.json MCP companion (ADR-1239)
Fold augment's runtime-literal conversion branches onto descriptor-driven
hostBehaviors and delete dead code:
- Site A (_applyRuntimeRewrites case 'augment'): the 4 ~/.augment dot-dir
  regexes now derive from getDirName('augment') via escapeRegExp (byte-
  identical; getDirName('augment')==='.augment') — no runtime literal.
- Site B (applyRuntimeContentRewritesForCommandsInPlace): the
  `if (runtime==='augment')` markdown-converter branch now reads
  runtime.hostBehaviors.commandBodyConverter and dispatches through a local
  COMMAND_BODY_CONVERTERS map (degrade-closed on unknown/absent name).
- Deleted dead `claudeToAugmentTools` map (zero refs; orphaned by ADR-1508
  single-sourcing) and the unreachable `else if (isAugment)` agent-conversion
  branch (augment ∈ _DESCRIPTOR_AGENTS_RUNTIMES → gated out upstream).
- Incidental orphan cleanup (no-defer): removed the equally-unreachable
  `else if (isTrae)` agent-conversion arm left behind by trae's already-merged
  migration #2094 (trae ∈ _DESCRIPTOR_AGENTS_RUNTIMES, same upstream gate).
  copilot/windsurf/codebuddy arms are removed by their own pending migrations.

UPGRADE 3 (transport:mcp): register the GSD companion MCP server in Augment's
settings.json under mcpServers.gsd (Augment hosts MCP in settings.json, not a
standalone file). mergeGsdMcpServerIntoSettings mutates the in-memory settings
object finishInstall already writes (gated on hostBehaviors.mcpCompanion===
'settings-json'); non-destructive + idempotent; symmetric uninstall removal.
settings.json is golden-excluded, so no golden change. UPGRADE 1 (named/
background dispatch) + UPGRADE 2 (settings-json hook bus, Claude dialect)
were already live in production — this adds tests exercising both.

Golden: byte-identical for all 16 runtimes (folds preserve regex behavior;
MCP lives in golden-excluded settings.json) — verified by a real double-install
tree diff. Tests: declarative-reference-augment (adapter/axes/fail-closed/
undocumented-sub-axes + source-grep guard scoped to conversion-logic branches)
+ augment-upgrades (dispatch negotiation, hook-bus live install, MCP add/
idempotent/preserve/uninstall). Matrix + connect-gsd-mcp-server + a stale
install-on-your-runtime hook-ownership claim corrected; changeset (Changed).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 04:47:56 -04:00
Tom Boucher
5695522d5f feat(#2096): migrate Antigravity onto EoS declarative adapter + permission-writer + MCP companion (ADR-1239)
Fold all antigravity literal branches into descriptor-driven reads:
getConfigDirFromHome (→ configHome.kind 'dot-home-nested'), projectLocalHookPrefix
(→ hostBehaviors.hookPathStyle 'raw'), applyAgentPathRewrites (→ noPathRewrite),
getProjectInstructionFile (→ projectInstructionFile 'GEMINI.md'); removed the dead
inline convertClaudeAgentToAntigravityAgent branch + dead isAntigravity
destructures (antigravity is already on the descriptor-agents path). subagentToolkit
flipped undocumented→full (Context7: antigravity.google/docs/cli/features);
namedDispatch/nested/maxDepth/backgroundDispatch stay undocumented. Byte-identical
golden parity for all 16 runtimes.

UPGRADE 1 (permission-writer): permissionWriter 'antigravity' + configureAntigravityPermissions
merges a scoped permissions.allow block (GSD's own tree + hooks) into Antigravity's
settings.json — non-destructive, idempotent, symmetric uninstall. Added to
VALID_PERMISSION_WRITERS + the FinishPermissionWriter union.
UPGRADE 2 (MCP companion): configureAntigravityMcpConfig writes mcp_config.json
registering the gsd-core companion MCP server (Gemini-successor mcpServers schema,
best-effort — raw schema unpublished). Both writers dispatch from finishInstall.
settings.json is golden-excluded (HOOK_CONFIG_FILES); mcp_config.json (portable,
no absolute paths) is golden-tracked → only antigravity.json changes.

Tests: declarative-reference-antigravity extended (source-grep guard across 4
modules, fail-closed for the 4 undocumented sub-axes, validator acceptance) +
antigravity-upgrades (permission-writer + mcp_config live-install, idempotency,
user-preservation). Matrix + ADR-1016 + capability-manifest + CONTEXT.md +
connect-gsd-mcp-server docs updated; changeset (Changed).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 03:07:01 -04:00
Tom Boucher
ab04916682 feat(#2095): migrate Kimi CLI onto EoS imperative adapter + native hook-bus + background dispatch (ADR-1239)
Fold all runtime==='kimi'/isKimi logic branches into descriptor-driven
hostBehaviors (localInstallDeferred, verificationStyle, agentManifestStyle,
reapplyCommand, doneBannerStyle) + add 'kimi' to _DESCRIPTOR_AGENTS_RUNTIMES.
Kimi's skills/kimi-agents dispatch was already descriptor-driven (converter-by-
name + kimi-agents kind). Zero isKimi/runtime==='kimi' branches remain.

UPGRADE 1 (native hook bus): new hooksSurface 'kimi-hooks-toml' + a marker-
delimited config.toml [[hooks]] emitter (buildKimiHooksTomlBlock/writeKimiHooksToml
in runtime-hooks-surface.cts; resolveKimiHooksTomlDir in runtime-homes.cts).
GSD's lifecycle hooks now wire into Kimi's native ~/.kimi/config.toml (Context7-
confirmed path) at SessionStart/PreToolUse/Stop/PreCompact/SubagentStart/
SubagentStop — kimi becomes a hooks/ consumer (the 3 && !isKimi exclusion guards
removed). config.toml holds absolute install paths so it's golden-excluded via
an exact relative-path (.kimi/config.toml), not a basename (which would blind
Codex's config.toml). New hooksSurface value added to the closed enum in
capability-validator + runtime-config-adapter-registry.
UPGRADE 2 (background dispatch): flip dispatch.backgroundDispatch true (Kimi's
Agent tool takes run_in_background; root agent already gets the Agent tool), so
negotiation no longer flattens dispatch. subagentToolkit stays 'undocumented'
per AC (coder/explore/plan have distinct tool policies).
MCP transport explicitly deferred (no installer-driven MCP for any runtime).

Golden: only kimi.json changes (hooks/ scripts now installed); all 15 others +
claude-local byte-identical (kilo/zcode keep their own exclusions). Tests:
kimi-imperative-reference (adapter/axes/fail-closed/hostBehaviors + source-grep
guard) + kimi-upgrades (config.toml [[hooks]] SessionStart + marker idempotency
+ backgroundDispatch negotiation). CONTEXT.md glossary + matrix + how-to updated;
changeset (Added).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 00:53:34 -04:00
Tom Boucher
f744635b3d feat(#2094): migrate Trae onto EoS imperative adapter + SOLO stage-metadata upgrade (ADR-1239)
Fold trae logic branches into descriptor-driven reads: skipSharedHooksInstall
gates (dropped && !isTrae), and the case 'trae' path-rewrite arm now computes
the self-alias from the descriptor-driven dirName (.trae). Dead isTrae bindings
removed from uninstall/writeManifest/finishInstall. trae's skills dispatch was
already descriptor-driven (converter-by-name). RUNTIME_CONTENT_DISPATCH.trae is
left as a runtime-keyed table registration (its regex/callback rewrites can't be
a byte-identical descriptor map — matches cursor/windsurf/cline). trae stays in
RUNTIME_FLAG_IDS: isTrae still gates the agents-converter selection (agents out
of scope; removal gated on the cross-runtime agents-dispatch migration).
Byte-identical golden parity for all 16 runtimes.

UPGRADE: SOLO stage/trigger metadata — emitted Trae SKILL.md now carries
stage: workflow (descriptor-gated via hostBehaviors.soloStageMetadata) so
Trae's SOLO Agent can auto-invoke GSD skills at the corresponding stage. Field
shape is best-effort/inferred (Trae publishes no formal schema). trae.json
golden regenerated.

Tests: trae-imperative-reference (adapter/axes/fail-closed shouldFlattenDispatch
+ no runtime==='trae' source-grep, isTrae exempted for agents) + trae-upgrades
(stage: workflow on installed SKILL.md, descriptor-gated). Matrix note +
changeset added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 21:31:10 -04:00
Tom Boucher
f014ec83bd feat(#2093): migrate Kilo onto EoS imperative adapter + hook-bus/model/MCP/dispatch upgrades (ADR-1239)
Fold remaining isKilo logic branches into descriptor-driven reads:
finishPermissionWriter (uninstall cleanup), skipSharedHooksInstall (hooks
copy), and a skills converter-name registry (the artifactLayout.converter
field is now load-bearing, not decorative). frontmatterDialect stays the
documented dispatch key for frontmatter (no descriptor field for it). Dead
isKilo destructure bindings removed. Byte-identical golden parity for all 16
runtimes (opencode, which shares kilo's combined-family path, verified clean).

UPGRADE 1 (hook bus): install .kilo/plugins/gsd-core.js native plugin +
extensionEvents:"kilo" + EXTENSION_EVENT_SURFACES.kilo (OpenCode-fork bus).
UPGRADE 2 (active model): populate runtimeTierDefaults.kilo + thread
modelOverride through convertClaudeToKiloFrontmatter — model no longer stripped
from agents. UPGRADE 3 (MCP): document the gsd-core MCP companion under kilo's
mcp config key. UPGRADE 4 (named dispatch): agents/*.md mode:subagent roster is
the Task-tool dispatch surface (tested); subagentToolkit stays 'undocumented'
per AC so dispatch degrades to 'degraded' by design.

Model-catalog single-source edit ripples the shared model-catalog.json hash
into all 16 golden fixtures (expected). Inline defect fixes (no-defer): stale-
bake-guard resolveAgentDir 'agent'->'agents' (was a silent no-op for opencode/
codex), hardcoded 'Removed OpenCode plugin' uninstall log -> generic, and the
connect-gsd-mcp-server.md OpenCode mcpServers->mcp doc error.

Tests: kilo-imperative-reference (adapter/axes/fail-closed/degradation/
hostBehaviors + widened isKilo source-grep across 4 modules) + kilo-upgrades
(plugin parity+load, model-override converter, agents dispatch surface, MCP
doc). Matrix + how-to + config docs updated; changeset added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 19:55:15 -04:00
Tom Boucher
c6ce110efa feat(#2092): migrate Qwen Code onto EoS imperative adapter + native subagents + SubagentStart (ADR-1239)
Fold all runtime==='qwen'/isQwen logic branches (skill-priority frontmatter,
branding/path rewrites, legacy commands/gsd cleanup, hyphen-namespace
normalization, RUNTIME_CONTENT_DISPATCH, hooks-surface label) into
descriptor-driven runtime.hostBehaviors on capabilities/qwen/capability.json,
read via _hostBehaviors(). Shared claude/qwen/hermes legacy-migration branches
in install-engine.cts folded to descriptor flags (claude+hermes descriptors
updated; FALLBACK_HOST_BEHAVIORS.claude floored). Byte-identical golden parity
for qwen/hermes/claude(global+local).

UPGRADE 1: native .qwen/agents/*.md subagent projection — new agents
artifact-layout kind + convertClaudeAgentToQwenAgent converter (name +
description + tools YAML block list; color/model dropped). qwen routed onto
the descriptor-driven agents path (_DESCRIPTOR_AGENTS_RUNTIMES).
UPGRADE 2: SubagentStart hook wired into extendedHookEvents + the
descriptor-gated hook-writer loop (activates only for qwen).

Tests: qwen-imperative-reference (adapter/axes/fail-closed/hostBehaviors +
no runtime==='qwen' source-grep across 4 files) + qwen-upgrades (agents file
validity + SubagentStart mirrors SubagentStop, descriptor-gated). Docs matrix
+ how-to updated; changeset added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 15:44:19 -04:00
github-actions[bot]
451e58e6d5 chore: sync next package version to 1.7.0-rc.5 2026-07-10 17:09:38 +00:00
Tom Boucher
f152e9a069 feat(#2091): migrate hermes onto EoS imperative adapter + extensionEvents dialect (ADR-1239)
- Fold 7 hardcoded isHermes/runtime === 'hermes' branches in bin/install.js into
  descriptor-driven _hostBehaviors lookups (skillFrontmatterVersion,
  skillsManifestPrefix, trackCategoryDescription, writeCategoryDescription,
  reportSkillsCount, legacyCommandsGsdCleanup, brandingRewrites)
- Add runtime.hostBehaviors block to capabilities/hermes/capability.json
- Register EXTENSION_EVENT_SURFACES.hermes (13 real plugin hook events) in
  src/host-integration.cts — replaces the borrowed hookEvents:'claude' 6-event
  surface that silently never fired on Hermes
- Add extensionEvents:'hermes' to the descriptor
- Add 'hermes' to VALID_EXTENSION_EVENTS in capability-validator.cjs
- Regenerate capability-registry.cjs
- Tests: hermes-imperative-reference (negotiation, axes, fail-closed, source-guard),
  hermes-dispatch-upgrade (dispatch posture, degradation, fail-closed)
- Changeset + docs update
2026-07-09 20:47:35 -04:00