Commit Graph

33 Commits

Author SHA1 Message Date
Tom Boucher
653f95e39f chore(#2801): remove the hostBehaviors.reviewerCli deprecated alias (#3272)
* test(#2801): failing-first suite for the hostBehaviors.reviewerCli alias removal

Inverts the Phase 5a rows that assert the derived legacy alias still
contributes a reviewer slug, and adds the removal-warning coverage the
alias's exit needs (ADR-2782 D9).

RED against unmodified production code, by design: the six shipped
manifests still declare the key and collectReviewerWarnings emits nothing
for hostBehaviors.

Refs #2801

* chore(#2801): remove the hostBehaviors.reviewerCli deprecated alias

ADR-2782 D9, Phase 7 — the final phase of epic #2782.

The derived legacy alias survived one release (Phase 5a shipped in 1.9.0;
1.9.1 and 1.10.0 have since gone out), so it goes. A declared reviewer
body is now the only route onto the reviewer roster.

- deriveReviewerSlugs no longer reads runtime.hostBehaviors.reviewerCli
- the key is stripped from the six manifests that carried it; each already
  declares a reviewer body whose slug equals its capability id, so the
  derived roster is unchanged at the same twelve slugs
- collectReviewerWarnings emits a presence-based, non-fatal removal notice
  for any manifest still declaring the key, reaching both the build-time
  registry generation and the third-party overlay load path. The check runs
  before the reviewer-body early-return, because the manifest it exists for
  is the alias-only one that has no body.
- hostBehaviors stays an open, unvalidated bag for its other 59 keys; this
  adds one keyed removal notice, not general validation

Refs #2801

* refactor(#2801): give the reviewer-warning channel a typed IR

Review finding: the new tests asserted with String#includes() on the
warning prose, which CONTRIBUTING.md's 'Prohibited: Raw Text Matching on
Test Outputs' bans in favor of a typed intermediate representation.

Adds the IR beside the renderer rather than replacing it, which is the
shape that section prescribes and bin/verify-reapply-patches.cjs already
models:

- REVIEWER_WARNING, a frozen code enum
- REMOVED_REVIEWER_CLI_FIELD, so the emitting site and its test share one
  symbol instead of duplicating a literal
- collectReviewerWarningRecords(cap), returning typed records

collectReviewerWarnings(cap) keeps its exact string[] contract as a thin
map over the records, so both production consumers are untouched. Every
section-K row now asserts on record.code/field/capId and none on the
rendered message. Locks the code surface, asserts the renderer stays
one-to-one with the records, and migrates the pre-existing Phase 2 test
on the same channel off prose matching.

Refs #2801

* test(#2801): invert the section F alias fall-through regression row

Caught by the remote runner: 2 unique failures on both Node lanes out of
31,692. tests/reviewer-lane-declarations.test.cjs section F — Phase 5a's
isolated-security-review regressions — asserted that a blank reviewer.slug
falls through to the hostBehaviors.reviewerCli alias rather than dropping
the lane. That is the direct inverse of this phase's contract.

The original rationale held only while the alias existed. With it gone
there is nothing to fall through to: a blank body is not a declaration,
and a declaration is the only route onto the roster.

Inverted rather than deleted — the row carries the adversarial-review
provenance for the slug trim, and removing a security regression guard to
make a change pass is backwards. The duplicate row added earlier in
section C is dropped instead; section F is its canonical home.

Also corrects two count strings Phase 5b left at eleven while asserting
twelve, which would misreport on failure.

Refs #2801

* docs(#2801): give the removed reviewerCli flag a migration path

The Reference edit alone satisfied CI — a file under docs/ moved, so
lint-docs-required.cjs was green — while the task-oriented quadrant said
nothing about the removal. A maintainer whose lane had just gone silent
would have found the field documented as removed and no page telling them
what to do about it.

Adds a migration section to the how-to: the symptom, the verbatim warning
they will see, the before/after manifest, and the note to keep the
reviewer slug equal to the capability id so existing
review.default_reviewers entries and --<slug> flags survive.

Refs #2801

* chore(#2801): backfill changeset pr number to 3272

* feat(#2801): close the runtime.hostBehaviors vocabulary

ADR-1016 closes twelve descriptor axes and rejects an open escape hatch
in the descriptor. It never mentioned runtime.hostBehaviors, and that
silence was read as permission: 59 keys across 18 manifests, 39 of them
set by a single capability, validated by nothing. The reference docs went
further and attributed the open seam to ADR-1016, which does not mention
the field at all.

KNOWN_HOST_BEHAVIORS enumerates the vocabulary. An undeclared key yields a
non-fatal UNKNOWN_HOST_BEHAVIOR record on the same D4.3 channel as the
alias removal notice, reaching both build-time generation and overlay
install.

Warning, never error, for the reason this phase exists: an error would
hard-break an out-of-tree descriptor carrying a bespoke key with no
deprecation window, which is what reviewerCli was given a release to
avoid. Escalation is a separate decision.

reviewerCli is excluded from the unknown-key sweep so it keeps its own
notice with the migration pointer rather than drawing two records.

A parity test binds the vocabulary to the shipped manifests in both
directions, and a second asserts no shipped capability draws a notice, so
the closure is provably inert in-tree.

Records the decision and the miscitation as an ADR-1016 amendment.

Refs #2801

* fix(#2801): bound and sanitize the unknown-key diagnostics

Two findings from an isolated adversarial review of the closure commit,
both proven by execution rather than asserted.

MAJOR, introduced by the closure: the new Object.keys(hostBehaviors) sweep
had no ceiling. An installed third-party manifest is bounded only by
MANIFEST_MAX_BYTES, and an 8.69MB manifest with 800,000 keys produced
800,000 records and ~139MB of message text, retained for the registry's
lifetime in OverlayMeta.diagnostics. Now capped at ten records plus a
summary carrying omittedCount, mirroring capability-loader's existing
slice(0,3) idiom. The same manifest now yields 11 records and 1748 chars.

MINOR, newly reachable: manifest-supplied key names were interpolated raw.
Unlike cap.id, which validateCapability gates on KEBAB_RE before these
diagnostics run, hostBehaviors keys have no grammar check anywhere, so
ANSI escapes and CRLF reached stderr and OverlayMeta.warnings intact. New
describeKey replaces C0/C1 controls and clips at 80 chars. The file
already had describeValue for this and applied it only to values.

Both fixes land on the pre-existing reviewer.* sweep too — it carried the
identical pair, and fixing only the new copy would leave the same defect
one screen from its own fix.

Refs #2801

---------

Co-authored-by: sim <sim@local>
2026-08-09 19:18:56 -04:00
github-actions[bot]
4f2a4932ce chore: sync next package version to 1.10.0 2026-08-08 05:07:32 +00:00
𝚌𝚕𝚎𝚣𝚌𝚘𝚍𝚒𝚗𝚐
88f6d9bd1b fix(#2644): deduplicate Cursor slash menu (#2812)
* fix(#2644): deduplicate Cursor slash menu

* fix: preserve installer executable mode

* chore: add changeset for PR #2812

* test(#2644): acknowledge Cursor emission changes

* test(#2644): drop spent emitted drift acknowledgments

* fix(#2644): remove retired Cursor command converter

---------

Co-authored-by: clezcoding <clezcoding@users.noreply.github.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-03 12:05:45 -04:00
github-actions[bot]
854c93533c chore: sync next package version to 1.9.1 2026-07-31 13:12:09 +00:00
github-actions[bot]
4232a79396 chore: sync next package version to 1.9.0 2026-07-31 03:14:04 +00:00
Tom Boucher
3f6b063fbb chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans

Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers
iterate declared lanes instead of hand-authored per-CLI bash.

Five additive descriptor amendments, each forced by a lane that ships today:
- LaneHandler gains 'opencode' — the lane rebuilds its review from assistant
  text parts of a --format json stream; a plain stdout copy re-breaks #1936.
- modelConfigKey — antigravity's key is review.models.agy, not .antigravity,
  so resolving by slug silently dropped a configured model.
- defaultHost/fallbackModel — Phase 4 federated every *_host with a default of
  empty string; the real fallback only existed in the bash.
- args becomes an argv template with a closed four-placeholder vocabulary.
  Positional splicing produced 'codex --model M -o F exec --ephemeral', which
  is not a valid invocation: codex injects in the middle, twice.
- kimi-code lane, with the bounded command-capability probe (needle
  --output-format) that tells Kimi Code from the legacy python kimi-cli.

Parity gate re-pointed: the workflow-text families it scanned are the text this
phase deletes, so they are replaced by descriptor-to-registry parity plus an
anti-parity check that no bespoke leg returns.

jq, curl and external timeout/gtimeout all drop out of the review path.

Refs #2782

* chore(#2799): add review-lane query surface and widen the manifest vocabulary

Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow
loops over, projects all twelve lanes into their capability manifests, and
widens capability-validator for the amendments.

opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's
own admission rule: one lane, justified by a documented upstream defect data
cannot express (#1936 — the agent can end its turn with zero output tokens and
--format default then drops the assistant text entirely).

Two bugs caught by an end-to-end stub run and fixed here:
- loadConfigResolved returns a provenance wrapper, not the config; using it
  directly resolved every key to undefined, which reads as 'nothing
  configured' and silently dropped every model override.
- hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced
  with a PATH scan that spawns nothing at all.

Refs #2782

* chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews

Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved
lanes, and renders REVIEWS.md sections from each lane's declared
reviewsSection instead of thirteen hardcoded headings. review.md drops from
1104 lines to 507 (61KB to 28.7KB).

Parity gate re-pointed, as agreed: the leg-marker and section-heading families
scanned exactly the text this phase deletes, so they are replaced by
descriptor-to-registry parity in both directions, plus an anti-parity check
that fires if a bespoke leg is ever re-added. Enum, emitting sites and the
Object.keys lock moved together.

The budget-trim helper is hoisted out of the Ollama leg: it was always
lane-agnostic, and any lane may now declare a promptBudgetKey.

Refs #2782

* feat(#2799): bind the consented egress host and re-verify it at invocation

Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3
but was not implemented: ConsentRecord had no host field and nothing in the
tree bound one, so this phase's rule-4 comparison had no baseline.

ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design:
isValidConsentRecord does not require it, so every record already on disk
stays valid and no re-consent storm fires (D4 rule 5). It is deliberately
excluded from disclosureSignature — the loader has no config resolver, so
folding a config-derived value in would make loader and lifecycle compute
different signatures for the same manifest and re-prompt forever.

Install resolves hostConfigKey (falling back to the lane's declared
defaultHost, which is what the invocation path uses) and records it.
Invocation re-resolves and blocks on mismatch rather than silently
redirecting. Absence allows: no record, or a record predating the field,
means nothing to compare — denying there would break every existing
local-model user on upgrade.

Refs #2782

* test(#2799): cover the resolver, runner and handlers; retarget the parity suites

Adds the golden invocation-plan table (one row per shipped lane, derived from
the bash legs rather than the descriptor types) plus runner coverage for the
probe, empty-output policy, the three handlers and the egress check.

Retargets the existing suites onto the new contract: descriptor-to-registry
parity, the anti-parity check, the opencode handler, and the twelfth lane.

Two corrections found by running them:
- modelConfigKey was required; that breaks D4 rule 2, since a reviewer
  manifest authored before this phase would fail validation on upgrade. It is
  optional, read as null when absent.
- the antigravity non-zero-exit test pre-seeded the transcript, which asserted
  that a STALE entry leaks through — the exact bug the watermark prevents. The
  spawn now appends, as the real tool does.

Refs #2782

* fix(#2799): restore agy --add-dir and the self-report prompt in the handler

Retargeting the three legacy reviewer suites off the deleted bash surfaced two
real regressions in the port, both #2176:

- --add-dir was dropped. Without it agy's permission context never receives the
  cwd repo, so the agent anchors on its own scratch dir and reviews the plan
  text in isolation — the exact failure the Review Instructions forbid. It is
  capability-probed, because an older agy rejects the unknown flag outright and
  a lane that fails to start is worse than one running on the prompt anchor.
- the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS
  self-report, which is what makes a blind review distinguishable from a
  grounded one. antigravity now builds its own prompt variant.

Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned
model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own
log is the only evidence that anything failed.

The three suites now assert against the plan and the handler instead of
matching fence text, so they no longer need allow-test-rule exemptions.

Refs #2782

* docs(#2799): document the declared lanes, the new flag, and dropped prerequisites

COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph,
which is now false: no lane requires jq, curl or an external timeout. Adds the
changed-egress-destination behavior, since a blocked lane is something a user
can hit.

CONFIGURATION.md records that the model config key is declared per lane rather
than derived from the flag — antigravity's is review.models.agy — and adds
review.models.kimi-code.

reviewer-instances.md now routes an instance through its lane's single
invocation seam instead of a copied per-adapter bash block, which is what lets
a cross-cutting fix reach instances for free. That required implementing the
--model/--agent/--as flags it documents; --model re-resolves through the lane's
argv template rather than splicing, so the flag lands where the lane declares
it rather than ahead of a subcommand.

CONTEXT.md glossary gains both new modules.

Refs #2782

* chore(#2799): drop the stale emitted-drift acknowledgment

The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md.
That file now shrinks by ~32KB and every emitted hash that moved is
attributable to this diff, so the ack no longer explains anything. Removing
the last entry means removing the file: its presence is the alarm, and an
empty one signals nothing.

Verified by deleting it and re-running the attribution and provenance gates
plus lint:ci — all green without it.

Refs #2782

* docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782

Five additive amendments, each forced by a lane that ships today, plus two
corrections the phase had to make rather than work around:

- D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so
  this phase's rule-4 comparison had no baseline. Recorded because an ADR
  asserting a rule was delivered is exactly what stops a later phase checking.
- The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families
  scanned the text this phase deletes.

Also records that D7's 'skip the probe where no bounding mechanism exists'
carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded
on every stock macOS host, which ships neither timeout nor gtimeout.

Refs #2782

* fix(#2799): close four defects found by adversarial review

Two confirmed bugs, both reproduced before fixing:

- resolveLanePlan was not total. An openai-http lane with a missing or
  non-object invoke dereferenced inv.hostConfigKey and threw, contradicting
  the module's own documented contract; the spawn branch guarded correctly and
  the http branch did not. The CLI seam resolves every selected lane in one
  map, so one malformed overlay manifest would have aborted the whole review
  rather than dropping its own lane. Guarded, plus a per-lane try/catch at the
  seam so a throw can never take down siblings.
- A reviewer-instance model was silently dropped for any lane declaring
  modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates
  that cli is a known slug but never that the slug accepts a model, so a user
  could configure one, get a clean run, and never learn a different model
  reviewed their plan. Now warns explicitly.

Two hardening fixes:

- The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in
  the resolver rather than inherited from a validator that does not run on this
  path — the module documents itself as the overlay-manifest trust boundary, so
  it should not depend on someone else having checked.
- normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses
  with an empty hostname, so it became 'localhost://11434' and was compared and
  requested as if real. An empty hostname now means not-a-URL.

Also documents the one gap that cannot be closed here: the antigravity
watermark is keyed by workspace, so two concurrent reviews of the same repo
share a transcript. agy exposes no per-invocation id to filter on, so the
handler now states which half of its never-stale guarantee actually holds.

Refs #2782

* test(#2799): retarget the remaining eight review.md-asserting suites

The remote runner found 37 failures the local sweep missed (it hit the shell's
two-minute cap before reaching these). All eight extract per-CLI bash from
review.md that this phase deletes; each protects a real invariant, so each is
retargeted onto the plan, the runner or the handler rather than removed.

Three real defects surfaced by doing so:

- effort args never reached ANY lane. model-resolver.cjs exports no
  resolveExecution, so effortFor silently returned [] every time. Restored by
  calling the same bounded resolve-execution query the bash legs used — and
  NOT with --raw, which prints the resolved effort rather than the picked
  field, so claude got 'low' instead of '--effort low'.
- the timeout guidance lost 'a silent empty output is a timeout kill, not a
  crash' — the operator note that exists because of the Codex 0xc0000142
  misdiagnosis. Restored.
- the opencode handler dropped EMPTY assistant text parts. The shipped jq was
  , and  only substitutes for false/null — an empty
  string is truthy in jq and contributed a blank line. Found by a property
  test shrinking to ['', ''].

The opencode property suite no longer spawns jq at all, which deletes the
#2099 hang mechanism it was architected around rather than mitigating it.

Refs #2782

* fix(#2799): register the two new generated modules, and untrack them

The remote runner caught build output committed to git. Both new modules
compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated
that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own
review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each
bin/lib/*.cjs is linted xor ignored according to migration state" failed.

Registered both in .gitignore and eslint.config.mjs alongside the Phase 1
module, and dropped them from the index. Nothing about the shipped behaviour
changes; the artifacts are rebuilt by build:lib.

This is the new-.cts-module registration ripple, and it is the one part of it I
had not completed - the CONTEXT.md glossary and the inventory manifest were
already done.

Refs #2782

* chore(#2799): backfill changeset pr number to 2861

* chore(#2799): backfill changeset pr number to 2861

---------

Co-authored-by: Test <test@example.com>
2026-07-30 12:48:06 -04:00
Tom Boucher
6a9babda69 chore(#2798): declare the eleven reviewer lanes as manifest data (#2837)
* chore(#2798): declare the eleven reviewer lanes as manifest data

Phase 5a of epic #2782, delivering ADR-2782 D9 (roster half) and D3.

- Five reviewers GSD never installs into become lane-only role:reviewer
  capabilities with no runtime body, no runtimeCompat and no install surface:
  gemini, coderabbit, ollama, lm-studio, llama-cpp. Before this they had no
  descriptor at all and lived as a hardcoded NON_RUNTIME_REVIEWER_SLUGS tail,
  which is now deleted outright.
- The six hosts that are ALSO reviewers gain a reviewer body alongside their
  runtime body. Their runtime bodies are byte-identical to next -- verified per
  capability against the git blob, not asserted -- so no install behaviour moves.
- KNOWN_REVIEWER_SLUGS derives from declared bodies via an exported
  deriveReviewerSlugs(registry). hostBehaviors.reviewerCli survives as a derived
  legacy alias for one release; where a capability carries both, the body wins
  and the slug appears once. Alias removal is Phase 7 (#2801).

THE KEYSTONE: the roster is the SAME ELEVEN SLUGS as before -- antigravity,
claude, coderabbit, codex, cursor, gemini, llama_cpp, lm_studio, ollama,
opencode, qwen. This phase changes HOW the roster is derived, not WHO is in it,
and the test asserts that literal list rather than a count.

kimi-code is deliberately NOT declared here. It is net-new with no
invoke_reviewers leg, so declaring it now would make it selectable but not
invocable -- present in --all, selected, emitting an empty section for the whole
5a-to-5b window -- and would break Phase 1's parity assertion. It lands in 5b
alongside the iteration that can run it. Legacy kimi (the Python CLI) is not a
reviewer at all and gains nothing.

The highest-value test is declaredManifestLanesMatchThePhase1Descriptor: it
deep-compares all eleven declared bodies against REVIEWER_LANES field-by-field,
including probe and invoke sub-fields. All eleven are byte-identical, key order
included. The epic's premise is that the manifest and the core descriptor
describe the same lane with NO translation layer, and Phase 2's review already
caught one divergence that every other test missed.

Two ADR corrections folded in, as Phases 1-3 each did:

1. PHASE ORDER. The ADR runs Phase 4 (federated config) before 5a and #2798
   claims a dependency on 4. That is inverted and makes Phase 4 unsatisfiable:
   D9 assigns review.<host>_host to lane capabilities that do not exist until
   THIS phase creates them, and a federated config slice must live inside
   capabilities/<id>/capability.json. Real graph: Phase 2 -> 5a -> 4.
2. #2798's INVENTORY acceptance item is vacuous. The inventory catalogs
   bin/lib/*.cjs modules, not capability directories -- antigravity, opencode
   and qwen appear zero times in it -- and gen-inventory-manifest --check passes
   with the five new dirs and no edit.

Also corrected a stale line in Phase 2's own ADR amendment: it recorded the slug
pattern as /^[a-z][a-z0-9_-]*$/, but Phase 2's security review widened the
shipped pattern to /^[a-z0-9][a-z0-9_-]*$/ to match Phase 1's exported
LANE_SLUG_RE. The prose had not followed the code.

Closes #2798

* fix(#2798): catalogue reviewer capabilities in the generated matrix

The capability matrix rendered exactly two tables, feature and runtime, via
renderTable(caps, role) filtering on c.role === role. ADR-2782 D3 added a THIRD
role, so every role:"reviewer" capability was silently dropped from the
first-party catalogue.

The drift guard did not catch it, and could not: --check compares generated
output against the committed file, and both omitted the five lanes identically,
so it reported "up to date" while five shipped capabilities were invisible in
the one document that is supposed to list what ships. A guard blind to an entire
role is not guarding.

This phase is what exposed it -- it ships the first role:"reviewer"
capabilities -- so it is fixed here rather than deferred (CLAUDE.md: a defect
found while working is fixed in the current change, which overrides
one-concern-per-PR).

Verified red-before-green: with a lane row deleted from the matrix, --check now
exits 1; restored, it exits 0. Before this fix the lanes were absent entirely, so
there was nothing for the guard to compare.

Phase 6 (#2800) still owns enriching the matrix with lane-specific detail
(slug/flag/transport columns) and the locale parity gate. This is the narrower
fix: the capabilities APPEAR at all.

* fix(#2798): close two hardening gaps and record three limits durably

Isolated security review (5 targets, no blockers) reproduced two gaps in the new
deriveReviewerSlugs. Both are unreachable through the checked-in registry -- it is
generated, JSON-sourced and code-reviewed -- but the function is EXPORTED for
reuse and carries no other validation, so it must not depend on its caller.

- A whitespace-only slug passed the length>0 test verbatim and occupied a roster
  entry it could never match. Slugs are now trimmed before the emptiness test. A
  blank body correctly falls through to the legacy alias rather than DROPPING the
  lane, which would have been worse than the blank slug.
- KNOWN_REVIEWER_SLUGS is computed at require() time, so an uncaught throw there
  breaks import for EVERY consumer rather than degrading selection. It is now
  guarded, yielding an empty roster on a malformed registry. That is a visible
  degradation, not a silent one: under D4 an explicitly requested reviewer that
  is unavailable is an ERROR, so /gsd:review --claude against an empty roster
  fails loudly. This also removes an asymmetry -- the sibling capability-trust
  module documents its collectors as TOTAL and wraps them for exactly this reason.

Also records three findings that previously existed ONLY in squash-merged PR
bodies, which is not a durable record:

- ADR-2782 D5 gains an implementation note explaining why the resolved host is
  deliberately EXCLUDED from the disclosure signature. Rule 1 says consent binds
  the resolved host; the loader has no config resolver, so folding it in would
  make the loader and lifecycle compute different signatures for one manifest and
  re-prompt forever. The binding is split: signature covers the SHA-pinned
  manifest fields, the consent record stores the resolved host, and Phase 5b
  re-resolves at invocation -- which is where rule 4 already puts the check. A
  reader comparing rule 1 to the code would otherwise conclude it is unimplemented.
- CONTEXT.md's capability-trust entry still described THREE executable surfaces.
  Phase 3 added the fourth and made that false; corrected here, since it is drift
  this epic introduced rather than Phase 6's new-glossary-term work.
- stableJson documents the NaN/Infinity/undefined -> null signature collision and
  why it is unreachable (JSON grammar has no such literal, so JSON.parse throws
  first). Reachability rests entirely on the ingest path staying JSON.parse-only,
  so the note lives where someone would break it.

* chore(#2798): backfill changeset pr number to 2837
2026-07-29 16:58:02 -04:00
Tom Boucher
6ad30f74b6 feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler.

harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run.

Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard.

Closes #2627

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:50:20 -04:00
Tom Boucher
ec7978c0b4 feat(#2584): add dispatch.isolation sub-field, descriptors, validator + negotiation (#2604) 2026-07-24 12:51:42 -04:00
Tom Boucher
d39ca0c004 Merge pull request #2490 from open-gsd/docs/2481-adr-1239-effort-axis
feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort
2026-07-21 20:14:27 -04:00
github-actions[bot]
8cf747724f chore: sync next package version to 1.8.0 2026-07-22 00:06:08 +00:00
Tom Boucher
09b535ac00 feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort
ADR-1239 gains a ninth negotiated axis, effortSurface (argv | none), declaring how
a host accepts reasoning effort. ADR-443 is amended in the same change because its
recorded deferral is what the axis resolves: its Unblock condition offered paths
(a) and (b) and stated the choice was 'a maintainer call this file records but does
not make'. Path (a) is selected and satisfied here.

Before this, effort reached a runtime only through install-time channels
(EFFORT_RENDERING's frontmatter/api), so reviewer CLIs spawned as subprocesses
silently inherited whatever effort sat in the user's own global CLI config. The
review lane now resolves one universal effort through the ADR-443 cascade and
renders it per host through the negotiated descriptor.

Every per-host value is documentation-sourced, never inferred:
- claude   argv  -- verified via 'claude --help' (--effort <level>)
- opencode argv  -- verified via 'opencode run --help' (--variant)
- codex    argv  -- codex-rs/exec/src/cli.rs: model_reasoning_effort is NOT a CLI
                    flag (config.toml key only), so the global -c override is the
                    only argv route
- 15 hosts undocumented -- their docs state no reasoning setting; the sentinel
                    fails closed rather than inheriting a profile baseline

No config-file vocabulary member: the only host that ever had one (Gemini CLI's
thinkingConfig) was removed as a sunset runtime by 8f2ebbe9b (#1928, PR #1996),
and neither Antigravity CLI nor ZCode documents a reasoning setting.

Closes #2481

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 19:10:22 -04:00
Tom Boucher
6b196ef638 fix(#2423): finalize job syncs next package.json after final release (#2437)
* fix(release): finalize job calls sync-next-version.cjs to bump next after final release (#2423)

The release pipeline's 'finalize' job shipped X.Y.0 to npm 'latest' and
merged to main, but never bumped 'next' to match. 'scripts/sync-next-version.cjs'
exists exactly for this — its docstring promises to run 'for every release
type (rc / hotfix / final)' — but it was wired only into the 'rc' job
(release.yml:479), not 'finalize'. As a result, after 1.7.0 shipped on
2026-07-15, 'next' stayed at 1.7.0-rc.6 and every npm script banner on
'next' (and feature branches cut from it) reported the stale rc version.

Regression of #1104 — closed incomplete (covered rc only, not final).

This patch:
  - adds a 'Sync next branch to the published release' step to the
    'finalize' job, mirroring the rc job's pattern at line 479 (uses
    VERSION from inputs.version rather than PRE_VERSION from steps.prerelease,
    since finalize does not run the prerelease step);
  - gates it on !inputs.dry_run + continue-on-error:true (matches rc);
  - adds tests/release-finalize-syncs-next-version.test.cjs — a structural
    YAML assertion that fails against the pre-fix workflow and passes after.
    Failing-first demonstrated during development; the test parses job blocks
    by indentation rather than grep so it stays valid as the file grows.

Root-cause diagnosis: scripts/sync-next-version.cjs:14 docstring admits
'used by release.yml's rc job, which has no next-targeting PR of its own'.
release.yml:506-685 (finalize job) had no sync-next-version step before
this patch. Every prior rc.N release has a matching 'chore: sync next
package version to 1.7.0-rc.N' commit; there is no such commit for 1.7.0.

* chore(release): sync next package version to 1.7.0 (#2423)

Replays the canonical 'chore: sync next package version to <v>' commit
that the release pipeline's rc job auto-produces via scripts/sync-next-version.cjs,
for the 1.7.0 final release that shipped on 2026-07-15 (commit dd4c90f82
'chore: finalize v1.7.0' on main). Without this, 'next' (and every feature
branch cut from it) carried 1.7.0-rc.6 indefinitely and reported it in
every npm script banner (e.g. 'lint:ci').

Bumps 43 synchronized manifests via the npm 'version' lifecycle hook
(scripts/sync-manifest-versions.cjs --stage + scripts/gen-capability-registry.cjs
--write), matching the file set of commit 27f69cc48 ('chore: sync next
package version to 1.7.0-rc.6') and commit dd4c90f82 ('chore: finalize
v1.7.0').

This is the immediate Layer-1 repair for #2423. Layer-2 (prevent recurrence)
is the workflow patch in the previous commit; Layer-3 (regression test)
ships with it. Future X.Y.0 final releases will produce this commit
automatically once the workflow fix lands.

* test(release): tighten #2423 dry-run gate assertion to the sync step

Code review of fix/2423 found that test #3 ('gates sync-next-version on
!inputs.dry_run') asserted too loosely: it scanned the entire finalize
block for any '!inputs.dry_run' line, so it would still pass if the
gate were stripped from the sync-next-version step specifically — the
exact regression the test name promises to catch. The finalize job has
multiple steps with their own !inputs.dry_run gates (e.g. Verify
publish), so the loose version masked the very bug it claimed to detect.

Tighten by extracting the specific YAML step block containing
'scripts/sync-next-version.cjs' and asserting the gate appears within
THAT step's lines, not anywhere in the job. Verified the tightened test:

  - PASSES against the post-fix workflow (sync step has its own gate)
  - FAILS when the sync step's gate is stripped (even when other steps
    in finalize retain their own !inputs.dry_run gates) — the exact
    regression that previously slipped through

Adds extractStepBlockContaining(jobBlock, marker) helper alongside the
existing extractJobBlock(text, jobName). Reuses the same indentation-
based parsing, so it stays valid as the file grows.

* chore(changeset): backfill pr:2437 in .changeset/sturdy-ibex-jump.md

CLAUDE.md changeset convention: 'Use placeholder pr:0 during initial commit.
Backfill immediately after gh api POST /pulls returns the real number.'

PR #2437 created from branch fix/2423-release-finalize-sync-next-version.

* fix(#2423): add see #2423 to allow-test-rule exemption per ADR-456

CI lint-allow-test-rule-refs failed on PR #2437: ADR-456 requires new
allow-test-rule exemptions added after the ADR's acceptance to include a
tracking issue number in the comment, in the form
  // allow-test-rule: <reason> (see #NNN)
The exemption added in commit 976c8b0a2 lacked this ref. Fixed.

Verified locally:
  node scripts/lint-allow-test-rule-refs.cjs
    → ok lint-allow-test-rule-refs: 173 grandfathered exemption(s) tracked, no novel untracked offenders
2026-07-19 14:33:19 -04:00
github-actions[bot]
27f69cc489 chore: sync next package version to 1.7.0-rc.6 2026-07-12 05:40:45 +00:00
Tom Boucher
c6ce110efa feat(#2092): migrate Qwen Code onto EoS imperative adapter + native subagents + SubagentStart (ADR-1239)
Fold all runtime==='qwen'/isQwen logic branches (skill-priority frontmatter,
branding/path rewrites, legacy commands/gsd cleanup, hyphen-namespace
normalization, RUNTIME_CONTENT_DISPATCH, hooks-surface label) into
descriptor-driven runtime.hostBehaviors on capabilities/qwen/capability.json,
read via _hostBehaviors(). Shared claude/qwen/hermes legacy-migration branches
in install-engine.cts folded to descriptor flags (claude+hermes descriptors
updated; FALLBACK_HOST_BEHAVIORS.claude floored). Byte-identical golden parity
for qwen/hermes/claude(global+local).

UPGRADE 1: native .qwen/agents/*.md subagent projection — new agents
artifact-layout kind + convertClaudeAgentToQwenAgent converter (name +
description + tools YAML block list; color/model dropped). qwen routed onto
the descriptor-driven agents path (_DESCRIPTOR_AGENTS_RUNTIMES).
UPGRADE 2: SubagentStart hook wired into extendedHookEvents + the
descriptor-gated hook-writer loop (activates only for qwen).

Tests: qwen-imperative-reference (adapter/axes/fail-closed/hostBehaviors +
no runtime==='qwen' source-grep across 4 files) + qwen-upgrades (agents file
validity + SubagentStart mirrors SubagentStop, descriptor-gated). Docs matrix
+ how-to updated; changeset added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 15:44:19 -04:00
github-actions[bot]
451e58e6d5 chore: sync next package version to 1.7.0-rc.5 2026-07-10 17:09:38 +00:00
Tom Boucher
b0d985ccb3 feat(#2089): migrate cursor onto imperative adapter + hook-bus/dispatch upgrades 2026-07-09 00:21:39 -04:00
github-actions[bot]
a6263bba8a chore: sync next package version to 1.7.0-rc.4 2026-07-07 06:08:27 +00:00
github-actions[bot]
a828d8207c chore: sync next package version to 1.7.0-rc.3 2026-07-06 02:19:50 +00:00
github-actions[bot]
8f791817f4 chore: sync next package version to 1.7.0-rc.2 2026-07-03 04:15:36 +00:00
github-actions[bot]
feb7036399 chore: sync next package version to 1.7.0-rc.1 2026-06-30 13:28:49 +00:00
Tom Boucher
ce62f2b68d refactor(#1763): ADR-1235 agent migration — cut over the trivial-converter group to the descriptor path (#1764)
ADR-1235 step 1: route the trivial-converter runtime group (cursor, windsurf, augment, trae, codebuddy) off the inline install() agent loop onto the descriptor-driven installRuntimeArtifacts path. Establishes the converter-context foundation (pre-converter cross-cutting + no agent-stamp). Agent install output is byte-identical for all 16 runtimes (golden-parity, global + verified local). cline deliberately excluded (local rules-only). Closes #1763.
2026-06-26 15:38:51 -04:00
Tom Boucher
a3d3c2a445 refactor(#1756): derive getDirName from a documented runtime.localConfigDir descriptor axis (#1757)
ADR-1239 Phase B (parent #1679). getDirName was a hand-maintained 15-branch
if-chain mapping each runtime to its local content-rewrite dot-dir. Relocate
those values into a documented runtime.localConfigDir descriptor field; derive
getDirName from registry.runtimes[id].runtime.localConfigDir (fallback .claude).

- 16 capability.json gain runtime.localConfigDir (byte-identical values)
- capability-validator.cjs requires it (non-empty dot-dir); registry regenerated
- docs/reference/capability-manifest.md documents the field + the three
  divergent values (copilot=.github, antigravity=.agents, kimi=.kimi-code)
- drift-guard test: golden value map + key-set equality both ways

Byte-identical install output for all 16 runtimes (golden-parity harness #1730).

Closes #1756

Co-authored-by: review-bot <review-bot@gsd>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 12:44:09 -04:00
Tom Boucher
cf2e66b39e feat(#1708): typed documentation-sourced #853 dispatch-flatten (ADR-1239 Phase B) (#1719)
* feat(#1708): typed documentation-sourced #853 dispatch-flatten

Graduate the #853 orchestrator-backgrounding decision from a scattered RUNTIME==='codex' prose check to a typed, documentation-sourced engine decision. Adds a backgroundDispatch dispatch sub-axis (sourced per host: codex+cursor documented true, 9 documented false, 5 undocumented), shouldFlattenDispatch(dispatch) (inline UNLESS background && backgroundDispatch, fail-closed), and a gsd_run query dispatch-should-flatten the plan/execute workflows call. Cursor is newly background-eligible per its docs (inline->background) — a documentation-justified behavior change. No RUNTIME-name residue for this decision.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1708): backgroundDispatch citations in matrix + CONTEXT note

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1708): address review findings on typed dispatch-flatten

Code/adversarial review: convert the manager.md/autonomous.md Compound Action preamble from hardcoded 'On Codex' to FLATTEN-based branching (the handlers already use the query; the preamble contradicted them and was wrong for cursor); make shouldFlattenDispatch null-safe + type-honest (accepts raw 'undocumented' registry values); make backgroundDispatch a required descriptor field (matching its siblings, all 16 carry it); strengthen the config.runtime behavioral test; update the bug-853 prose-pin test + comment. Security review clean; Codex confirmed no fail-open.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1708): backfill backgroundDispatch in role:runtime test fixtures

Making backgroundDispatch a required descriptor field broke role:runtime fixtures in capability-manifest-version/capability-registry/host-integration-descriptors tests that build a dispatch object without it (caught by full gsd-test, not scoped npm test). Backfill backgroundDispatch:false into the well-formed fixtures; the deliberately-malformed 'required-field' test fixture is left malformed by design.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1708): update fix-1521 dispatch-gating assertion to the FLATTEN gate

fix-1521 pinned the codex-specific run_in_background prose that #1708 graduated to the typed dispatch-should-flatten/FLATTEN gate. Update its assertions to verify FLATTEN=false gating (not a runtime name) + that the old RUNTIME===codex gate is gone. Caught by full gsd-test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1708): add changeset for typed dispatch-flatten

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1708): remove stray temp PR-body file

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1708): add issue ref to bug-853 allow-test-rule annotations

ADR-456 requires every allow-test-rule exemption to carry a see #NNN reference; the source-text-is-the-product annotations added when migrating the prose assertions lacked it (lint-tests CI gate). Add (see #1708).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 15:37:33 -04:00
Tom Boucher
30d4b85de5 feat(#1684): negotiated host-integration interface (ADR-1239 Phase A) (#1690)
* feat(#1684): add negotiated host-integration interface module

ADR-1239 Phase A: a pure, additive, no-I/O module exposing PROTOCOL_VERSION, the 8-axis HOST_INTEGRATION_AXES closed vocabulary, the UNDOCUMENTED fail-closed sentinel, negotiateHostCapabilities (effective subset of host-declared and engine-known), a typed degradation ladder, and host-capability profiles.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1684): validate and document host-integration axes (16 runtimes)

Extend validateRuntimeBody to validate the 8 hostIntegration axes (closed enums + undocumented sentinel + dispatch struct + reserved-key guards) and the widened runtime vocabulary; author a documentation-sourced hostIntegration block in all 16 runtime descriptors; regenerate the registry. Every per-CLI value is documented (cited) or the explicit undocumented sentinel.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1684): add host-integration capability matrix and adr amendment

New per-CLI, per-axis citation reference (value/source/evidence for all 16 CLIs); ADR-1239 Phase-A-implemented amendment; CONTEXT.md glossary seam entry.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1684): harden dispatch negotiation edge cases

Code-review hardening: treat NaN/Infinity maxDepth as missing (fail-closed, +warning); reset nested/background when namedDispatch collapses to false (struct consistency); SAFE_DEFAULTS dispatch floor to read-only; warn on non-finite protocolVersion; symmetric undocumented warnings for dispatch fields. Pure module — no consumers; behaviour fail-closed throughout.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1684): register host-integration.cjs in lint-ignore and inventory

New tsc-generated bin/lib artifact: add to the eslint ignore list (ADR-457 — lint the .cts source), regenerate docs/INVENTORY-MANIFEST.json, and add the docs/INVENTORY.md CLI-modules row. Fixes the 3 gsd-test failures (551-eslint-bin-lib-coverage x2 + inventory-manifest-sync).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1684): add changeset fragment for host-integration interface

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1684): add how-to for sourcing a host's integration axes

Diataxis how-to guide for adding/updating a host's runtime.hostIntegration axes from authoritative docs, the undocumented-sentinel rule, validation, and extending the closed vocabulary. Completes the Step-5 doc quadrants (reference + explanation + how-to). Indexed in docs/README.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 11:33:06 -04:00
github-actions[bot]
548a986ffc chore: sync next package version to 1.6.0 2026-06-24 23:09:44 +00:00
github-actions[bot]
2e406e8e05 chore: sync next package version to 1.6.0-rc.3 2026-06-24 03:52:20 +00:00
github-actions[bot]
f20dc691c1 chore: sync next package version to 1.6.0-rc.2 2026-06-22 05:22:23 +00:00
github-actions[bot]
4441b20b4a chore: sync next package version to 1.6.0-rc.1 2026-06-20 19:58:40 +00:00
Tom Boucher
2421cf1b4a feat(#1430): versioned capability manifest + native stamping (ADR-1244 Phase 1) (#1436)
* feat(#1430): versioned capability manifest + native stamping (ADR-1244 Phase 1)

Make the capability manifest versioned — the data substrate the Capability
Ecosystem (ADR-1244) keys off:

- capability.json gains a REQUIRED semver `version` plus the optional
  ecosystem envelope (`engines.gsd`, `compatVersions`, `integrity`,
  `provenance`); the build-time conformance validator enforces them via a new
  `validateVersionEnvelope()` (exported for the Phase 2 runtime overlay).
- All 32 native capabilities stamped with `version` (= package version,
  lockstep) + `engines.gsd`; `sync-manifest-versions.cjs` gains a glob sweep
  that keeps them in sync, and the issue-844 regression guard is extended.
- Strict SemVer 2.0.0 grammar blocks metacharacter/space/unicode smuggling in
  version strings; range/integrity fields are shape-validated (satisfaction
  and the load-time gate are deferred to Phase 2/4).
- Capability rel-paths emitted forward-slash for cross-platform git correctness.

Closes #1430

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1430): add changeset for versioned capability manifest

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 10:45:53 -04:00
Tom Boucher
94872662e9 feat(#1082): complete phase 5 — descriptor-drive all install surfaces + materialize the InstallPlan — ADR-857/1016/58 (#1080)
* feat(#1077): phase 5f-2 — drive the hookEvents dialect (PostToolUse/AfterTool) from the descriptor

postToolEvent (bin/install.js) and preToolEvent (applySettingsJsonHooks in
runtime-hooks-surface.cts) now select the event-name dialect from
registry.runtimes[id].runtime.hookEvents instead of the hardcoded
(runtime === 'gemini' || runtime === 'antigravity') check: hookEvents === 'gemini'
→ AfterTool/BeforeTool; else → PostToolUse/PreToolUse. hookEvents threaded into the
applySettingsJsonHooks opts bag. Equivalence-preserving (Codex-verified): hookEvents
'gemini' is exactly {gemini, antigravity}, 'claude' the rest; undefined → claude
dialect (matches the old else).

The per-event SET guards (isQwen||claude → SubagentStop/Stop/PreCompact;
runtime==='claude' → FileChanged; isGemini → Gemini agent-events) stay HARDCODED —
hookEvents (2-value) is too coarse to drive them (the event set differs within
hookEvents='claude'); per-event-set drive tracked in #1076.

Registry-parity test (enh-1077): asserts BOTH post-tool (AfterTool/PostToolUse) AND
pre-tool (BeforeTool/PreToolUse) dialects are a pure function of hookEvents, for
gemini/antigravity/claude/augment — non-vacuous (catches a broken hookEvents thread).

Closes #1077

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1077): build hooks/dist in before() so dialect-drive test passes in scoped CI

hooks/dist is gitignored and absent in scoped/windows CI jobs that do not
pre-run build:hooks. Without it, install() finds no hook files and all
AfterTool/BeforeTool/PostToolUse/PreToolUse event arrays come back empty,
failing every hook-presence assertion. Added an idempotent ensureHooksDist()
called in a top-level before() — mirrors the pattern from
bug-376-claude-js-hook-gsd-rewriter.test.cjs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1055): add installSurface/writesSharedSettings/permissionWriter/extendedHookEvents to runtime descriptors

Purely additive: four new fields on all 16 runtime capability.json descriptors,
validator extended with three new closed-vocab sets, registry regenerated.
Test fixtures (VALID_RUNTIME_CAP and makeRuntimeCap) updated to include the new
required fields so all 255 capability-registry tests continue to pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1076): drive per-event hook guards from extendedHookEvents descriptor

Replace hardcoded runtime-name checks (isQwen||runtime==='claude',
runtime==='claude', isGemini) in applySettingsJsonHooks with a single
descriptor-driven extendedEvents array derived from the new opts field.
Remove isQwen and isGemini derivations (no remaining uses after the three
guard blocks are migrated). Wire extendedHookEvents from the capability
registry in bin/install.js call site. Add behavioral regression test
(enh-1076-extended-hook-events-drive.test.cjs) confirming the drive is
purely descriptor-based and runtime-name-agnostic.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1055): drive resolveRuntimeConfigIntent from the runtime descriptor; retire hand-kept REGISTRY

- Rewrites src/runtime-config-adapter-registry.cts to require capability-registry.cjs
  and read installSurface / writesSharedSettings / permissionWriter from
  runtimes[id].runtime; deletes the hand-kept REGISTRY const (ADR-857 phase 5g drive 2).
- ALLOWED_CONFIG_RUNTIMES is now derived from descriptor entries that have installSurface.
- Fixes the configFormat parity gate in scripts/gen-capability-registry.cjs to read
  installSurface directly from capMap descriptor bodies, breaking the require cycle
  (adapter now requires the generated registry; gen-script must not require the adapter).
- Adds golden-master test tests/enh-1055-config-intent-descriptor-drive.test.cjs (41 tests)
  pinning all 16 runtimes' return shapes and the TypeError-on-unknown contract.
- Updates scripts/lint-test-file-count.allowlist.json (config module, +1 file).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1076): make hooksSurface descriptor load-bearing for the settings-json hook-skip

- Adds hooksSurface?: string to ApplySettingsJsonHooksOpts and destructuring in
  applySettingsJsonHooks (src/runtime-hooks-surface.cts).
- Replaces the hardcoded !isOpencode && !isKilo hook-skip guard with
  hooksSurface !== 'none'; removes the now-unused isOpencode/isKilo derivations
  (ADR-857 phase 5g drive 3).
- Passes hooksSurface from the runtime descriptor at the applySettingsJsonHooks
  call site in bin/install.js using the established
  _capabilityRegistry?.runtimes?.[runtime]?.runtime?.hooksSurface idiom.
- Extends tests/enh-1076-extended-hook-events-drive.test.cjs with two new suites
  proving: (a) hooksSurface:'none' writes no hooks regardless of runtime name;
  (b) hooksSurface:'settings-json' writes hooks even for 'opencode' (previously
  hardcoded to skip).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: record installSurface/writesSharedSettings/permissionWriter/extendedHookEvents descriptor axes in ADR-1016

Add Decision 7a documenting the four axes added in the 5f-completion pass,
update axis counts from "six" to "twelve", note 5f-completion drives as done
in Decision 8's ladder, update Out of scope to reflect #1055/#1076 are done
and 5g (InstallPlan capstone) remains the only open phase.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1055): parity gate must fire on configFormat↔installSurface mismatch (read installSurface at the descriptor level)

The test fixture makeRuntimeCapMap did not include installSurface in the runtime object,
so the gate's typeof r.installSurface !== 'string' guard always skipped the entry and never threw.
Added installSurface as an optional third parameter to makeRuntimeCapMap and passed the correct
installSurface values ('settings-json' for claude, 'codex-toml' for codex) to the two THROWS tests.
The gate implementation already reads r.installSurface correctly from the descriptor level.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1076): add installSurface↔hooksSurface + extendedHookEvents↔hookEvents consistency gates with rejection tests

GATE A: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES map in validateRuntimeBody enforces that a
runtime's hooksSurface is valid for its installSurface (e.g. profile-marker-only only allows none,
codex-toml only allows codex-hooks-json). Derived from the 16 real runtime descriptors.

GATE B: validateRuntimeBody checks that if extendedHookEvents contains Gemini agent-events
(BeforeAgent/AfterAgent/BeforeModel), hookEvents must be 'gemini'; if it contains Claude-family
events (SubagentStop/Stop/PreCompact/FileChanged), hookEvents must be 'claude'.

Added 10 rejection tests in suite 27 covering each gate + each new field validator.
All 16 real runtimes satisfy both gates (verified before coding).
Exports: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES, VALID_INSTALL_SURFACES,
VALID_EXTENDED_HOOK_EVENTS, VALID_PERMISSION_WRITERS, validateRuntimeBody.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1076): strengthen hooksSurface-drive assertions; defensive hooksSurface fallback; drop vacuous dup

1. bin/install.js: add explicit literal fallback for hooksSurface when the committed
   capability registry fails to load (opencode/kilo → 'none', all others → 'settings-json').
   The descriptor is always the source of truth in normal operation.

2. enh-1076 Suite 7: change SessionStart assertion from key-presence (hasOwnProperty)
   to at least-one-command (hasHooksFor), so the test fails if hooks are initialized-but-empty.
   ensureHooksDist() in before() guarantees hook files exist.

3. enh-1055 Test 2: remove vacuous duplicate suite that re-asserted intent.runtime === row.runtime
   already fully covered by Test 1's deepStrictEqual over all four fields.

4. capability-registry.test.cjs: fix stale comments in the grok-skip test that claimed the
   parity gate uses the adapter registry; gate reads purely from the descriptor (installSurface
   absent → typeof r.installSurface !== 'string' → soft-skip).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1082): materialize the InstallPlan — collect install-level descriptor axes into resolveInstallPlan; route install()/finishInstall() through it (ADR-58/5g)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: record 5g InstallPlan materialization (ADR-58 Accepted, ADR-1016 phase-5 complete)

ADR-1016 Decision 8 step 7 updated to DONE: resolveInstallPlan(runtime) in
runtime-config-adapter-registry collects install-level descriptor axes into
the typed InstallPlan consumed by install()/finishInstall(). Out-of-scope
section updated: 5g capstone is complete, phase 5 fully materialized.
ADR-1016 line ~20 updated: InstallPlan IS now materialized (both halves).
ADR-58 Implementation note added (2026-06-11): realized in
runtime-config-adapter-registry (co-located with adapter-selection).
CONTEXT.md Runtime Config Adapter Registry entry extended to document
resolveInstallPlan and both-halves realization.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1082): update install drift guard to the resolveInstallPlan seam (5g)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 22:36:47 -04:00
Tom Boucher
58ed55683e feat(#1049): phase 5d — drive artifactLayout from the runtime descriptor (retire the 128-LOC switch) (#1053)
resolveRuntimeArtifactLayout now builds Layout from
registry.runtimes[id].runtime.artifactLayout[scope] — a loop dispatching each
ArtifactKind through the SAME 5 builders (commandsKind/agentsKind/skillsKind/
convertedCommandsKind/kimiAgentsKind, unchanged) by (kind, converter, nesting) —
replacing the hardcoded switch(runtime). Equivalence-preserving for all 16 runtimes
× {global, local} (Codex-verified, no divergence). -43 LOC; bin/install.js + the
converters + the install loop untouched. getInstallExports()[converterName]
resolution, configDir threading, scope default, unknown-runtime guard all preserved.

Driving the local scope surfaced a 5a gap: the old switch had no scope branch for 13
runtimes (cursor/gemini/codex/copilot/antigravity/windsurf/augment/trae/qwen/hermes/
codebuddy/opencode/kilo) → local == global for them, but 5a authored local:[].
Backfilled local=global for those 13 (descriptor-faithful; a fall-through shim would
wrongly give cline/kimi local=global). claude/cline/kimi scope-gating untouched.

validateArtifactKindEntry tightened: destSubpath/prefix/nesting/converter required
(ConverterName enum still open — 5e). New 39-case deep-equal golden equivalence test.

Closes #1049

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 14:08:46 -04:00
Tom Boucher
fbd62cd84f feat(#1035): phase 5a — author 16 role:runtime capability descriptors (registry-only) (#1039)
Author capabilities/<runtime>/capability.json for all 16 runtimes (role:runtime),
populating the registry's runtimes index ({} -> 16). Each carries the 6 ADR-1016
axes (configHome structured, configFormat, artifactLayout structured, commandStyle,
hooksSurface + hookEvents, sandboxTier, supportTier), extracted from the live
modules (runtime-homes/runtime-slash/runtime-artifact-layout/runtime-config-adapter
+ CODEX_AGENT_SANDBOX). validateRuntimeBody tightened to enforce the closed
vocabularies + structured configHome/artifactLayout (rejects old string configHome;
env required; skillsHome recursively validated).

Registry-only: bin/install.js + src/*.cts untouched, zero behavior change. The 5b-5f
drive steps consume these one axis at a time (ADR-1016 §8).

Codex accuracy review caught + fixed: opencode/kilo register ZERO lifecycle hooks
(install.js skips the whole block) -> hooksSurface:none (configFormat stays
settings-json); kimi probe selects on <root>/skills existence -> probeExists:"skills".
ArtifactKind required-field strictness + ConverterName closing deferred to 5d/5e.
ADR-1016 (Proposed) amended to match (Decisions 1 + 5).

Closes #1035

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 09:34:07 -04:00