Commit Graph

121 Commits

Author SHA1 Message Date
Cody Anderson
98e4233ce9 fix(#2176): ground the Antigravity reviewer in the repo under review (#2184)
* fix(#2176): ground the Antigravity reviewer in the repo under review

- capability-probe --add-dir (mirrors the Codex bypass-flag probe) and pass
  the repo root on both invocation arms
- anchor _AGY_PROMPT to the absolute repo root; mandate a
  REVIEWED-WITHOUT-REPO-ACCESS self-report when the repo is unreadable
- stamp a [reviewed-without-repo-access] marker on self-reported or
  scratch-anchored output; Consensus Summary down-weights marked reviews
- apply the same absolute-root anchor to the cursor-agent prompt (AC5)

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* docs(#2176): changeset fragment for PR #2184

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2176): review fixes — size baseline, cursor root anchor, anchored blind tells

- regenerate tests/workflow-size-baseline.json for review.md's growth
- cursor anchor uses git rev-parse --show-toplevel (bare pwd resolved the
  wrong root from a repo subdirectory)
- blind-review tells anchored: self-report to the first lines of output,
  scratch tell to a workspace-declaration phrasing — a grounded review
  quoting either string is no longer mis-stamped

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2176): round-2 review fixes — scratch-tell bridge, behavioral test, changeset

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test: regenerate golden-install-parity fixtures for the review.md change

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test(#2176): pass the transcript path to bash with forward slashes

The behavioral detection test substitutes a mkdtemp path into the bash
compound; on Windows runners that path contains backslashes, which bash
strips, so the transcript is never found and the first assertion fails
(windows-latest/24 lane). Git Bash accepts D:/-style paths.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2176): use /gsd:review namespace syntax in workflow comment

The slash-command namespace invariant (#3443) bans retired /gsd-<cmd>
references in Claude-facing sources; a cursor-anchor comment used
/gsd-review. Size baseline + golden fixtures regenerated for the byte
change.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test(#2176): derive the POSIX path via path.sep, not a hardcoded separator

Review finding: out.replaceAll('\\', '/') hardcodes both separators;
use the separator-safe out.split(path.sep).join(path.posix.sep) idiom.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test(#2176): use the merged toPosixPath seam for the bash path

Per maintainer note: #2247's shell-command-projection now centralizes
running-OS → POSIX path conversion; import it instead of the inline
split/join idiom.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
2026-07-13 15:54:19 -04:00
Adnan
3592697bed fix(#2107): orchestrator honors gate="blocking-human" checkpoints in auto-mode (#2113)
* fix(execute-phase): honor gate="blocking-human" in auto-mode checkpoint handling

The package-legitimacy gate (#2827) spans two layers. gsd-executor refuses to
auto-approve a gate="blocking-human" checkpoint and escalates it so a human can
vet the package. execute-phase's checkpoint_handling step then dispatched purely
on checkpoint *type* and never read gate -- so under --auto/--chain it
auto-approved the checkpoint the executor had just refused to auto-approve.

Net effect: the slopsquatting defence was inert in exactly the unattended mode
where it matters. An [ASSUMED]/[SUS] package reached install with no human ever
seeing the prompt.

- gsd-core/workflows/execute-phase.md: carve out gate="blocking-human" (and the
  package-legitimacy what-built markers) ahead of every auto-mode branch.
- gsd-core/references/checkpoints.md: document the gate attribute and its two
  values. blocking-human previously appeared nowhere outside gsd-executor.md,
  so no planner had a documented way to author a non-auto-approvable checkpoint.
- tests/package-legitimacy-gate.test.cjs: the existing regression test asserted
  the executor half only, which is why it stayed green while the gate was open.
  Now asserts the orchestrator half too.

* chore(changeset): link to issue #2107

* chore(changeset): backfill PR number 2113

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa

* test(#2107): refresh golden-install-parity hashes for edited gsd-core files

The golden fixtures pin content hashes for gsd-core/references/checkpoints.md
and gsd-core/workflows/execute-phase.md, both edited by this fix. Regenerated
via UPDATE_GOLDEN=1; only those two keys change across all 17 runtime fixtures.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa

* fix(#2107): keep the carve-out inside the ADR-857 host-loop budget

The ADR-857 phase-6 ratchet pins execute-phase.md below 93600 LF bytes so
optional-feature logic keeps migrating out of the host loop. The carve-out
first landed 623 bytes over that ceiling.

Move the two-layer rationale (why gsd-executor escalates these checkpoints)
into references/checkpoints.md, where the gate is now documented, and reduce
the workflow to the operative rule. execute-phase.md is 93589 bytes, under
the ceiling; the gate token and both <what-built> marker strings are kept
because the orchestrator matches on them.

Refresh the two baselines the edit invalidates: golden-install-parity
fixtures (only the checkpoints.md and execute-phase.md hashes move) and
workflow-size-baseline.json (one line). The ADR-857 ceiling itself is
untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa

* fix(#2107): executor honors blocking-human on the decision branch + gate transport

Review found the fix incomplete one layer down. Two executor-layer gaps:

1. Blocker — agents/gsd-executor.md auto-mode dispatch gated
   checkpoint:human-verify on gate="blocking-human" but the checkpoint:decision
   branch below auto-selected the first option with no gate check. The executor
   resolves a decision itself (auto-selects and continues) without returning it,
   so the orchestrator carve-out never runs for it. A planner following the new
   checkpoints.md rule 6 ("gate a decision whose default would be wrong to
   assume") would have it silently auto-selected under --auto/--chain — the exact
   #2107 harm, one checkpoint type over. The decision branch now STOPs and
   returns for an explicit human decision when gate="blocking-human".

2. Major (transport) — checkpoint_return_format carried no field conveying the
   gate to the freshly-spawned orchestrator, so recognition of the proactive
   pre-install checkpoint rested on freeform prose. Added a **Gate:** field to
   the return format and re-pointed the execute-phase carve-out at it
   ("If the returned Gate: is blocking-human"). Net byte-negative: execute-phase.md
   drops 93589 -> 93583, widening ADR-857 headroom from 11 to 17 bytes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2107): cover decision carve-out + gate transport, de-vacuum conditional tests

- New: 'auto mode does not auto-select a blocking-human decision checkpoint'
  asserts the executor decision branch STOPs on blocking-human. Verified red on
  the pre-fix executor (2 fail), green with the fix (27 pass).
- New: 'checkpoint_return_format transports the gate ...' asserts the **Gate:**
  field carries blocking-human across the executor->orchestrator boundary.
- New: 'auto-select rule for decision is conditional' — orchestrator-side mirror
  of the human-verify conditional test, for the execute-phase decision branch.
- Fix vacuous test: both conditional tests now assert the anchor matched
  (length > 0) before iterating, so anchor drift can no longer pass with zero
  assertions.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2107): refresh golden + size baselines for executor + execute-phase edits

Regenerated via UPDATE_GOLDEN=1 and update-size-baseline.cjs. Only the
gsd-executor.md and gsd-core/workflows/execute-phase.md hashes move across the
runtime fixtures (35 ins / 35 del, no keys added or removed); checkpoints.md is
unchanged this round. Size baselines: gsd-executor.md 43607 -> 43973,
execute-phase.md 93589 -> 93583 (still under the ADR-857 ceiling).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-13 13:43:29 -04:00
Cody Anderson
ad7111e50b feat(#2161): opt-in absolute token count on the statusline context meter (#2174)
* feat(#2161): opt-in absolute token count on the statusline context meter

New statusline.show_context_tokens config (default false). When enabled,
the context meter shows the absolute token total after the percentage,
e.g. "████░░░░░░ 46% (156k)" — summing input, cache-creation, cache-read,
and output tokens from context_window.current_usage (matching /context).

Default output is byte-for-byte unchanged when the flag is absent or
false. The .planning config is now read once per render and shared with
the last-command/position block instead of being re-read.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* docs(#2161): changeset fragment for PR #2174

* fix(#2161): review fixes — k-to-M threshold, boundary tests, changeset format

- formatTokens promotes to the M branch when k-rounding reaches 1000
  (999,500-999,999 rendered "1000k" instead of "1.0M")
- boundary tests at 999499/999500/999999/1000000/1000001
- Number() guards on the four usage fields (silent string-concat gap)
- changeset body ends with the (#2161) citation per house convention

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2161): round-2 review fixes — config-set coverage, precision claim, exports style

- config-set accept/reject tests for statusline.show_context_tokens
  (mirrors the post-planning-gaps precedent the issue scope names)
- changeset + docs no longer claim parity with /context: the suffix sums
  four fields while the meter %% derives from used_percentage (three), so
  the figures can diverge slightly
- module.exports one entry per line

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test: regenerate golden-install-parity fixtures for the statusline hook change

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-13 13:25:47 -04:00
Tom Boucher
b5ce72f729 fix(#2119): single SECURITY.md writer — auditor is return-only (#2154)
* fix(2119): single SECURITY.md writer — auditor is return-only

The gsd-security-auditor held Write/Edit and was instructed to write
SECURITY.md (no <N>- prefix, no template frontmatter), while the
orchestrator's Step 6 also wrote the correct padded <N>-SECURITY.md
from templates/SECURITY.md. Two writers, two naming conventions, two
shapes — the auditor's unprefixed file was invisible to the workflow's
*-SECURITY.md glob detector and unparseable for the threats_open gate.

Fix (option 1 from the issue): make the auditor return-only.
- Remove Write/Edit from auditor's tools
- Rewrite all 'Write SECURITY.md' instructions to 'Return structured
  verdict' with threats_open count
- Add explicit constraint in workflow Step 5 spawn prompt
- Update existing test (was asserting Write in tools — now asserts absence)
- Add new regression test for single-writer contract
- Update docs/AGENTS.md stale Tools/Produces rows
- Regenerate golden fixtures + agent size baseline

* docs(changeset): backfill PR number (#2154)

* chore(#2119): regenerate pi/qwen golden fixtures after next merge

The single-writer change edits gsd-core/workflows/secure-phase.md and
agents/gsd-security-auditor.md; pi.json (added on next) and qwen.json (merge
straggler) were the only runtime fixtures still holding pre-change hashes for
those files. All other runtimes already reflect the change. Regenerated via
the sanctioned gen-golden-install-parity script.

* merge origin/next — regenerate goldens + baseline for merged state

* fix slash-command syntax: /gsd-secure-phase → /gsd:secure-phase (#2154 CI fix)
2026-07-13 00:47:15 -04:00
Tom Boucher
4bb846b67a fix(#2112): scope commit to --files pathspec, not entire index (#2148)
* fix(2112): scope commit to --files pathspec, not entire index

cmdCommit/cmdCommitToSubrepo/cmdPrSubrepo staged exactly the files
named in --files but then ran a bare 'git commit' with no pathspec,
absorbing anything else in the index into a commit whose message
described only the named files (#2112).

Fix: append '-- ...stagedPaths' to the commit args when the caller
declared a scope. Three guards are load-bearing:
- stagedPaths (not filesToStage) excludes skipped missing files (#2014)
- explicitFiles gate keeps the default .planning/ path byte-identical
- MERGE_HEAD check via 'git rev-parse' falls back to bare commit during merge
- --amend is left without pathspec (different operation)

cmdPrSubrepo pathspec uses changedFiles (old+new for renames) so the
full rename is captured atomically.

Also fixes workflow markdown in spec-phase.md and add-tests.md.

All-files-missing now short-circuits to nothing_to_commit instead of
absorbing the entire index under a message describing files that
were not committed.

* docs(changeset): backfill PR number (#2148)

* test: update golden-install-parity fixtures for workflow markdown changes (#2112)

* test: update golden fixtures + workflow baselines for #2112 changes

- claude-local.json golden fixture (now generated via gen script)
- workflow-size-baseline.json (add-tests.md +16, spec-phase.md +42 bytes)
- Extended gen-golden-install-parity-zcode.cjs to also regenerate the
  claude local-layout fixture
2026-07-13 00:21:46 -04:00
Tom Boucher
481f00e38a fix(#2118): honor --dry-run in milestone complete with zero-mutation preview (#2155)
* fix(2118): honor --dry-run in milestone complete with zero-mutation preview

milestone complete treated --dry-run as a no-op: the flag was neither
parsed nor rejected, so a caller who expected a preview instead
triggered the full destructive mutation (archive phases → move audit
artifacts → rewrite STATE.md) with no way to back out.

Fix (option 2 from the issue): add dryRun to MilestoneCompleteOptions,
parse --dry-run in the dispatcher, and return a JSON preview plan
(would_archive, would_update) after the read-only stats gathering but
before any mutations. Also gated platformEnsureDir on !dryRun so the
archive directory is not created during preview.

3 regression tests: no-mutation happy path, --no-archive-phases combo,
and --force bypass combo.

* docs(changeset): backfill PR number (#2155)

* chore(#2118): regenerate pi/qwen golden fixtures after next merge

The milestone --dry-run fix changes gsd-core/bin/gsd-tools.cjs; qwen.json
(missed at authoring) and pi.json (added on next, never carried the fix)
were the only two runtime fixtures still holding the pre-fix hash. All
other runtimes already reflect the change. Regenerated via the sanctioned
gen-golden-install-parity script.

* fix(#2118): surface accomplishments in dry-run preview; fix --dry-run --raw

Orthogonal review findings on the --dry-run preview:
- The preview omitted the already-computed accomplishments (the primary
  MILESTONES.md content a real run writes); surface it as a top-level field,
  mirroring the real-run result.
- `--dry-run --raw` discarded the structured payload and printed the literal
  string "dry-run"; drop the raw-value arg so --raw emits the full preview
  JSON, matching the real-run output() call.
Adds tests: --dry-run --raw is parseable JSON, preview includes accomplishments,
and --dry-run --force is proven zero-mutation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 00:02:52 -04:00
Tom Boucher
f8c5c1590f fix(#2196): declare the debug session-manager spawn foreground + no-TaskOutput + recovery (#2227) 2026-07-12 19:47:42 -04:00
Tom Boucher
24c324cbfa fix(#2194): add Bash timeout guidance for prompt-fed reviewers in review.md
The Gemini, Claude, and Codex reviewer blocks invoked the CLIs with no explicit
timeout, so each inherited the host default (~2 min on Claude Code). A source-
grounded review of a large plan set takes ~570s (Codex xhigh) / ~525s (headless
Claude) — both exceed that window, so the lane is killed mid-review, its output
is empty, and the cross-AI review silently proceeds with fewer lanes. CodeRabbit
and OpenCode already documented a timeout; the four main lanes did not.

Add a shared timeout-guidance note directing a high Bash timeout (>= 900000;
1200000 for Codex xhigh / headless Claude), referencing BASH_MAX_TIMEOUT_MS for
the Claude Code host cap, and framing a slow-lane empty output as a timeout kill
(not the 0xc0000142 crash it gets misdiagnosed as) so operators re-run with more
time instead of diagnosing a CLI failure.

Closes #2194

Recaptures the 18 golden-install-parity fixtures + workflow-size baseline (only
the review.md entry changed in each; review is LARGE-tier, 47168 < 61440).
2026-07-12 18:53:49 -04:00
Tom Boucher
924ff6822e fix(#2138): surface track_shipping push failures (review)
Review (MEDIUM): `git push ... 2>&1` did not check exit code, so a silent push
failure (auth-token expiry, network blip, non-fast-forward) would proceed to the
report step and declare success — silently reproducing the exact #2138 defect.
Add a fallback warning naming the rerun command so a failed push is visible
(best-effort: the PR already exists, so we still report it rather than abort).
Recaptures goldens + size baseline.
2026-07-12 15:55:39 -04:00
Tom Boucher
aaf74878f7 fix(#2138): push the track_shipping ship-note onto the PR branch [ci skip]
track_shipping committed the STATE ship-note ('Phase N shipped — PR #N') AFTER
create_pr and never pushed it, so the commit stayed local-only. When the GitHub
PR merged (especially fast/auto-merge) the ship-note was not in the source branch
and never reached the default branch — STATE's ship-status was silently lost,
recoverable only by STATE self-heal on the next /gsd-start.

Push the ship-note commit onto the PR branch with a [ci skip] trailer. GitHub
honors [ci skip]/[skip ci], so this lands the note on merge without triggering a
redundant pipeline, and preserves the PR number in STATE.

Recaptures the 18 golden-install-parity fixtures + the workflow-size baseline
(only the ship.md entry changed in each).

Closes #2138
2026-07-12 15:55:39 -04:00
Tom Boucher
a5110c4b8c Merge branch 'next' into fix/2133-fast-md-log-to-state-schema-gate 2026-07-12 12:42:15 -04:00
Tom Boucher
c6b4304bab test(#2133): regenerate golden + workflow-size baselines for fast.md
The fast.md log_to_state fix changes an installed workflow file, so the per-
runtime golden-install-parity fixtures (18 runtimes) and the workflow-size
baseline are recaptured. Diff is exactly one entry per fixture (the fast.md
hash) and one baseline byte count — no spurious drift.
2026-07-12 10:48:43 -04:00
Tom Boucher
1153eefedd Merge branch 'next' into fix/2116-surface-bare-require 2026-07-12 10:33:02 -04:00
Tom Boucher
64c4a105f6 fix(#2116): use resolvable paths in surface.md require + correct package name
Four require() examples in commands/gsd/surface.md used bare
'gsd-core/...' specifiers that Node cannot resolve (wrong package
name + runtime-mirror layout off module path). Now derives the path
from runtimeConfigDir. Also fixes reinstall hint from
'npm i -g gsd-core' to 'npm i -g @opengsd/gsd-core'.
2026-07-12 01:32:12 -04:00
Tom Boucher
14b755d928 Merge branch 'next' into fix/2117-audit-milestone-not-validated 2026-07-12 00:03:52 -04:00
Tom Boucher
a0fafedfa0 feat(#2103): drive VS Code through the Embeddable Orchestration System (ADR-1239)
VS Code is a net-new EoS runtime that — unlike every prior migration — is NOT
CLI-installed (Marketplace/VSIX extension). It has zero runtime==='vscode'
branches in bin/install.js and stays that way (regression-guarded); it is driven
entirely through the negotiated imperative Host-Integration adapter.

Registry + validator (the hard part):
- capabilities/vscode/capability.json (role:runtime): full hostIntegration block
  (imperative / palette / active vscode.lm model / engine hook bus /
  sandboxed-storage / mcp transport / sandboxed-web runtime; dispatch nested,
  maxDepth 5 per VS Code's documented subagent depth).
- capability-validator.cjs extended so a role:runtime capability can legitimately
  declare "extension-distributed, no config directory": new configHome.kind:'none'
  + installSurface:'none' (+ GATE-A pairing + the parity maps), with localConfigDir
  and configHome.name made conditional on kind!=='none'. All 18 runtimes still
  validate; getDirName returns a distinct sentinel (not '.claude') for a no-config
  runtime.
- The add-a-registry-runtime tax: NON_INSTALLABLE_RUNTIMES exemption in the
  runtime-flags drift guard, vscode added to global-config-home SPECIAL_CASED,
  EXPECTED_PROFILES.vscode='ide', and the config-adapter/derivation/pin-count
  guards updated. No golden-install fixture, model-catalog, or CONFIGURATION rows
  (vscode never enters allRuntimes).

Dispatch + extension surface:
- Fixed vscode/extension.js's createHub()-no-args bug (every dispatch was
  UnknownCommand, masked by a vacuous reachability test) — now reuses the shared
  dispatchGsdCommand subprocess-shim (Node/desktop); the reachability test is
  tightened to assert real dispatch.
- Promoted the #1933 host binding to a shipped vscode/host-binding.js; activate()
  now composes the model/hookBus/stateIO seams through it. Corrected the model
  seam to VS Code's real API (vscode.lm.selectChatModels() -> model.sendRequest();
  vscode.lm.sendRequest does not exist) so the binding actually composes on real
  desktop VS Code instead of throwing.
- New vscode/browser.js Web Extension entry with ZERO Node APIs (the engine's
  config/capability loading is Node-bound, so the web entry registers the surface
  and directs full dispatch to the native MCP server — honestly documented).
- UPGRADE 1: GSD skills as native Language Model Tools (contributes.languageModelTools
  + vscode.lm.registerTool), invoke() dispatching through the hub.
- UPGRADE 2: native subagent dispatch wired onto #runSubagent /
  chat.subagents.allowInvocationsFromSubagents (fail-soft on API availability,
  maxDepth 5 enforced).
- vscode/package.json: browser entry, engines.vscode ^1.105, chatParticipants +
  languageModelTools contributions; fixed a stale activationPoints->activationEvents
  manifest key. Added "vscode" to the package files array.

Docs (## vscode matrix section) + changeset (Added).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 23:24:42 -04:00
Tom Boucher
b55a2e7655 fix(#2117): distinguish not-yet-validated phase from validated failure in audit-milestone
audit-milestone's Nyquist scan classified a phase from `nyquist_compliant`
alone, so a phase seeded by plan-phase but never run through validate-phase
read PARTIAL — identical to a phase that validated and genuinely failed. The
template's `status` field could discriminate the two, but no workflow ever
promoted it off `draft`, so it was dead.

Make `status` live and read it:
- validate-phase.md §6: set `status: validated` in both the create (State B)
  and update (State A) VALIDATION.md paths.
- audit-milestone.md §5.5: parse `status`; add a distinct NOT-VALIDATED bucket
  keyed on `status: draft`, gate COMPLIANT/PARTIAL on `status: validated`, and
  report `not_validated_phases` in the audit YAML.
- VALIDATION.md template: document the draft → validated lifecycle.

Tests & generated artifacts:
- Regression test folded into policy-138 (owning workflow-contract file);
  fail-first verified vs origin/next (0 matches pre-fix).
- Regenerate golden-install-parity fixtures cleanly: adds the previously-missed
  qwen.json and removes a contaminated `settings.local.json` entry that had
  leaked into claude-local.json (the harness excludes hook-config files).
- Correct a stale validate-phase.md workflow-size-baseline entry.

Closes #2117

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 22:47:51 -04:00
Tom Boucher
79d7657eff feat(#2102): make pi a first-class installable runtime + fix its dispatch (ADR-1239)
Net-new EoS/pi installable runtime — purely additive (no prior runtime==='pi'
branches). pi is a bun-runtime programmatic-CLI whose /gsd command is registered
by a native ExtensionAPI extension and dispatches through the embedded engine.

Stage 1 (install plumbing):
- capabilities/pi/capability.json: full hostIntegration descriptor (imperative /
  slash-programmatic / active-model / native-extension / bun) + hostBehaviors
  {nativePlugin, pluginOnlyInstall}.
- --pi flag + interactive-menu renumber (All 17->18); pi added to RUNTIME_FLAG_IDS,
  RUNTIME_LABELS, RUNTIME_META, allRuntimes/runtimeMap, model-catalog defaults.
- Install mirrors OpenCode: pi installs the gsd.cjs extension + the shared engine
  payload (gsd-core + scripts + config markers) + the shared hooks bundle (spawned
  by the extension at lifecycle events, like OpenCode's plugin). pluginOnlyInstall
  EXCLUDES declarative command/agent/skill markdown, which pi has no host-read
  surface for (its /gsd is programmatic). _installNativePluginIfDeclared (extracted
  from the opencode-family path) copies pi/gsd.cjs -> ~/.pi/agent/extensions/gsd.cjs
  (global) / .pi/extensions/ (local). pi added to package.json files.
- Golden: new pi.json (320 files: extension + engine + 27-file hooks bundle, no
  markdown); the 16 other fixtures + claude-local change only by the shared
  model-catalog hash line.

Stage 2 (real dispatch + upgrades):
- Shared dispatchGsdCommand() (shell-command-projection): bounded, no-throw
  subprocess-shim to gsd-tools.cjs (the only full-surface dispatch path; no
  in-process full-hub factory exists). Fixes pi/gsd.cjs's createHub()-no-args bug
  (every dispatch was UnknownCommand) AND the identical bug in mcp-server.cts's
  gsd_invoke_command, which a vacuous unknown-family-only test had masked (now has
  a real dispatch regression test).
- pi/gsd.cjs: /gsd handler now (args, ctx) - tokenizes (quote-aware, via the
  shipped hooks/lib/git-cmd.js) + dispatches real family/subcommand (not hardcoded
  query/help); gsd_invoke gets a TypeBox (JSON-schema-fallback) parameters schema +
  consumes params; getArgumentCompletions; before_provider_request active-model
  steering (fail-open on null resolution); functional session_start /
  before_agent_start / session_before_compact hook bridges (spawn the shipped GSD
  hook scripts).
- EXTENSION_EVENT_SURFACES.pi expanded from ['tool_call'] to the full 30-event
  vocabulary.

Docs (host-integration matrix + how-to) + changeset (Added).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 21:07:38 -04:00
Tom Boucher
bd613566cb feat(#2100): drive Windsurf through the EoS descriptor + wire Cascade's blocking hook bus (ADR-1239)
Fold all 10 residual isWindsurf branches in bin/install.js onto descriptor-driven
hostBehaviors (byte-parity — no fold changes any install output):
- 2 dead destructures dropped (uninstall, finishInstall); the dead
  `else if (isWindsurf)` legacy agent-loop arm removed (windsurf ∈
  _DESCRIPTOR_AGENTS_RUNTIMES → unreachable).
- skipSharedHooksInstall:true folds the two `!isWindsurf` shared-hooks exclusions.
- legacyDevinSkillsCleanup:true folds the `.devin`→`.windsurf` one-time cleanup gate.
- installsCommandBodiesForWorkflowDelegation:true folds the #1629 command-body copy
  (workflow-delegation target — load-bearing; local-install verified intact).
- verificationStyle:"windsurf-workflows" folds the workflow-count report.
- Corrected stale _LEGACY_SCAN_SUBDIR_NAMES + hooks-json manifest comments (cursor + windsurf).
Zero live runtime==='windsurf'/isWindsurf branches remain across bin/install.js,
install-engine.cts, surface.cts, runtime-artifact-conversion.cts (AC2 guard scans all four).

UPGRADE (Cascade hook bus): wire GSD's write/command safety guards into Windsurf's
native hook bus. New hooksSurface 'windsurf-hooks-json' (VALID_HOOKS_SURFACES 7→8, GATE A
profile-marker-only allowlist, the HooksSurface union) + writeWindsurfHooksJson
(Cursor-templated, Cascade's flat {hooks:{<event>:[{command}]}} shape) writing
.windsurf/hooks.json with two BLOCKING pre-hooks:
- pre_write_code → gsd-windsurf-pre-write.js: blocks writes to a file outside the
  active git worktree / into .git internals.
- pre_run_command → gsd-windsurf-pre-command.js: conservative destructive-command
  deny-list (rm -rf of root/home incl. sudo/env/path-prefixed forms; fork bombs;
  force-push refspec forms — HEAD:main, +main, --force/-f — to main/master/next).
Both use Cascade's protocol (stdin JSON, exit 2 + stderr to block, exit 0 to allow,
fail-open on error/timeout). Tokenize-based classifier (no catastrophic-backtracking regex;
4096-char cap) with the fail-closed false-positives fixed post-review.

The 4 advisory GSD guards + pre_mcp_tool_use + 5 post_* logging events are deliberately
NOT wired: Cascade has no context-injection channel for advisory hooks and GSD has no MCP
guard — porting them would be non-functional padding (documented; codebuddy #2098 / copilot
#2099 faithful-subset precedent). extendedHookEvents stays [].

Golden: the 2 guard scripts ship in the shared hook bundle (HOOKS_TO_COPY + the shared
managed-hooks-registry), exactly like cursor's 6 gsd-cursor-*.js scripts — so the 8
shared-bundle runtimes' fixtures gain the 2 inert windsurf scripts + the registry hash
(functionally inert for non-windsurf; the established cursor pattern). No install-output
change beyond that (the folds are byte-parity; skip-bundle runtimes untouched). New scripts
registered in managed-hooks-registry + build-hooks + INVENTORY. Tests: declarative-reference-
windsurf (adapter/axes/fail-closed + AC2 guard) + windsurf-hooks-bridge (live exit-2 blocking
+ allow/fail-open + ReDoS-bound + writer/reconcile/remove idempotency); VALID_HOOKS_SURFACES
pin updated to 8. Matrix hookBus delta + changeset (Changed). capability-registry regenerated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 16:04:24 -04:00
Tom Boucher
d1e9491fef feat(#2099): drive GitHub Copilot through the EoS descriptor + multi-event hook bus (ADR-1239)
Fold Copilot's residual runtime-literal branches onto descriptor-driven
hostBehaviors. Several issue premises were inaccurate (verified via research)
and deliberately NOT followed: writesSharedSettings/"legacy exclusion list"
(per-runtime descriptor data, copilot's false is correct); RUNTIME_CONTENT_DISPATCH.copilot
+ installSurface==='copilot-instructions' (already descriptor-driven);
extendedHookEvents (closed Claude/Gemini enum — hooks extended in code instead);
reapply ternary (already folded, kimi #2095).

Real folds (all byte-parity — golden byte-identical for every runtime):
- src/install-engine.cts + src/surface.cts: the two `.agent.md` filename cutovers
  (_copyStaged + _syncGsdDir) unified onto hostBehaviors.agentFileExtension via a
  new exported agentFileExtensionFor() accessor (kills the two-mechanism divergence).
- src/runtime-artifact-conversion.cts: applyAgentPathRewrites' copilot skip →
  hostBehaviors.noPathRewrite:true (antigravity #2096 precedent).
- bin/install.js uninstall: the two isCopilot cleanup branches → installSurface===
  'copilot-instructions' gate (symmetric with install-time).
- bin/install.js: `!isCopilot` in the two skipSharedHooksInstall checks →
  hostBehaviors.skipSharedHooksInstall:true (copilot has no shared gsd-*.js hooks).
- bin/install.js: three dead legacy inline-agent-loop isCopilot refs removed
  (copilot ∈ _DESCRIPTOR_AGENTS_RUNTIMES → unreachable; byte-parity proven by clean
  golden + real reachable-runtime install diffs). isCopilot dropped from 4 destructures.
Zero live `runtime==='copilot'`/`isCopilot` branches remain in bin/install.js,
install-engine.cts, surface.cts, or runtime-artifact-conversion.cts (AC2 guard scans all four).

UPGRADE 1 (multi-event hook bus): buildCopilotHookConfig() now emits preToolUse/
postToolUse/userPromptSubmitted/sessionEnd advisory handlers alongside sessionStart
(static inline bash/powershell — deterministic, golden-trackable). Only copilot.json's
gsd-session.json hash changes.

UPGRADE 2 (background dispatch): surfaced via the negotiated contract only —
dispatch.background:true exceeds the declarative-cli baseline and survives negotiation
with no downgrade warning. NO .agent.md frontmatter field (copilot has none). MCP
companion out of scope (AC4 names only 2 upgrades).

Tests: declarative-reference-copilot (adapter/axes/fail-closed + AC2 4-file source-grep
guard) + copilot-upgrades (live 5-event hook wiring; dispatch.background negotiation).
Matrix EoS note + how-to; changeset (Changed). capability-registry regenerated.

Incidental flaky-test RE-ARCHITECTURE (no-defer, maintainer-directed):
tests/opencode-review-reconstruction.property.test.cjs spawned ~600 synchronous
execFileSync('jq') subprocesses (numRuns:200 × 3 fast-check properties, one jq per
generated stream); a single jq freezing on a contended macos-22 CI runner hung the whole
unit-test chunk to its 600s kill (this PR's CI). --test-force-exit can't interrupt a
synchronous execFileSync, so the cure is to stop spawning per case, not just time-bound
it. Re-architected to run the SHIPPED jq program over the whole fast-check corpus in ONE
jq process: each generated stream is one compact-JSON array per line in a temp file,
`jq -c <PROGRAM>` (no -s) applies PROGRAM to each array (`.` == the array, exactly what
production's `jq -rs <file>` sees after slurping) and emits one result per line —
empirically byte-identical to the per-stream form across embedded-newline/empty/quote/
unicode/null-drop cases, and file-input (like production) so there's no stdin pipe to
deadlock on large I/O. ~600 spawns → 6; coverage unchanged (200-case corpus per property,
deterministic seeds) plus explicit boundary/diagnostic example batches. Still property-
tests the real shipped jq (no JS reimplementation). Per-call jq timeout retained as a
belt-and-suspenders bound.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 13:10:31 -04:00
Tom Boucher
5695522d5f feat(#2096): migrate Antigravity onto EoS declarative adapter + permission-writer + MCP companion (ADR-1239)
Fold all antigravity literal branches into descriptor-driven reads:
getConfigDirFromHome (→ configHome.kind 'dot-home-nested'), projectLocalHookPrefix
(→ hostBehaviors.hookPathStyle 'raw'), applyAgentPathRewrites (→ noPathRewrite),
getProjectInstructionFile (→ projectInstructionFile 'GEMINI.md'); removed the dead
inline convertClaudeAgentToAntigravityAgent branch + dead isAntigravity
destructures (antigravity is already on the descriptor-agents path). subagentToolkit
flipped undocumented→full (Context7: antigravity.google/docs/cli/features);
namedDispatch/nested/maxDepth/backgroundDispatch stay undocumented. Byte-identical
golden parity for all 16 runtimes.

UPGRADE 1 (permission-writer): permissionWriter 'antigravity' + configureAntigravityPermissions
merges a scoped permissions.allow block (GSD's own tree + hooks) into Antigravity's
settings.json — non-destructive, idempotent, symmetric uninstall. Added to
VALID_PERMISSION_WRITERS + the FinishPermissionWriter union.
UPGRADE 2 (MCP companion): configureAntigravityMcpConfig writes mcp_config.json
registering the gsd-core companion MCP server (Gemini-successor mcpServers schema,
best-effort — raw schema unpublished). Both writers dispatch from finishInstall.
settings.json is golden-excluded (HOOK_CONFIG_FILES); mcp_config.json (portable,
no absolute paths) is golden-tracked → only antigravity.json changes.

Tests: declarative-reference-antigravity extended (source-grep guard across 4
modules, fail-closed for the 4 undocumented sub-axes, validator acceptance) +
antigravity-upgrades (permission-writer + mcp_config live-install, idempotency,
user-preservation). Matrix + ADR-1016 + capability-manifest + CONTEXT.md +
connect-gsd-mcp-server docs updated; changeset (Changed).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 03:07:01 -04:00
Tom Boucher
ab04916682 feat(#2095): migrate Kimi CLI onto EoS imperative adapter + native hook-bus + background dispatch (ADR-1239)
Fold all runtime==='kimi'/isKimi logic branches into descriptor-driven
hostBehaviors (localInstallDeferred, verificationStyle, agentManifestStyle,
reapplyCommand, doneBannerStyle) + add 'kimi' to _DESCRIPTOR_AGENTS_RUNTIMES.
Kimi's skills/kimi-agents dispatch was already descriptor-driven (converter-by-
name + kimi-agents kind). Zero isKimi/runtime==='kimi' branches remain.

UPGRADE 1 (native hook bus): new hooksSurface 'kimi-hooks-toml' + a marker-
delimited config.toml [[hooks]] emitter (buildKimiHooksTomlBlock/writeKimiHooksToml
in runtime-hooks-surface.cts; resolveKimiHooksTomlDir in runtime-homes.cts).
GSD's lifecycle hooks now wire into Kimi's native ~/.kimi/config.toml (Context7-
confirmed path) at SessionStart/PreToolUse/Stop/PreCompact/SubagentStart/
SubagentStop — kimi becomes a hooks/ consumer (the 3 && !isKimi exclusion guards
removed). config.toml holds absolute install paths so it's golden-excluded via
an exact relative-path (.kimi/config.toml), not a basename (which would blind
Codex's config.toml). New hooksSurface value added to the closed enum in
capability-validator + runtime-config-adapter-registry.
UPGRADE 2 (background dispatch): flip dispatch.backgroundDispatch true (Kimi's
Agent tool takes run_in_background; root agent already gets the Agent tool), so
negotiation no longer flattens dispatch. subagentToolkit stays 'undocumented'
per AC (coder/explore/plan have distinct tool policies).
MCP transport explicitly deferred (no installer-driven MCP for any runtime).

Golden: only kimi.json changes (hooks/ scripts now installed); all 15 others +
claude-local byte-identical (kilo/zcode keep their own exclusions). Tests:
kimi-imperative-reference (adapter/axes/fail-closed/hostBehaviors + source-grep
guard) + kimi-upgrades (config.toml [[hooks]] SessionStart + marker idempotency
+ backgroundDispatch negotiation). CONTEXT.md glossary + matrix + how-to updated;
changeset (Added).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 00:53:34 -04:00
Tom Boucher
f744635b3d feat(#2094): migrate Trae onto EoS imperative adapter + SOLO stage-metadata upgrade (ADR-1239)
Fold trae logic branches into descriptor-driven reads: skipSharedHooksInstall
gates (dropped && !isTrae), and the case 'trae' path-rewrite arm now computes
the self-alias from the descriptor-driven dirName (.trae). Dead isTrae bindings
removed from uninstall/writeManifest/finishInstall. trae's skills dispatch was
already descriptor-driven (converter-by-name). RUNTIME_CONTENT_DISPATCH.trae is
left as a runtime-keyed table registration (its regex/callback rewrites can't be
a byte-identical descriptor map — matches cursor/windsurf/cline). trae stays in
RUNTIME_FLAG_IDS: isTrae still gates the agents-converter selection (agents out
of scope; removal gated on the cross-runtime agents-dispatch migration).
Byte-identical golden parity for all 16 runtimes.

UPGRADE: SOLO stage/trigger metadata — emitted Trae SKILL.md now carries
stage: workflow (descriptor-gated via hostBehaviors.soloStageMetadata) so
Trae's SOLO Agent can auto-invoke GSD skills at the corresponding stage. Field
shape is best-effort/inferred (Trae publishes no formal schema). trae.json
golden regenerated.

Tests: trae-imperative-reference (adapter/axes/fail-closed shouldFlattenDispatch
+ no runtime==='trae' source-grep, isTrae exempted for agents) + trae-upgrades
(stage: workflow on installed SKILL.md, descriptor-gated). Matrix note +
changeset added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 21:31:10 -04:00
Tom Boucher
f014ec83bd feat(#2093): migrate Kilo onto EoS imperative adapter + hook-bus/model/MCP/dispatch upgrades (ADR-1239)
Fold remaining isKilo logic branches into descriptor-driven reads:
finishPermissionWriter (uninstall cleanup), skipSharedHooksInstall (hooks
copy), and a skills converter-name registry (the artifactLayout.converter
field is now load-bearing, not decorative). frontmatterDialect stays the
documented dispatch key for frontmatter (no descriptor field for it). Dead
isKilo destructure bindings removed. Byte-identical golden parity for all 16
runtimes (opencode, which shares kilo's combined-family path, verified clean).

UPGRADE 1 (hook bus): install .kilo/plugins/gsd-core.js native plugin +
extensionEvents:"kilo" + EXTENSION_EVENT_SURFACES.kilo (OpenCode-fork bus).
UPGRADE 2 (active model): populate runtimeTierDefaults.kilo + thread
modelOverride through convertClaudeToKiloFrontmatter — model no longer stripped
from agents. UPGRADE 3 (MCP): document the gsd-core MCP companion under kilo's
mcp config key. UPGRADE 4 (named dispatch): agents/*.md mode:subagent roster is
the Task-tool dispatch surface (tested); subagentToolkit stays 'undocumented'
per AC so dispatch degrades to 'degraded' by design.

Model-catalog single-source edit ripples the shared model-catalog.json hash
into all 16 golden fixtures (expected). Inline defect fixes (no-defer): stale-
bake-guard resolveAgentDir 'agent'->'agents' (was a silent no-op for opencode/
codex), hardcoded 'Removed OpenCode plugin' uninstall log -> generic, and the
connect-gsd-mcp-server.md OpenCode mcpServers->mcp doc error.

Tests: kilo-imperative-reference (adapter/axes/fail-closed/degradation/
hostBehaviors + widened isKilo source-grep across 4 modules) + kilo-upgrades
(plugin parity+load, model-override converter, agents dispatch surface, MCP
doc). Matrix + how-to + config docs updated; changeset added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 19:55:15 -04:00
Tom Boucher
c6ce110efa feat(#2092): migrate Qwen Code onto EoS imperative adapter + native subagents + SubagentStart (ADR-1239)
Fold all runtime==='qwen'/isQwen logic branches (skill-priority frontmatter,
branding/path rewrites, legacy commands/gsd cleanup, hyphen-namespace
normalization, RUNTIME_CONTENT_DISPATCH, hooks-surface label) into
descriptor-driven runtime.hostBehaviors on capabilities/qwen/capability.json,
read via _hostBehaviors(). Shared claude/qwen/hermes legacy-migration branches
in install-engine.cts folded to descriptor flags (claude+hermes descriptors
updated; FALLBACK_HOST_BEHAVIORS.claude floored). Byte-identical golden parity
for qwen/hermes/claude(global+local).

UPGRADE 1: native .qwen/agents/*.md subagent projection — new agents
artifact-layout kind + convertClaudeAgentToQwenAgent converter (name +
description + tools YAML block list; color/model dropped). qwen routed onto
the descriptor-driven agents path (_DESCRIPTOR_AGENTS_RUNTIMES).
UPGRADE 2: SubagentStart hook wired into extendedHookEvents + the
descriptor-gated hook-writer loop (activates only for qwen).

Tests: qwen-imperative-reference (adapter/axes/fail-closed/hostBehaviors +
no runtime==='qwen' source-grep across 4 files) + qwen-upgrades (agents file
validity + SubagentStart mirrors SubagentStop, descriptor-gated). Docs matrix
+ how-to updated; changeset added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 15:44:19 -04:00
Dave
a8ff8fbb30 docs(#1867): specify propose-then-confirm --auto behavior in ui-phase Step 9.5 (review Minor) 2026-07-10 14:41:12 -04:00
Dave
31a500b970 docs(#1867): replace stale plan-phase.md:921 line-pointer with section-name reference (review #6) 2026-07-10 14:40:03 -04:00
Dave
20608cba90 chore(#1867): regen goldens + size baselines after merge onto next
next advanced to a3b9cbae (incl. #1820 specless-rail merge); regenerate
golden-install-parity fixtures + workflow-size-baseline authoritatively.
Manifest + agent baseline already in sync. plan-phase.md conflict
hand-merged to keep both EDGE_ABSENT/PROHIB_ABSENT and UI Considerations.
2026-07-10 14:40:03 -04:00
Tom Boucher
185abe2d66 feat(codex): advance Codex/OpenAI model defaults to GPT-5.6 (Sol/Terra/Luna)
Update runtimeTierDefaults.codex and providerPresets.openai in
model-catalog.json to the GPT-5.6 family (gpt-5.6-sol/terra/luna),
advancing from the superseded GPT-5.4/5.5 generation.

Model IDs verified against OpenAI developer API docs:
- gpt-5.6-sol: flagship, /, reasoning xhigh
- gpt-5.6-terra: balanced, .50/, reasoning medium
- gpt-5.6-luna: fast/cheap, /, reasoning medium

Tier mapping is 1:1 (Sol↔flagship, Terra↔balanced, Luna↔fast),
so profile semantics are unchanged — only the underlying IDs advance.

Updates: catalog JSON, test assertions (catalog defaults), docs
(CONFIGURATION.md + zh-CN/pt-BR translations, workflow settings),
and changeset.

Closes #2122
2026-07-10 12:11:55 -04:00
Tom Boucher
ae1bd14691 Merge branch 'next' into fix/2073-antigravity-reviewer-block 2026-07-09 18:55:02 -04:00
Tom Boucher
c68bca55cf test(#2089): regenerate all golden fixtures for managed-hooks-registry + cursor hook additions 2026-07-09 01:19:25 -04:00
Tom Boucher
303a796579 docs(changeset): #2089 cursor host-integration migration + golden fixture 2026-07-09 00:21:51 -04:00
Tom Boucher
6e773d97df feat(#2088): migrate Codex onto the Embeddable Orchestration System (ADR-1239)
Drive Codex install/uninstall through the descriptor-driven Host-Integration
Interface (declarative embedding adapter → engine surface dispatch) and fold
every positive `runtime === 'codex'` / `isCodex` projection into descriptor-driven
`runtime.hostBehaviors`. Install/uninstall output stays byte-parity-gated
(tests/fixtures/golden-install-parity/codex.json); no other runtime changes.

Three Context7-verified upgrades, each with a test on the user-reachable surface:
- Skill root → canonical $HOME/.agents/skills via a skills-kind `home` override,
  with pre-move migration cleanup (stale ~/.codex/skills/gsd-* removed on install
  and uninstall; user content preserved). Fixes getGlobalSkillsBase, writeManifest,
  and the skill-manifest inventory to honor the override so --skills-root /
  sync-skills / the manifest report the real location.
- Six new hooks.json lifecycle events (PreToolUse, PermissionRequest, PreCompact,
  PostCompact, SubagentStop, UserPromptSubmit) shared by install + uninstall;
  extendedHookEvents reconciled [] -> the schema-valid wired subset.
- Explicit `[agents] max_depth = 1` in the managed config.toml block, pinning the
  negotiated dispatch.maxDepth:1 axis. validateCodexConfigSchema now permits a
  known-scalar-only bare `[agents]` AgentsToml table (still rejects [[agents]] and
  unknown-key break-forms, #2760); mergeCodexConfig preserves the user's own
  AgentsToml scalars (max_threads etc.) instead of dropping them.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 21:42:11 -04:00
Tom Boucher
88d008553d Merge remote-tracking branch 'origin/next' into fix/2073-antigravity-reviewer-block 2026-07-08 20:00:25 -04:00
Tom Boucher
dbc730d8de fix(#2073): capability-probe external killer (timeout/gtimeout) for macOS
Code review (HIGH): a hardcoded 'timeout 600 agy' fails with rc 127 on stock
macOS (no GNU timeout/gtimeout), silently losing the agy reviewer. Probe for
'timeout'/'gtimeout' via command -v and fall back to agy's native --print-timeout
alone when neither exists (mirrors scripts/base64-scan.sh). External cap (600s)
stays >= --print-timeout (540s) so it only backstops a pre-session stall. Factor
the prompt into _AGY_PROMPT to avoid duplicating the long -p string across both
branches. Update the agy + #687 tests to assert the probe + bound + fallback,
regen the 17 goldens + size baseline, refresh the maintainer-note version stamp
to 1.0.16.
2026-07-08 17:53:28 -04:00
Tom Boucher
396f44bd0b feat(architecture): [EoS/opencode] Migrate OpenCode onto the Embeddable Orchestration System (ADR-1239, #2087)
Route OpenCode (and its Kilo sibling) through the public Host-Integration Interface and
land two Context7-verified capability upgrades. Byte-identical install output for all 16
runtimes (golden parity asserted).

Through the interface (AC2):
- OpenCode/Kilo's bespoke commands+skills+plugin install (the inline
  `else if (isOpencode || isKilo)` block) moves into the engine
  (installOpencodeFamilyCommands/Artifacts in src/install-engine.cts), dispatched by
  installRuntimeArtifacts when the descriptor declares hostBehaviors.combinedFamilyInstall.
  opencode/kilo now flow CLI -> _runtimeAdapter -> installRuntimeArtifacts like the skills
  runtimes. _isSkillsRuntime no longer excludes them; the bespoke block + dead
  copyFlattenedCommands are removed.
- Every hardcoded `runtime === 'opencode'`/`isOpencode` branch is folded into
  descriptor-driven runtime.hostBehaviors. ZERO `runtime === 'opencode'`/`'kilo'`
  string-equality remain in bin/install.js / install-engine.cts / runtime-artifact-conversion.cts.

Upgrades (AC4):
- Background dispatch: OpenCode shipped experimental background subagents in v1.15 and
  made them default-on in v1.17 -> dispatch.background/backgroundDispatch flip to true;
  shouldFlattenDispatch(opencode) now returns false (behavioral change; type: Changed).
- Expanded event surface: the OpenCode plugin subscribes permission.asked/replied +
  session.error.

Tests: opencode-imperative-reference (adapter/profile, shouldFlattenDispatch pin,
fail-closed negotiate, hostBehaviors, AC2 source-guard) + extended plugin surface test.
Docs: capability matrix v1.15/v1.17 citations. Changeset (Changed). gitignore .memdb//.memtrace/.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 17:17:05 -04:00
Tom Boucher
e0f245a6b2 fix(#2073): regen claude-local golden; reword changeset (no product parenthetical) 2026-07-08 17:15:48 -04:00
Tom Boucher
6ae23ffc21 fix(#2073): supersede #687 contract; regen golden install-parity baselines
#687 encoded 'agy bounded ONLY by --print-timeout, no external killer' and
'inline -p "$(cat)"'. Documentation since then (see PR description) shows:
  * agy's own print-mode guidance pairs --print-timeout with an external
    terminal 'timeout' (it cannot fire pre-session);
  * agy gained --model in ~1.0.3 (#3782's 'no --model' note was correct then,
    stale now);
  * inline "$(cat)" overflows the exec arg list on a large review prompt.
Rewrite the #687 describe block to the new contract (file-reference prompt +
--print-timeout PAIRED with a >= external timeout + --model + discard-on-
nonzero), and regenerate the 16 golden install-parity fixtures (only the
review.md hash line changed per runtime).
2026-07-08 16:47:02 -04:00
Tom Boucher
ff6e928530 fix(#2086): normalize realpath temp root in golden parity manifest (macOS /private)
The claude LOCAL install resolves its config dir via realpath, which on macOS
prepends /private to the temp root and embeds it in projected agents/commands/
workflows (@ references). buildParityManifest normalized only `root` (/var/folders/…),
leaving the /private prefix on macOS while Linux has none — so the mac-generated
claude-local fixture failed the Linux CI leg (198 files). Normalize the realpath
form too; no-op for the global fixtures (literal --config-dir, never realpath-resolved).
Regenerated claude-local.json now matches the Linux hashes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 14:49:58 -04:00
Tom Boucher
fdd5e401eb feat(architecture): [EoS/claude] drive claude through the imperative adapter + descriptor-driven hostBehaviors (#2086)
Fold Claude Code's install/uninstall onto the Embeddable Orchestration System
(ADR-1239 Phase D). claude is GSD's tier-1 reference host, but its install path
was still driven by 13 hardcoded `runtime === 'claude'` string-equality branches
scattered across bin/install.js rather than the public Host-Integration Interface.

- Route install()/uninstall() through `createImperativeAdapter({runtime})` — the
  adapter delegates to the SAME installRuntimeArtifacts/uninstallRuntimeArtifacts
  engine calls, so output is byte-identical (proven pre/post, both scopes).
- Replace all 13 `runtime === 'claude'` / `runtime !== 'claude'` branches with
  descriptor-driven `runtime.hostBehaviors` lookups on capabilities/claude/
  capability.json (attributionSource, authorsCanonicalWorkflow, localInstallStyle,
  permissionsSchema, settingsFileByScope, sourceMarkerFile, agentFrontmatterExtensions,
  ownsClaudePaths, nativeModelAliases, skillsGlobalOnboarding). Behavior is
  identical; the brittle string-equality coupling (the add-a-host tax) is gone.
- Single-source the scattered literal 'claude' defaults/rosters behind DEFAULT_RUNTIME.
- Extend golden-install-parity to assert the claude LOCAL legacy layout is
  byte-identical too (AC1 "both scopes"); exclude the platform-varying
  settings.local.json (same reason settings.json is excluded).
- New tests/claude-imperative-reference.test.cjs: adapter kind, programmatic-cli
  profile, fail-closed negotiation on a corrupted/partial descriptor, and an AC2
  source guard that no `runtime === 'claude'` branch remains.

No user-visible install-output change (internal architecture only).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 13:31:53 -04:00
Alex V.
f15c25867d docs(#1578): rebase B2+B3 agent prompt changes 2026-07-08 12:57:49 +03:00
Tom Boucher
1ee00320c7 Merge branch 'next' into fix/2072-thread-model-into-assumptions-and-review-spawns 2026-07-07 22:03:47 -04:00
Tom Boucher
4483300253 fix(#2072): thread resolved model into routed-agent spawns (assumptions-analyzer, code-reviewer, code-fixer)
model_overrides / models.<phaseType> were silently inert for gsd-assumptions-analyzer,
gsd-code-reviewer, and gsd-code-fixer on Claude Code: resolveModelInternal honors them,
but the workflows spawned these agents with no model= param, so the resolved value
never reached the Agent tool and the agents inherited the session model — no warning.

Fix — thread each agent's resolved model at every spawn site (the established
plan-phase pattern; the architecture-consistent Claude mechanism, since 13 other
agents already thread their model):
- discuss-phase-assumptions.md: `resolve-model gsd-assumptions-analyzer --raw`
  → ANALYZER_MODEL, threaded.
- code-review.md + code-review-fix.md (re-review): `resolve-model gsd-code-reviewer --raw`
  → REVIEWER_MODEL, threaded.
- code-review-fix.md (both fixer spawns): `resolve-model gsd-code-fixer --raw`
  → FIXER_MODEL, threaded (same silently-inert bug, same file — folded in per review).
- quick.md review step: was reusing `{executor_model}` for gsd-code-reviewer (so the
  reviewer's own override was ignored); init.quick now resolves `reviewer_model`
  (gsd-code-reviewer) and the spawn threads it.

resolve-model --raw returns the bare model string (resolve-execution --raw would
return effort — wrong). The resolver maps these agents to phaseType discuss /
verification / execution, so models.<phaseType> apply too.

Scope: the three agents reachable from the two issue-named workflows + quick.md. The
wider systemic class (other agents in UNTOUCHED workflows with the same pattern) stays
documented on the issue for a maintainer-scoped structural decision (thread-at-source
vs embed-at-install like #2256), not widened here.

Docs: the stale "discuss — reserved, no subagent today" model-profile tables now list
gsd-assumptions-analyzer and the verification row includes gsd-code-reviewer, across
the English docs, the shipped gsd-core/references/model-profiles.md reference, and the
ja-JP / zh-CN / ko-KR / pt-BR locale mirrors.

Tests:
- tests/model-resolver.test.cjs: #2072 acceptance — model_overrides and
  models.discuss/verification/execution resolve for all three agents.
- tests/model-routing-spawn-threading.test.cjs: every spawn of the three agents threads
  a resolved model (fails pre-fix); a header-precise parity guard fails the suite if a
  new un-threaded spawn of any of them regresses.
All 16 golden-install-parity fixtures + the workflow size baseline regenerated for the
changed shipped files (4 workflows + the reference doc); bin/lib is excluded from parity.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 21:22:48 -04:00
Tom Boucher
6addeccd19 feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates
an external API/SDK/service can no longer seal without a decided coverage matrix.

- src/api-coverage.cts: deterministic detector (compound verb+noun signal +
  <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix
  parse/validate/render with field-length caps.
- check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as
  a token under .planning/phases/ only (traversal-neutralized); validates
  COVERAGE.md or blocks iff a strong integration signal is detected and no
  matrix exists; fail-closed when phases tree exists but phase unresolvable.
- capabilities/ai-integration: workflow.api_coverage_gate config key (default
  true), plan:pre contribution, blocking verify:pre gate. Data-driven.
- gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch.
- Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e.
  Code+security review findings fixed (stopword FP, scope containment, pipe/cap
  rejection, prompt-injection message hygiene).
- Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs.

Closes #1562
2026-07-07 15:11:12 -04:00
Tom Boucher
12cc1955b2 Merge branch 'next' into fix/1857-test-gate-watch-mode-timeout 2026-07-07 12:58:57 -04:00
Tom Boucher
603593d41d fix(#1857): test gates normalize to one-shot + bounded timeout (no watch-mode hang)
A GSD verification gate resolves a project's test command and runs it. vitest
defaults to WATCH mode in an interactive TTY — exactly where a user runs
`gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never
exited and the orchestrator waited indefinitely. Recovery needed the user to
manually prompt "something blocking?".

Fix — one shared helper + a bounded, surfacing timeout on the test-command gates:
- New pure module src/normalize-test-command.cts + `gsd-tools query
  normalize-test-command` verb: rewrites a resolved command to a best-effort
  one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`;
  a package-manager `test` script whose package.json runner is watch-vitest →
  `CI=true` prefix; handles `--dir`; already-one-shot commands unchanged —
  never double-flagged). Named `normalize-test-command` (not `test-*`) so the
  file does not match node --test's default `test-*` discovery glob.
- The three gates that HUNG or silently-continued route through that ONE helper
  and bound execution with `timeout $(config-get workflow.test_gate_timeout)`
  (new config key, default 600s): the regression gate (extracted to
  execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen —
  it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the
  audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124.
- verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so
  it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint
  on 124, staying under its frozen 40960-byte tier cap.

Security hardening (review): the normalizer only rewrites a runner named as a
standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never
mangled), is length-capped and uses only linear-time split-based scanning (no
super-linear backtracking on an adversarial `workflow.test_command`), and reads
package.json only when it is a regular file (never blocks on a FIFO via `--dir`).

Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the
schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors
workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores,
inventory manifest/index. All 16 golden-install-parity fixtures + workflow size
baseline regenerated for the changed shipped files; bin/lib is excluded from the
parity manifest.

Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security
hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route
through the shared helper + configured timeout + exit-124 watch-mode hint;
verify-phase asserted as normalize-only/already-bounded).
tests/execute-phase-active-flags.test.cjs repointed at the extracted step;
tests/planner-language-regression.test.cjs allowlist comment updated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 12:14:43 -04:00
Tom Boucher
2b806904e2 fix(#1821): stop copying dead hook scripts for Kilo and ZCode (hooksSurface:none)
Kilo and ZCode both declare `hooksSurface: 'none'` and have no plugin surface,
so the GSD installer staged lifecycle hook scripts (hooks/*.js, hooks/*.sh,
hooks/lib/) plus a `{"type":"commonjs"}` package.json marker into their config
dirs where nothing ever invokes them — dead weight (#1821).

The installer's two hook-copy guards at bin/install.js were still on the legacy
hardcoded runtime-name list and never excluded Kilo or ZCode. Add
`&& !isKilo && !isZcode` to both (isZcode added to the install() runtimeFlags
destructure).

OpenCode — which #1821 also named — is deliberately NOT excluded: since the
issue was filed, #1914 shipped a native OpenCode plugin (plugins/gsd-core.js)
that spawns those exact staged hooks via OpenCode's event bus and requires both
the hook scripts and the CommonJS package.json marker. Excluding OpenCode would
regress #1914, so its hooks stay live. The genuinely-dead cases are Kilo & ZCode.

- bin/install.js: add `&& !isKilo && !isZcode` to the hooks/dist copy guard and
  the hooks/lib copy guard; document the OpenCode-vs-Kilo/ZCode split.
- tests/install-minimal-hooks.test.cjs: regression test asserting Kilo and ZCode
  receive no gsd-*.js/.sh hooks or hooks/lib, while OpenCode keeps its hooks +
  #1914 plugin and Claude keeps its hooks (over-exclusion guard).
- tests/fixtures/golden-install-parity/{kilo,zcode}.json: drop the 21 hooks/*
  entries and the package.json marker they no longer receive (opencode unchanged).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 09:40:01 -04:00
Tom Boucher
847de596b8 fix: third-party capability skills surface correctly (#2045)
A skills-only role:feature third-party capability installed 'active' but its skills never reached the runtime surface, capability enable/set rejected it as 'unknown capability', and capability list disagreed with capability state. Three defects, fix shape 1b (teach resolveSurface, no on-disk linking):

D1 (materialization): resolveSurface built the surfaced skills Set only from the on-disk manifest; third-party cap skills live at ~/.gsd/capabilities/<id>/skills/ and never entered the Set -> surfaced:false. Fix: union registry.capabilityClusters values into the Set in the full-profile branch (idempotent for first-party, additive for third-party; prototype-pollution guarded).

D2 (enable/set unknown): setCapabilityState validated the capId against the static first-party registry instead of the composed overlay-aware registry. Fix: validate against loadRegistry({includeInstalled}) like capability-state.cts does.

D3 (list vs state): capability list derived status purely from ledger existence, never surface composition. Fix: add a 'surfaced' field to list rows sourced from the same resolver capability state uses, so list and state agree.

Regression coverage: tests/issue-2045-third-party-skills-surface.test.cjs asserts all five acceptance criteria (resolveSurface union, state surfaced:true, enable/set not-unknown, list/state agreement, first-party + unknown-id regression). gsd-test: 23916/23916 green (linux-node22 + linux-node24).
2026-07-07 08:25:33 -04:00
Tom Boucher
ea4378063f Merge branch 'next' into feat/1820-specless-predicate-rail 2026-07-07 07:49:07 -04:00
Tom Boucher
0d7c15badc Merge branch 'next' into fix/2003-capability-state-runtime-flag 2026-07-07 00:24:27 -04:00