Commit Graph

285 Commits

Author SHA1 Message Date
github-actions[bot]
451e58e6d5 chore: sync next package version to 1.7.0-rc.5 2026-07-10 17:09:38 +00:00
Tom Boucher
185abe2d66 feat(codex): advance Codex/OpenAI model defaults to GPT-5.6 (Sol/Terra/Luna)
Update runtimeTierDefaults.codex and providerPresets.openai in
model-catalog.json to the GPT-5.6 family (gpt-5.6-sol/terra/luna),
advancing from the superseded GPT-5.4/5.5 generation.

Model IDs verified against OpenAI developer API docs:
- gpt-5.6-sol: flagship, /, reasoning xhigh
- gpt-5.6-terra: balanced, .50/, reasoning medium
- gpt-5.6-luna: fast/cheap, /, reasoning medium

Tier mapping is 1:1 (Sol↔flagship, Terra↔balanced, Luna↔fast),
so profile semantics are unchanged — only the underlying IDs advance.

Updates: catalog JSON, test assertions (catalog defaults), docs
(CONFIGURATION.md + zh-CN/pt-BR translations, workflow settings),
and changeset.

Closes #2122
2026-07-10 12:11:55 -04:00
Tom Boucher
b321bc04f4 fix(#2128): bound the sibling bracket-prefix clause — complete the ReDoS fix
Review caught that the prior commit bounded only the paren tag clause and left
the SIBLING bracket-prefix `(?:\[[^\]]+\]\s*)?` (same host regexes, before Phase)
UNBOUNDED — the identical quadratic reachable via a `[...]` run (measured ~16s at
1.7MB). Bound `[^\]]+`/`[^\]]*` -> {1,200}/{0,200} across all 19 phase/milestone
heading prefixes. Comprehensive re-measurement now shows EVERY vector linear
(bracket/paren/id/name/milestone all ~2-44ms at 2.45MB; bracket scaling
2k->2ms, 4k->5ms, 8k->10ms). Also: update the #1729 literal-mirror parity test
off its stale unbounded constant, and add limit-1 (199) boundary coverage.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 09:44:49 -04:00
Tom Boucher
07dcd85c98 fix(#2091): regenerate capability-registry from clean origin/next (includes cline hostBehaviors from #2090)
The previous registry was generated from a stale main-repo state that was
missing cline's hostBehaviors (merged in #2090). CI's lint:generated-sync
detected the staleness. Regenerated from clean origin/next + hermes changes.
2026-07-09 23:19:44 -04:00
Tom Boucher
f152e9a069 feat(#2091): migrate hermes onto EoS imperative adapter + extensionEvents dialect (ADR-1239)
- Fold 7 hardcoded isHermes/runtime === 'hermes' branches in bin/install.js into
  descriptor-driven _hostBehaviors lookups (skillFrontmatterVersion,
  skillsManifestPrefix, trackCategoryDescription, writeCategoryDescription,
  reportSkillsCount, legacyCommandsGsdCleanup, brandingRewrites)
- Add runtime.hostBehaviors block to capabilities/hermes/capability.json
- Register EXTENSION_EVENT_SURFACES.hermes (13 real plugin hook events) in
  src/host-integration.cts — replaces the borrowed hookEvents:'claude' 6-event
  surface that silently never fired on Hermes
- Add extensionEvents:'hermes' to the descriptor
- Add 'hermes' to VALID_EXTENSION_EVENTS in capability-validator.cjs
- Regenerate capability-registry.cjs
- Tests: hermes-imperative-reference (negotiation, axes, fail-closed, source-guard),
  hermes-dispatch-upgrade (dispatch posture, degradation, fail-closed)
- Changeset + docs update
2026-07-09 20:47:35 -04:00
Tom Boucher
760eb71b6b feat(#2090): migrate cline onto imperative adapter + beforeTool/createAgentModel upgrades
Fold all hardcoded runtime === 'cline' / isCline branches in bin/install.js
into descriptor-driven runtime.hostBehaviors lookups (reapplyCommand,
frontmatterDialect, skipSharedHooksInstall, localTargetIsProjectRoot,
clineRulesSurface, localCommandsViaRules). Add cline-sdk-binding adapter
(ADR-1239 Phase D): UPGRADE 1 re-implements the .clinerules/hooks/PreToolUse
guard as a real AgentPlugin.hooks.beforeTool handler (fail-open, same
semantics); UPGRADE 2 wires DefaultGateway.createAgentModel params from
model_overrides/model_profile_overrides resolution (modelMode: active).
Install output is byte-identical (golden parity asserted for cline +
claude/cursor/codex/opencode).
2026-07-09 19:38:27 -04:00
Tom Boucher
ae1bd14691 Merge branch 'next' into fix/2073-antigravity-reviewer-block 2026-07-09 18:55:02 -04:00
Tom Boucher
b0d985ccb3 feat(#2089): migrate cursor onto imperative adapter + hook-bus/dispatch upgrades 2026-07-09 00:21:39 -04:00
Tom Boucher
6e773d97df feat(#2088): migrate Codex onto the Embeddable Orchestration System (ADR-1239)
Drive Codex install/uninstall through the descriptor-driven Host-Integration
Interface (declarative embedding adapter → engine surface dispatch) and fold
every positive `runtime === 'codex'` / `isCodex` projection into descriptor-driven
`runtime.hostBehaviors`. Install/uninstall output stays byte-parity-gated
(tests/fixtures/golden-install-parity/codex.json); no other runtime changes.

Three Context7-verified upgrades, each with a test on the user-reachable surface:
- Skill root → canonical $HOME/.agents/skills via a skills-kind `home` override,
  with pre-move migration cleanup (stale ~/.codex/skills/gsd-* removed on install
  and uninstall; user content preserved). Fixes getGlobalSkillsBase, writeManifest,
  and the skill-manifest inventory to honor the override so --skills-root /
  sync-skills / the manifest report the real location.
- Six new hooks.json lifecycle events (PreToolUse, PermissionRequest, PreCompact,
  PostCompact, SubagentStop, UserPromptSubmit) shared by install + uninstall;
  extendedHookEvents reconciled [] -> the schema-valid wired subset.
- Explicit `[agents] max_depth = 1` in the managed config.toml block, pinning the
  negotiated dispatch.maxDepth:1 axis. validateCodexConfigSchema now permits a
  known-scalar-only bare `[agents]` AgentsToml table (still rejects [[agents]] and
  unknown-key break-forms, #2760); mergeCodexConfig preserves the user's own
  AgentsToml scalars (max_threads etc.) instead of dropping them.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 21:42:11 -04:00
Tom Boucher
88d008553d Merge remote-tracking branch 'origin/next' into fix/2073-antigravity-reviewer-block 2026-07-08 20:00:25 -04:00
Tom Boucher
dbc730d8de fix(#2073): capability-probe external killer (timeout/gtimeout) for macOS
Code review (HIGH): a hardcoded 'timeout 600 agy' fails with rc 127 on stock
macOS (no GNU timeout/gtimeout), silently losing the agy reviewer. Probe for
'timeout'/'gtimeout' via command -v and fall back to agy's native --print-timeout
alone when neither exists (mirrors scripts/base64-scan.sh). External cap (600s)
stays >= --print-timeout (540s) so it only backstops a pre-session stall. Factor
the prompt into _AGY_PROMPT to avoid duplicating the long -p string across both
branches. Update the agy + #687 tests to assert the probe + bound + fallback,
regen the 17 goldens + size baseline, refresh the maintainer-note version stamp
to 1.0.16.
2026-07-08 17:53:28 -04:00
Tom Boucher
396f44bd0b feat(architecture): [EoS/opencode] Migrate OpenCode onto the Embeddable Orchestration System (ADR-1239, #2087)
Route OpenCode (and its Kilo sibling) through the public Host-Integration Interface and
land two Context7-verified capability upgrades. Byte-identical install output for all 16
runtimes (golden parity asserted).

Through the interface (AC2):
- OpenCode/Kilo's bespoke commands+skills+plugin install (the inline
  `else if (isOpencode || isKilo)` block) moves into the engine
  (installOpencodeFamilyCommands/Artifacts in src/install-engine.cts), dispatched by
  installRuntimeArtifacts when the descriptor declares hostBehaviors.combinedFamilyInstall.
  opencode/kilo now flow CLI -> _runtimeAdapter -> installRuntimeArtifacts like the skills
  runtimes. _isSkillsRuntime no longer excludes them; the bespoke block + dead
  copyFlattenedCommands are removed.
- Every hardcoded `runtime === 'opencode'`/`isOpencode` branch is folded into
  descriptor-driven runtime.hostBehaviors. ZERO `runtime === 'opencode'`/`'kilo'`
  string-equality remain in bin/install.js / install-engine.cts / runtime-artifact-conversion.cts.

Upgrades (AC4):
- Background dispatch: OpenCode shipped experimental background subagents in v1.15 and
  made them default-on in v1.17 -> dispatch.background/backgroundDispatch flip to true;
  shouldFlattenDispatch(opencode) now returns false (behavioral change; type: Changed).
- Expanded event surface: the OpenCode plugin subscribes permission.asked/replied +
  session.error.

Tests: opencode-imperative-reference (adapter/profile, shouldFlattenDispatch pin,
fail-closed negotiate, hostBehaviors, AC2 source-guard) + extended plugin surface test.
Docs: capability matrix v1.15/v1.17 citations. Changeset (Changed). gitignore .memdb//.memtrace/.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 17:17:05 -04:00
Tom Boucher
af9069f865 fix(#2073): harden agy reviewer block (arg overflow, 404, pre-session stall)
Three failure modes on agy 1.0.16, all fixed by mirroring the Cursor block's
invocation discipline:
  * file-reference prompt instead of inline "$(cat)" — a large review prompt
    overflowed the exec arg list (rc 126).
  * external 'timeout 600' wrapper — --print-timeout cannot fire before agy
    creates a session, so a pre-session stall hung unbounded.
  * --model from review.models.agy when set — escape hatch for a pinned model
    that 404s (exit 0, empty stdout + transcript).
  * stdin </dev/null so agy never blocks on a tty.
Also enrich the Step 3 empty-output stub to grep agy cli.log for a
model-availability diagnostic, and correct the stale 'no --model flag' note
plus the 'review.models.agy reserved for future' comment (the config key was
already read but never passed through).
2026-07-08 16:47:02 -04:00
Tom Boucher
fdd5e401eb feat(architecture): [EoS/claude] drive claude through the imperative adapter + descriptor-driven hostBehaviors (#2086)
Fold Claude Code's install/uninstall onto the Embeddable Orchestration System
(ADR-1239 Phase D). claude is GSD's tier-1 reference host, but its install path
was still driven by 13 hardcoded `runtime === 'claude'` string-equality branches
scattered across bin/install.js rather than the public Host-Integration Interface.

- Route install()/uninstall() through `createImperativeAdapter({runtime})` — the
  adapter delegates to the SAME installRuntimeArtifacts/uninstallRuntimeArtifacts
  engine calls, so output is byte-identical (proven pre/post, both scopes).
- Replace all 13 `runtime === 'claude'` / `runtime !== 'claude'` branches with
  descriptor-driven `runtime.hostBehaviors` lookups on capabilities/claude/
  capability.json (attributionSource, authorsCanonicalWorkflow, localInstallStyle,
  permissionsSchema, settingsFileByScope, sourceMarkerFile, agentFrontmatterExtensions,
  ownsClaudePaths, nativeModelAliases, skillsGlobalOnboarding). Behavior is
  identical; the brittle string-equality coupling (the add-a-host tax) is gone.
- Single-source the scattered literal 'claude' defaults/rosters behind DEFAULT_RUNTIME.
- Extend golden-install-parity to assert the claude LOCAL legacy layout is
  byte-identical too (AC1 "both scopes"); exclude the platform-varying
  settings.local.json (same reason settings.json is excluded).
- New tests/claude-imperative-reference.test.cjs: adapter kind, programmatic-cli
  profile, fail-closed negotiation on a corrupted/partial descriptor, and an AC2
  source guard that no `runtime === 'claude'` branch remains.

No user-visible install-output change (internal architecture only).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 13:31:53 -04:00
Tom Boucher
1ee00320c7 Merge branch 'next' into fix/2072-thread-model-into-assumptions-and-review-spawns 2026-07-07 22:03:47 -04:00
Tom Boucher
4483300253 fix(#2072): thread resolved model into routed-agent spawns (assumptions-analyzer, code-reviewer, code-fixer)
model_overrides / models.<phaseType> were silently inert for gsd-assumptions-analyzer,
gsd-code-reviewer, and gsd-code-fixer on Claude Code: resolveModelInternal honors them,
but the workflows spawned these agents with no model= param, so the resolved value
never reached the Agent tool and the agents inherited the session model — no warning.

Fix — thread each agent's resolved model at every spawn site (the established
plan-phase pattern; the architecture-consistent Claude mechanism, since 13 other
agents already thread their model):
- discuss-phase-assumptions.md: `resolve-model gsd-assumptions-analyzer --raw`
  → ANALYZER_MODEL, threaded.
- code-review.md + code-review-fix.md (re-review): `resolve-model gsd-code-reviewer --raw`
  → REVIEWER_MODEL, threaded.
- code-review-fix.md (both fixer spawns): `resolve-model gsd-code-fixer --raw`
  → FIXER_MODEL, threaded (same silently-inert bug, same file — folded in per review).
- quick.md review step: was reusing `{executor_model}` for gsd-code-reviewer (so the
  reviewer's own override was ignored); init.quick now resolves `reviewer_model`
  (gsd-code-reviewer) and the spawn threads it.

resolve-model --raw returns the bare model string (resolve-execution --raw would
return effort — wrong). The resolver maps these agents to phaseType discuss /
verification / execution, so models.<phaseType> apply too.

Scope: the three agents reachable from the two issue-named workflows + quick.md. The
wider systemic class (other agents in UNTOUCHED workflows with the same pattern) stays
documented on the issue for a maintainer-scoped structural decision (thread-at-source
vs embed-at-install like #2256), not widened here.

Docs: the stale "discuss — reserved, no subagent today" model-profile tables now list
gsd-assumptions-analyzer and the verification row includes gsd-code-reviewer, across
the English docs, the shipped gsd-core/references/model-profiles.md reference, and the
ja-JP / zh-CN / ko-KR / pt-BR locale mirrors.

Tests:
- tests/model-resolver.test.cjs: #2072 acceptance — model_overrides and
  models.discuss/verification/execution resolve for all three agents.
- tests/model-routing-spawn-threading.test.cjs: every spawn of the three agents threads
  a resolved model (fails pre-fix); a header-precise parity guard fails the suite if a
  new un-threaded spawn of any of them regresses.
All 16 golden-install-parity fixtures + the workflow size baseline regenerated for the
changed shipped files (4 workflows + the reference doc); bin/lib is excluded from parity.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 21:22:48 -04:00
Tom Boucher
6addeccd19 feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates
an external API/SDK/service can no longer seal without a decided coverage matrix.

- src/api-coverage.cts: deterministic detector (compound verb+noun signal +
  <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix
  parse/validate/render with field-length caps.
- check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as
  a token under .planning/phases/ only (traversal-neutralized); validates
  COVERAGE.md or blocks iff a strong integration signal is detected and no
  matrix exists; fail-closed when phases tree exists but phase unresolvable.
- capabilities/ai-integration: workflow.api_coverage_gate config key (default
  true), plan:pre contribution, blocking verify:pre gate. Data-driven.
- gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch.
- Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e.
  Code+security review findings fixed (stopword FP, scope containment, pipe/cap
  rejection, prompt-injection message hygiene).
- Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs.

Closes #1562
2026-07-07 15:11:12 -04:00
Tom Boucher
603593d41d fix(#1857): test gates normalize to one-shot + bounded timeout (no watch-mode hang)
A GSD verification gate resolves a project's test command and runs it. vitest
defaults to WATCH mode in an interactive TTY — exactly where a user runs
`gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never
exited and the orchestrator waited indefinitely. Recovery needed the user to
manually prompt "something blocking?".

Fix — one shared helper + a bounded, surfacing timeout on the test-command gates:
- New pure module src/normalize-test-command.cts + `gsd-tools query
  normalize-test-command` verb: rewrites a resolved command to a best-effort
  one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`;
  a package-manager `test` script whose package.json runner is watch-vitest →
  `CI=true` prefix; handles `--dir`; already-one-shot commands unchanged —
  never double-flagged). Named `normalize-test-command` (not `test-*`) so the
  file does not match node --test's default `test-*` discovery glob.
- The three gates that HUNG or silently-continued route through that ONE helper
  and bound execution with `timeout $(config-get workflow.test_gate_timeout)`
  (new config key, default 600s): the regression gate (extracted to
  execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen —
  it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the
  audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124.
- verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so
  it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint
  on 124, staying under its frozen 40960-byte tier cap.

Security hardening (review): the normalizer only rewrites a runner named as a
standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never
mangled), is length-capped and uses only linear-time split-based scanning (no
super-linear backtracking on an adversarial `workflow.test_command`), and reads
package.json only when it is a regular file (never blocks on a FIFO via `--dir`).

Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the
schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors
workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores,
inventory manifest/index. All 16 golden-install-parity fixtures + workflow size
baseline regenerated for the changed shipped files; bin/lib is excluded from the
parity manifest.

Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security
hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route
through the shared helper + configured timeout + exit-124 watch-mode hint;
verify-phase asserted as normalize-only/already-bounded).
tests/execute-phase-active-flags.test.cjs repointed at the extracted step;
tests/planner-language-regression.test.cjs allowlist comment updated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 12:14:43 -04:00
Tom Boucher
847de596b8 fix: third-party capability skills surface correctly (#2045)
A skills-only role:feature third-party capability installed 'active' but its skills never reached the runtime surface, capability enable/set rejected it as 'unknown capability', and capability list disagreed with capability state. Three defects, fix shape 1b (teach resolveSurface, no on-disk linking):

D1 (materialization): resolveSurface built the surfaced skills Set only from the on-disk manifest; third-party cap skills live at ~/.gsd/capabilities/<id>/skills/ and never entered the Set -> surfaced:false. Fix: union registry.capabilityClusters values into the Set in the full-profile branch (idempotent for first-party, additive for third-party; prototype-pollution guarded).

D2 (enable/set unknown): setCapabilityState validated the capId against the static first-party registry instead of the composed overlay-aware registry. Fix: validate against loadRegistry({includeInstalled}) like capability-state.cts does.

D3 (list vs state): capability list derived status purely from ledger existence, never surface composition. Fix: add a 'surfaced' field to list rows sourced from the same resolver capability state uses, so list and state agree.

Regression coverage: tests/issue-2045-third-party-skills-surface.test.cjs asserts all five acceptance criteria (resolveSurface union, state surfaced:true, enable/set not-unknown, list/state agreement, first-party + unknown-id regression). gsd-test: 23916/23916 green (linux-node22 + linux-node24).
2026-07-07 08:25:33 -04:00
Tom Boucher
ea4378063f Merge branch 'next' into feat/1820-specless-predicate-rail 2026-07-07 07:49:07 -04:00
github-actions[bot]
a6263bba8a chore: sync next package version to 1.7.0-rc.4 2026-07-07 06:08:27 +00:00
Tom Boucher
0d7c15badc Merge branch 'next' into fix/2003-capability-state-runtime-flag 2026-07-07 00:24:27 -04:00
Tom Boucher
8a935a08a3 Merge branch 'next' into fix/2003-capability-state-runtime-flag 2026-07-07 00:19:08 -04:00
Tom Boucher
d7129222c0 Merge branch 'next' into codex/gsd-onboard 2026-07-06 23:49:49 -04:00
Tom Boucher
9b34fd5c08 Merge branch 'next' into fix/1941-quick-worktree-stale-base 2026-07-06 23:37:23 -04:00
Tom Boucher
ab82e73af3 fix(#2003): add --runtime override to capability state + loop render-hooks
resolveCapabilityRuntimeState derived the config dir from resolveRuntime(cwd)
(GSD_RUNTIME -> config.runtime -> 'claude') when no --config-dir was passed,
so a repo with persisted runtime:'codex' resolved the config dir to ~/.codex
where the Claude skill isn't installed -> surfaced:false / hooks silently
no-op when the operator drove from Claude Code. capability state and loop
render-hooks parsed only --config-dir, never --runtime, so there was no way
to assert the actually-active runtime.

Add a runtimeOverride param to resolveCapabilityRuntimeState (canonicalized
via runtime-name-policy so aliases like codex-app work); when present it
short-circuits the persisted-runtime fallback and resolves getGlobalConfigDir
for the explicit runtime. Thread --runtime through cmdCapabilityState and
cmdLoopRenderHooks, and parse it in gsd-tools.cjs for both commands (dual
--runtime X / --runtime=X form, mirroring --config-dir and the existing
capability-set --runtime precedent). Help text updated.

Without the override, behavior is byte-identical to today (regression-guarded).
2026-07-06 22:57:05 -04:00
Tom Boucher
080bacdb4b Merge branch 'next' into codex/gsd-onboard 2026-07-06 22:43:32 -04:00
Tom Boucher
f433db8b88 fix(#1143): address adversarial review — full semver precedence, docs/reality alignment
- compareSemver: implement full SemVer 2.0.0 §11 pre-release identifier
  comparison (two pre-releases of the same triple now order correctly; was 0).
- capability description + fragment: scope the plan-checker/verifier claim
  (this capability delivers the parallel-execution backend; those gates remain
  inline until separately wired). Correct the 'each wave is one barrier' prose
  (a wave splits into multiple sequential parallel() barriers on files_modified
  overlap). Frame detect-backend CLI as a simulation harness; the pure function
  with the live host descriptor is the real detection seam.
- partitionStages docstring: 'near-minimal via greedy first-fit' (not 'fewest');
  document empty-files_modified behavior.
2026-07-06 15:41:19 -04:00
Tom Boucher
e3262d94d3 feat(capabilities): add claude-orchestration capability (Workflow backend) (#1143)
Default-off, BETA, claude-only capability adopting Claude Code's Workflow tool
(/effort ultracode, Agent SDK >= v0.3.149) as an optional parallel-execution
backend for the GSD loop. Restores the wave parallelism + plan-checker + verifier
that #853 forces inline on Claude Code, and folds gsd-ultraplan-phase under one
runtime gate.

- Pure fail-closed core (src/claude-orchestration.cts): detectWorkflowBackend
  (gate ladder: enabled -> Claude -> backend != inline -> nested+background host
  -> valid Agent SDK -> SDK >= floor; every miss degrades to inline) and
  emitWorkflowScript (waves -> parallel() barriers, plans -> gsd-executor +
  worktree, files_modified overlap -> separate stages, resumeFromRunId, budget).
  All interpolated identifiers validated script-safe; briefs JSON-quoted.
- claude-orchestration command family (gsd-tools claude-orchestration
  detect-backend|emit-workflow) for orchestrator invocation.
- Two gated loop contributions at wired points (execute:wave:post, plan:post);
  federated config keys (enabled/execution_backend/min_agent_sdk_version).
- ADR-1143 implementation amendment; CONTEXT.md glossary entry; explanation doc.

On any runtime lacking the Workflow tool, behaviour is byte-identical to today.

closes #1143
2026-07-06 15:18:23 -04:00
Dave
7ef834cabc feat(#1820): spec-optional predicate rail — author probe predicates into must_haves when SPEC omits them 2026-07-06 14:40:38 -04:00
Joe Slitzker
3fe9a81428 fix(#1941): degrade /gsd-quick worktree dispatch when fork base is stale
Claude Code's isolation="worktree" forks new worktrees from origin/HEAD, not
the live local HEAD. When prior local commits (e.g. an earlier quick task in
the same session, or this task's own Step 5.6 pre-dispatch plan commit)
advance local HEAD without an intervening push, origin/HEAD stays pinned to a
stale ancestor and the executor's worktree_branch_check guard halts with a
base-mismatch fatal that can be many commits behind, not just one.

Port the worktree.base-check auto-degrade pattern already used by
execute-phase (#683/#1369) into quick.md's single-dispatch path, run
immediately before EXPECTED_BASE is captured in Step 6.
2026-07-06 13:01:47 -05:00
Tom Boucher
d671171698 feat(#1575): complete agent-converter descriptor cutover for copilot/antigravity + surface path parity
- Teach applySurface to build agentCtx (pathPrefix + attribution) and pass it
  to kind.stage() for agents kind, mirroring createRuntimeArtifactInstallPlan
  (ADR-1235 §1). Surface-path agents now receive path-rewrite + attribution +
  converter + normalize, matching install output byte-for-byte.

- Pass skills:'*' sentinel for agents staging when no surface state modifications
  exist, so ALL agents are staged (not just those referenced by _calls_agents_).

- Declare converted agents kind in copilot and antigravity capability.json;
  add to _DESCRIPTOR_AGENTS_RUNTIMES in bin/install.js.

- Handle copilot .agent.md filename rename in both _copyStaged (install path)
  and _syncGsdDir (surface path).

- Ship golden-parity harness (ADR-1235 §0): tests/issue-1575-agent-descriptor-
  parity.test.cjs asserts applySurface output is byte-identical to
  installRuntimeArtifacts for all 7 descriptor-driven runtimes, plus stale-
  cleanup convergence and prune data-loss coverage.

- Update ADR-1235 with cutover progress.

Cline remains deferred (rules-only local branch + local/global complication).
2026-07-06 11:34:32 -04:00
Tom Boucher
7c4494bddb chore(#1925): bump zcode descriptor version to 1.7.0-rc.3 (post-rebase sync) 2026-07-06 08:33:02 -04:00
Tom Boucher
7acbdc71b8 test(#1925): register zcode across installer surfaces + fluidify count pins
Resolve the zcode test cascade exposed by gsd-test:
- model-catalog.json: add zcode runtimeTierDefaults (null tiers, matching the
  other profile-marker-only runtimes) so KNOWN_RUNTIMES stays parity with allRuntimes.
- runtime-name-policy: add zcode to GLOBAL_CONFIG_HOME_FRAGMENTS (~/.zcode) so
  getGlobalConfigHomeFragment stops falling through to the .claude default.
- installer-migration-install: add the zcode fresh-install contract (flat-skills,
  no settings, no package.json — same shape as trae).
- golden-install-parity: capture the zcode fixture via a standalone generator
  script (not node --test — the gate stays gsd-test).
- capability-matrix.md: regenerate so zcode appears as a first-party row.

Fluidify the remaining count-pinned guards so adding a runtime no longer trips a
hand-pinned snapshot: gemini-runtime-removed (flag count derives from the
registry), non-claude-runtimes-registry-derivation (golden list derived from the
registry).
2026-07-06 08:31:52 -04:00
Tom Boucher
69b309e4e0 feat(#1925): add ZCode (Z.ai) as a pluggable runtime descriptor
Add ZCode as a first-party runtime via a declarative capability descriptor
(capabilities/zcode/capability.json) with zero hardcoded runtime === 'zcode'
branches — exercising the de-hardcoded, data-driven runtime path that 1.7.0
(ADR-1016 / ADR-1239) enables.

Descriptor (all axes sourced verbatim from zcode.z.ai docs):
- configHome ~/.zcode; nested skills + flat commands/agents; profile-marker install
- Claude-shaped skill format reuses convertClaudeCommandToClaudeSkill (no new converter)
- hostIntegration: declarative / slash-file / mcp / electron; dispatch background=false
  (foreground-only per docs); nested+maxDepth undocumented; passive model mode

Installer registration (data, not branches): --zcode flag, allRuntimes, runtimeMap,
interactive menu, --all list. getGlobalConfigDir + resolveRuntimeArtifactLayout +
ALLOWED_CONFIG_RUNTIMES + resolveInstallPlan all derive from the descriptor.

Revamped the brittle per-runtime golden-master tests to be count-agnostic,
descriptor-derived property tests (1.7.0 makes runtimes pluggable data, so pinning
frozen '15 runtime' snapshots is the wrong invariant): getdirname, label-policy,
config-adapter-registry (intent + install-plan golden master), capability-registry,
host-integration-descriptors (counts derive from curated maps). Adding a runtime
descriptor now extends coverage with zero edits to those suites.

Welcome banner, --zcode help, supported-runtimes how-to, and the host-integration
capability matrix (every axis cited) updated.
2026-07-06 08:31:51 -04:00
github-actions[bot]
a828d8207c chore: sync next package version to 1.7.0-rc.3 2026-07-06 02:19:50 +00:00
Tom Boucher
19fd9343f6 feat(#2002): self-healing runtime build for plugin-marketplace installs (#2036)
The compiled gsd-core/bin/lib/*.cjs modules are gitignored build artifacts
(ADR-457) shipped prebuilt in the npm tarball. A Claude Code plugin-marketplace
/ git-clone install materializes the repo tree directly and never runs
`npm run build:lib`, so those ~148 files are absent and every CLI command dies
at load with `Cannot find module './lib/cli-exit.cjs'`.

gsd-tools.cjs now calls ensureRuntimeBuild() before its ./lib requires: a new
committed bin helper that compiles the tree once on demand (lock-guarded,
incremental-cache-cleared, portable `node <tsc>` invocation) when the sentinel
cli-exit.cjs is missing, and surfaces an actionable error when TypeScript is
unavailable. The already-built npm path is a single fs.existsSync no-op.
ADR-457 (gitignored, build-at-publish) is preserved; no release-pipeline change.

Verified against a simulated marketplace checkout (148 build outputs removed):
the CLI self-heals and the result is byte-identical to `npm run build:lib`.

Closes #2002

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 20:56:12 -04:00
Tom Boucher
9f0d785b61 fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows (#2027)
* fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows

gsd-executor.md referenced sdk/src/query/QUERY-HANDLERS.md and reapply-patches.md
referenced sdk/dist/cli.js — both retired with the SDK (ADR-0174). AI runtimes
that locate doc refs via filesystem search ran find /, which on Git Bash for
Windows traverses the whole drive (14h+, orphaned find.exe, 4M+ handles).

- agents/gsd-executor.md: drop dead QUERY-HANDLERS.md ref.
- workflows/reapply-patches.md: drop dead sdk/dist/cli.js clause.
- tests/no-dead-sdk-refs.test.cjs: regression guard — no sdk/src|dist|handlers
  file refs in agents/workflows/references markdown.

Closes #2020

* docs(#2020): backfill changeset pr 2027
2026-07-05 16:43:26 -04:00
Tom Boucher
f272983d90 fix(#2019): planning-config.md global-learnings path ~/.gsd/learnings → ~/.gsd/knowledge (#2026)
* fix(#2019): planning-config.md global-learnings path ~/.gsd/learnings → ~/.gsd/knowledge

The features.global_learnings row pointed at ~/.gsd/learnings/ but the
implementation (src/learnings.cts, execute-phase.md) uses ~/.gsd/knowledge/.
Docs now match the code.

Closes #2019

* docs(#2019): backfill changeset pr 2026
2026-07-05 15:59:36 -04:00
Tom Boucher
1bfadec2d0 fix(#1921): preserve verify-work state across gap-closure + defer follow-ups (#2025)
* fix(#1921): preserve verify-work state across gap-closure + defer follow-ups

Resuming /gsd:verify-work after /gsd:execute-phase --gaps-only lost the
verification state: UAT ## Gaps still read status: failed even after their
fix plans executed, so verify-work re-diagnosed them as fresh blockers,
spawned a new gap plan, and reported only the new plan verified. A new
state contract links each gap to its fix plan so fixed gaps are recognized.

- verify-work.md gap YAML: stable gap_id (G-{phase}-{N}) per gap.
- plan_gap_closure planner prompt: each *-PLAN.md tags the gap_ids it
  addresses in its frontmatter (gap_closure: true, gap_ids: [...]).
- new reconcile_gaps step (run at resume_from_file entry): marks a gap
  status: resolved when its plan has a matching *-SUMMARY.md — fixed gaps
  are not re-diagnosed and do not spawn new gap plans; a re-reported break
  is treated as a fresh regression with a new gap_id.
- deferred-follow-up branch: a future-work idea (signals: 'later',
  'next version', 'out of scope', ...) is captured to UAT ## Deferred
  Follow-Ups instead of becoming a blocking gap/plan.
- workflow-size baseline recaptured (verify-work LARGE tier, 38221/61440).

Closes #1921

* docs(#1921): backfill changeset pr 2025

* fix(#1921): balance <step> tags in verify-work.md via 4-backtick outer fence

The plan_gap_closure step nested a ```yaml example inside a ``` block;
same-length fences made stripFencedCode close the outer block early,
swallowing the step's </step> (14 opens / 13 closes). Widen the outer
fence to 4 backticks so the nested yaml is contained. Regenerate golden
fixtures + size baselines.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 15:42:10 -04:00
Codesmith
425d2f2c8b fix(#1990): keep onboard resolver delegation in sync with canonical launcher
onboard.md delegates the gsd_run preamble to
gsd-core/references/gsd-run-resolver.md via @-include, but two guards
regressed once the canonical launcher snippet advanced to
${CLAUDE_CONFIG_DIR:-$HOME/.claude} (#2024):

- runtime-launcher-parity (B2): the resolver reference still shipped the
  old $HOME/.claude arm. references/ is not covered by
  sync-runtime-launcher.cjs, so refresh the resolver bash block to be
  byte-equal to _runtime-launcher.snippet.sh.
- /gsd:onboard command contract: sync-runtime-launcher.cjs had re-inlined
  the preamble into onboard.md (a delegating file). Teach the sync
  transform to strip-but-never-inline files that @-include the resolver,
  mirroring the exemption already in the parity test (B/B2).

Regenerate golden install fixtures and the workflow size baseline for the
smaller onboard.md.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-05 19:16:10 +00:00
Codesmith
4c673e51f3 chore(#1990): resync runtime launcher and regenerate artifacts after rebase onto next
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-05 19:16:10 +00:00
jeremymcs
7734dee051 fix(onboard): resolve Tests-lane failures for onboard workflow
- runtime-launcher-parity: recognize workflows that delegate gsd_run to references/gsd-run-resolver.md (onboard.md) and exempt them from the inline-preamble checks; add a compensating byte-equality guard (B2) asserting the reference bash block matches _runtime-launcher.snippet.sh. - onboard.md: document the TEXT_MODE plain-text/numbered-list fallback for AskUserQuestion on non-Claude runtimes (fixes ask-user-questions-fallback, #2012). - Regenerate golden-install-parity fixtures, docs/INVENTORY-MANIFEST.json (add gsd-run-resolver.md + onboard-projection.cjs), and tests/workflow-size-baseline.json.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-05 19:16:09 +00:00
Cursor Agent
83da2e1ca9 fix(onboard): enforce complete map gate in fast mode and add skip-ingest rerun
Remove the projectExists guard so fast mode with a partial map still routes
to complete-map-before-new-project even after project planning exists.

Add the missing onboard rerun instruction to the skip-mapping docs-ingest
handoff so the onboarding loop can continue after ingest.
2026-07-05 19:16:09 +00:00
Cursor Agent
ee2ddbae53 fix(onboard): add onboard rerun handoffs and correct summary next step
Add missing 'Then rerun onboard' instructions after new-project handoffs
so brownfield onboarding returns to create SUMMARY.md per REQ-ONBOARD-05.

Use handoff_commands.manager instead of next_action.reason in the summary
template so persisted SUMMARY.md recommends the correct post-onboarding step.
2026-07-05 19:16:09 +00:00
Jeremy McSpadden
2f2d33aea7 no-mistakes(review): Format onboard handoffs by runtime 2026-07-05 19:16:09 +00:00
Jeremy McSpadden
a845a7d231 no-mistakes(review): Preserve partial-planning skip guard 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
c9d964243a no-mistakes(review): Fix onboard fast-map and skip routing 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
cd5971fd3f no-mistakes(review): Fix onboard skip handoffs 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
a5298c1fc0 refactor: project onboard routing in init 2026-07-05 19:16:08 +00:00