The targeted CI lane runs changed files UNSHARDED; #2088 touched 13 install-heavy
test files that all landed in one chunk, blowing the 600s per-chunk backstop on
the slow Windows runner (pure slowness, not a leak — per run-tests.cjs's own
comment). Weight install*/codex-* files (~10x a unit file) toward the per-chunk
budget so they spread across chunks instead of clustering; light-file chunking is
unchanged (weight 1). Adds harness regression tests (heavy split vs light control).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Drive Codex install/uninstall through the descriptor-driven Host-Integration
Interface (declarative embedding adapter → engine surface dispatch) and fold
every positive `runtime === 'codex'` / `isCodex` projection into descriptor-driven
`runtime.hostBehaviors`. Install/uninstall output stays byte-parity-gated
(tests/fixtures/golden-install-parity/codex.json); no other runtime changes.
Three Context7-verified upgrades, each with a test on the user-reachable surface:
- Skill root → canonical $HOME/.agents/skills via a skills-kind `home` override,
with pre-move migration cleanup (stale ~/.codex/skills/gsd-* removed on install
and uninstall; user content preserved). Fixes getGlobalSkillsBase, writeManifest,
and the skill-manifest inventory to honor the override so --skills-root /
sync-skills / the manifest report the real location.
- Six new hooks.json lifecycle events (PreToolUse, PermissionRequest, PreCompact,
PostCompact, SubagentStop, UserPromptSubmit) shared by install + uninstall;
extendedHookEvents reconciled [] -> the schema-valid wired subset.
- Explicit `[agents] max_depth = 1` in the managed config.toml block, pinning the
negotiated dispatch.maxDepth:1 axis. validateCodexConfigSchema now permits a
known-scalar-only bare `[agents]` AgentsToml table (still rejects [[agents]] and
unknown-key break-forms, #2760); mergeCodexConfig preserves the user's own
AgentsToml scalars (max_threads etc.) instead of dropping them.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Route OpenCode (and its Kilo sibling) through the public Host-Integration Interface and
land two Context7-verified capability upgrades. Byte-identical install output for all 16
runtimes (golden parity asserted).
Through the interface (AC2):
- OpenCode/Kilo's bespoke commands+skills+plugin install (the inline
`else if (isOpencode || isKilo)` block) moves into the engine
(installOpencodeFamilyCommands/Artifacts in src/install-engine.cts), dispatched by
installRuntimeArtifacts when the descriptor declares hostBehaviors.combinedFamilyInstall.
opencode/kilo now flow CLI -> _runtimeAdapter -> installRuntimeArtifacts like the skills
runtimes. _isSkillsRuntime no longer excludes them; the bespoke block + dead
copyFlattenedCommands are removed.
- Every hardcoded `runtime === 'opencode'`/`isOpencode` branch is folded into
descriptor-driven runtime.hostBehaviors. ZERO `runtime === 'opencode'`/`'kilo'`
string-equality remain in bin/install.js / install-engine.cts / runtime-artifact-conversion.cts.
Upgrades (AC4):
- Background dispatch: OpenCode shipped experimental background subagents in v1.15 and
made them default-on in v1.17 -> dispatch.background/backgroundDispatch flip to true;
shouldFlattenDispatch(opencode) now returns false (behavioral change; type: Changed).
- Expanded event surface: the OpenCode plugin subscribes permission.asked/replied +
session.error.
Tests: opencode-imperative-reference (adapter/profile, shouldFlattenDispatch pin,
fail-closed negotiate, hostBehaviors, AC2 source-guard) + extended plugin surface test.
Docs: capability matrix v1.15/v1.17 citations. Changeset (Changed). gitignore .memdb//.memtrace/.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The #338 fail-safe commit added a comment containing the literal `runtime === 'claude'`
(explaining what the data lookup is NOT), which the AC2 source-grep test matched as a
false positive (the test read the whole file, prose included). Strip block/line comments
+ backtick spans before matching so the guard flags only LIVE code, and reword the
comment. CRLF-safe line-comment strip.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Reviewer (PR #2106, elevated): if capability-registry.cjs fails to load,
_hostBehaviors('claude') returned {} — silently routing a claude LOCAL install to
the repo-shared settings.json instead of the gitignored settings.local.json (#338),
skipping mergeClaudePermissions + the .gsd-source marker. The migration is what
introduced that registry dependency (pre-PR the path had none).
Add FALLBACK_HOST_BEHAVIORS (keyed by runtime id — a data lookup, not a
runtime==='claude' branch) mirroring the reference host's #338-privacy-critical keys
(settingsFileByScope, permissionsSchema, sourceMarkerFile), consulted only when the
registry (or the descriptor) is unavailable. Behavior degrades CLOSED, never open;
the live descriptor stays the source of truth. Normal (registry-present) output is
unchanged (golden parity preserved). Pinned by tests via a registry-injected
_resolveHostBehaviors helper.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The claude LOCAL install resolves its config dir via realpath, which on macOS
prepends /private to the temp root and embeds it in projected agents/commands/
workflows (@ references). buildParityManifest normalized only `root` (/var/folders/…),
leaving the /private prefix on macOS while Linux has none — so the mac-generated
claude-local fixture failed the Linux CI leg (198 files). Normalize the realpath
form too; no-op for the global fixtures (literal --config-dir, never realpath-resolved).
Regenerated claude-local.json now matches the Linux hashes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Fold Claude Code's install/uninstall onto the Embeddable Orchestration System
(ADR-1239 Phase D). claude is GSD's tier-1 reference host, but its install path
was still driven by 13 hardcoded `runtime === 'claude'` string-equality branches
scattered across bin/install.js rather than the public Host-Integration Interface.
- Route install()/uninstall() through `createImperativeAdapter({runtime})` — the
adapter delegates to the SAME installRuntimeArtifacts/uninstallRuntimeArtifacts
engine calls, so output is byte-identical (proven pre/post, both scopes).
- Replace all 13 `runtime === 'claude'` / `runtime !== 'claude'` branches with
descriptor-driven `runtime.hostBehaviors` lookups on capabilities/claude/
capability.json (attributionSource, authorsCanonicalWorkflow, localInstallStyle,
permissionsSchema, settingsFileByScope, sourceMarkerFile, agentFrontmatterExtensions,
ownsClaudePaths, nativeModelAliases, skillsGlobalOnboarding). Behavior is
identical; the brittle string-equality coupling (the add-a-host tax) is gone.
- Single-source the scattered literal 'claude' defaults/rosters behind DEFAULT_RUNTIME.
- Extend golden-install-parity to assert the claude LOCAL legacy layout is
byte-identical too (AC1 "both scopes"); exclude the platform-varying
settings.local.json (same reason settings.json is excluded).
- New tests/claude-imperative-reference.test.cjs: adapter kind, programmatic-cli
profile, fail-closed negotiation on a corrupted/partial descriptor, and an AC2
source guard that no `runtime === 'claude'` branch remains.
No user-visible install-output change (internal architecture only).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The checkbox regex in cmdPhaseComplete used a greedy .* between ] and
'Phase N', so completing an already-checked phase (idempotent re-run)
matched a LATER phase whose description merely mentioned the target phase
number — checking the wrong phase's box. Restrict the gap to whitespace /
optional markdown bold emphasis, mirroring the tight pattern already used
by phase-insert.
The checkbox regex in cmdPhaseComplete uses a greedy .* between ] and
'Phase N'. Completing an already-checked phase (idempotent re-run) wrongly
matches a later phase whose description mentions the target phase. These
tests encode the exact repro from #2067; they fail on next @ origin/next.
Worker A (lock holder) used holdMs=1000 as a safety cap on its
Atomics.wait for the writer's contention signal. Under container load
(full suite, node24) Worker B's spawn + require('state.cjs') + stub
installation can exceed 1s, so A's wait timed out and removed the lock
before B ever contended. B's first atomic-create then succeeded
(lockAttempts:1), failing the retry-path witness and red-flagging the
gsd-test gate on otherwise-green branches.
The handshake is the real release trigger; holdMs is only a safety cap
for a dead/hung writer, so it must be large enough to never elapse
during B's spawn+init. Bump to 30s (bounded; afterEach terminate()s A
on the normal path, so no added latency) and widen the per-test timeout
to 15s for spawn headroom under heavy parallel load.
`gsd-tools effort sync` crashed in every installed runtime (e.g. ~/.claude/gsd-core/)
with `Cannot find module '../../../bin/install.js'`: cmdEffortSync (src/commands.cts)
required the package-root bin/install.js for its install-time effort resolvers, but the
installer only copies the gsd-core/ subtree into a runtime home — bin/install.js is never
present there. So `effort` config changes silently never reached installed agents without
a full reinstall (exactly the gap #488 was meant to close). 4th instance of the recurring
"runtime code under gsd-core/ requires a file outside the shipped subtree via ../../../"
anti-pattern (#1223/#1920/#1383 were the prior three, all already mitigated).
Fix (ADR-457 direction — extract, single source): move readGsdEffectiveEffortConfig +
resolveInstallTimeEffort (with their _getGsdEffortCatalog + _readGsdConfigFile helpers)
out of the hand-authored bin/install.js into a new src/install-effort-resolver.cts that
compiles into the shipped gsd-core/bin/lib/install-effort-resolver.cjs. commands.cts now
requires it as a sibling (`./install-effort-resolver.cjs`) — always present in the
installed tree — instead of `../../../bin/install.js`. bin/install.js imports the same four
symbols back from the new module (it still calls them + re-exports them), so there is one
source of truth and no duplication/drift. The lazy manifest read is repointed from the
package-root layout (`.., gsd-core, bin, shared`) to the bin/lib layout (`.., shared`).
Scope note: this is one of four instances of the anti-pattern; the other three are already
shipped/guarded. A build-time guard rejecting new cross-boundary requires whose target isn't
in the installer copy manifest (to prevent instance #5) is recommended on the issue but kept
out of this fix.
Tests: tests/effort-sync-installed-runtime.test.cjs does a real minimal install into a temp
home (the golden-parity helper) and runs the issue's exact repro
(`gsd-tools effort sync --config-dir <temp>`), asserting no MODULE_NOT_FOUND for
bin/install.js. Fail-first verified: against pristine next the same test throws
`Cannot find module '../../../bin/install.js'` at cmdEffortSync; post-fix it syncs cleanly.
New module registered in .gitignore (ADR-457), eslint ignores, docs/INVENTORY.md +
INVENTORY-MANIFEST.json. bin/install.js is not shipped and the new module is under bin/lib
(excluded from golden parity), so no golden fixtures change.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
model_overrides / models.<phaseType> were silently inert for gsd-assumptions-analyzer,
gsd-code-reviewer, and gsd-code-fixer on Claude Code: resolveModelInternal honors them,
but the workflows spawned these agents with no model= param, so the resolved value
never reached the Agent tool and the agents inherited the session model — no warning.
Fix — thread each agent's resolved model at every spawn site (the established
plan-phase pattern; the architecture-consistent Claude mechanism, since 13 other
agents already thread their model):
- discuss-phase-assumptions.md: `resolve-model gsd-assumptions-analyzer --raw`
→ ANALYZER_MODEL, threaded.
- code-review.md + code-review-fix.md (re-review): `resolve-model gsd-code-reviewer --raw`
→ REVIEWER_MODEL, threaded.
- code-review-fix.md (both fixer spawns): `resolve-model gsd-code-fixer --raw`
→ FIXER_MODEL, threaded (same silently-inert bug, same file — folded in per review).
- quick.md review step: was reusing `{executor_model}` for gsd-code-reviewer (so the
reviewer's own override was ignored); init.quick now resolves `reviewer_model`
(gsd-code-reviewer) and the spawn threads it.
resolve-model --raw returns the bare model string (resolve-execution --raw would
return effort — wrong). The resolver maps these agents to phaseType discuss /
verification / execution, so models.<phaseType> apply too.
Scope: the three agents reachable from the two issue-named workflows + quick.md. The
wider systemic class (other agents in UNTOUCHED workflows with the same pattern) stays
documented on the issue for a maintainer-scoped structural decision (thread-at-source
vs embed-at-install like #2256), not widened here.
Docs: the stale "discuss — reserved, no subagent today" model-profile tables now list
gsd-assumptions-analyzer and the verification row includes gsd-code-reviewer, across
the English docs, the shipped gsd-core/references/model-profiles.md reference, and the
ja-JP / zh-CN / ko-KR / pt-BR locale mirrors.
Tests:
- tests/model-resolver.test.cjs: #2072 acceptance — model_overrides and
models.discuss/verification/execution resolve for all three agents.
- tests/model-routing-spawn-threading.test.cjs: every spawn of the three agents threads
a resolved model (fails pre-fix); a header-precise parity guard fails the suite if a
new un-threaded spawn of any of them regresses.
All 16 golden-install-parity fixtures + the workflow size baseline regenerated for the
changed shipped files (4 workflows + the reference doc); bin/lib is excluded from parity.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The perf-316 state-lock test released the held lock on a fixed 1000ms timer, then
asserted the SUT had contended (lockAttempts >= 2). On a slow/loaded runner the
SUT worker's spawn+init exceeded the timer, so it acquired the lock on the first
try (lockAttempts:1) and the test failed intermittently (observed on linux-node22
while linux-node24 passed). The holder now releases only when the writer signals
its first FAILED lock attempt (via a shared SharedArrayBuffer + Atomics), so a
retry is guaranteed regardless of worker-spawn latency; holdMs becomes a safety
cap. Surfaced by the #2009 gsd-test run.
Previously a capability that failed to LOAD (e.g. incompatible engines.gsd) but
declared a gate-kind loop hook caused the loop resolver to inject a BLOCKING
synthetic gate (blocking:true, onError:halt) at every declared point, halting
every ship:pre / verify:post project-wide over an unrelated load error, with no
remediation surfaced.
Per maintainer decision (#2009) it now fails OPEN: no gate is injected (the loop
proceeds; --active-cap correctly reports the failed cap inactive) and a loud
warning is emitted — to stderr (the channel host workflows/agents actually see)
and in the envelope 'warnings' array — naming the load reason and the exact
'gsd capability remove <id>' remediation. The loader still records blockedGates;
only the consequence changes from block to warn.
Security (review): capId and reason originate from a third-party manifest /
directory name. capId is validated against the canonical kebab-case id shape
before it is placed in the runnable remediation command (withheld otherwise);
reason is stripped of control chars and backticks. This closes an argument/
prompt-injection vector in the surfaced message.
Docs updated to the fail-open-warning posture (ARCHITECTURE, INVENTORY,
CONFIGURATION, README, capability-overlay-model). Also removes a dead 'before'
import surfaced by lint in the issue-2045 test.
Asserts the loop resolver injects NO gate for a load-failed overlay capability
(fail open) and instead emits a loud warning — to stderr and the envelope
'warnings' array — carrying the 'gsd capability remove <id>' remediation, and
that an invalid (non-kebab) capability id is withheld from the runnable command.
Fails against origin/next (which injects a blocking fail-closed gate).
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates
an external API/SDK/service can no longer seal without a decided coverage matrix.
- src/api-coverage.cts: deterministic detector (compound verb+noun signal +
<Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix
parse/validate/render with field-length caps.
- check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as
a token under .planning/phases/ only (traversal-neutralized); validates
COVERAGE.md or blocks iff a strong integration signal is detected and no
matrix exists; fail-closed when phases tree exists but phase unresolvable.
- capabilities/ai-integration: workflow.api_coverage_gate config key (default
true), plan:pre contribution, blocking verify:pre gate. Data-driven.
- gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch.
- Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e.
Code+security review findings fixed (stopword FP, scope containment, pipe/cap
rejection, prompt-injection message hygiene).
- Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs.
Closes#1562
Two code-confirmed defects in `gsd-tools phase complete` (re-verified against
next; the three severe corruption paths the issue filed are superseded by the
ADR-1769 Transition Module migration + #2012, so this is the confirmed remainder).
1. Milestone-end mislabel (isLastPhase). The milestone-end determination only
cleared isLastPhase when a HIGHER-numbered phase existed, so completing the
numerically-highest phase out of order (e.g. Phase 10 before Phase 9) stamped
STATE.md `Status: Milestone complete` while a lower phase was still outstanding.
Added a lower-phase check: after the existing higher-phase scans, if any earlier
phase in the current milestone has an unchecked roadmap checkbox (`[ ]`),
isLastPhase becomes false AND next_phase/next_phase_name point at the LOWEST
outstanding lower phase — so STATE.md advances to the real gap instead of
parking on the just-completed phase. A completed phase always has `[x]`
(phase.complete sets it), so all-lower-complete still reports milestone-end;
heading-only roadmaps (no checkboxes) retain prior behavior. The checkbox regex
mirrors the sibling phasePattern's anchoring (whitespace/bold + required `:`) so
unrelated checklist lines mentioning "Phase N" don't match.
2. Workstream root-fallback (no guard). cmdPhaseComplete resolves every path via
planningDir(cwd); with a `workstreams/` dir present but no active workstream and
no --ws, that returns root `.planning`, so phase.complete wrote STATE.md/
ROADMAP.md (and the mislabel) into the shared root other workstreams read.
Added the same #1912 fail-safe guard init.progress got: refuse (asking for
`--ws`/active workstream) instead of silently writing root. Resolution itself
was already wired globally (resolveActiveWorkstream: --ws > GSD_WORKSTREAM >
pointer, set in bin/gsd-tools.cjs), so only the refusal guard was missing.
The workstream-mode detection (`listAvailableWorkstreams`) is extracted into
planning-workspace.cts as the single source of truth and consumed by BOTH
init.progress and phase.complete, so the two fail-safe paths cannot drift.
Tests (tests/phase.test.cjs, new #2028 describe): out-of-order completion becomes
`Ready to plan` with is_last_phase=false, next_phase pointing at the outstanding
phase and Current Phase advancing to it (not the completed phase); all-lower-
complete still reports milestone-end; workstream-mode-no-active refuses with an
`--ws` hint; `--ws` completes in the workstream leaving root untouched; flat mode
unaffected. Fail-first verified locally via direct gsd-tools invocation.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>