Fold remaining isKilo logic branches into descriptor-driven reads:
finishPermissionWriter (uninstall cleanup), skipSharedHooksInstall (hooks
copy), and a skills converter-name registry (the artifactLayout.converter
field is now load-bearing, not decorative). frontmatterDialect stays the
documented dispatch key for frontmatter (no descriptor field for it). Dead
isKilo destructure bindings removed. Byte-identical golden parity for all 16
runtimes (opencode, which shares kilo's combined-family path, verified clean).
UPGRADE 1 (hook bus): install .kilo/plugins/gsd-core.js native plugin +
extensionEvents:"kilo" + EXTENSION_EVENT_SURFACES.kilo (OpenCode-fork bus).
UPGRADE 2 (active model): populate runtimeTierDefaults.kilo + thread
modelOverride through convertClaudeToKiloFrontmatter — model no longer stripped
from agents. UPGRADE 3 (MCP): document the gsd-core MCP companion under kilo's
mcp config key. UPGRADE 4 (named dispatch): agents/*.md mode:subagent roster is
the Task-tool dispatch surface (tested); subagentToolkit stays 'undocumented'
per AC so dispatch degrades to 'degraded' by design.
Model-catalog single-source edit ripples the shared model-catalog.json hash
into all 16 golden fixtures (expected). Inline defect fixes (no-defer): stale-
bake-guard resolveAgentDir 'agent'->'agents' (was a silent no-op for opencode/
codex), hardcoded 'Removed OpenCode plugin' uninstall log -> generic, and the
connect-gsd-mcp-server.md OpenCode mcpServers->mcp doc error.
Tests: kilo-imperative-reference (adapter/axes/fail-closed/degradation/
hostBehaviors + widened isKilo source-grep across 4 modules) + kilo-upgrades
(plugin parity+load, model-override converter, agents dispatch surface, MCP
doc). Matrix + how-to + config docs updated; changeset added.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Update runtimeTierDefaults.codex and providerPresets.openai in
model-catalog.json to the GPT-5.6 family (gpt-5.6-sol/terra/luna),
advancing from the superseded GPT-5.4/5.5 generation.
Model IDs verified against OpenAI developer API docs:
- gpt-5.6-sol: flagship, /, reasoning xhigh
- gpt-5.6-terra: balanced, .50/, reasoning medium
- gpt-5.6-luna: fast/cheap, /, reasoning medium
Tier mapping is 1:1 (Sol↔flagship, Terra↔balanced, Luna↔fast),
so profile semantics are unchanged — only the underlying IDs advance.
Updates: catalog JSON, test assertions (catalog defaults), docs
(CONFIGURATION.md + zh-CN/pt-BR translations, workflow settings),
and changeset.
Closes#2122
Drive Codex install/uninstall through the descriptor-driven Host-Integration
Interface (declarative embedding adapter → engine surface dispatch) and fold
every positive `runtime === 'codex'` / `isCodex` projection into descriptor-driven
`runtime.hostBehaviors`. Install/uninstall output stays byte-parity-gated
(tests/fixtures/golden-install-parity/codex.json); no other runtime changes.
Three Context7-verified upgrades, each with a test on the user-reachable surface:
- Skill root → canonical $HOME/.agents/skills via a skills-kind `home` override,
with pre-move migration cleanup (stale ~/.codex/skills/gsd-* removed on install
and uninstall; user content preserved). Fixes getGlobalSkillsBase, writeManifest,
and the skill-manifest inventory to honor the override so --skills-root /
sync-skills / the manifest report the real location.
- Six new hooks.json lifecycle events (PreToolUse, PermissionRequest, PreCompact,
PostCompact, SubagentStop, UserPromptSubmit) shared by install + uninstall;
extendedHookEvents reconciled [] -> the schema-valid wired subset.
- Explicit `[agents] max_depth = 1` in the managed config.toml block, pinning the
negotiated dispatch.maxDepth:1 axis. validateCodexConfigSchema now permits a
known-scalar-only bare `[agents]` AgentsToml table (still rejects [[agents]] and
unknown-key break-forms, #2760); mergeCodexConfig preserves the user's own
AgentsToml scalars (max_threads etc.) instead of dropping them.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Code review (HIGH): a hardcoded 'timeout 600 agy' fails with rc 127 on stock
macOS (no GNU timeout/gtimeout), silently losing the agy reviewer. Probe for
'timeout'/'gtimeout' via command -v and fall back to agy's native --print-timeout
alone when neither exists (mirrors scripts/base64-scan.sh). External cap (600s)
stays >= --print-timeout (540s) so it only backstops a pre-session stall. Factor
the prompt into _AGY_PROMPT to avoid duplicating the long -p string across both
branches. Update the agy + #687 tests to assert the probe + bound + fallback,
regen the 17 goldens + size baseline, refresh the maintainer-note version stamp
to 1.0.16.
Route OpenCode (and its Kilo sibling) through the public Host-Integration Interface and
land two Context7-verified capability upgrades. Byte-identical install output for all 16
runtimes (golden parity asserted).
Through the interface (AC2):
- OpenCode/Kilo's bespoke commands+skills+plugin install (the inline
`else if (isOpencode || isKilo)` block) moves into the engine
(installOpencodeFamilyCommands/Artifacts in src/install-engine.cts), dispatched by
installRuntimeArtifacts when the descriptor declares hostBehaviors.combinedFamilyInstall.
opencode/kilo now flow CLI -> _runtimeAdapter -> installRuntimeArtifacts like the skills
runtimes. _isSkillsRuntime no longer excludes them; the bespoke block + dead
copyFlattenedCommands are removed.
- Every hardcoded `runtime === 'opencode'`/`isOpencode` branch is folded into
descriptor-driven runtime.hostBehaviors. ZERO `runtime === 'opencode'`/`'kilo'`
string-equality remain in bin/install.js / install-engine.cts / runtime-artifact-conversion.cts.
Upgrades (AC4):
- Background dispatch: OpenCode shipped experimental background subagents in v1.15 and
made them default-on in v1.17 -> dispatch.background/backgroundDispatch flip to true;
shouldFlattenDispatch(opencode) now returns false (behavioral change; type: Changed).
- Expanded event surface: the OpenCode plugin subscribes permission.asked/replied +
session.error.
Tests: opencode-imperative-reference (adapter/profile, shouldFlattenDispatch pin,
fail-closed negotiate, hostBehaviors, AC2 source-guard) + extended plugin surface test.
Docs: capability matrix v1.15/v1.17 citations. Changeset (Changed). gitignore .memdb//.memtrace/.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
#687 encoded 'agy bounded ONLY by --print-timeout, no external killer' and
'inline -p "$(cat)"'. Documentation since then (see PR description) shows:
* agy's own print-mode guidance pairs --print-timeout with an external
terminal 'timeout' (it cannot fire pre-session);
* agy gained --model in ~1.0.3 (#3782's 'no --model' note was correct then,
stale now);
* inline "$(cat)" overflows the exec arg list on a large review prompt.
Rewrite the #687 describe block to the new contract (file-reference prompt +
--print-timeout PAIRED with a >= external timeout + --model + discard-on-
nonzero), and regenerate the 16 golden install-parity fixtures (only the
review.md hash line changed per runtime).
The claude LOCAL install resolves its config dir via realpath, which on macOS
prepends /private to the temp root and embeds it in projected agents/commands/
workflows (@ references). buildParityManifest normalized only `root` (/var/folders/…),
leaving the /private prefix on macOS while Linux has none — so the mac-generated
claude-local fixture failed the Linux CI leg (198 files). Normalize the realpath
form too; no-op for the global fixtures (literal --config-dir, never realpath-resolved).
Regenerated claude-local.json now matches the Linux hashes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Fold Claude Code's install/uninstall onto the Embeddable Orchestration System
(ADR-1239 Phase D). claude is GSD's tier-1 reference host, but its install path
was still driven by 13 hardcoded `runtime === 'claude'` string-equality branches
scattered across bin/install.js rather than the public Host-Integration Interface.
- Route install()/uninstall() through `createImperativeAdapter({runtime})` — the
adapter delegates to the SAME installRuntimeArtifacts/uninstallRuntimeArtifacts
engine calls, so output is byte-identical (proven pre/post, both scopes).
- Replace all 13 `runtime === 'claude'` / `runtime !== 'claude'` branches with
descriptor-driven `runtime.hostBehaviors` lookups on capabilities/claude/
capability.json (attributionSource, authorsCanonicalWorkflow, localInstallStyle,
permissionsSchema, settingsFileByScope, sourceMarkerFile, agentFrontmatterExtensions,
ownsClaudePaths, nativeModelAliases, skillsGlobalOnboarding). Behavior is
identical; the brittle string-equality coupling (the add-a-host tax) is gone.
- Single-source the scattered literal 'claude' defaults/rosters behind DEFAULT_RUNTIME.
- Extend golden-install-parity to assert the claude LOCAL legacy layout is
byte-identical too (AC1 "both scopes"); exclude the platform-varying
settings.local.json (same reason settings.json is excluded).
- New tests/claude-imperative-reference.test.cjs: adapter kind, programmatic-cli
profile, fail-closed negotiation on a corrupted/partial descriptor, and an AC2
source guard that no `runtime === 'claude'` branch remains.
No user-visible install-output change (internal architecture only).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
model_overrides / models.<phaseType> were silently inert for gsd-assumptions-analyzer,
gsd-code-reviewer, and gsd-code-fixer on Claude Code: resolveModelInternal honors them,
but the workflows spawned these agents with no model= param, so the resolved value
never reached the Agent tool and the agents inherited the session model — no warning.
Fix — thread each agent's resolved model at every spawn site (the established
plan-phase pattern; the architecture-consistent Claude mechanism, since 13 other
agents already thread their model):
- discuss-phase-assumptions.md: `resolve-model gsd-assumptions-analyzer --raw`
→ ANALYZER_MODEL, threaded.
- code-review.md + code-review-fix.md (re-review): `resolve-model gsd-code-reviewer --raw`
→ REVIEWER_MODEL, threaded.
- code-review-fix.md (both fixer spawns): `resolve-model gsd-code-fixer --raw`
→ FIXER_MODEL, threaded (same silently-inert bug, same file — folded in per review).
- quick.md review step: was reusing `{executor_model}` for gsd-code-reviewer (so the
reviewer's own override was ignored); init.quick now resolves `reviewer_model`
(gsd-code-reviewer) and the spawn threads it.
resolve-model --raw returns the bare model string (resolve-execution --raw would
return effort — wrong). The resolver maps these agents to phaseType discuss /
verification / execution, so models.<phaseType> apply too.
Scope: the three agents reachable from the two issue-named workflows + quick.md. The
wider systemic class (other agents in UNTOUCHED workflows with the same pattern) stays
documented on the issue for a maintainer-scoped structural decision (thread-at-source
vs embed-at-install like #2256), not widened here.
Docs: the stale "discuss — reserved, no subagent today" model-profile tables now list
gsd-assumptions-analyzer and the verification row includes gsd-code-reviewer, across
the English docs, the shipped gsd-core/references/model-profiles.md reference, and the
ja-JP / zh-CN / ko-KR / pt-BR locale mirrors.
Tests:
- tests/model-resolver.test.cjs: #2072 acceptance — model_overrides and
models.discuss/verification/execution resolve for all three agents.
- tests/model-routing-spawn-threading.test.cjs: every spawn of the three agents threads
a resolved model (fails pre-fix); a header-precise parity guard fails the suite if a
new un-threaded spawn of any of them regresses.
All 16 golden-install-parity fixtures + the workflow size baseline regenerated for the
changed shipped files (4 workflows + the reference doc); bin/lib is excluded from parity.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates
an external API/SDK/service can no longer seal without a decided coverage matrix.
- src/api-coverage.cts: deterministic detector (compound verb+noun signal +
<Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix
parse/validate/render with field-length caps.
- check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as
a token under .planning/phases/ only (traversal-neutralized); validates
COVERAGE.md or blocks iff a strong integration signal is detected and no
matrix exists; fail-closed when phases tree exists but phase unresolvable.
- capabilities/ai-integration: workflow.api_coverage_gate config key (default
true), plan:pre contribution, blocking verify:pre gate. Data-driven.
- gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch.
- Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e.
Code+security review findings fixed (stopword FP, scope containment, pipe/cap
rejection, prompt-injection message hygiene).
- Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs.
Closes#1562
A GSD verification gate resolves a project's test command and runs it. vitest
defaults to WATCH mode in an interactive TTY — exactly where a user runs
`gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never
exited and the orchestrator waited indefinitely. Recovery needed the user to
manually prompt "something blocking?".
Fix — one shared helper + a bounded, surfacing timeout on the test-command gates:
- New pure module src/normalize-test-command.cts + `gsd-tools query
normalize-test-command` verb: rewrites a resolved command to a best-effort
one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`;
a package-manager `test` script whose package.json runner is watch-vitest →
`CI=true` prefix; handles `--dir`; already-one-shot commands unchanged —
never double-flagged). Named `normalize-test-command` (not `test-*`) so the
file does not match node --test's default `test-*` discovery glob.
- The three gates that HUNG or silently-continued route through that ONE helper
and bound execution with `timeout $(config-get workflow.test_gate_timeout)`
(new config key, default 600s): the regression gate (extracted to
execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen —
it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the
audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124.
- verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so
it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint
on 124, staying under its frozen 40960-byte tier cap.
Security hardening (review): the normalizer only rewrites a runner named as a
standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never
mangled), is length-capped and uses only linear-time split-based scanning (no
super-linear backtracking on an adversarial `workflow.test_command`), and reads
package.json only when it is a regular file (never blocks on a FIFO via `--dir`).
Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the
schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors
workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores,
inventory manifest/index. All 16 golden-install-parity fixtures + workflow size
baseline regenerated for the changed shipped files; bin/lib is excluded from the
parity manifest.
Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security
hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route
through the shared helper + configured timeout + exit-124 watch-mode hint;
verify-phase asserted as normalize-only/already-bounded).
tests/execute-phase-active-flags.test.cjs repointed at the extracted step;
tests/planner-language-regression.test.cjs allowlist comment updated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Kilo and ZCode both declare `hooksSurface: 'none'` and have no plugin surface,
so the GSD installer staged lifecycle hook scripts (hooks/*.js, hooks/*.sh,
hooks/lib/) plus a `{"type":"commonjs"}` package.json marker into their config
dirs where nothing ever invokes them — dead weight (#1821).
The installer's two hook-copy guards at bin/install.js were still on the legacy
hardcoded runtime-name list and never excluded Kilo or ZCode. Add
`&& !isKilo && !isZcode` to both (isZcode added to the install() runtimeFlags
destructure).
OpenCode — which #1821 also named — is deliberately NOT excluded: since the
issue was filed, #1914 shipped a native OpenCode plugin (plugins/gsd-core.js)
that spawns those exact staged hooks via OpenCode's event bus and requires both
the hook scripts and the CommonJS package.json marker. Excluding OpenCode would
regress #1914, so its hooks stay live. The genuinely-dead cases are Kilo & ZCode.
- bin/install.js: add `&& !isKilo && !isZcode` to the hooks/dist copy guard and
the hooks/lib copy guard; document the OpenCode-vs-Kilo/ZCode split.
- tests/install-minimal-hooks.test.cjs: regression test asserting Kilo and ZCode
receive no gsd-*.js/.sh hooks or hooks/lib, while OpenCode keeps its hooks +
#1914 plugin and Claude keeps its hooks (over-exclusion guard).
- tests/fixtures/golden-install-parity/{kilo,zcode}.json: drop the 21 hooks/*
entries and the package.json marker they no longer receive (opencode unchanged).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A skills-only role:feature third-party capability installed 'active' but its skills never reached the runtime surface, capability enable/set rejected it as 'unknown capability', and capability list disagreed with capability state. Three defects, fix shape 1b (teach resolveSurface, no on-disk linking):
D1 (materialization): resolveSurface built the surfaced skills Set only from the on-disk manifest; third-party cap skills live at ~/.gsd/capabilities/<id>/skills/ and never entered the Set -> surfaced:false. Fix: union registry.capabilityClusters values into the Set in the full-profile branch (idempotent for first-party, additive for third-party; prototype-pollution guarded).
D2 (enable/set unknown): setCapabilityState validated the capId against the static first-party registry instead of the composed overlay-aware registry. Fix: validate against loadRegistry({includeInstalled}) like capability-state.cts does.
D3 (list vs state): capability list derived status purely from ledger existence, never surface composition. Fix: add a 'surfaced' field to list rows sourced from the same resolver capability state uses, so list and state agree.
Regression coverage: tests/issue-2045-third-party-skills-surface.test.cjs asserts all five acceptance criteria (resolveSurface union, state surfaced:true, enable/set not-unknown, list/state agreement, first-party + unknown-id regression). gsd-test: 23916/23916 green (linux-node22 + linux-node24).
The --runtime parsing + help-text edit to gsd-core/bin/gsd-tools.cjs changes
the installed file's content (gsd-tools.cjs is installed and compared by the
golden snapshot, unlike gsd-core/bin/lib/ which is excluded). Regenerated via
UPDATE_GOLDEN=1; every runtime's manifest updates exactly one line (the
gsd-tools.cjs hash).
next gained a new zcode runtime (#2039) since the last rebase, adding
tests/fixtures/golden-install-parity/zcode.json and shifting every other
runtime's install hashes again. Regenerated on Linux (WSL); full suite
(2830 tests) passes.
The rebase onto origin/next pulled in the runtime-launcher preamble resync
(applied repo-wide on next) alongside this branch's quick.md change; both
together shift every runtime's install hashes and workflow sizes, so the
fixtures from the pre-rebase regen were stale.
The quick.md fix in this PR changes the file content, so its per-runtime
content hash in the golden-install-parity snapshot (tests/golden-install-parity.test.cjs,
Linux/macOS-only) is stale. Regenerated with UPDATE_GOLDEN=1 on Linux
(WSL); only the gsd-core/workflows/quick.md hash line changed in each of
the 16 runtime fixtures — no other drift.
The legacy _isSkillsRuntime gate was a hardcoded isCodex || isCopilot || ...
roster; zcode was absent so the layout-driven skills-install path was skipped
(commands + agents installed, but no skills). Adding || isZcode would violate
AC#3 (no runtime === 'zcode' branches), so the gate is now derived from the
descriptor: a runtime takes the skills path iff its scoped artifactLayout
declares a skills kind. Behavior-preserving for all 15 existing runtimes
(verified by their install contracts); includes zcode (and any future skills
runtime) with zero per-runtime branches. opencode/kilo keep their specialized
combined path.
Regenerate the golden-install-parity fixtures: the non-zcode runtimes shift only
the shared model-catalog.json hash (zcode was added to the catalog); zcode now
includes its skills tree.
Resolve the zcode test cascade exposed by gsd-test:
- model-catalog.json: add zcode runtimeTierDefaults (null tiers, matching the
other profile-marker-only runtimes) so KNOWN_RUNTIMES stays parity with allRuntimes.
- runtime-name-policy: add zcode to GLOBAL_CONFIG_HOME_FRAGMENTS (~/.zcode) so
getGlobalConfigHomeFragment stops falling through to the .claude default.
- installer-migration-install: add the zcode fresh-install contract (flat-skills,
no settings, no package.json — same shape as trae).
- golden-install-parity: capture the zcode fixture via a standalone generator
script (not node --test — the gate stays gsd-test).
- capability-matrix.md: regenerate so zcode appears as a first-party row.
Fluidify the remaining count-pinned guards so adding a runtime no longer trips a
hand-pinned snapshot: gemini-runtime-removed (flag count derives from the
registry), non-claude-runtimes-registry-derivation (golden list derived from the
registry).
The compiled gsd-core/bin/lib/*.cjs modules are gitignored build artifacts
(ADR-457) shipped prebuilt in the npm tarball. A Claude Code plugin-marketplace
/ git-clone install materializes the repo tree directly and never runs
`npm run build:lib`, so those ~148 files are absent and every CLI command dies
at load with `Cannot find module './lib/cli-exit.cjs'`.
gsd-tools.cjs now calls ensureRuntimeBuild() before its ./lib requires: a new
committed bin helper that compiles the tree once on demand (lock-guarded,
incremental-cache-cleared, portable `node <tsc>` invocation) when the sentinel
cli-exit.cjs is missing, and surfaces an actionable error when TypeScript is
unavailable. The already-built npm path is a single fs.existsSync no-op.
ADR-457 (gitignored, build-at-publish) is preserved; no release-pipeline change.
Verified against a simulated marketplace checkout (148 build outputs removed):
the CLI self-heals and the result is byte-identical to `npm run build:lib`.
Closes#2002
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2017: grant mcp__plugin_context7_context7__* for plugin-marketplace context7
The 8 context7-using agents granted only mcp__context7__* (standalone server
form). Claude Code's plugin-marketplace context7 install names tools
mcp__plugin_context7_context7__*, so the grant never matched and every
researcher/planner/executor silently lost doc lookup (fell back to WebSearch).
- 8 agents: add mcp__plugin_context7_context7__* alongside mcp__context7__*.
- scripts/research-profiles.cjs: update the researcher profile tools to match.
- tests/context7-plugin-grant-parity.test.cjs: regression guard — no agent
grants the standalone form without the plugin form.
Closes#2017
* docs(#2017): backfill changeset pr 2029
* fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows
gsd-executor.md referenced sdk/src/query/QUERY-HANDLERS.md and reapply-patches.md
referenced sdk/dist/cli.js — both retired with the SDK (ADR-0174). AI runtimes
that locate doc refs via filesystem search ran find /, which on Git Bash for
Windows traverses the whole drive (14h+, orphaned find.exe, 4M+ handles).
- agents/gsd-executor.md: drop dead QUERY-HANDLERS.md ref.
- workflows/reapply-patches.md: drop dead sdk/dist/cli.js clause.
- tests/no-dead-sdk-refs.test.cjs: regression guard — no sdk/src|dist|handlers
file refs in agents/workflows/references markdown.
Closes#2020
* docs(#2020): backfill changeset pr 2027
* fix(#2019): planning-config.md global-learnings path ~/.gsd/learnings → ~/.gsd/knowledge
The features.global_learnings row pointed at ~/.gsd/learnings/ but the
implementation (src/learnings.cts, execute-phase.md) uses ~/.gsd/knowledge/.
Docs now match the code.
Closes#2019
* docs(#2019): backfill changeset pr 2026
* fix(#1921): preserve verify-work state across gap-closure + defer follow-ups
Resuming /gsd:verify-work after /gsd:execute-phase --gaps-only lost the
verification state: UAT ## Gaps still read status: failed even after their
fix plans executed, so verify-work re-diagnosed them as fresh blockers,
spawned a new gap plan, and reported only the new plan verified. A new
state contract links each gap to its fix plan so fixed gaps are recognized.
- verify-work.md gap YAML: stable gap_id (G-{phase}-{N}) per gap.
- plan_gap_closure planner prompt: each *-PLAN.md tags the gap_ids it
addresses in its frontmatter (gap_closure: true, gap_ids: [...]).
- new reconcile_gaps step (run at resume_from_file entry): marks a gap
status: resolved when its plan has a matching *-SUMMARY.md — fixed gaps
are not re-diagnosed and do not spawn new gap plans; a re-reported break
is treated as a fresh regression with a new gap_id.
- deferred-follow-up branch: a future-work idea (signals: 'later',
'next version', 'out of scope', ...) is captured to UAT ## Deferred
Follow-Ups instead of becoming a blocking gap/plan.
- workflow-size baseline recaptured (verify-work LARGE tier, 38221/61440).
Closes#1921
* docs(#1921): backfill changeset pr 2025
* fix(#1921): balance <step> tags in verify-work.md via 4-backtick outer fence
The plan_gap_closure step nested a ```yaml example inside a ``` block;
same-length fences made stripFencedCode close the outer block early,
swallowing the step's </step> (14 opens / 13 closes). Widen the outer
fence to 4 backticks so the nested yaml is contained. Regenerate golden
fixtures + size baselines.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Wire the two forward-declared mempalace.memory_mode modes so they actually
route recall/capture instead of silently behaving as `augment`:
- kg_backend: the palace temporal KG is the primary knowledge-graph source;
native .planning/graphs/ is the fallback. Non-KG drawer recall stays additive.
- replace: recall resolves through the palace as the source of truth; native
artifacts are the fallback.
Every mode stays onError:skip and default-resilient — an unreachable palace
degrades to native memory and GSD keeps writing .planning/graphs/, so no memory
is lost. Cross-mode .planning/graphs/ migration remains a documented open
question (PRD/ADR §17), out of scope here.
Surfaces updated (instruction-only contract): recall/capture commands (+ generated
skills), discuss/wave fragments, curator agent, capability.json schema. Docs:
how-to Step 3, CONFIGURATION, FEATURES, CONTEXT glossary. Regenerated
capability-registry, golden install-parity fixtures (mempalace hashes only),
agent-size-baseline. Added a routing-contract + cross-surface parity test.
Incidental (folded per no-defer rule): removed pre-existing unused imports
(spawnSync in capability-registry.test.cjs; fs in issue-498-package-identity.test.cjs)
that eslint flagged in/alongside the touched files.
Closes#2007
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1825): configurable graphify graph location (graphify.graph_path)
Add a graphify.graph_path config key (.planning/config.json) that overrides
where /gsd-graphify query|status|diff read the knowledge graph, so one curated
umbrella-level cross-repo graph can serve multiple sibling projects without N
drifting ~5 MB mirror copies. Previously the graph location was hardcoded to
<cwd>/.planning/graphs/.
- src/graphify.cts: resolveGraphLocation(cwd, planningDir) honors the key
(resolved relative to project root; absolute paths honored via path.resolve);
falls back to the historical .planning/graphs/graph.json when unset/blank/
non-string (byte-identical). Wired into graphifyQuery, graphifyStatus,
graphifyDiff (snapshot travels with the configured graph via dirname), and
writeSnapshot. Configured-but-missing -> actionable error naming the path.
Build stays project-scoped (skill hardcodes the cp dest); umbrella graph is
built in the umbrella project, sub-projects only READ it.
- config-schema.manifest.json: register graphify.graph_path in validKeys.
- tests/graphify-graph-path.test.cjs: boundary matrix (unset byte-identical,
set+present reads configured graph not default, set+missing actionable error,
relative resolved vs project root, blank treated as unset, snapshot alongside
configured graph, diff from configured dir, build project-scoped) +
VALID_CONFIG_KEYS registration.
- docs: CONFIGURATION.md row, FEATURES.md REQ-GRAPH-06, CONTEXT.md module note,
.changeset (Added).
Closes#1825
* docs(#1825): backfill changeset pr number 2013
* fix(#1864): balance <step> tags in settings-advanced.md §8 Model Policy
§8 Model Policy ended with </step> but had no matching opening tag (5 opens /
6 closes), leaving it as loose inter-step content. Add the missing
<step name="model_policy"> opener so the section is a proper step.
- gsd-core/workflows/settings-advanced.md: add <step name="model_policy">
- tests/workflow-step-tag-balance.test.cjs: regression guard — every top-level
workflow must have balanced <step>/</step> (fenced code stripped), plus a
focused assertion that §8 is wrapped in model_policy.
- goldens + workflow-size baseline recaptured.
Closes#1864
* docs(#1864): backfill changeset pr 2014
* fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR
The gsd_run preamble resolved the Claude global install only at
$HOME/.claude/gsd-core/bin/, but the installer honors CLAUDE_CONFIG_DIR —
so a global install redirected via CLAUDE_CONFIG_DIR was invisible to every
gsd_run call (every command failed with 'gsd-tools.cjs not found').
The Claude resolver arm now uses ${CLAUDE_CONFIG_DIR:-$HOME/.claude},
matching the installer + the other runtimes' ${VAR:-default} pattern.
Default $HOME/.claude behavior is unchanged.
- _runtime-launcher.snippet.sh: Claude arm honors CLAUDE_CONFIG_DIR.
- sync-runtime-launcher.cjs re-run: 95 workflows/agents re-synced.
- review.md / discuss-phase.md: trimmed to stay under their byte budgets.
- runtime-launcher-parity.test.cjs: (A) substring updated for the new form
+ explicit #1865 assertion that the snippet honors CLAUDE_CONFIG_DIR.
- goldens + size baselines recaptured.
Closes#1865
* docs(#1865): backfill changeset pr 2024
* fix(#1936): reconstruct OpenCode review from JSON events; diagnosable empty-output stub
On a large review prompt, OpenCode's default `build` agent runs a few read
tool calls then ends its turn with zero output tokens (reason:"stop",
output:0), so `opencode run --format default` emits empty stdout. The reviewer
block redirected stderr to /dev/null and wrote a generic "failed or returned
empty output" stub — so the phase silently lost its second independent reviewer
with no diagnostic and no timeout.
Rewrite the OpenCode reviewer block to invoke `--format json` as the primary
call and reconstruct the review from the assistant `text` parts (jq). Capture
stderr to a `.err` sidecar (mirrors the Codex block). When the agent emits no
text, surface the stop reason, output-token count, and stderr so the failure is
diagnosable. Gate the stub on the extracted CONTENT, not the output file size —
an empty jq extraction still prints a lone newline that a `[ -s file ]` check
would treat as populated. Document the wall-clock timeout as a Bash-tool param
(macOS lacks GNU timeout; opencode has no native timeout flag).
review.md was already at the DEFAULT size-tier ceiling (40956/40960), so the
fix cannot fit without reclassifying it into the LARGE tier (it is a
multi-reviewer orchestration file that outgrew "focused single-purpose"; 43.4 KB
sits well under the LARGE high-water mark). Recapture the 16 golden-install
fixtures — the diff is exactly one review.md hash per runtime. Regression block
folded into review-default-reviewers-workflow.test.cjs (new bug-NNNN test files
are not accepted).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1936): add changeset
* test(#1936): property-test the OpenCode review jq reconstruction
Address the re-review's one actionable finding: the jq JSON-event → text
reconstruction had no fast-check property test.
Add tests/opencode-review-reconstruction.property.test.cjs. It extracts the two
shipped jq programs (OPENCODE_REVIEW, OPENCODE_DIAG) verbatim from
gsd-core/workflows/review.md and runs the real jq — not a reimplementation — so
the shipped logic is what gets tested. Properties: the reconstructed review
equals the newline-join of every assistant text part (order preserved); a stream
with no text part reconstructs to empty (drives the #1936 stub); null/absent text
parts are dropped, never rendered as "null". Plus example-based coverage of the
diagnostic edges the reviewer cited: missing .tokens.output and no step_finish
degrade to "?"; non-JSON stdout makes jq fail rather than masquerade as a review.
Verified the invariant has teeth (a comma-join jq fails the property).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1936): skip jq reconstruction property test when jq is absent
The property test shells out to `jq`, which GitHub's windows-latest runners do
not ship (macOS/Linux runners do). `execFileSync('jq')` therefore ENOENT-failed
the whole file on `test (windows-latest, *)`. Probe `jq --version` at load and
skip the suite when jq is not on PATH — the reconstruction logic is
platform-independent, so the assertions still run in full on every jq-present
runner (mirrors how golden-install-parity skips on win32).
Verified: jq present → 7 pass; jq removed from PATH → 7 skipped, 0 fail.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1936): skip jq reconstruction property test on Windows, not just when jq is absent
The prior guard skipped only when `jq` was absent from PATH — but the
windows-latest runners DO ship jq, so the suite still ran there and failed with
`jq: parse error: Invalid numeric literal` (confirmed from the CI job log). Root
cause is Node's child_process argument quoting mangling the jq program (it embeds
double quotes) on Windows, not the shipped review.md logic — the macOS/Linux legs
pass. Gate the suite on `process.platform === 'win32'` (still also skipping when
jq is absent), mirroring golden-install-parity's win32 skip. Logic is
platform-independent and fully asserted on every macOS/Linux CI leg.
Verified: macOS → 7 pass; simulated win32 → skips.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>