- restructure _loadFlatCommandsGsdManifest try/catch to wrap read+parse+set
together (mirrors loadSkillsManifest exactly), so a thrown parser degrades
both keys to [] — closes the latent catch-scope parity drift [Nit-1]
- add boundary test: gsd-.md (empty stem) skipped, gsd-x.md single-char stem
kept (slice(4,-3) boundary) [Low-1]
- add unreadable-file test (POSIX-gated): both keys degrade to [] [Low-2]
- strengthen parity test to also compare requires + _calls_agents_ VALUES,
not just the stem set [Nit-2]
Both orthogonal reviews returned APPROVE with no Critical/High findings.
Security review confirmed no new trust-boundary crossing, prototype-pollution
immune (Map/Set throughout), and symlink/path-traversal surface identical to
the pre-existing nested loader (not a regression).
_resolveManifest only recognized the nested source layout (commands/gsd/*.md)
and the installed-runtime skills layout (skills/gsd-<stem>/SKILL.md). A flat
source install (Claude local project shape: commands/gsd-<stem>.md, no
commands/gsd/ subdir) matched neither branch, so the manifest came back empty
and resolveSurface materialized the full profile to an empty Set — silently
reporting every skill-bearing capability as surfaced:false / enabled:false /
active:false. The nyquist/code-review/security/ui verify:post and execute:post
hooks never fired even with their workflow.* toggles on.
Add a third branch: when commandsGsdDir is absent, scan dirname(commandsGsdDir)
for gsd-<stem>.md files, strip the gsd- prefix, and build the same Map shape
the nested loader produces (requires via shared parseRequires, companion
_calls_agents_<stem> via shared parseCallsAgents). Falls through to the
installed-skills branch when the flat dir has no gsd-*.md files (precedence:
nested > flat-source > installed).
Also export parseCallsAgents from install-profiles so capability-state reuses
the SAME parser the nested loader uses (no drift; mirrors the existing
parseRequires export+reuse pattern).
Mirrors the #1160 installed-layout tests for the flat source layout
(<repo>/commands/gsd-<stem>.md, no commands/gsd/ subdir). Covers stem
extraction (strip gsd- prefix), requires parsing via shared parseRequires,
companion _calls_agents_ key parity, _resolveManifest flat-branch detection,
precedence (flat-empty falls through to installed), and a generative-parity
assertion that the flat loader and nested loader produce identical stem sets
for the real command tree.
Expected RED against unfixed capability-state.cts (_resolveManifest has no
flat branch; _loadFlatCommandsGsdManifest not exported).
- add typeof guard so a non-string override passes through verbatim instead of
crashing on .startsWith (preserves pre-fix no-crash behaviour) [LOW-1]
- use Object.hasOwn() for the alias lookup so __proto__/constructor cannot
return a truthy non-string from the plain object literal [LOW-D3]
- cap the unmappable-override stderr warning at 64 chars so an oversized or
secret-shaped value cannot leak in full to stderr/logs [LOW-D4]
- remove the unused mapClaudeOverrideForRuntime export (helpers are covered
behaviourally via resolveModelInternal/resolveModelForTier) [NIT]
- add resolveModelForTier unmappable-override fall-through test (closes the
mutation-score gap) [MEDIUM-1]
- add case-sensitivity contract test (Claude-Sonnet-5 passes through verbatim) [LOW-2]
Both orthogonal reviews returned APPROVE with no Critical/High findings.
model_overrides values that are full Claude model IDs (claude-sonnet-5,
claude-opus-4-8, claude-haiku-4-5, claude-fable-5) were returned verbatim on
the claude runtime and handed to the Claude Agent tool, whose typed model
parameter documents only tier aliases (opus/sonnet/haiku/fable). The
model_policy path already mapped full IDs -> aliases via
CLAUDE_POLICY_ID_TO_ALIAS (#1144); model_overrides skipped that mapping, so
the two resolver paths produced different shapes for the same underlying
Claude model. The fix mirrors #1144 on the override path via a shared
mapClaudeOverrideForRuntime helper used by both resolveModelInternal and
resolveModelForTier. Bare aliases pass through verbatim; non-Claude runtimes
and non-Claude custom/vendor values keep full IDs verbatim (parity). An
unmappable Claude ID (e.g. claude-opus-4-5) warns once to stderr and falls
through to tier resolution, exactly as the model_policy path already does.
Alias mapping is also the documented best practice (prevents staleness when
new model versions ship).
Mirrors the #1133 model_policy alias-mapping tests for the model_overrides
path. Covers AC1-AC6: mappable Claude full IDs (claude-sonnet-5/opus-4-8/
haiku-4-5/fable-5) resolve to aliases on runtime:claude; bare aliases pass
through; non-claude runtimes keep full IDs verbatim; unmappable Claude IDs
warn-once + fall through; resolveModelForTier escalation path also maps;
non-Claude custom/vendor values pass through verbatim (regression guards).
Expected RED against unfixed model-resolver.cts (override short-circuit at
lines 162-167 / 288-290 returns override verbatim with no alias mapping).
- compareSemver: implement full SemVer 2.0.0 §11 pre-release identifier
comparison (two pre-releases of the same triple now order correctly; was 0).
- capability description + fragment: scope the plan-checker/verifier claim
(this capability delivers the parallel-execution backend; those gates remain
inline until separately wired). Correct the 'each wave is one barrier' prose
(a wave splits into multiple sequential parallel() barriers on files_modified
overlap). Frame detect-backend CLI as a simulation harness; the pure function
with the live host descriptor is the real detection seam.
- partitionStages docstring: 'near-minimal via greedy first-fit' (not 'fewest');
document empty-files_modified behavior.
Shard 2/3 chunk 2 (~80 files including state.test.cjs, perf-*, worktree-cleanup)
exceeded the 600s per-chunk timeout on macOS Node 22. Reducing the cap from 90
to 60 splits this into two ~40-file chunks, each well within the 600s budget.
Three chunks at ~5 min each = ~15 min, safely under the 20m job cap.
AGENTS.md is legitimately created by gsd install copilot (issue #786) when
the installer runs inside a repo checkout. The file may exist on disk — that
is expected. The test must verify it is not git-tracked, not that it is absent
from the working tree. Uses git ls-files --error-unmatch instead of fs.existsSync.
- Teach applySurface to build agentCtx (pathPrefix + attribution) and pass it
to kind.stage() for agents kind, mirroring createRuntimeArtifactInstallPlan
(ADR-1235 §1). Surface-path agents now receive path-rewrite + attribution +
converter + normalize, matching install output byte-for-byte.
- Pass skills:'*' sentinel for agents staging when no surface state modifications
exist, so ALL agents are staged (not just those referenced by _calls_agents_).
- Declare converted agents kind in copilot and antigravity capability.json;
add to _DESCRIPTOR_AGENTS_RUNTIMES in bin/install.js.
- Handle copilot .agent.md filename rename in both _copyStaged (install path)
and _syncGsdDir (surface path).
- Ship golden-parity harness (ADR-1235 §0): tests/issue-1575-agent-descriptor-
parity.test.cjs asserts applySurface output is byte-identical to
installRuntimeArtifacts for all 7 descriptor-driven runtimes, plus stale-
cleanup convergence and prune data-loss coverage.
- Update ADR-1235 with cutover progress.
Cline remains deferred (rules-only local branch + local/global complication).
Support regenerating a single runtime (node scripts/gen-golden-install-parity-zcode.cjs zcode) or all runtimes (no args). Needed when a shared payload file changes and every fixture must be recaptured.
The descriptor-driven gate must treat a runtime as layout-driven when its scoped
artifactLayout is non-empty (any kind), not only when it has a skills kind —
windsurf's global layout is agents-only and is a legitimate layout-driven runtime.
Preserve the three legacy special-cased paths (opencode/kilo combined path,
claude-local copyWithPathReplacement).
zcode falls into the default package.json+hooks path (not in the CommonJS-mode
exclusion roster), so its install contract is packageJson: true — matching
observed install output. No runtime === 'zcode' branch is added (AC#3): zcode
gets the default by not being excluded.
The legacy _isSkillsRuntime gate was a hardcoded isCodex || isCopilot || ...
roster; zcode was absent so the layout-driven skills-install path was skipped
(commands + agents installed, but no skills). Adding || isZcode would violate
AC#3 (no runtime === 'zcode' branches), so the gate is now derived from the
descriptor: a runtime takes the skills path iff its scoped artifactLayout
declares a skills kind. Behavior-preserving for all 15 existing runtimes
(verified by their install contracts); includes zcode (and any future skills
runtime) with zero per-runtime branches. opencode/kilo keep their specialized
combined path.
Regenerate the golden-install-parity fixtures: the non-zcode runtimes shift only
the shared model-catalog.json hash (zcode was added to the catalog); zcode now
includes its skills tree.
Resolve the zcode test cascade exposed by gsd-test:
- model-catalog.json: add zcode runtimeTierDefaults (null tiers, matching the
other profile-marker-only runtimes) so KNOWN_RUNTIMES stays parity with allRuntimes.
- runtime-name-policy: add zcode to GLOBAL_CONFIG_HOME_FRAGMENTS (~/.zcode) so
getGlobalConfigHomeFragment stops falling through to the .claude default.
- installer-migration-install: add the zcode fresh-install contract (flat-skills,
no settings, no package.json — same shape as trae).
- golden-install-parity: capture the zcode fixture via a standalone generator
script (not node --test — the gate stays gsd-test).
- capability-matrix.md: regenerate so zcode appears as a first-party row.
Fluidify the remaining count-pinned guards so adding a runtime no longer trips a
hand-pinned snapshot: gemini-runtime-removed (flag count derives from the
registry), non-claude-runtimes-registry-derivation (golden list derived from the
registry).
Add ZCode as a first-party runtime via a declarative capability descriptor
(capabilities/zcode/capability.json) with zero hardcoded runtime === 'zcode'
branches — exercising the de-hardcoded, data-driven runtime path that 1.7.0
(ADR-1016 / ADR-1239) enables.
Descriptor (all axes sourced verbatim from zcode.z.ai docs):
- configHome ~/.zcode; nested skills + flat commands/agents; profile-marker install
- Claude-shaped skill format reuses convertClaudeCommandToClaudeSkill (no new converter)
- hostIntegration: declarative / slash-file / mcp / electron; dispatch background=false
(foreground-only per docs); nested+maxDepth undocumented; passive model mode
Installer registration (data, not branches): --zcode flag, allRuntimes, runtimeMap,
interactive menu, --all list. getGlobalConfigDir + resolveRuntimeArtifactLayout +
ALLOWED_CONFIG_RUNTIMES + resolveInstallPlan all derive from the descriptor.
Revamped the brittle per-runtime golden-master tests to be count-agnostic,
descriptor-derived property tests (1.7.0 makes runtimes pluggable data, so pinning
frozen '15 runtime' snapshots is the wrong invariant): getdirname, label-policy,
config-adapter-registry (intent + install-plan golden master), capability-registry,
host-integration-descriptors (counts derive from curated maps). Adding a runtime
descriptor now extends coverage with zero edits to those suites.
Welcome banner, --zcode help, supported-runtimes how-to, and the host-integration
capability matrix (every axis cited) updated.
The compiled gsd-core/bin/lib/*.cjs modules are gitignored build artifacts
(ADR-457) shipped prebuilt in the npm tarball. A Claude Code plugin-marketplace
/ git-clone install materializes the repo tree directly and never runs
`npm run build:lib`, so those ~148 files are absent and every CLI command dies
at load with `Cannot find module './lib/cli-exit.cjs'`.
gsd-tools.cjs now calls ensureRuntimeBuild() before its ./lib requires: a new
committed bin helper that compiles the tree once on demand (lock-guarded,
incremental-cache-cleared, portable `node <tsc>` invocation) when the sentinel
cli-exit.cjs is missing, and surfaces an actionable error when TypeScript is
unavailable. The already-built npm path is a single fs.existsSync no-op.
ADR-457 (gitignored, build-at-publish) is preserved; no release-pipeline change.
Verified against a simulated marketplace checkout (148 build outputs removed):
the CLI self-heals and the result is byte-identical to `npm run build:lib`.
Closes#2002
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Leaked working-notes/triage file (Started 2026-05-16, Grok+user /gsd-inbox
session) accidentally committed to the root in 05316369a. References resolved
issues; not referenced by any code, test, doc, manifest, or golden fixture.
Closes#2034
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The #2018 regression test used raw fs.rmSync() in afterEach, tripping the
local/no-raw-rmsync-in-tests eslint rule (Windows-EBUSY retry budget). CI's
lint-tests job (no eslint cache) caught it on next; local --cache runs had
false-greened it. Switch to helpers.cleanup(). Also drop two unused
destructured imports (evaluateCommandExitZero, interpolate) in the
gate-predicate-evaluator test that emitted no-unused-vars warnings.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2012): scope Progress-row regex to ## Progress section (was binding to earlier table)
The Progress-row writer used a non-global regex that matched ANY table row
starting with the phase number. When an earlier table (e.g. Requirements
coverage | Phase | Requirements | Count |) preceded ## Progress, the regex
bound to the wrong row (3-column), no-op'd, and never reached the real
Progress row. roadmap_updated stayed true (it's existsSync), masking the
failure.
- src/phase.cts: scope the tableRowPattern regex to the ## Progress section
(indexOf + slice) so it only matches Progress-table rows.
- tests/phase.test.cjs: regression test — ROADMAP with a phase-numbered
Requirements table before ## Progress → Progress row updated, Requirements
row untouched.
Closes#2012
* docs(#2012): backfill changeset pr 2032
* fix(#2018): applySurface empty manifest no longer deletes gsd-* agents
The agent-prune loop in _syncGsdDir deleted any gsd-*.md not in the staged
set. When the manifest resolved empty (null/non-object, no entries, no files
key, or unresolvable install source root), the staged set was empty → every
gsd-* agent was deleted. Skills were guarded by pruneSkillDirs' manifest-
membership check; agents had no equivalent.
- src/surface.cts: skip the agent-prune loop when the manifest is empty/absent
(copy still runs — genuinely new agents are added).
- tests/surface-empty-manifest-agents.test.cjs: boundary matrix — empty
manifest preserves agents (Map() + undefined); non-empty manifest + empty
staged still prunes (boundary); empty manifest + new staged copies new +
preserves existing.
Closes#2018
* docs(#2018): backfill changeset pr 2031
* fix(#2017: grant mcp__plugin_context7_context7__* for plugin-marketplace context7
The 8 context7-using agents granted only mcp__context7__* (standalone server
form). Claude Code's plugin-marketplace context7 install names tools
mcp__plugin_context7_context7__*, so the grant never matched and every
researcher/planner/executor silently lost doc lookup (fell back to WebSearch).
- 8 agents: add mcp__plugin_context7_context7__* alongside mcp__context7__*.
- scripts/research-profiles.cjs: update the researcher profile tools to match.
- tests/context7-plugin-grant-parity.test.cjs: regression guard — no agent
grants the standalone form without the plugin form.
Closes#2017
* docs(#2017): backfill changeset pr 2029
* fix(#2022): gate roadmap update-plan-progress checkbox on verification passed
cmdRoadmapUpdatePlanProgress stamped the phase checkbox + completion date
the moment all summaries landed, with NO verification gate — unlike
cmdPhaseComplete (phase.cts:1436) which checks readVerificationStatus. Since
update-plan-progress is called after every wave and every plan, the checkbox
fired before gsd-verifier confirmed the phase.
- src/roadmap.cts: isComplete now requires summaryCount >= planCount AND
readVerificationStatus(phaseDir).status === 'passed'.
- tests/roadmap.test.cjs: 2 regression tests (no VERIFICATION.md → not
complete; gaps_found → not complete) + updated 3 existing complete tests
to include a passed VERIFICATION.md.
Closes#2022
* docs(#2022): backfill changeset pr 2030
* fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows
gsd-executor.md referenced sdk/src/query/QUERY-HANDLERS.md and reapply-patches.md
referenced sdk/dist/cli.js — both retired with the SDK (ADR-0174). AI runtimes
that locate doc refs via filesystem search ran find /, which on Git Bash for
Windows traverses the whole drive (14h+, orphaned find.exe, 4M+ handles).
- agents/gsd-executor.md: drop dead QUERY-HANDLERS.md ref.
- workflows/reapply-patches.md: drop dead sdk/dist/cli.js clause.
- tests/no-dead-sdk-refs.test.cjs: regression guard — no sdk/src|dist|handlers
file refs in agents/workflows/references markdown.
Closes#2020
* docs(#2020): backfill changeset pr 2027
* fix(#2019): planning-config.md global-learnings path ~/.gsd/learnings → ~/.gsd/knowledge
The features.global_learnings row pointed at ~/.gsd/learnings/ but the
implementation (src/learnings.cts, execute-phase.md) uses ~/.gsd/knowledge/.
Docs now match the code.
Closes#2019
* docs(#2019): backfill changeset pr 2026
* fix(#1921): preserve verify-work state across gap-closure + defer follow-ups
Resuming /gsd:verify-work after /gsd:execute-phase --gaps-only lost the
verification state: UAT ## Gaps still read status: failed even after their
fix plans executed, so verify-work re-diagnosed them as fresh blockers,
spawned a new gap plan, and reported only the new plan verified. A new
state contract links each gap to its fix plan so fixed gaps are recognized.
- verify-work.md gap YAML: stable gap_id (G-{phase}-{N}) per gap.
- plan_gap_closure planner prompt: each *-PLAN.md tags the gap_ids it
addresses in its frontmatter (gap_closure: true, gap_ids: [...]).
- new reconcile_gaps step (run at resume_from_file entry): marks a gap
status: resolved when its plan has a matching *-SUMMARY.md — fixed gaps
are not re-diagnosed and do not spawn new gap plans; a re-reported break
is treated as a fresh regression with a new gap_id.
- deferred-follow-up branch: a future-work idea (signals: 'later',
'next version', 'out of scope', ...) is captured to UAT ## Deferred
Follow-Ups instead of becoming a blocking gap/plan.
- workflow-size baseline recaptured (verify-work LARGE tier, 38221/61440).
Closes#1921
* docs(#1921): backfill changeset pr 2025
* fix(#1921): balance <step> tags in verify-work.md via 4-backtick outer fence
The plan_gap_closure step nested a ```yaml example inside a ``` block;
same-length fences made stripFencedCode close the outer block early,
swallowing the step's </step> (14 opens / 13 closes). Widen the outer
fence to 4 backticks so the nested yaml is contained. Regenerate golden
fixtures + size baselines.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
onboard.md delegates the gsd_run preamble to
gsd-core/references/gsd-run-resolver.md via @-include, but two guards
regressed once the canonical launcher snippet advanced to
${CLAUDE_CONFIG_DIR:-$HOME/.claude} (#2024):
- runtime-launcher-parity (B2): the resolver reference still shipped the
old $HOME/.claude arm. references/ is not covered by
sync-runtime-launcher.cjs, so refresh the resolver bash block to be
byte-equal to _runtime-launcher.snippet.sh.
- /gsd:onboard command contract: sync-runtime-launcher.cjs had re-inlined
the preamble into onboard.md (a delegating file). Teach the sync
transform to strip-but-never-inline files that @-include the resolver,
mirroring the exemption already in the parity test (B/B2).
Regenerate golden install fixtures and the workflow size baseline for the
smaller onboard.md.
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>