Two gen-time validation tightenings (validation-only; serialized registry content
unchanged; bin/install.js + adapter + descriptors untouched):
Part B: validateArtifactKindEntry now requires artifactLayout[].converter ∈
VALID_CONVERTER_NAMES (15 names, all exported by install.js) ∪ {null} — a typo'd
converter fails at gen time instead of silently → installExports[name]===undefined
at install time.
Part A: a HARD buildRegistry parity gate asserts each runtime descriptor's
configFormat agrees with the adapter registry's installSurface via a fixed mapping
(cursor-hooks-json/profile-marker-only→none, codex-toml→toml, copilot-instructions→
markdown, cline-rules→markdown-dir, settings-json→settings-json) — keeps configFormat
from drifting; prerequisite-validation for the deferred full drive (#1055).
The full config-writing drive (retire resolveRuntimeConfigIntent) is deferred to
#1055: configFormat is lossy vs installSurface (cursor vs profile-marker both → none;
opencode/kilo permissionWriter has no descriptor field) → needs schema extension.
Closes#1056
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The `full test (windows-latest, 22)` job intermittently got CANCELLED at its
20m wall-clock cap with no failed test step — a false-negative gate (recurrence
of #869). Root cause: a unit test leaves an open event-loop handle, so the
chunk's `node --test` child hangs ~150s on Windows after its last test prints;
two such stalls push the already-~13m job past 20m.
Fix (defense in depth):
- run-tests.cjs: pass --test-force-exit (Node >=22; engines requires >=22.0.0)
so the runner exits once all tests finish regardless of lingering handles —
the durable backstop. Account for the flag in the argv-length ceiling.
- run-tests.cjs: add a per-chunk execFileSync timeout (default 600000ms, env
RUN_TESTS_CHUNK_TIMEOUT_MS) that fails loudly with a diagnostic naming the
chunk's files, so a hung chunk can never silently eat the job budget.
- perf-316 test: terminate both Worker threads on all paths (afterEach +
finally) so they cannot outlive the test.
- locking-bugs test: kill spawned children in a finally that wraps the whole
spawn -> waitFor -> barrier-release -> Promise.all sequence, so a barrier
timeout no longer leaks live child processes.
- Refresh the stale synckit comment (synckit/SDK bridge was removed).
Regression tests in run-tests-harness: a hung chunk hits the per-chunk timeout
and fails with a clear message; force-exit lets a chunk with a leaked handle
exit cleanly.
Closes#1051
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
resolveRuntimeArtifactLayout now builds Layout from
registry.runtimes[id].runtime.artifactLayout[scope] — a loop dispatching each
ArtifactKind through the SAME 5 builders (commandsKind/agentsKind/skillsKind/
convertedCommandsKind/kimiAgentsKind, unchanged) by (kind, converter, nesting) —
replacing the hardcoded switch(runtime). Equivalence-preserving for all 16 runtimes
× {global, local} (Codex-verified, no divergence). -43 LOC; bin/install.js + the
converters + the install loop untouched. getInstallExports()[converterName]
resolution, configDir threading, scope default, unknown-runtime guard all preserved.
Driving the local scope surfaced a 5a gap: the old switch had no scope branch for 13
runtimes (cursor/gemini/codex/copilot/antigravity/windsurf/augment/trae/qwen/hermes/
codebuddy/opencode/kilo) → local == global for them, but 5a authored local:[].
Backfilled local=global for those 13 (descriptor-faithful; a fall-through shim would
wrongly give cline/kimi local=global). claude/cline/kimi scope-gating untouched.
validateArtifactKindEntry tightened: destSubpath/prefix/nesting/converter required
(ConverterName enum still open — 5e). New 39-case deep-equal golden equivalence test.
Closes#1049
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
`expandRunsOn` enumerated a multi-axis `strategy.matrix` one key at a time,
producing partial realization contexts. A true `os × shell` matrix therefore
left `${{ matrix.shell }}` unresolvable against any `{ os: ... }`-only context,
firing spurious `UNRESOLVABLE_MATRIX` violations and leaving shell-pinning
coverage incomplete on Cartesian jobs.
Enumerate the full GitHub Actions cross-product of all base-list matrix keys
(every `matrix.<k>` array, excluding the `include`/`exclude` control keys) via
a named `cartesianProduct` helper. Each realization's context now carries a
value for every matrix key, so `${{ matrix.<key> }}` resolves per realization.
Single-axis matrices keep byte-for-byte identical output; only multi-axis
matrices change shape. The `include` and `exclude` blocks are unchanged
(full tuple-aware exclude is a documented out-of-scope follow-up).
Tests: updated the `os × shell` test to assert post-fix behavior (4 step
realizations, 2 WRONG_SHELL_FOR_OS, 0 UNRESOLVABLE_MATRIX); added a compliant
`os × node-version` cross-product test; added a fast-check property test that
the realization count equals the product of axis lengths and that no axis key
is dropped from any context.
Closes#435
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver
Source agents/*.md (gsd-planner, gsd-executor, gsd-verifier, gsd-plan-checker,
gsd-intel-updater, gsd-debugger, …) called bare "gsd-tools …" in shell blocks.
On a shim-only install — where gsd-tools.cjs exists under the runtime home but
gsd-tools is NOT on PATH — those calls fail with "command not found" and the
agent silently skips init/state/validate/commit ceremony, deferring to the
orchestrator or bypassing GSD bookkeeping entirely.
were never migrated, so it persisted on Claude Code and every other runtime that
consumes the source agents directly. Only gsd-phase-researcher.md carried a
resolver — and a stale, claude-only truncated one.
Fix (all runtimes):
- Inject the canonical multi-runtime gsd_run preamble (byte-equal to
_runtime-launcher.snippet.sh — claude/codex/cursor/gemini/copilot/windsurf/
augment/trae/qwen/cline/opencode/kilo/hermes/antigravity homes) at the top of
the first gsd_run block of all 12 gsd-tools-calling agents, and rewrite every
command-position bare gsd-tools to gsd_run.
- Upgrade gsd-phase-researcher.md's stale resolver to the canonical one.
- Extend scripts/sync-runtime-launcher.cjs to maintain agents/ in parity (the
sync caught and corrected a mis-placed preamble during development).
- Extend the bare-gsd-tools (#2851) and launcher-parity (#373) regression guards
to agents/ so no runtime can silently regress.
Closes#1041
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1041): backfill changeset PR number to 1045
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Author capabilities/<runtime>/capability.json for all 16 runtimes (role:runtime),
populating the registry's runtimes index ({} -> 16). Each carries the 6 ADR-1016
axes (configHome structured, configFormat, artifactLayout structured, commandStyle,
hooksSurface + hookEvents, sandboxTier, supportTier), extracted from the live
modules (runtime-homes/runtime-slash/runtime-artifact-layout/runtime-config-adapter
+ CODEX_AGENT_SANDBOX). validateRuntimeBody tightened to enforce the closed
vocabularies + structured configHome/artifactLayout (rejects old string configHome;
env required; skillsHome recursively validated).
Registry-only: bin/install.js + src/*.cts untouched, zero behavior change. The 5b-5f
drive steps consume these one axis at a time (ADR-1016 §8).
Codex accuracy review caught + fixed: opencode/kilo register ZERO lifecycle hooks
(install.js skips the whole block) -> hooksSurface:none (configFormat stays
settings-json); kimi probe selects on <root>/skills existence -> probeExists:"skills".
ArtifactKind required-field strictness + ConverterName closing deferred to 5d/5e.
ADR-1016 (Proposed) amended to match (Decisions 1 + 5).
Closes#1035
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Replace the inlined ui-review invocation in autonomous.md §3d.5 with a
loop.render-hooks verify:post dispatch — the first workflow to consume
render-hooks and fire a skill from it (closes the #1018 live-execution residual
as real wiring). Capability-driven, equivalence-preserving for the current
registry (only ui-review at verify:post, default on): fires gsd-ui-review under
the same precondition (UI-SPEC exists via consumes-gate + workflow.ui_review).
Gate findings (real pattern issues, fixed so every future cutover inherits them):
- bug-2643 static "Skill() references a real skill" check vs templated
Skill(skill="gsd-${ref.skill}") dispatch → skip ${...}-templated names.
- Coverage moved, not lost: gen-capability-registry now validates
steps[].ref.skill in skills + ref.agent in agents + rejects gsd- double-prefix.
- Tightened §3d.5 tests; markdown clarity (consumes rule, LLM-native JSON read,
UI-REVIEW.md score hint).
gsd-ui-review skill + §3a.5/ui-phase untouched. §5.6/ui-phase cutover deferred
(#1022 step-can-halt-vs-gate model question).
Closes#1023
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1006): harden render --preview against fragment parse failures
`render --preview` wrote `report.preview` unconditionally. When a `.changeset`
fragment fails to parse, `cmdRender` early-returns with `{exitCode:1, report:
{failures}}` and NO `preview` key, so `process.stdout.write(undefined)` threw
ERR_INVALID_ARG_TYPE and the rc release job's "Preview CHANGELOG" step died with
a cryptic TypeError that masked the real cause.
Guard the preview write on `typeof report.preview === 'string'` (ADR-227: shape,
not just type); when absent, fall through to the existing failure reporter that
names the offending fragment and exits non-zero — identical to a non-preview
render. Also backfills the stray placeholder `pr: 0` -> `pr: 939` in
.changeset/936-convergence-inline-plan-phase.md that triggered the live failure.
Regression test (red-then-green verified) added at the render --preview seam.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1006): validate changeset fragment content at the Changeset Required gate
The `Changeset Required` gate (scripts/changeset/lint.cjs) only checked that a
`.changeset/*.md` fragment EXISTS in the PR diff; it never validated the
fragment's contents. So a malformed fragment (e.g. an un-backfilled `pr: 0`
placeholder) silently merged to `next` and only detonated later in the rc
release job. This is the upstream prevention for #1006 — the crash hardening
turns the failure into a clear message, this stops the bad fragment ever
reaching the release path.
evaluateLint now accepts `fragmentFailures` and fails with the typed reason
`fail_invalid_fragment` (naming each offending file) before the existence/
opt-out checks — a malformed fragment beats `no-changelog`, since it will break
the render regardless. main() reads + parseFragment()s every changed fragment:
a deleted fragment (not on disk) is skipped, a present-but-unreadable one fails
closed. Tests assert on the typed LINT_REASON enum (no raw-text matching), a
precedence case over the opt-out label, and an end-to-end suite that drives the
real main() against a temp git repo (malformed -> fail, valid -> pass, deleted
-> skipped) so the wiring is regression-proof.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1006): assert the typed --json report in the preview regression test
Code review flagged the preview parse-failure regression test for positive
raw-text matching on CLI output (`combined.includes('bad-fragment.md')` /
`'invalid_pr'`), which this repo's testing standards forbid. Keep the non-json
`runRenderRaw` call for the negative crash proof (the ERR_INVALID_ARG_TYPE
crash lives only on the non-json stdout.write path), and add a `--json`
invocation that asserts the offending fragment + typed `invalid_pr` reason via
the structured `report.failures[]` surface instead of rendered prose.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#969): make bug-969 hardening tests hermetic and move build tsbuildinfo out of shipped tree (regression from #996)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#969): self-heal legacy bin-local tsbuildinfo and make sentinel test hermetic (adversarial-review follow-ups)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(changeset): set pr number to 1002
* docs(#1001): record DEFECT.SHARED-ARTIFACT-MUTATION-IN-CONCURRENT-TEST anti-pattern in CONTEXT.md
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#983): rewrite bare .claude paths in Trae/Windsurf converters (Codex/Cline parity)
Both convertClaudeToWindsurfMarkdown and convertClaudeToTraeMarkdown only
handled trailing-slash .claude/ forms; bare ~/.claude and $HOME/.claude
references (e.g. configDir = ~/.claude, RUNTIME_CONFIG_DIR=".../$HOME/.claude")
survived conversion and pointed users at the wrong config dir.
Fix: add bare-form replacements using negative lookahead (?![\w-]) to protect
.claude-plugin and .claudeignore, mirroring Cline (#782) and Codex (#570) precedent.
Also adds CLAUDE_CONFIG_DIR -> WINDSURF_CONFIG_DIR / TRAE_CONFIG_DIR rewrite.
_applyRuntimeRewrites windsurf case gets matching \b-anchored bare-form lines,
mirroring the existing trae case.
Closes#983
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(#983): backfill changeset pr number (995)
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#974): error on graphify --budget with missing/non-numeric value instead of silent NaN no-op
When `--budget` was the last arg or followed by a non-numeric token,
parseInt(undefined/NaN-string, 10) produced NaN. NaN is falsy so both
the router check and applyBudget gate silently skipped budget trimming.
The query ran unbounded with no warning.
Fix: guard in graphify-command-router.cts — if args[budgetIdx+1] is
absent or parses to NaN, emit ERROR_REASON.USAGE and return early.
Defensive fix in graphify.cts: tighten `if (!budgetTokens)` →
`if (budgetTokens == null)` and `if (options.budget)` →
`if (options.budget != null)` so a real 0/NaN caller is handled
predictably by both independent guards.
Regression tests: 16 cases (unit/mock, subprocess, property-based)
covering boundary inputs: missing value, non-numeric, valid integers,
and fast-check properties over the budget parse contract.
Closes#974
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(#974): backfill changeset pr number (986)
* fix(#974): convert static property (b) to real fc.assert property with constrained generator + deterministic seed
The test named "property: --budget as last arg always produces usage error"
was a static test with no fc.assert — it only checked a single hardcoded
term ("someterm") and could never flake or produce a fast-check path. This
is a generator/property bug (case b): the test was mislabeled as a property
test but lacked parameterization.
Root-cause of CI path "46:0:0" / seed 42 failure: a naive parameterized
version without the !startsWith('--') filter could feed term='--budget',
causing args.indexOf('--budget') to hit index 2 (the term slot) rather than
index 3 (the flag slot), placing the router in a different code path. The
property still holds — NaN detection fires on rawBudget='--budget' — but
the assertion text referenced the wrong invariant, making the failure appear
spurious. Fix: constrain the generator to non-flag terms (filter out
strings starting with '--') and pass explicit { seed: 42, numRuns: 200 } to
fc.assert so the test is fully deterministic in CI regardless of GSD_FC_SEED.
No change to src/graphify-command-router.cts (router is correct).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#974): constrain non-numeric budget property generator to genuinely-NaN values and pin fast-check seeds (CI determinism)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#974): use a strict valid-term generator and pin seeds so budget property tests are deterministic
Replace fc.string({minLength:1}).filter(!startsWith('--')) term generators in
properties (b) and (d) with a shared validTerm = fc.stringMatching(/^[A-Za-z0-9][A-Za-z0-9_.-]{0,29}$/)
that is alphanumeric-leading and contains no whitespace, flags, or sign-numerics.
This eliminates the class of CI failures where the old generator produced out-of-
contract inputs (" ", "+5", empty) that the router legitimately rejects for reasons
outside the budget-parse contract under test. Pin distinct seeds (1001–1004) on
every fc.assert for CI determinism. Stress-tested at numRuns=100000 per property
and across 10 seeds (1,2,7,13,42,43,44,99,12345,999999) at numRuns=2000 — all pass.
No src change: gsd-core/bin/lib/graphify-command-router.cjs is correct.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#974): replace flaky term-fuzzing properties with deterministic examples; keep value-fuzzing properties
The validTerm regex /^[A-Za-z0-9][A-Za-z0-9_.-]{0,29}$/ admitted single-digit
strings like "0" which are falsy; the router's `if (!term)` guard fires before
the budget-missing-value path, producing a spurious errFn call. Properties (b)
and (d), which test the BUDGET contract (not term handling), are replaced with
deterministic example loops over fixed valid terms. Properties (a) and (c),
which fuzz the BUDGET VALUE with a fixed term, are kept unchanged (seed+numRuns
pinned). The validTerm generator is fully removed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#978): parse --force in milestone complete dispatcher so the guard's documented override works
The dispatcher built `{ name, archivePhases }` but never parsed `--force`, so
`options.force` was always `undefined` and the guard inside `cmdMilestoneComplete`
(which tells users to "Re-run with --force to override") could never be bypassed.
Add `const force = args.includes('--force')` and pass it into the options object.
The guard already honors `options.force` — no changes to milestone.cts needed.
Closes#978
* chore(#978): backfill changeset pr number (982)
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4)
Migrate the intel CLI command family from a hardcoded gsd-tools.cjs case arm to
a registry-dispatched Capability (commandFamilies mechanism, #961), mirroring
the graphify (#972) and audit (#984) cutovers. New src/intel-command-router.cts
exports routeIntelCommand reproducing all 9 subcommands verbatim (incl. the
status timeAgo non-raw post-processing), lazily requiring intel.cjs inside the
route fn. capabilities/intel/capability.json declares the intel command family;
commands-only (skills:[]), declares the existing intel.enabled gate (default
false — behavior unchanged).
Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing intel.test.cjs passes unchanged. Completes the first-party command-
family cutover sequence (graphify/audit/intel).
Closes#985
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#985): compute expected planningDir via path.join in intel cutover unit tests (Windows CI)
The intel-command-cutover unit-mock assertions hardcoded a POSIX `/.planning`
expectation while the router builds it with path.join(cwd, '.planning') →
backslashes on Windows, so the planningDir-arg assertions (query/status/diff/
snapshot/validate/update/api-surface) failed only on windows-latest CI. Compute
the expectation with path.join (cross-platform); production router unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Migrate the audit-uat + audit-open CLI commands from hardcoded gsd-tools.cjs
case arms to a registry-dispatched Capability (commandFamilies mechanism, #961),
mirroring the graphify cutover (#972). New src/audit-command-router.cts exports
routeAuditUat/routeAuditOpen, each lazily requiring only its backing module
(uat.cjs/audit.cjs) inside the route fn — matching the old per-case lazy loads.
capabilities/audit/capability.json declares the two command families;
commands-only (skills:[]), no config gate, audit_review cluster untouched.
Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing audit regression tests pass unchanged.
Closes#981
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Migrate graphify into a Capability that owns its `graphify` command family,
dispatched via the registry (#961 mechanism) instead of a hardcoded case.
graphify is now an enable/disable plug-in.
- src/graphify-command-router.cts: routeGraphifyCommand (standard route*Command),
reproduces the removed case EXACTLY (query +--budget, status, diff, build,
hidden build snapshot, usage/unknown errors); injectable _graphify test seam.
- capabilities/graphify/capability.json: role feature, tier:full, skills:[graphify],
config:{graphify.enabled default false}, commands:[{family:graphify, module,
router:routeGraphifyCommand}].
- removed case 'graphify' from gsd-tools.cjs; graphify now flows
default -> dispatchCapabilityCommand -> commandFamilies.graphify -> router.
- regenerated registry (commandFamilies/bySkill/configSchema/profileMembership/
capabilityClusters for graphify); tier:full keeps 4c install/surface a no-op.
Equivalence-proven: 8 recording-mock unit tests assert the exact fn+args per
subcommand (budget, snapshot-vs-build); 9 subprocess tests assert distinguishing
output shapes; existing graphify tests pass unchanged through the new path.
Surfaced (not silently accepted): the pre-existing --budget-no-value NaN no-op
quirk, preserved for equivalence, filed separately.
Closes#972
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#961): capability command mechanism — commandFamilies index + default-case dispatch (ADR-857 phase 4d-impl-1)
Build the capability command mechanism per ADR-959: the `commands` declaration
field on the feature role, the registry commandFamilies index, and a real
dispatchCapabilityCommand consulted in runCommand's default case (replacing the
dead _dispatchNonFamily shim's role).
The registry DISCOVERS a standard route*Command (no rebuilt handler table). On a
default-case command, dispatch enforces a bare-.cjs-basename, resolves the module
under gsd-core/bin/lib/, asserts confinement, requires the resolved path, and
own-property-guards the router export before calling it. A require.main===module
guard makes gsd-tools.cjs importable for tests; the CLI path is unchanged.
Additive: commandFamilies is empty today, so the default case is behavior-
preserving for every command; the 10 dead _dispatchNonFamily sites are untouched
(future migration markers). The graphify cutover is the separate next step.
Closes#961
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#961): surface capability router failures structurally + enforce sync contract
Review follow-up: dispatchCapabilityCommand now wraps the router invocation so an
unexpected (non-ExitError) throw is converted to a structured, attributed
error(msg, SDK_FAIL_FAST) — honoring --json-errors — instead of escaping as a raw
stack trace; an intentional ExitError propagates unchanged. An async router
(returns a thenable) is rejected loudly with a structured error (the contract is
synchronous, like the 12 host routers). Not shipped as "consistent with existing
behavior": the host's pre-existing version of this gap is filed as #965.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#851): correct Codex adapter for generic multi_agent_v1 schema
The Codex skill adapter header in getCodexSkillAdapterHeader() documented
typed spawn_agent(agent_type=...) as a direct, unconditional mapping for all
Task()/Agent() calls. In sessions exposing only the generic multi_agent_v1
schema (message/items/fork_context — no agent_type field), this mapping is
silently invalid: the orchestrator cannot natively dispatch typed gsd-planner/
gsd-executor agents and may fall back to inline execution or produce errors.
Fix: Section C now requires schema detection before spawning. It documents the
typed mapping as conditional on the agent_type-capable schema (e.g. multi_agent_v2)
and introduces an explicitly-labeled generic-agent workaround for multi_agent_v1
sessions — read the agent TOML, inject its instructions as a role-preamble, and
call spawn_agent(message=...) — clearly marking the result as NOT equivalent to
typed gsd-planner/gsd-executor execution.
Regression test: tests/bug-851-codex-quick-adapter-agent-type-fallback.test.cjs
asserts schema-awareness language, the multi_agent_v1 fallback, the workaround
label, and backward compat with the existing bug-279 typed-spawn contract.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: add changeset for PR #958
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#851): keep Codex adapter block consistent with materialized skill surface
The `~/.codex/agents/<agent-name>.toml` literal introduced in the #851 prose
was being rewritten to the real install path by `_applyRuntimeRewrites` (the
`~/.codex/` → pathPrefix substitution) before the SKILL.md was written to
disk. `getCodexSkillAdapterHeader()` still returned `~/.codex/agents/...` so
the test assertion (exact match between builder output and materialized file)
always failed.
Fix: replace the `~/.codex/agents/` literal with the runtime-neutral form
`agents/<agent-name>.toml` plus a parenthetical naming `$CODEX_HOME/` — which
is not matched by any rewrite pattern and survives the path-substitution step
unchanged. The #851 schema-detection + generic-subagent-fallback intent is
fully preserved.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#851): resolve active Codex config root in fallback; strengthen tests (adversarial review)
- Rewrites the generic-agent workaround step 1 to explicitly describe
active config root resolution (priority: $CODEX_HOME → --config-dir →
--local .codex → default global dir) without the literal ~/.codex/
substring that _applyRuntimeRewrites replaces, preventing bug-3582
divergence.
- Replaces OR/loose-includes test assertions in bug-851 with AND-logic
checks covering all four required elements: (a) schema-detection step,
(b) active-config-root resolution for the TOML path including all three
override mechanisms, (c) NOT-equivalent-to-typed-gsd-planner/gsd-executor
label, and (d) fail-closed rule when typed dispatch is mandatory.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#851): register bug-851/947/948/950 in lint-regression-test-names allowlist
The ratchet (622e4be) bans NEW top-level bug-NNNN test files; the four
sibling PRs (#851, #947, #948, #950) landed AFTER the baseline was cut,
so their test files were not yet grandfathered. Add all four to the
identity allowlist so lint-regression-test-names passes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* test(#948): add regression tests for no-op write guard and record-session auto-create (#944)
Red before fix: 11/15 tests fail. Green after: 15/15.
Covers zero-match patch byte-identity, milestone_name preservation,
stopped_at frontmatter-wins, record-session auto-create fallback, and
adversarial fixtures (CRLF, empty body, non-canonical labels).
Also registers bug-948-state-noop-write-guard.test.cjs in the state
bucket of lint-test-file-count.allowlist.json.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#948): guard STATE.md no-op writes; preserve milestone_name/stopped_at (closes#944)
Shared root cause: `readModifyWriteStateMd` wrote STATE.md unconditionally
even when the transform produced no change, and `syncStateFrontmatter`
re-derived frontmatter from the possibly-stale body on every write.
Three coordinated fixes in src/state.cts:
1. readModifyWriteStateMd: add no-op guard — when transform result ===
input content, skip the write entirely (no platformWriteSync, no
last_updated bump, no frontmatter re-derive). Fixes#948 zero-match
phantom write and the #944 phantom last_updated bump.
2. syncStateFrontmatter: extend existing-frontmatter preserve logic —
fall back to existingFm['milestone_name'] / existingFm['milestone']
when the derived value is the template placeholder 'milestone'
(getMilestoneInfo returns this literal when it cannot match the
version in ROADMAP.md); prefer existingFm['stopped_at'] /
existingFm['paused_at'] over a body-derived value (the frontmatter
value, written by the canonical record-session path, wins over stale
historical body lines). Mirrors the fallback already in cmdStateJson.
3. cmdStateRecordSession: when --stopped-at / --resume-file are supplied
but body labels are absent, DWIM auto-create a canonical ## Session
section (mirroring how add-decision / add-blocker / record-metric
auto-create their sections). Never return a silent recorded:false when
the caller supplied values.
SDK check: no sdk/src/state.ts exists in this repo (the comment in
cmdStateSnapshot references a sibling concern in the TypeScript SDK
codebase, which is a separate repo not present here).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: add changeset for PR #952 (fix #948/#944)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#948): correct stopped_at preserve rule; adjust test for sync behaviour
The "always prefer frontmatter stopped_at" rule in syncStateFrontmatter
was too aggressive — it broke phase.complete which intentionally updates
stopped_at in the body and expects syncStateFrontmatter to pick it up.
The primary fix (no-op guard in readModifyWriteStateMd) already prevents
the stale-body-overwrites-frontmatter scenario from #948: the file is not
written when the transform produces no change, so syncStateFrontmatter
never runs on a zero-match patch. The body-derived value can only win when
an actual write occurs, which means the body was legitimately updated.
Reverted to the original #905 rule for stopped_at/paused_at: fall back to
existing frontmatter only when the derived value is absent (empty/null).
Also adjusted the sync-suite test to assert what state sync actually does
(milestone_name preservation) rather than a stopped_at-wins property that
state sync does not have by design.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#944): update existing session block in place (adversarial review)
HIGH finding: the DWIM auto-create in cmdStateRecordSession was appending
a second ## Session block unconditionally, even when one already existed
with non-canonical content (e.g. a markdown table). Both
buildStateFrontmatter and cmdStateSnapshot read only the FIRST ## Session
block via regex, so the newly-written Stopped at / Resume file values
landed in the second, invisible block — frontmatter stopped_at stayed
stale and state-snapshot returned nulls.
Fix: check for an existing ## Session heading. When one is present,
normalize that section in place by replacing its body with canonical
**Last session:** / **Stopped at:** / **Resume file:** bold-label lines.
Only append a brand-new section when NO ## Session heading exists.
LOW finding: the auto-create scaffold emits **Last session:** but
cmdStateSnapshot only matched **Last Date:**, so session.last_date was
null after auto-create despite a valid timestamp being written.
Fix: extend the lastDateMatch regex in cmdStateSnapshot to also accept
**Last session:** / Last session: (the form the scaffold writes).
Tests: 3 new tests added to bug-948-state-noop-write-guard.test.cjs that
confirmed failure against the previous HEAD and pass after this fix:
- exactly one ## Session block after record-session with non-canonical existing block
- state-snapshot sees correct stopped_at via first Session block (not a duplicate)
- state-snapshot session.last_date is non-null after auto-create on body-less file
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#944): improve in-place section replace to cleanly remove old body content
The previous regex `/(^## Session[ \t]*$)([\s\S]*?)(?=\n^## |\n*$)/im`
with a lazy match consumed nothing after the heading, so old non-canonical
body content (e.g. table rows) remained after the new canonical lines.
While functionally correct (parsers found the canonical lines first in the
FIRST ## Session block), it left stale content in the section. Replace with
a negative-lookahead per-line pattern that consumes all content from the
heading up to (but not including) the next ## heading, producing a clean
section with only the canonical bold-label lines.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The three shell scanners exempt their own adversarial test fixtures by exact
filename; the *.security.test.cjs renames broke those entries, so the PR diff
scan flagged the scanners' own test payloads. Verified locally with all three
scanners in --diff origin/next mode (0 findings) and the security suite
(207/207). The .sh files were missed in the original reference sweep because
the rename grep filtered to .cjs/.yml/.json/.md extensions.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
13 bug-* files landed upstream between the audit baseline and this branch's
rebase; they predate the ratchet policy, so they are grandfathered.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
244 one-off bug-* files (~38% of the suite) are grandfathered in
lint-regression-test-names.allowlist.json; new ones fail lint with
fold-into-module guidance, and deletions force allowlist pruning so the
baseline only shrinks. Wired into npm run lint:ci (new single entry point
for every CI lint). Policy documented in docs/TESTING-SUITES.md.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Renames (git mv) with all references updated (ci-test-scope RULES,
windows-parity allowlist, test-file-count allowlist, docs in 6 locales):
- 5 scanner tests -> *.security.test.cjs — the 'Run security tests' CI step
ran zero files since the suite taxonomy landed; it is now honest.
- graphify-auto-update -> *.slow.test.cjs (36s, slowest file in the suite;
e2e gsd-tools spawns) — runs on full-matrix lanes and push to next.
- installer-migration-install-integration -> *.integration.test.cjs
(13s; an integration test by its own name).
Coverage gate measured after retags: 88.55% lines (gate 70%).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
full_matrix fired on 15/15 sampled PRs because any tests/** change forced it,
costing ~25 runner-minutes each. A changed test file now always joins the
scoped windows lane (covering the #482 OS-specific failure class per-file)
and still runs on ubuntu 22/24 via targeted_tests; the residual macOS /
windows-node-22 cross-product is covered on every push to next.
Also narrows WINDOWS_HINTS from 6 substrings (102/633 files, a ~10-minute
scoped lane) to windows/win32/shell/path — the dropped hints (workflow,
install, hook) are either platform-independent lint tests or already covered
by fullMatrix rules.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Establish `tier` as the source of install-profile + cluster membership
(ADR-894 §4). The registry now derives two views: capabilityClusters
(capId → its skills) and profileMembership (capId → {tier, profiles}, where
profiles is the PROFILE_RANK suffix from the capability's tier). A consistency
gate cross-checks them against the hand-authored PROFILES/CLUSTERS: HARD (throws)
on a capId-matching-a-cluster-name with a different skill set; SOFT
(pending-reconciliation stderr warning, not serialized) when a capability skill
isn't yet in the closure-resolved hand-authored profile.
The SOFT gate loads the real skills manifest and resolves each profile's closure
once so transitively-included skills don't false-warn; warnings are de-duped to
one per (capability, skill); both derived views are scoped to skill-owning
capabilities; serialized with a global capId sort for determinism; reserved-name
guards at every write site; lazy requires of the built constants.
Behavior-preserving: install/surface untouched; derived views consumed by
nothing. Full generation rides along with the phase-6 migration.
Closes#942
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
- bin/install.js now copies scripts/changeset/ and scripts/lib/ into
<configDir>/scripts/ so $GSD_DIR/scripts/changeset/cli.cjs resolves
at runtime; aborts install with an explicit failure if the source
directory is missing from the package.
- gsd-core/workflows/update.md: corrected path from
gsd-core/scripts/changeset/cli.cjs to scripts/changeset/cli.cjs;
added an explicit [ ! -f ] guard so a missing CLI surfaces a clear
message rather than silently swallowing the error; stderr captured
via 2>&1 sentinel so node errors are visible in the preview output.
- release.yml's changeset-CLI invocations (node scripts/changeset/cli.cjs)
remain at the repo-root path and are unaffected by this change.
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Build the federated config merge: each capability owns its config-key slice
(ADR-857 decision 3 / ADR-894), and loadConfig merges them defensively. The
registry now emits a full configSchema index ({key:{owner,type,default,
description}}, generator-validated); a new src/federated-config.cts resolves
federated keys defensively (skip central keys -> pending-migration warning,
skip malformed slices -> warning never throw, else type-checked user override
?? default, with nested dotted-path lookup and enum validation); and loadConfig
applies the overlay on every return path.
Wired as a provably-empty no-op channel: every UI-pilot key is still central,
so validKeys is empty and loadConfig returns byte-identical output on all paths
(identity return when the overlay is empty; shared CONFIG_DEFAULTS never
mutated). Registry-only; no key is cut over; nothing in the live loop changes.
Closes#910
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
syncStateFrontmatter was silently dropping current_phase, current_phase_name,
current_plan, and progress when body annotations were absent (e.g. after an
agent or tool rewrote the body). These scalars can only be derived from body
annotations — when absent, buildStateFrontmatter returns nothing for those
keys. Added existingFm fallbacks mirroring the same pattern already applied in
cmdStateJson, so every writeStateMd call preserves the existing values instead
of stripping them. Also extended cmdStateJson with the same fallbacks for the
three non-progress scalars.
Adds regression test (7 cases) + lint-test-file-count allowlist entry.
Closes#905
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
buildRoadmapPhaseVariants() only matched heading-style phases (## Phase N:),
silently skipping the supported checklist format (- [x] **Phase N: name**).
This caused W007 false-positives for every on-disk phase dir when the project
uses a checklist ROADMAP. Fix adds a second regex pass (mirroring the existing
buildNotStartedPhaseVariants() approach). Also refactors the duplicate
inline heading-only regex in cmdValidateConsistency() to delegate to
buildRoadmapPhaseVariants() (DRY). Regression test in
tests/bug-892-validate-checklist-roadmap-phases.test.cjs covers both paths.
Closes#892
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Replace the inline LOOP_HOST_CONTRACT constant in the Capability Registry
generator with a generated-from-workflows contract (ADR-894 §3). The contract
is now derived from inert `<!-- gsd:loop-host ... -->` marker blocks in the five
step workflows, emitted as the committed gsd-core/bin/lib/loop-host-contract.cjs,
and required by gen-capability-registry.cjs — one source of truth, no drift.
Drift guards in gen-loop-host-contract.cjs: per-step point ownership (each step
must declare exactly its canonical loop points), multiple-block + duplicate-key
hard errors, and a word-boundary agent-role cross-check. Contract content is
byte-identical to the former constant; registry-only, nothing wired into the
live loop.
Closes#903
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#896): Capability Registry generator + UI pilot (ADR-857 phase 3a-impl)
First phase-3 code: the Capability Registry generation pipeline, built against
the ADR-894 contract and NOT wired into the live loop (registry-only, per the
staged-cutover design).
- capabilities/ui/capability.json — the UI pilot (ADR-894 worked example): 2
skills, 2 agents, 3 config keys, 2 steps + 1 gate, with `when` activation.
- scripts/gen-capability-registry.cjs — --write/--check generator. Hand-rolled
schema validation (envelope + role-typed feature/runtime bodies + typed
steps/contributions/gates + when + gate-check variants); cross-capability
invariants (single ownership; requires exist+acyclic+tier-monotone; config-key
ownership exclusive, collision-vs-central as a pending-migration warning);
hooks validated against an inline LOOP_HOST_CONTRACT (3a-impl-2 swaps its
source to the generated-from-workflows contract); GLOBAL point-ordered
consumes-satisfiability; materialized byLoopPoint ordering (produces/consumes
topo-sort); emits gsd-core/bin/lib/capability-registry.cjs (role-partitioned
indexes + requiresClosure). Prototype-pollution guards (Object.create(null) +
inline literal key checks) + fragment.path traversal guard.
- gsd-core/bin/lib/capability-registry.cjs — committed generated artifact
(mirrors package-identity.cjs: script-generated, tracked, linted, regenerated
on build, drift-tested), wired via the new `gen:capability-registry` build step.
- tests/capability-registry.test.cjs — 72 tests: schema + invariant + hook +
ordering + adversarial (path-traversal, proto-pollution, runtime body,
self-consume, cycles, collisions) + committed-file staleness guard.
New-CLI-module checklist (INVENTORY 97->98, MANIFEST, ARCHITECTURE), CONTEXT.md
"Capability Registry" un-[Planned]'d. Nothing wired into install/surface/loop.
Gates: lint, code-review (4 bugs fixed), security-review (path-traversal +
prototype-pollution fixed), codex adversarial-review ×3 (8+ findings fixed,
confirmed sound), clean-build docker 13190 pass / 0 fail.
Closes#896
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#896): CRLF-agnostic --check for capability-registry staleness (Windows)
The committed capability-registry.cjs staleness guard failed on Windows CI only:
git checks out the committed .cjs as CRLF (autocrlf, no .gitattributes) while the
generator emits LF, so the byte-for-byte --check comparison mismatched. Normalize
line endings on both sides of the --check comparison (no .gitattributes change,
no change to the LF the generator writes). Adds a regression test simulating the
Windows CRLF checkout.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Emit the 6 gsd-ns-* routers as the only top-level skill bundles and nest
the ~61 concrete skills under <router>/skills/<name>/SKILL.md on runtimes
with confirmed non-recursive skill loaders (claude global, cline, qwen,
hermes, augment, trae, antigravity). Router bodies rewrite their routing
tables from Skill-tool dispatch to a Read skills/<name>/SKILL.md pattern.
Recursive/unconfirmed loaders (cursor, codex, copilot, windsurf, codebuddy,
opencode, kilo) keep the flat layout. Completes the v1.40 namespace
architecture (#2792) so the eager skill listing drops to ~6 entries.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
extractCurrentMilestone reads STATE.md via planningDir(cwd), which is
workstream-aware (honours GSD_PROJECT/GSD_WORKSTREAM). The fixtures write
STATE.md to the plain <tmp>/.planning/STATE.md, so a developer shell inside a
GSD workstream (GSD_WORKSTREAM exported) redirected the read to a non-existent
workstream subdir -> version=null -> closed milestone sections leaked into the
slice and assertions failed. Clean CI/Docker env never hit it. Not a Node-26
regex bug; reproduces identically on any Node with GSD_WORKSTREAM set.
- scripts/run-tests.cjs: strip GSD_PROJECT/GSD_WORKSTREAM before spawning test
children so the local runner env matches clean CI/Docker.
- tests/roadmap-phase-fallback.test.cjs: file-level beforeEach/afterEach
save/delete/restore of both vars; new regression test pinning workstream-aware
STATE.md resolution.
- tests/run-tests-harness.test.cjs: guard asserting the runner strips both vars
(so removing the deletion fails clean CI).
Closes#872
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The PR Gate workflow's only job, size-check, labeled every PR with
size/S–size/XL based on lines changed. Those labels aren't used in any
review, triage, or automation flow, so the workflow was pure noise.
- Delete .github/workflows/pr-gate.yml
- Drop size-check from required status checks in both rulesets so PRs
don't block forever on a check that never reports
- Remove pr-gate.yml from INERT_WORKFLOWS (ci-test-scope.cjs) and the
knownInert list (ci-test-scope.test.cjs)
- Remove "PR Gate / size-check" from setup-branch-protection.sh
Closes#846
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#844): sync runtime manifest versions on npm version bump
The release workflow bumps package.json via `npm version` but never
stamped the runtime-integration manifests that must track it
(.claude-plugin/plugin.json #766, gemini-extension.json #775), so the
first RC/finalize whose version diverged from the -dev stream failed the
test suite before tagging/publishing.
Add scripts/sync-manifest-versions.cjs (single VERSIONED_MANIFESTS
registry) wired to a `version` npm lifecycle hook that stamps + stages
the manifests on every `npm version` — covering all four release bump
sites and local bumps with no workflow edits. A regression guard test
fails if any repo JSON whose version matches package.json is not
registered, forcing future version-bearing manifests into the sync.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#844): add changeset for manifest version sync fix
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#770): register Claude Code lifecycle hooks (SubagentStop/Stop/PreCompact/FileChanged)
Wire three new context-tracking events (SubagentStop, Stop, PreCompact) to
gsd-context-monitor so context-headroom warnings surface at model-stop and
subagent-finalisation moments — not just on PostToolUse. Add a new
FileChanged hook (gsd-config-reload.js) that hot-reloads .planning/config.json
context mid-session when the user edits it, injecting a config summary as
hookSpecificOutput.additionalContext. Updates plugin manifest hooks.json,
managed-hooks-registry, installer-migration-report allowlist, and
shell-command-projection cleanup tables. Tests: 21 new assertions in
enh-770-claude-hook-events.test.cjs; enh-788 and issue-766 test suites updated.
Closes#770
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#770): document newly-registered Claude Code lifecycle hooks
Add a Hook coverage table to the Claude Code npm installer section of
docs/how-to/install-on-your-runtime.md describing SubagentStop, Stop,
PreCompact, and the new FileChanged (gsd-config-reload.js) hook that
hot-reloads .planning/config.json mid-session. Also fixes the changeset
frontmatter (adds type: Added + pr: 821) so docs-lint can consume the
fragment.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#770): add gsd-config-reload.js to INVENTORY.md and regenerate manifest
The feat commit added hooks/gsd-config-reload.js but did not bump the
Hooks count in docs/INVENTORY.md (14→15) or add the new row, and did not
regenerate docs/INVENTORY-MANIFEST.json. Both inventory-counts and
inventory-manifest-sync tests failed across the full CI matrix.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#770): make lifecycle-hook tests deterministic on scoped runner
Replace the shared hooks/dist/ ensemble setup (ensureHooksDist /
teardownHooksDist) in the Claude hook tests with per-test isolation:
pre-populate each test's own tmpDir/.claude/hooks/ with stub files and
pass installerMigrations:[] to install() so the first-time-baseline
migration does not remove the stubs before the copy step can run.
Root cause: hooks/dist/ is gitignored and absent on a fresh npm ci.
ensureHooksDist() created it and teardownHooksDist() deleted it, but
with --test-concurrency=4 both test files ran concurrently as separate
Node.js worker processes sharing the same filesystem. One file's
afterEach teardown deleted hooks/dist/ while the other file's install()
was copying from it, producing an ENOENT (reproduced 2/10 runs locally).
The additional issue: even with pre-placed stubs surviving the copy race,
the 000-first-time-baseline migration classified hooks/gsd-*.js as
bundled-gsd-hook artifacts, auto-removed them, and the copy step never
re-ran (hooks/dist/ absent) — leaving contextMonitorFile missing and all
hook registrations silently skipped (the 'got: []' symptom).
Fix: pre-populate targetDir/hooks/ per-test (isolated temp dir) AND pass
installerMigrations:[] so the baseline scan is skipped. The Qwen suites
already used this pattern correctly; the Claude suites are aligned to it.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#770): ship gsd-config-reload.js by adding it to build-hooks HOOKS_TO_COPY
The #770 feature added hooks/gsd-config-reload.js and registered it in
MANAGED_HOOKS, the installer, INVENTORY, and the test EXPECTED_ALL_HOOKS
list — but never added it to scripts/build-hooks.js HOOKS_TO_COPY. As a
result the hook was never copied into hooks/dist/ during the build, so:
- the hook would never ship to users (real production bug — the
FileChanged config-reload feature was dead-on-arrival), and
- install-minimal-hooks.test.cjs #1755 ("all expected hooks are copied
from hooks/dist/ to target", ".js hooks are executable after copy",
"manifest contains .js hook entries") failed on any environment with
a clean checkout (no pre-existing hooks/dist/): coverage, full test
macos-22/macos-24, test ubuntu-24.
The failures were masked locally only by a stale hooks/dist/ left from a
prior build (build-hooks copies into dist without clearing it). On CI's
fresh `npm ci` there is no dist, so the omission surfaced.
Fix: add 'gsd-config-reload.js' to HOOKS_TO_COPY so build-hooks stages it
into hooks/dist/ alongside the other JS hooks. Verified by removing
hooks/dist/ and rerunning the full suite green (0 fail).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#770): make config prototype-pollution beforeEach deterministic on scoped runner
Root cause: the #663 and alert-#26 prototype-pollution describe blocks
seeded .planning/config.json in beforeEach via a bare
runGsdTools('config-ensure-section') whose result was discarded. That
command runs in a spawned gsd-tools child; on the scoped CI lane
(--test-concurrency=4, config.test.cjs scheduled alongside the heavy
install/tarball suites that #770 pulled into the targeted set) the child
can be transiently killed under resource pressure (non-zero exit, empty
stderr — an OS-level kill, not an app error). The swallowed failure left
config.json absent, so the first subtest's readConfig() threw ENOENT
opening <tmp>/.planning/config.json. Only 1 of 4 subtests failed,
confirming a per-invocation transient, not a deterministic miss; the full
suite schedules files differently so config.test.cjs did not collide with
those heavy neighbors → passed there.
Fix: add ensureConfigReady(tmpDir) which retries config-ensure-section on
ANY failure or missing file and throws a clear diagnostic if it still
cannot create config.json, then use it in both prototype-pollution
beforeEach blocks. Setup is now deterministic under load; the #663/alert-#26
security assertions are unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#836): no-LLM duplicate-issue detection + challenge + 1-day auto-close
Adds a deterministic (no-LLM) duplicate-issue governance lifecycle:
- scripts/issue-dedupe.cjs: pure, unit-tested module (tokenize, Sørensen–Dice
title similarity, scoreCandidates, renderChallengeComment, shouldClose) with
fail-safe destructive-action guards.
- duplicate-check.yml (issues:opened): scores new-issue title against open
issues, posts a challenge comment + applies the pending `possible-duplicate`
label on a clear match.
- duplicate-sweep.yml (daily cron): closes possible-duplicate issues whose
challenge comment is >24h old with no human reply and no 👎 veto; honors
exempt labels; re-checks the label immediately before close (TOCTOU guard);
strips the label on close to avoid reopen loops.
- remove-duplicate-label.yml (issue_comment:created): clears the label and
applies needs-maintainer-review when any human responds.
- bug_report.yml / docs_issue.yml: add the required "I searched existing
issues" preflight checkbox so all five forms force a pre-search attestation.
- docs/agents/triage-labels.md: document the label + lifecycle.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#836): add changeset fragment for duplicate-issue detection
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
CI test-scope detection diffed changed files with a two-dot
`git diff --name-only base head`, where base is the moving tip of `next`.
A PR branch cut from a slightly older `next` surfaced every product file
`next` had gained since the merge-base, flipping product_changed/full_matrix
and running the full Windows/macOS matrix + coverage on docs-only PRs.
Switch to a three-dot `git diff --name-only base...head` (vs the merge-base),
matching GitHub's PR "Files changed" semantics. Add a regression test that
builds a stale-base topology, plus a guard test pinning `fetch-depth: 0` on
the `changes` job (required for the merge-base to be locally available).
Closes#837
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#777): register Cursor-native hooks (.cursor/hooks.json) for session-start/post-tool parity
- Add gsd-cursor-session-start.js: injects STATE.md presence reminder (or
new-project nudge) into Cursor sessions via the sessionStart hook event
- Add gsd-cursor-post-tool.js: emits an additional_context nudge when
write-class tool calls touch .planning/ files (postToolUse hook event)
- Add 'cursor-hooks-json' installSurface to runtime-config-adapter-registry;
writeCursorHooksJson/reconcileCursorHooksJson write the canonical
{ version: 1, hooks: { sessionStart, postToolUse } } JSON shape with
idempotent reconciliation that preserves user-owned hook entries
- Hook scripts are copied with /gsd:→gsd- rewrite so installed files
contain no colon-form slash-command refs (bug-376 invariant)
- 20 new tests in tests/cursor-hooks.test.cjs cover all reconciler paths,
entry helpers, removal, runtime adapter surface, and hook script behavior
- Update CONTEXT.md, ARCHITECTURE.md, installer-migrations.md, and
000-first-time-baseline.cts to include Cursor hooks.json surface
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#777): build hooks/dist on demand in bug-376 test for scoped/windows CI
hooks/dist is gitignored and only produced by `npm run build:hooks`.
The CI scoped (ubuntu-latest/node-22) and windows (windows-latest/node-24)
test jobs do NOT run build:hooks before executing tests, so bug-376's
prerequisite suite was failing with "hooks/dist not found" on both legs.
Add ensureHooksDist() helper (mirrors bug-3357 pattern) that builds
hooks/dist on demand in the before() hooks of prerequisite and Suite 3.
Also add ensureHooksDist() call to Suite 3's before() so the snapshot
step is also hermetic.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#776): Gemini hook events (BeforeAgent/AfterAgent/BeforeModel) + hooksConfig.enabled check
Register three new Gemini-CLI hook events on install:
- BeforeAgent: fires before agent planning; wired to gsd-context-monitor
- AfterAgent: fires after final response generation; wired to gsd-context-monitor
- BeforeModel: fires before each LLM call (per-turn); wired to gsd-context-monitor
All three reuse gsd-context-monitor.js (no new hook files). Uninstall cleanup
loop extended to remove the new events. Non-array guard added for robustness
against malformed settings.
Also detect hooksConfig.enabled:false in Gemini settings and emit a clear
warning — without this check, all registered hooks silently do nothing.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: update changeset pr: 829
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#776): document Gemini hook events
Add hook coverage table to the Gemini CLI section of install-on-your-runtime.md,
covering the three new events (BeforeAgent/AfterAgent/BeforeModel wired to
gsd-context-monitor) plus a callout for the hooksConfig.enabled:false silent
failure mode detected by the installer.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#771): convert agent color: hex/magenta values to documented named colors
Claude Code's sub-agent `color:` field documents only 8 named colors
(red, blue, green, yellow, purple, orange, pink, cyan). Twelve agent
files used hex values and two used the undocumented `magenta`; convert
each to the nearest documented named color so the intended per-agent
TUI color differentiation is spec-compliant.
- agents/*.md: 14 color values hex/magenta -> nearest named color
- scripts/research-profiles.cjs: update the 3 generated research-agent
profiles (source of truth) so gen-research-agents stays in sync
- docs/AGENTS.md: update documented colors; add missing Color rows for
gsd-nyquist-auditor, gsd-project-researcher, gsd-phase-researcher
- tests/agent-frontmatter.test.cjs: add regression guard asserting every
agent color: is in the documented named-color set
Closes#771
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#771): add changeset
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
test.yml had no paths filter and the ci-test-scope classifier treated docs/
and every .github/workflows/* as code_changed, so documentation edits and
product-irrelevant automation tweaks still spun up the full Linux/Windows/macOS
matrix. Narrow the heavy matrix to changes that can actually affect the product
or the test pipeline.
- ci-test-scope.cjs: drop docs/ from code_changed (docs-only -> full skip; the
required-tests fan-in still reports green). Add src/ to code_changed (it was
missing -> a source-only PR previously skipped all tests). Add INERT_WORKFLOWS
allowlist + isInertCi() + an "inert CI" rule, and a product_changed output that
gates the heavy test/coverage jobs. Fail-safe: any workflow not on the inert
allowlist defaults to the full matrix. A module-load assertion throws if a
PROTECTED_WORKFLOWS entry (test/install-smoke/mutation/security-scan/release)
is ever added to the inert set, so a weakening edit fails CI loudly.
- test.yml: keep the static 3-lane matrix (so the H1 shell-policy linter can
still statically verify the Windows lane), gate test/coverage on
product_changed, add a lightweight ubuntu-only test-inert job, and branch the
required-tests fan-in on product_changed.
- docs-required.yml: run docs-parity-live-registry (gated on docs/ changes) so
pure-docs PRs still catch live-registry drift without the matrix.
- tests: cover docs-only, inert-only, src/, pipeline, unknown-workflow fail-safe,
mixed escalation, the code_changed=false -> no-lanes invariant, and protected-
workflow tamper-evidence.
Closes#764
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The rc action publishes a release candidate to @next for testing but
never surfaces the curated CHANGELOG section for the version under test —
render only runs destructively at finalize (#715), so there was no safe
way to preview the upcoming notes during the RC window.
Add a --preview mode to scripts/changeset/cli.cjs cmdRender: it renders
the dated release section to stdout via the existing renderChangelog/
serializeChangelog path (with priorChangelog: null, so only the new
section is emitted), reuses the shared injectEmptyPlaceholder helper for
zero-fragment releases, and returns WITHOUT writing CHANGELOG.md or
deleting any .changeset fragment. Wire a "Preview CHANGELOG" step into
the rc job that renders to a file (standalone command, so a malformed
fragment fails the step) and cats it to the job summary and log.
Closes#759
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
`extractCurrentMilestone()` scoped the current-milestone window to its
`## Phases` checklist subsection and terminated at the milestone's own
`## Milestone … (Phase Details)` heading, so the `### Phase N:` detail
headers fell outside scope. Every parser-backed command — `init.phase-op`
(and thus `/gsd:discuss-phase`, `/gsd:plan-phase`), `state`, `roadmap list`,
and `validate health` (W006) — therefore could not resolve phases of any
milestone after the first until a `.planning/phases/` directory already
existed, blocking discuss/plan.
The parser now additionally includes the current milestone's `(Phase Details)`
section in scope, located via the already-computed version matches and anchored
(boundary-aware) to the selected milestone's version token so sibling
sub-milestones sharing a version prefix do not cross-pollinate. The existing
heading selection and primary window are unchanged.
Adds tests/bug-730-milestone-phase-details-scope.test.cjs covering the
two-milestone reproduction, first-milestone non-regression, direct
getRoadmapPhaseInternal resolution, validate-health W006 visibility, a
three-milestone roadmap, and the closed-sibling sub-milestone case.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Part 1 of 2 of the n/no-process-exit cleanup (umbrella #738): convert every
process.exit() call in standalone scripts/** CLIs to the rule-compliant pattern.
- New shared helper scripts/lib/cli-exit.cjs: ExitError(code,message) + runMain()
which translates a thrown ExitError / returned number into process.exitCode
(never process.exit()), flushing output and still firing process.on('exit').
- main()-based entrypoints: throw new ExitError(code) for errors, return <code>
for verdicts; invoked via runMain(main). Child exit codes preserved via return.
- top-level-only scripts: imperative body extracted into main() so mid-flow
aborts (throw ExitError) actually halt; pure consts/helpers stay at module scope.
- diff-touches-shipped-paths.cjs: stdin event handling restructured to an async
read so the whole flow runs under runMain; uncaughtException/unhandledRejection
nets replaced by an in-band catch that preserves EXIT_ERROR=2.
Exit codes verified unchanged for every converted script (success/error/help and
the 0/1/2 semantic codes in diff-touches). Rule stays warn here; flipped to error
in part 2 (#738) once gsd-core/bin/** is also clean.
Refs #739
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Pay down pre-existing error→warn lint debt. Removes dead imports/vars, unused functions, redundant regex/string escapes, and stale eslint-disable directives; converts unused `catch (_e)` to optional catch binding (src/*.cts).
No behavior change. Lint 345→125 warnings (0 errors); deferred categories (n/no-process-exit, test-sleeps, control-regex) tracked in #732 for follow-up. Full test suite green (0 failures); code-review verified all removals unused and all escape fixes semantics-preserving.
Closes#732
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>