Commit Graph

177 Commits

Author SHA1 Message Date
Tom Boucher
1fab2e10ba fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver (#1045)
* fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver

Source agents/*.md (gsd-planner, gsd-executor, gsd-verifier, gsd-plan-checker,
gsd-intel-updater, gsd-debugger, …) called bare "gsd-tools …" in shell blocks.
On a shim-only install — where gsd-tools.cjs exists under the runtime home but
gsd-tools is NOT on PATH — those calls fail with "command not found" and the
agent silently skips init/state/validate/commit ceremony, deferring to the
orchestrator or bypassing GSD bookkeeping entirely.

were never migrated, so it persisted on Claude Code and every other runtime that
consumes the source agents directly. Only gsd-phase-researcher.md carried a
resolver — and a stale, claude-only truncated one.

Fix (all runtimes):
- Inject the canonical multi-runtime gsd_run preamble (byte-equal to
  _runtime-launcher.snippet.sh — claude/codex/cursor/gemini/copilot/windsurf/
  augment/trae/qwen/cline/opencode/kilo/hermes/antigravity homes) at the top of
  the first gsd_run block of all 12 gsd-tools-calling agents, and rewrite every
  command-position bare gsd-tools to gsd_run.
- Upgrade gsd-phase-researcher.md's stale resolver to the canonical one.
- Extend scripts/sync-runtime-launcher.cjs to maintain agents/ in parity (the
  sync caught and corrected a mis-placed preamble during development).
- Extend the bare-gsd-tools (#2851) and launcher-parity (#373) regression guards
  to agents/ so no runtime can silently regress.

Closes #1041

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1041): backfill changeset PR number to 1045

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:42:59 -04:00
Tom Boucher
fbd62cd84f feat(#1035): phase 5a — author 16 role:runtime capability descriptors (registry-only) (#1039)
Author capabilities/<runtime>/capability.json for all 16 runtimes (role:runtime),
populating the registry's runtimes index ({} -> 16). Each carries the 6 ADR-1016
axes (configHome structured, configFormat, artifactLayout structured, commandStyle,
hooksSurface + hookEvents, sandboxTier, supportTier), extracted from the live
modules (runtime-homes/runtime-slash/runtime-artifact-layout/runtime-config-adapter
+ CODEX_AGENT_SANDBOX). validateRuntimeBody tightened to enforce the closed
vocabularies + structured configHome/artifactLayout (rejects old string configHome;
env required; skillsHome recursively validated).

Registry-only: bin/install.js + src/*.cts untouched, zero behavior change. The 5b-5f
drive steps consume these one axis at a time (ADR-1016 §8).

Codex accuracy review caught + fixed: opencode/kilo register ZERO lifecycle hooks
(install.js skips the whole block) -> hooksSurface:none (configFormat stays
settings-json); kimi probe selects on <root>/skills existence -> probeExists:"skills".
ArtifactKind required-field strictness + ConverterName closing deferred to 5d/5e.
ADR-1016 (Proposed) amended to match (Decisions 1 + 5).

Closes #1035

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 09:34:07 -04:00
Tom Boucher
2ac6592096 feat(#1023): first phase-6 cutover — ui-review (verify:post) inline → loop.render-hooks dispatch (#1024)
Replace the inlined ui-review invocation in autonomous.md §3d.5 with a
loop.render-hooks verify:post dispatch — the first workflow to consume
render-hooks and fire a skill from it (closes the #1018 live-execution residual
as real wiring). Capability-driven, equivalence-preserving for the current
registry (only ui-review at verify:post, default on): fires gsd-ui-review under
the same precondition (UI-SPEC exists via consumes-gate + workflow.ui_review).

Gate findings (real pattern issues, fixed so every future cutover inherits them):
- bug-2643 static "Skill() references a real skill" check vs templated
  Skill(skill="gsd-${ref.skill}") dispatch → skip ${...}-templated names.
- Coverage moved, not lost: gen-capability-registry now validates
  steps[].ref.skill in skills + ref.agent in agents + rejects gsd- double-prefix.
- Tightened §3d.5 tests; markdown clarity (consumes rule, LLM-native JSON read,
  UI-REVIEW.md score hint).

gsd-ui-review skill + §3a.5/ui-phase untouched. §5.6/ui-phase cutover deferred
(#1022 step-can-halt-vs-gate model question).

Closes #1023

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 22:40:58 -04:00
Tom Boucher
19edab21da fix(#1006): rc CHANGELOG preview crash on malformed changeset fragment + validate fragment content at the gate (#1007)
* fix(#1006): harden render --preview against fragment parse failures

`render --preview` wrote `report.preview` unconditionally. When a `.changeset`
fragment fails to parse, `cmdRender` early-returns with `{exitCode:1, report:
{failures}}` and NO `preview` key, so `process.stdout.write(undefined)` threw
ERR_INVALID_ARG_TYPE and the rc release job's "Preview CHANGELOG" step died with
a cryptic TypeError that masked the real cause.

Guard the preview write on `typeof report.preview === 'string'` (ADR-227: shape,
not just type); when absent, fall through to the existing failure reporter that
names the offending fragment and exits non-zero — identical to a non-preview
render. Also backfills the stray placeholder `pr: 0` -> `pr: 939` in
.changeset/936-convergence-inline-plan-phase.md that triggered the live failure.

Regression test (red-then-green verified) added at the render --preview seam.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1006): validate changeset fragment content at the Changeset Required gate

The `Changeset Required` gate (scripts/changeset/lint.cjs) only checked that a
`.changeset/*.md` fragment EXISTS in the PR diff; it never validated the
fragment's contents. So a malformed fragment (e.g. an un-backfilled `pr: 0`
placeholder) silently merged to `next` and only detonated later in the rc
release job. This is the upstream prevention for #1006 — the crash hardening
turns the failure into a clear message, this stops the bad fragment ever
reaching the release path.

evaluateLint now accepts `fragmentFailures` and fails with the typed reason
`fail_invalid_fragment` (naming each offending file) before the existence/
opt-out checks — a malformed fragment beats `no-changelog`, since it will break
the render regardless. main() reads + parseFragment()s every changed fragment:
a deleted fragment (not on disk) is skipped, a present-but-unreadable one fails
closed. Tests assert on the typed LINT_REASON enum (no raw-text matching), a
precedence case over the opt-out label, and an end-to-end suite that drives the
real main() against a temp git repo (malformed -> fail, valid -> pass, deleted
-> skipped) so the wiring is regression-proof.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1006): assert the typed --json report in the preview regression test

Code review flagged the preview parse-failure regression test for positive
raw-text matching on CLI output (`combined.includes('bad-fragment.md')` /
`'invalid_pr'`), which this repo's testing standards forbid. Keep the non-json
`runRenderRaw` call for the negative crash proof (the ERR_INVALID_ARG_TYPE
crash lives only on the non-json stdout.write path), and add a `--json`
invocation that asserts the offending fragment + typed `invalid_pr` reason via
the structured `report.failures[]` surface instead of rendered prose.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 15:55:31 -04:00
Tom Boucher
adaf3e17d8 fix(#1001): make bug-969 hardening tests hermetic + move build tsbuildinfo out of shipped tree (regression from #996) (#1002)
* fix(#969): make bug-969 hardening tests hermetic and move build tsbuildinfo out of shipped tree (regression from #996)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#969): self-heal legacy bin-local tsbuildinfo and make sentinel test hermetic (adversarial-review follow-ups)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(changeset): set pr number to 1002

* docs(#1001): record DEFECT.SHARED-ARTIFACT-MUTATION-IN-CONCURRENT-TEST anti-pattern in CONTEXT.md

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 13:58:56 -04:00
Tom Boucher
88e30d5342 test(#969): fix stale-build flake (incremental + re-emit-on-missing) and make runGsdTools retry-once before surfacing subprocess kills (#996)
Closes #969

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 12:04:59 -04:00
Tom Boucher
1fd5c86a1e fix(#983): rewrite bare .claude paths in Trae/Windsurf converters (Codex/Cline parity) (#995)
* fix(#983): rewrite bare .claude paths in Trae/Windsurf converters (Codex/Cline parity)

Both convertClaudeToWindsurfMarkdown and convertClaudeToTraeMarkdown only
handled trailing-slash .claude/ forms; bare ~/.claude and $HOME/.claude
references (e.g. configDir = ~/.claude, RUNTIME_CONFIG_DIR=".../$HOME/.claude")
survived conversion and pointed users at the wrong config dir.

Fix: add bare-form replacements using negative lookahead (?![\w-]) to protect
.claude-plugin and .claudeignore, mirroring Cline (#782) and Codex (#570) precedent.
Also adds CLAUDE_CONFIG_DIR -> WINDSURF_CONFIG_DIR / TRAE_CONFIG_DIR rewrite.
_applyRuntimeRewrites windsurf case gets matching \b-anchored bare-form lines,
mirroring the existing trae case.

Closes #983

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#983): backfill changeset pr number (995)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 11:43:31 -04:00
Tom Boucher
36b68ac81d fix(#977): map ephemeral fnm multishell execPath to a stable fnm alias in normalizeNodePath (#992)
* fix(#977): map ephemeral fnm multishell execPath to a stable fnm alias in normalizeNodePath

Closes #977

* chore(#977): backfill changeset pr number (992)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 11:16:13 -04:00
Tom Boucher
972a41a528 fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:) (#990)
* fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:)

Closes #967

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#967): backfill changeset pr number (990)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 11:10:40 -04:00
Tom Boucher
921a7cd618 fix(#974): error on graphify --budget with missing/non-numeric value instead of silent NaN no-op (#986)
* fix(#974): error on graphify --budget with missing/non-numeric value instead of silent NaN no-op

When `--budget` was the last arg or followed by a non-numeric token,
parseInt(undefined/NaN-string, 10) produced NaN. NaN is falsy so both
the router check and applyBudget gate silently skipped budget trimming.
The query ran unbounded with no warning.

Fix: guard in graphify-command-router.cts — if args[budgetIdx+1] is
absent or parses to NaN, emit ERROR_REASON.USAGE and return early.
Defensive fix in graphify.cts: tighten `if (!budgetTokens)` →
`if (budgetTokens == null)` and `if (options.budget)` →
`if (options.budget != null)` so a real 0/NaN caller is handled
predictably by both independent guards.

Regression tests: 16 cases (unit/mock, subprocess, property-based)
covering boundary inputs: missing value, non-numeric, valid integers,
and fast-check properties over the budget parse contract.

Closes #974

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#974): backfill changeset pr number (986)

* fix(#974): convert static property (b) to real fc.assert property with constrained generator + deterministic seed

The test named "property: --budget as last arg always produces usage error"
was a static test with no fc.assert — it only checked a single hardcoded
term ("someterm") and could never flake or produce a fast-check path. This
is a generator/property bug (case b): the test was mislabeled as a property
test but lacked parameterization.

Root-cause of CI path "46:0:0" / seed 42 failure: a naive parameterized
version without the !startsWith('--') filter could feed term='--budget',
causing args.indexOf('--budget') to hit index 2 (the term slot) rather than
index 3 (the flag slot), placing the router in a different code path. The
property still holds — NaN detection fires on rawBudget='--budget' — but
the assertion text referenced the wrong invariant, making the failure appear
spurious. Fix: constrain the generator to non-flag terms (filter out
strings starting with '--') and pass explicit { seed: 42, numRuns: 200 } to
fc.assert so the test is fully deterministic in CI regardless of GSD_FC_SEED.

No change to src/graphify-command-router.cts (router is correct).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#974): constrain non-numeric budget property generator to genuinely-NaN values and pin fast-check seeds (CI determinism)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#974): use a strict valid-term generator and pin seeds so budget property tests are deterministic

Replace fc.string({minLength:1}).filter(!startsWith('--')) term generators in
properties (b) and (d) with a shared validTerm = fc.stringMatching(/^[A-Za-z0-9][A-Za-z0-9_.-]{0,29}$/)
that is alphanumeric-leading and contains no whitespace, flags, or sign-numerics.
This eliminates the class of CI failures where the old generator produced out-of-
contract inputs (" ", "+5", empty) that the router legitimately rejects for reasons
outside the budget-parse contract under test. Pin distinct seeds (1001–1004) on
every fc.assert for CI determinism. Stress-tested at numRuns=100000 per property
and across 10 seeds (1,2,7,13,42,43,44,99,12345,999999) at numRuns=2000 — all pass.
No src change: gsd-core/bin/lib/graphify-command-router.cjs is correct.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#974): replace flaky term-fuzzing properties with deterministic examples; keep value-fuzzing properties

The validTerm regex /^[A-Za-z0-9][A-Za-z0-9_.-]{0,29}$/ admitted single-digit
strings like "0" which are falsy; the router's `if (!term)` guard fires before
the budget-missing-value path, producing a spurious errFn call. Properties (b)
and (d), which test the BUDGET contract (not term handling), are replaced with
deterministic example loops over fixed valid terms. Properties (a) and (c),
which fuzz the BUDGET VALUE with a fixed term, are kept unchanged (seed+numRuns
pinned). The validTerm generator is fully removed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 10:58:42 -04:00
Tom Boucher
626575cbc5 fix(#978): parse --force in milestone complete dispatcher so the guard's documented override works (#982)
* fix(#978): parse --force in milestone complete dispatcher so the guard's documented override works

The dispatcher built `{ name, archivePhases }` but never parsed `--force`, so
`options.force` was always `undefined` and the guard inside `cmdMilestoneComplete`
(which tells users to "Re-run with --force to override") could never be bypassed.

Add `const force = args.includes('--force')` and pass it into the options object.
The guard already honors `options.force` — no changes to milestone.cts needed.

Closes #978

* chore(#978): backfill changeset pr number (982)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 10:40:06 -04:00
Tom Boucher
caca4d255c feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4) (#988)
* feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4)

Migrate the intel CLI command family from a hardcoded gsd-tools.cjs case arm to
a registry-dispatched Capability (commandFamilies mechanism, #961), mirroring
the graphify (#972) and audit (#984) cutovers. New src/intel-command-router.cts
exports routeIntelCommand reproducing all 9 subcommands verbatim (incl. the
status timeAgo non-raw post-processing), lazily requiring intel.cjs inside the
route fn. capabilities/intel/capability.json declares the intel command family;
commands-only (skills:[]), declares the existing intel.enabled gate (default
false — behavior unchanged).

Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing intel.test.cjs passes unchanged. Completes the first-party command-
family cutover sequence (graphify/audit/intel).

Closes #985

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#985): compute expected planningDir via path.join in intel cutover unit tests (Windows CI)

The intel-command-cutover unit-mock assertions hardcoded a POSIX `/.planning`
expectation while the router builds it with path.join(cwd, '.planning') →
backslashes on Windows, so the planningDir-arg assertions (query/status/diff/
snapshot/validate/update/api-surface) failed only on windows-latest CI. Compute
the expectation with path.join (cross-platform); production router unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 10:00:30 -04:00
Tom Boucher
9a03539c2d feat(#981): audit-uat + audit-open command cutover — commands-only capability (ADR-857 phase 4d-impl-3) (#984)
Migrate the audit-uat + audit-open CLI commands from hardcoded gsd-tools.cjs
case arms to a registry-dispatched Capability (commandFamilies mechanism, #961),
mirroring the graphify cutover (#972). New src/audit-command-router.cts exports
routeAuditUat/routeAuditOpen, each lazily requiring only its backing module
(uat.cjs/audit.cjs) inside the route fn — matching the old per-case lazy loads.
capabilities/audit/capability.json declares the two command families;
commands-only (skills:[]), no config gate, audit_review cluster untouched.

Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing audit regression tests pass unchanged.

Closes #981

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 08:34:49 -04:00
Tom Boucher
77bfd943dd feat(#972): graphify command cutover — first capability owning a command family (ADR-857 phase 4d-impl-2) (#975)
Migrate graphify into a Capability that owns its `graphify` command family,
dispatched via the registry (#961 mechanism) instead of a hardcoded case.
graphify is now an enable/disable plug-in.

- src/graphify-command-router.cts: routeGraphifyCommand (standard route*Command),
  reproduces the removed case EXACTLY (query +--budget, status, diff, build,
  hidden build snapshot, usage/unknown errors); injectable _graphify test seam.
- capabilities/graphify/capability.json: role feature, tier:full, skills:[graphify],
  config:{graphify.enabled default false}, commands:[{family:graphify, module,
  router:routeGraphifyCommand}].
- removed case 'graphify' from gsd-tools.cjs; graphify now flows
  default -> dispatchCapabilityCommand -> commandFamilies.graphify -> router.
- regenerated registry (commandFamilies/bySkill/configSchema/profileMembership/
  capabilityClusters for graphify); tier:full keeps 4c install/surface a no-op.

Equivalence-proven: 8 recording-mock unit tests assert the exact fn+args per
subcommand (budget, snapshot-vs-build); 9 subprocess tests assert distinguishing
output shapes; existing graphify tests pass unchanged through the new path.

Surfaced (not silently accepted): the pre-existing --budget-no-value NaN no-op
quirk, preserved for equivalence, filed separately.

Closes #972

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 07:12:57 -04:00
Colin Johnson
354e0e1b94 fix(ratchet): add --update drift repair + inherited-drift guidance to regression-name lint (#971) 2026-06-10 01:05:24 -04:00
Tom Boucher
0c567c17e4 feat(#961): capability command mechanism (commandFamilies index + default-case dispatch) — ADR-857 phase 4d-impl-1 (#964)
* feat(#961): capability command mechanism — commandFamilies index + default-case dispatch (ADR-857 phase 4d-impl-1)

Build the capability command mechanism per ADR-959: the `commands` declaration
field on the feature role, the registry commandFamilies index, and a real
dispatchCapabilityCommand consulted in runCommand's default case (replacing the
dead _dispatchNonFamily shim's role).

The registry DISCOVERS a standard route*Command (no rebuilt handler table). On a
default-case command, dispatch enforces a bare-.cjs-basename, resolves the module
under gsd-core/bin/lib/, asserts confinement, requires the resolved path, and
own-property-guards the router export before calling it. A require.main===module
guard makes gsd-tools.cjs importable for tests; the CLI path is unchanged.

Additive: commandFamilies is empty today, so the default case is behavior-
preserving for every command; the 10 dead _dispatchNonFamily sites are untouched
(future migration markers). The graphify cutover is the separate next step.

Closes #961

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#961): surface capability router failures structurally + enforce sync contract

Review follow-up: dispatchCapabilityCommand now wraps the router invocation so an
unexpected (non-ExitError) throw is converted to a structured, attributed
error(msg, SDK_FAIL_FAST) — honoring --json-errors — instead of escaping as a raw
stack trace; an intentional ExitError propagates unchanged. An async router
(returns a thenable) is rejected loudly with a structured error (the contract is
synchronous, like the 12 host routers). Not shipped as "consistent with existing
behavior": the host's pre-existing version of this gap is filed as #965.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 01:03:46 -04:00
Tom Boucher
e8cfb560b2 fix(#851): correct Codex quick adapter for generic multi_agent_v1 schema (#958)
* fix(#851): correct Codex adapter for generic multi_agent_v1 schema

The Codex skill adapter header in getCodexSkillAdapterHeader() documented
typed spawn_agent(agent_type=...) as a direct, unconditional mapping for all
Task()/Agent() calls. In sessions exposing only the generic multi_agent_v1
schema (message/items/fork_context — no agent_type field), this mapping is
silently invalid: the orchestrator cannot natively dispatch typed gsd-planner/
gsd-executor agents and may fall back to inline execution or produce errors.

Fix: Section C now requires schema detection before spawning. It documents the
typed mapping as conditional on the agent_type-capable schema (e.g. multi_agent_v2)
and introduces an explicitly-labeled generic-agent workaround for multi_agent_v1
sessions — read the agent TOML, inject its instructions as a role-preamble, and
call spawn_agent(message=...) — clearly marking the result as NOT equivalent to
typed gsd-planner/gsd-executor execution.

Regression test: tests/bug-851-codex-quick-adapter-agent-type-fallback.test.cjs
asserts schema-awareness language, the multi_agent_v1 fallback, the workaround
label, and backward compat with the existing bug-279 typed-spawn contract.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: add changeset for PR #958

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#851): keep Codex adapter block consistent with materialized skill surface

The `~/.codex/agents/<agent-name>.toml` literal introduced in the #851 prose
was being rewritten to the real install path by `_applyRuntimeRewrites` (the
`~/.codex/` → pathPrefix substitution) before the SKILL.md was written to
disk.  `getCodexSkillAdapterHeader()` still returned `~/.codex/agents/...` so
the test assertion (exact match between builder output and materialized file)
always failed.

Fix: replace the `~/.codex/agents/` literal with the runtime-neutral form
`agents/<agent-name>.toml` plus a parenthetical naming `$CODEX_HOME/` — which
is not matched by any rewrite pattern and survives the path-substitution step
unchanged.  The #851 schema-detection + generic-subagent-fallback intent is
fully preserved.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#851): resolve active Codex config root in fallback; strengthen tests (adversarial review)

- Rewrites the generic-agent workaround step 1 to explicitly describe
  active config root resolution (priority: $CODEX_HOME → --config-dir →
  --local .codex → default global dir) without the literal ~/.codex/
  substring that _applyRuntimeRewrites replaces, preventing bug-3582
  divergence.
- Replaces OR/loose-includes test assertions in bug-851 with AND-logic
  checks covering all four required elements: (a) schema-detection step,
  (b) active-config-root resolution for the TOML path including all three
  override mechanisms, (c) NOT-equivalent-to-typed-gsd-planner/gsd-executor
  label, and (d) fail-closed rule when typed dispatch is mandatory.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#851): register bug-851/947/948/950 in lint-regression-test-names allowlist

The ratchet (622e4be) bans NEW top-level bug-NNNN test files; the four
sibling PRs (#851, #947, #948, #950) landed AFTER the baseline was cut,
so their test files were not yet grandfathered. Add all four to the
identity allowlist so lint-regression-test-names passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 00:45:27 -04:00
Tom Boucher
46967baae8 fix(#948): guard STATE.md no-op writes; preserve milestone_name/stopped_at (closes #944) (#952)
* test(#948): add regression tests for no-op write guard and record-session auto-create (#944)

Red before fix: 11/15 tests fail. Green after: 15/15.
Covers zero-match patch byte-identity, milestone_name preservation,
stopped_at frontmatter-wins, record-session auto-create fallback, and
adversarial fixtures (CRLF, empty body, non-canonical labels).

Also registers bug-948-state-noop-write-guard.test.cjs in the state
bucket of lint-test-file-count.allowlist.json.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#948): guard STATE.md no-op writes; preserve milestone_name/stopped_at (closes #944)

Shared root cause: `readModifyWriteStateMd` wrote STATE.md unconditionally
even when the transform produced no change, and `syncStateFrontmatter`
re-derived frontmatter from the possibly-stale body on every write.

Three coordinated fixes in src/state.cts:

1. readModifyWriteStateMd: add no-op guard — when transform result ===
   input content, skip the write entirely (no platformWriteSync, no
   last_updated bump, no frontmatter re-derive). Fixes #948 zero-match
   phantom write and the #944 phantom last_updated bump.

2. syncStateFrontmatter: extend existing-frontmatter preserve logic —
   fall back to existingFm['milestone_name'] / existingFm['milestone']
   when the derived value is the template placeholder 'milestone'
   (getMilestoneInfo returns this literal when it cannot match the
   version in ROADMAP.md); prefer existingFm['stopped_at'] /
   existingFm['paused_at'] over a body-derived value (the frontmatter
   value, written by the canonical record-session path, wins over stale
   historical body lines). Mirrors the fallback already in cmdStateJson.

3. cmdStateRecordSession: when --stopped-at / --resume-file are supplied
   but body labels are absent, DWIM auto-create a canonical ## Session
   section (mirroring how add-decision / add-blocker / record-metric
   auto-create their sections). Never return a silent recorded:false when
   the caller supplied values.

SDK check: no sdk/src/state.ts exists in this repo (the comment in
cmdStateSnapshot references a sibling concern in the TypeScript SDK
codebase, which is a separate repo not present here).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: add changeset for PR #952 (fix #948/#944)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#948): correct stopped_at preserve rule; adjust test for sync behaviour

The "always prefer frontmatter stopped_at" rule in syncStateFrontmatter
was too aggressive — it broke phase.complete which intentionally updates
stopped_at in the body and expects syncStateFrontmatter to pick it up.

The primary fix (no-op guard in readModifyWriteStateMd) already prevents
the stale-body-overwrites-frontmatter scenario from #948: the file is not
written when the transform produces no change, so syncStateFrontmatter
never runs on a zero-match patch. The body-derived value can only win when
an actual write occurs, which means the body was legitimately updated.

Reverted to the original #905 rule for stopped_at/paused_at: fall back to
existing frontmatter only when the derived value is absent (empty/null).

Also adjusted the sync-suite test to assert what state sync actually does
(milestone_name preservation) rather than a stopped_at-wins property that
state sync does not have by design.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#944): update existing session block in place (adversarial review)

HIGH finding: the DWIM auto-create in cmdStateRecordSession was appending
a second ## Session block unconditionally, even when one already existed
with non-canonical content (e.g. a markdown table). Both
buildStateFrontmatter and cmdStateSnapshot read only the FIRST ## Session
block via regex, so the newly-written Stopped at / Resume file values
landed in the second, invisible block — frontmatter stopped_at stayed
stale and state-snapshot returned nulls.

Fix: check for an existing ## Session heading. When one is present,
normalize that section in place by replacing its body with canonical
**Last session:** / **Stopped at:** / **Resume file:** bold-label lines.
Only append a brand-new section when NO ## Session heading exists.

LOW finding: the auto-create scaffold emits **Last session:** but
cmdStateSnapshot only matched **Last Date:**, so session.last_date was
null after auto-create despite a valid timestamp being written.

Fix: extend the lastDateMatch regex in cmdStateSnapshot to also accept
**Last session:** / Last session: (the form the scaffold writes).

Tests: 3 new tests added to bug-948-state-noop-write-guard.test.cjs that
confirmed failure against the previous HEAD and pass after this fix:
- exactly one ## Session block after record-session with non-canonical existing block
- state-snapshot sees correct stopped_at via first Session block (not a duplicate)
- state-snapshot session.last_date is non-null after auto-create on body-less file

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#944): improve in-place section replace to cleanly remove old body content

The previous regex `/(^## Session[ \t]*$)([\s\S]*?)(?=\n^## |\n*$)/im`
with a lazy match consumed nothing after the heading, so old non-canonical
body content (e.g. table rows) remained after the new canonical lines.

While functionally correct (parsers found the canonical lines first in the
FIRST ## Session block), it left stale content in the section. Replace with
a negative-lookahead per-line pattern that consumes all content from the
heading up to (but not including) the next ## heading, producing a clean
section with only the canonical bold-label lines.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 00:22:29 -04:00
Colin
533b518553 fix(security-scan): update scanner self-exemption allowlists for renamed suite files
The three shell scanners exempt their own adversarial test fixtures by exact
filename; the *.security.test.cjs renames broke those entries, so the PR diff
scan flagged the scanners' own test payloads. Verified locally with all three
scanners in --diff origin/next mode (0 findings) and the security suite
(207/207). The .sh files were missed in the original reference sweep because
the rename grep filtered to .cjs/.yml/.json/.md extensions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 00:15:02 -04:00
Colin
8bb6784009 chore(ratchet): regenerate bug-* allowlist after rebase onto next (244 -> 257)
13 bug-* files landed upstream between the audit baseline and this branch's
rebase; they predate the ratchet policy, so they are grandfathered.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:56:57 -04:00
Colin
622e4be6d8 test(ratchet): ban new top-level bug-NNNN test files via identity allowlist
244 one-off bug-* files (~38% of the suite) are grandfathered in
lint-regression-test-names.allowlist.json; new ones fail lint with
fold-into-module guidance, and deletions force allowlist pruning so the
baseline only shrinks. Wired into npm run lint:ci (new single entry point
for every CI lint). Policy documented in docs/TESTING-SUITES.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:50:41 -04:00
Colin
cd5db1f8db test(suites): seed security/slow/integration suites via measured retags
Renames (git mv) with all references updated (ci-test-scope RULES,
windows-parity allowlist, test-file-count allowlist, docs in 6 locales):

- 5 scanner tests -> *.security.test.cjs — the 'Run security tests' CI step
  ran zero files since the suite taxonomy landed; it is now honest.
- graphify-auto-update -> *.slow.test.cjs (36s, slowest file in the suite;
  e2e gsd-tools spawns) — runs on full-matrix lanes and push to next.
- installer-migration-install-integration -> *.integration.test.cjs
  (13s; an integration test by its own name).

Coverage gate measured after retags: 88.55% lines (gate 70%).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:50:41 -04:00
Colin
a647053dcf ci(scope): narrow #494 invariant — changed tests join windows lane, not full matrix
full_matrix fired on 15/15 sampled PRs because any tests/** change forced it,
costing ~25 runner-minutes each. A changed test file now always joins the
scoped windows lane (covering the #482 OS-specific failure class per-file)
and still runs on ubuntu 22/24 via targeted_tests; the residual macOS /
windows-node-22 cross-product is covered on every push to next.

Also narrows WINDOWS_HINTS from 6 substrings (102/633 files, a ~10-minute
scoped lane) to windows/win32/shell/path — the dropped hints (workflow,
install, hook) are either platform-independent lint tests or already covered
by fullMatrix rules.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:50:41 -04:00
Tom Boucher
ed467cd8f2 feat(#942): tier→profile/cluster derivation + consistency gate (ADR-857 phase 4a) (#943)
Establish `tier` as the source of install-profile + cluster membership
(ADR-894 §4). The registry now derives two views: capabilityClusters
(capId → its skills) and profileMembership (capId → {tier, profiles}, where
profiles is the PROFILE_RANK suffix from the capability's tier). A consistency
gate cross-checks them against the hand-authored PROFILES/CLUSTERS: HARD (throws)
on a capId-matching-a-cluster-name with a different skill set; SOFT
(pending-reconciliation stderr warning, not serialized) when a capability skill
isn't yet in the closure-resolved hand-authored profile.

The SOFT gate loads the real skills manifest and resolves each profile's closure
once so transitively-included skills don't false-warn; warnings are de-duped to
one per (capability, skill); both derived views are scoped to skill-owning
capabilities; serialized with a global capId sort for determinism; reserved-name
guards at every write site; lazy requires of the built constants.

Behavior-preserving: install/surface untouched; derived views consumed by
nothing. Full generation rides along with the phase-6 migration.

Closes #942

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 15:11:42 -04:00
Tom Boucher
dcb0d8a28d fix(#935): install changeset CLI so /gsd-update changelog preview works (#938)
- bin/install.js now copies scripts/changeset/ and scripts/lib/ into
  <configDir>/scripts/ so $GSD_DIR/scripts/changeset/cli.cjs resolves
  at runtime; aborts install with an explicit failure if the source
  directory is missing from the package.
- gsd-core/workflows/update.md: corrected path from
  gsd-core/scripts/changeset/cli.cjs to scripts/changeset/cli.cjs;
  added an explicit [ ! -f ] guard so a missing CLI surfaces a clear
  message rather than silently swallowing the error; stderr captured
  via 2>&1 sentinel so node errors are visible in the preview output.
- release.yml's changeset-CLI invocations (node scripts/changeset/cli.cjs)
  remain at the repo-root path and are unaffected by this change.

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 12:37:26 -04:00
Tom Boucher
6dbd895028 feat(#910): federated config merge in config-loader (ADR-857 phase 3b) (#914)
Build the federated config merge: each capability owns its config-key slice
(ADR-857 decision 3 / ADR-894), and loadConfig merges them defensively. The
registry now emits a full configSchema index ({key:{owner,type,default,
description}}, generator-validated); a new src/federated-config.cts resolves
federated keys defensively (skip central keys -> pending-migration warning,
skip malformed slices -> warning never throw, else type-checked user override
?? default, with nested dotted-path lookup and enum validation); and loadConfig
applies the overlay on every return path.

Wired as a provably-empty no-op channel: every UI-pilot key is still central,
so validKeys is empty and loadConfig returns byte-identical output on all paths
(identity return when the overlay is empty; shared CONFIG_DEFAULTS never
mutated). Registry-only; no key is cut over; nothing in the live loop changes.

Closes #910

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 23:37:41 -04:00
Tom Boucher
d12809985e fix(#905): preserve STATE.md frontmatter scalars in syncStateFrontmatter (#907)
syncStateFrontmatter was silently dropping current_phase, current_phase_name,
current_plan, and progress when body annotations were absent (e.g. after an
agent or tool rewrote the body). These scalars can only be derived from body
annotations — when absent, buildStateFrontmatter returns nothing for those
keys. Added existingFm fallbacks mirroring the same pattern already applied in
cmdStateJson, so every writeStateMd call preserves the existing values instead
of stripping them. Also extended cmdStateJson with the same fallbacks for the
three non-progress scalars.

Adds regression test (7 cases) + lint-test-file-count allowlist entry.

Closes #905

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 22:52:43 -04:00
Tom Boucher
808df9110c fix(#892): parse checklist-style roadmap phases in validate/verify (#908)
buildRoadmapPhaseVariants() only matched heading-style phases (## Phase N:),
silently skipping the supported checklist format (- [x] **Phase N: name**).
This caused W007 false-positives for every on-disk phase dir when the project
uses a checklist ROADMAP. Fix adds a second regex pass (mirroring the existing
buildNotStartedPhaseVariants() approach). Also refactors the duplicate
inline heading-only regex in cmdValidateConsistency() to delegate to
buildRoadmapPhaseVariants() (DRY). Regression test in
tests/bug-892-validate-checklist-roadmap-phases.test.cjs covers both paths.

Closes #892

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 22:52:19 -04:00
Tom Boucher
48cc27bd84 feat(#903): generate Loop Host Contract from workflow markers (ADR-857 phase 3a-impl-2) (#906)
Replace the inline LOOP_HOST_CONTRACT constant in the Capability Registry
generator with a generated-from-workflows contract (ADR-894 §3). The contract
is now derived from inert `<!-- gsd:loop-host ... -->` marker blocks in the five
step workflows, emitted as the committed gsd-core/bin/lib/loop-host-contract.cjs,
and required by gen-capability-registry.cjs — one source of truth, no drift.

Drift guards in gen-loop-host-contract.cjs: per-step point ownership (each step
must declare exactly its canonical loop points), multiple-block + duplicate-key
hard errors, and a word-boundary agent-role cross-check. Contract content is
byte-identical to the former constant; registry-only, nothing wired into the
live loop.

Closes #903

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 22:16:58 -04:00
Tom Boucher
ad754ca6cd feat(#896): Capability Registry generator + UI pilot (ADR-857 phase 3a-impl) (#902)
* feat(#896): Capability Registry generator + UI pilot (ADR-857 phase 3a-impl)

First phase-3 code: the Capability Registry generation pipeline, built against
the ADR-894 contract and NOT wired into the live loop (registry-only, per the
staged-cutover design).

- capabilities/ui/capability.json — the UI pilot (ADR-894 worked example): 2
  skills, 2 agents, 3 config keys, 2 steps + 1 gate, with `when` activation.
- scripts/gen-capability-registry.cjs — --write/--check generator. Hand-rolled
  schema validation (envelope + role-typed feature/runtime bodies + typed
  steps/contributions/gates + when + gate-check variants); cross-capability
  invariants (single ownership; requires exist+acyclic+tier-monotone; config-key
  ownership exclusive, collision-vs-central as a pending-migration warning);
  hooks validated against an inline LOOP_HOST_CONTRACT (3a-impl-2 swaps its
  source to the generated-from-workflows contract); GLOBAL point-ordered
  consumes-satisfiability; materialized byLoopPoint ordering (produces/consumes
  topo-sort); emits gsd-core/bin/lib/capability-registry.cjs (role-partitioned
  indexes + requiresClosure). Prototype-pollution guards (Object.create(null) +
  inline literal key checks) + fragment.path traversal guard.
- gsd-core/bin/lib/capability-registry.cjs — committed generated artifact
  (mirrors package-identity.cjs: script-generated, tracked, linted, regenerated
  on build, drift-tested), wired via the new `gen:capability-registry` build step.
- tests/capability-registry.test.cjs — 72 tests: schema + invariant + hook +
  ordering + adversarial (path-traversal, proto-pollution, runtime body,
  self-consume, cycles, collisions) + committed-file staleness guard.

New-CLI-module checklist (INVENTORY 97->98, MANIFEST, ARCHITECTURE), CONTEXT.md
"Capability Registry" un-[Planned]'d. Nothing wired into install/surface/loop.

Gates: lint, code-review (4 bugs fixed), security-review (path-traversal +
prototype-pollution fixed), codex adversarial-review ×3 (8+ findings fixed,
confirmed sound), clean-build docker 13190 pass / 0 fail.

Closes #896

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#896): CRLF-agnostic --check for capability-registry staleness (Windows)

The committed capability-registry.cjs staleness guard failed on Windows CI only:
git checks out the committed .cjs as CRLF (autocrlf, no .gitattributes) while the
generator emits LF, so the byte-for-byte --check comparison mismatched. Normalize
line endings on both sides of the --check comparison (no .gitattributes change,
no change to the LF the generator writes). Adds a regression test simulating the
Windows CRLF checkout.

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 21:15:41 -04:00
Tom Boucher
0a11d361ca feat(#69): nest concrete skills under namespace routers at install (#883)
Emit the 6 gsd-ns-* routers as the only top-level skill bundles and nest
the ~61 concrete skills under <router>/skills/<name>/SKILL.md on runtimes
with confirmed non-recursive skill loaders (claude global, cline, qwen,
hermes, augment, trae, antigravity). Router bodies rewrite their routing
tables from Skill-tool dispatch to a Read skills/<name>/SKILL.md pattern.
Recursive/unconfirmed loaders (cursor, codex, copilot, windsurf, codebuddy,
opencode, kilo) keep the flat layout. Completes the v1.40 namespace
architecture (#2792) so the eager skill listing drops to ~6 entries.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 15:09:34 -04:00
Tom Boucher
a480510f54 fix(#872): make roadmap-phase-fallback tests hermetic against ambient GSD env (#873)
extractCurrentMilestone reads STATE.md via planningDir(cwd), which is
workstream-aware (honours GSD_PROJECT/GSD_WORKSTREAM). The fixtures write
STATE.md to the plain <tmp>/.planning/STATE.md, so a developer shell inside a
GSD workstream (GSD_WORKSTREAM exported) redirected the read to a non-existent
workstream subdir -> version=null -> closed milestone sections leaked into the
slice and assertions failed. Clean CI/Docker env never hit it. Not a Node-26
regex bug; reproduces identically on any Node with GSD_WORKSTREAM set.

- scripts/run-tests.cjs: strip GSD_PROJECT/GSD_WORKSTREAM before spawning test
  children so the local runner env matches clean CI/Docker.
- tests/roadmap-phase-fallback.test.cjs: file-level beforeEach/afterEach
  save/delete/restore of both vars; new regression test pinning workstream-aware
  STATE.md resolution.
- tests/run-tests-harness.test.cjs: guard asserting the runner strips both vars
  (so removing the deletion fails clean CI).

Closes #872

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 12:04:26 -04:00
Tom Boucher
606363c416 chore(#846): remove unused PR-size labeler (size/S–XL) workflow (#848)
The PR Gate workflow's only job, size-check, labeled every PR with
size/S–size/XL based on lines changed. Those labels aren't used in any
review, triage, or automation flow, so the workflow was pure noise.

- Delete .github/workflows/pr-gate.yml
- Drop size-check from required status checks in both rulesets so PRs
  don't block forever on a check that never reports
- Remove pr-gate.yml from INERT_WORKFLOWS (ci-test-scope.cjs) and the
  knownInert list (ci-test-scope.test.cjs)
- Remove "PR Gate / size-check" from setup-branch-protection.sh

Closes #846

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 22:03:48 -04:00
Tom Boucher
543e51e71f fix(#844): sync runtime manifest versions on npm version bump (#845)
* fix(#844): sync runtime manifest versions on npm version bump

The release workflow bumps package.json via `npm version` but never
stamped the runtime-integration manifests that must track it
(.claude-plugin/plugin.json #766, gemini-extension.json #775), so the
first RC/finalize whose version diverged from the -dev stream failed the
test suite before tagging/publishing.

Add scripts/sync-manifest-versions.cjs (single VERSIONED_MANIFESTS
registry) wired to a `version` npm lifecycle hook that stamps + stages
the manifests on every `npm version` — covering all four release bump
sites and local bumps with no workflow edits. A regression guard test
fails if any repo JSON whose version matches package.json is not
registered, forcing future version-bearing manifests into the sync.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#844): add changeset for manifest version sync fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 21:43:30 -04:00
Tom Boucher
1b6bd66f2c feat(#770): register Claude Code lifecycle hooks (SubagentStop/Stop/PreCompact/FileChanged) (#821)
* feat(#770): register Claude Code lifecycle hooks (SubagentStop/Stop/PreCompact/FileChanged)

Wire three new context-tracking events (SubagentStop, Stop, PreCompact) to
gsd-context-monitor so context-headroom warnings surface at model-stop and
subagent-finalisation moments — not just on PostToolUse.  Add a new
FileChanged hook (gsd-config-reload.js) that hot-reloads .planning/config.json
context mid-session when the user edits it, injecting a config summary as
hookSpecificOutput.additionalContext.  Updates plugin manifest hooks.json,
managed-hooks-registry, installer-migration-report allowlist, and
shell-command-projection cleanup tables.  Tests: 21 new assertions in
enh-770-claude-hook-events.test.cjs; enh-788 and issue-766 test suites updated.

Closes #770

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#770): document newly-registered Claude Code lifecycle hooks

Add a Hook coverage table to the Claude Code npm installer section of
docs/how-to/install-on-your-runtime.md describing SubagentStop, Stop,
PreCompact, and the new FileChanged (gsd-config-reload.js) hook that
hot-reloads .planning/config.json mid-session. Also fixes the changeset
frontmatter (adds type: Added + pr: 821) so docs-lint can consume the
fragment.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#770): add gsd-config-reload.js to INVENTORY.md and regenerate manifest

The feat commit added hooks/gsd-config-reload.js but did not bump the
Hooks count in docs/INVENTORY.md (14→15) or add the new row, and did not
regenerate docs/INVENTORY-MANIFEST.json. Both inventory-counts and
inventory-manifest-sync tests failed across the full CI matrix.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#770): make lifecycle-hook tests deterministic on scoped runner

Replace the shared hooks/dist/ ensemble setup (ensureHooksDist /
teardownHooksDist) in the Claude hook tests with per-test isolation:
pre-populate each test's own tmpDir/.claude/hooks/ with stub files and
pass installerMigrations:[] to install() so the first-time-baseline
migration does not remove the stubs before the copy step can run.

Root cause: hooks/dist/ is gitignored and absent on a fresh npm ci.
ensureHooksDist() created it and teardownHooksDist() deleted it, but
with --test-concurrency=4 both test files ran concurrently as separate
Node.js worker processes sharing the same filesystem.  One file's
afterEach teardown deleted hooks/dist/ while the other file's install()
was copying from it, producing an ENOENT (reproduced 2/10 runs locally).

The additional issue: even with pre-placed stubs surviving the copy race,
the 000-first-time-baseline migration classified hooks/gsd-*.js as
bundled-gsd-hook artifacts, auto-removed them, and the copy step never
re-ran (hooks/dist/ absent) — leaving contextMonitorFile missing and all
hook registrations silently skipped (the 'got: []' symptom).

Fix: pre-populate targetDir/hooks/ per-test (isolated temp dir) AND pass
installerMigrations:[] so the baseline scan is skipped.  The Qwen suites
already used this pattern correctly; the Claude suites are aligned to it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#770): ship gsd-config-reload.js by adding it to build-hooks HOOKS_TO_COPY

The #770 feature added hooks/gsd-config-reload.js and registered it in
MANAGED_HOOKS, the installer, INVENTORY, and the test EXPECTED_ALL_HOOKS
list — but never added it to scripts/build-hooks.js HOOKS_TO_COPY. As a
result the hook was never copied into hooks/dist/ during the build, so:

  - the hook would never ship to users (real production bug — the
    FileChanged config-reload feature was dead-on-arrival), and
  - install-minimal-hooks.test.cjs #1755 ("all expected hooks are copied
    from hooks/dist/ to target", ".js hooks are executable after copy",
    "manifest contains .js hook entries") failed on any environment with
    a clean checkout (no pre-existing hooks/dist/): coverage, full test
    macos-22/macos-24, test ubuntu-24.

The failures were masked locally only by a stale hooks/dist/ left from a
prior build (build-hooks copies into dist without clearing it). On CI's
fresh `npm ci` there is no dist, so the omission surfaced.

Fix: add 'gsd-config-reload.js' to HOOKS_TO_COPY so build-hooks stages it
into hooks/dist/ alongside the other JS hooks. Verified by removing
hooks/dist/ and rerunning the full suite green (0 fail).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#770): make config prototype-pollution beforeEach deterministic on scoped runner

Root cause: the #663 and alert-#26 prototype-pollution describe blocks
seeded .planning/config.json in beforeEach via a bare
runGsdTools('config-ensure-section') whose result was discarded. That
command runs in a spawned gsd-tools child; on the scoped CI lane
(--test-concurrency=4, config.test.cjs scheduled alongside the heavy
install/tarball suites that #770 pulled into the targeted set) the child
can be transiently killed under resource pressure (non-zero exit, empty
stderr — an OS-level kill, not an app error). The swallowed failure left
config.json absent, so the first subtest's readConfig() threw ENOENT
opening <tmp>/.planning/config.json. Only 1 of 4 subtests failed,
confirming a per-invocation transient, not a deterministic miss; the full
suite schedules files differently so config.test.cjs did not collide with
those heavy neighbors → passed there.

Fix: add ensureConfigReady(tmpDir) which retries config-ensure-section on
ANY failure or missing file and throws a clear diagnostic if it still
cannot create config.json, then use it in both prototype-pollution
beforeEach blocks. Setup is now deterministic under load; the #663/alert-#26
security assertions are unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 20:36:11 -04:00
Tom Boucher
d32b8db635 feat(#836): no-LLM duplicate-issue detection + challenge + 1-day auto-close (#843)
* feat(#836): no-LLM duplicate-issue detection + challenge + 1-day auto-close

Adds a deterministic (no-LLM) duplicate-issue governance lifecycle:

- scripts/issue-dedupe.cjs: pure, unit-tested module (tokenize, Sørensen–Dice
  title similarity, scoreCandidates, renderChallengeComment, shouldClose) with
  fail-safe destructive-action guards.
- duplicate-check.yml (issues:opened): scores new-issue title against open
  issues, posts a challenge comment + applies the pending `possible-duplicate`
  label on a clear match.
- duplicate-sweep.yml (daily cron): closes possible-duplicate issues whose
  challenge comment is >24h old with no human reply and no 👎 veto; honors
  exempt labels; re-checks the label immediately before close (TOCTOU guard);
  strips the label on close to avoid reopen loops.
- remove-duplicate-label.yml (issue_comment:created): clears the label and
  applies needs-maintainer-review when any human responds.
- bug_report.yml / docs_issue.yml: add the required "I searched existing
  issues" preflight checkbox so all five forms force a pre-search attestation.
- docs/agents/triage-labels.md: document the label + lifecycle.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#836): add changeset fragment for duplicate-issue detection

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 20:35:43 -04:00
Tom Boucher
988024c1a3 fix(#837): three-dot diff in ci-test-scope so docs-only PRs skip the heavy matrix (#841)
CI test-scope detection diffed changed files with a two-dot
`git diff --name-only base head`, where base is the moving tip of `next`.
A PR branch cut from a slightly older `next` surfaced every product file
`next` had gained since the merge-base, flipping product_changed/full_matrix
and running the full Windows/macOS matrix + coverage on docs-only PRs.

Switch to a three-dot `git diff --name-only base...head` (vs the merge-base),
matching GitHub's PR "Files changed" semantics. Add a regression test that
builds a stale-base topology, plus a guard test pinning `fetch-depth: 0` on
the `changes` job (required for the merge-base to be locally available).

Closes #837

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 20:28:45 -04:00
Tom Boucher
1040fb792e feat(#777): register Cursor-native hooks (.cursor/hooks.json) for session-start/post-tool parity (#831)
* feat(#777): register Cursor-native hooks (.cursor/hooks.json) for session-start/post-tool parity

- Add gsd-cursor-session-start.js: injects STATE.md presence reminder (or
  new-project nudge) into Cursor sessions via the sessionStart hook event
- Add gsd-cursor-post-tool.js: emits an additional_context nudge when
  write-class tool calls touch .planning/ files (postToolUse hook event)
- Add 'cursor-hooks-json' installSurface to runtime-config-adapter-registry;
  writeCursorHooksJson/reconcileCursorHooksJson write the canonical
  { version: 1, hooks: { sessionStart, postToolUse } } JSON shape with
  idempotent reconciliation that preserves user-owned hook entries
- Hook scripts are copied with /gsd:→gsd- rewrite so installed files
  contain no colon-form slash-command refs (bug-376 invariant)
- 20 new tests in tests/cursor-hooks.test.cjs cover all reconciler paths,
  entry helpers, removal, runtime adapter surface, and hook script behavior
- Update CONTEXT.md, ARCHITECTURE.md, installer-migrations.md, and
  000-first-time-baseline.cts to include Cursor hooks.json surface

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#777): build hooks/dist on demand in bug-376 test for scoped/windows CI

hooks/dist is gitignored and only produced by `npm run build:hooks`.
The CI scoped (ubuntu-latest/node-22) and windows (windows-latest/node-24)
test jobs do NOT run build:hooks before executing tests, so bug-376's
prerequisite suite was failing with "hooks/dist not found" on both legs.

Add ensureHooksDist() helper (mirrors bug-3357 pattern) that builds
hooks/dist on demand in the before() hooks of prerequisite and Suite 3.
Also add ensureHooksDist() call to Suite 3's before() so the snapshot
step is also hermetic.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 19:22:48 -04:00
Tom Boucher
e04e757672 feat(#776): Gemini hook events (BeforeAgent/AfterAgent/BeforeModel) + hooksConfig.enabled check (#829)
* feat(#776): Gemini hook events (BeforeAgent/AfterAgent/BeforeModel) + hooksConfig.enabled check

Register three new Gemini-CLI hook events on install:
  - BeforeAgent: fires before agent planning; wired to gsd-context-monitor
  - AfterAgent: fires after final response generation; wired to gsd-context-monitor
  - BeforeModel: fires before each LLM call (per-turn); wired to gsd-context-monitor

All three reuse gsd-context-monitor.js (no new hook files). Uninstall cleanup
loop extended to remove the new events. Non-array guard added for robustness
against malformed settings.

Also detect hooksConfig.enabled:false in Gemini settings and emit a clear
warning — without this check, all registered hooks silently do nothing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: update changeset pr: 829

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#776): document Gemini hook events

Add hook coverage table to the Gemini CLI section of install-on-your-runtime.md,
covering the three new events (BeforeAgent/AfterAgent/BeforeModel wired to
gsd-context-monitor) plus a callout for the hooksConfig.enabled:false silent
failure mode detected by the installer.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 19:16:32 -04:00
Tom Boucher
3aed02822d chore(#771): convert agent color: hex/magenta values to documented named colors (#823)
* chore(#771): convert agent color: hex/magenta values to documented named colors

Claude Code's sub-agent `color:` field documents only 8 named colors
(red, blue, green, yellow, purple, orange, pink, cyan). Twelve agent
files used hex values and two used the undocumented `magenta`; convert
each to the nearest documented named color so the intended per-agent
TUI color differentiation is spec-compliant.

- agents/*.md: 14 color values hex/magenta -> nearest named color
- scripts/research-profiles.cjs: update the 3 generated research-agent
  profiles (source of truth) so gen-research-agents stays in sync
- docs/AGENTS.md: update documented colors; add missing Color rows for
  gsd-nyquist-auditor, gsd-project-researcher, gsd-phase-researcher
- tests/agent-frontmatter.test.cjs: add regression guard asserting every
  agent color: is in the documented named-color set

Closes #771

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#771): add changeset

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 18:50:50 -04:00
Tom Boucher
571d7b5a1c feat(#764): skip cross-platform test matrix for docs-only and inert-CI PRs (#798)
test.yml had no paths filter and the ci-test-scope classifier treated docs/
and every .github/workflows/* as code_changed, so documentation edits and
product-irrelevant automation tweaks still spun up the full Linux/Windows/macOS
matrix. Narrow the heavy matrix to changes that can actually affect the product
or the test pipeline.

- ci-test-scope.cjs: drop docs/ from code_changed (docs-only -> full skip; the
  required-tests fan-in still reports green). Add src/ to code_changed (it was
  missing -> a source-only PR previously skipped all tests). Add INERT_WORKFLOWS
  allowlist + isInertCi() + an "inert CI" rule, and a product_changed output that
  gates the heavy test/coverage jobs. Fail-safe: any workflow not on the inert
  allowlist defaults to the full matrix. A module-load assertion throws if a
  PROTECTED_WORKFLOWS entry (test/install-smoke/mutation/security-scan/release)
  is ever added to the inert set, so a weakening edit fails CI loudly.
- test.yml: keep the static 3-lane matrix (so the H1 shell-policy linter can
  still statically verify the Windows lane), gate test/coverage on
  product_changed, add a lightweight ubuntu-only test-inert job, and branch the
  required-tests fan-in on product_changed.
- docs-required.yml: run docs-parity-live-registry (gated on docs/ changes) so
  pure-docs PRs still catch live-registry drift without the matrix.
- tests: cover docs-only, inert-only, src/, pipeline, unknown-workflow fail-safe,
  mixed escalation, the code_changed=false -> no-lanes invariant, and protected-
  workflow tamper-evidence.

Closes #764

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 11:46:48 -04:00
Solvely-Colin
dd75400801 fix discord release changelog announcements 2026-06-07 10:27:20 -04:00
Tom Boucher
5237fae537 feat(#759): non-destructive CHANGELOG preview in the rc release job (#763)
The rc action publishes a release candidate to @next for testing but
never surfaces the curated CHANGELOG section for the version under test —
render only runs destructively at finalize (#715), so there was no safe
way to preview the upcoming notes during the RC window.

Add a --preview mode to scripts/changeset/cli.cjs cmdRender: it renders
the dated release section to stdout via the existing renderChangelog/
serializeChangelog path (with priorChangelog: null, so only the new
section is emitted), reuses the shared injectEmptyPlaceholder helper for
zero-fragment releases, and returns WITHOUT writing CHANGELOG.md or
deleting any .changeset fragment. Wire a "Preview CHANGELOG" step into
the rc job that renders to a file (standalone command, so a malformed
fragment fails the step) and cats it to the job summary and log.

Closes #759

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 10:05:21 -04:00
Tom Boucher
edd3c6c332 fix(#730): scope current-milestone Phase Details section in roadmap parser (#748)
`extractCurrentMilestone()` scoped the current-milestone window to its
`## Phases` checklist subsection and terminated at the milestone's own
`## Milestone … (Phase Details)` heading, so the `### Phase N:` detail
headers fell outside scope. Every parser-backed command — `init.phase-op`
(and thus `/gsd:discuss-phase`, `/gsd:plan-phase`), `state`, `roadmap list`,
and `validate health` (W006) — therefore could not resolve phases of any
milestone after the first until a `.planning/phases/` directory already
existed, blocking discuss/plan.

The parser now additionally includes the current milestone's `(Phase Details)`
section in scope, located via the already-computed version matches and anchored
(boundary-aware) to the selected milestone's version token so sibling
sub-milestones sharing a version prefix do not cross-pollinate. The existing
heading selection and primary window are unchanged.

Adds tests/bug-730-milestone-phase-details-scope.test.cjs covering the
two-milestone reproduction, first-milestone non-regression, direct
getRoadmapPhaseInternal resolution, validate-health W006 visibility, a
three-milestone roadmap, and the closed-sibling sub-milestone case.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 22:13:20 -04:00
Tom Boucher
f729101eec refactor(scripts): replace process.exit() with ExitError + runMain handler (#739) (#740)
Part 1 of 2 of the n/no-process-exit cleanup (umbrella #738): convert every
process.exit() call in standalone scripts/** CLIs to the rule-compliant pattern.

- New shared helper scripts/lib/cli-exit.cjs: ExitError(code,message) + runMain()
  which translates a thrown ExitError / returned number into process.exitCode
  (never process.exit()), flushing output and still firing process.on('exit').
- main()-based entrypoints: throw new ExitError(code) for errors, return <code>
  for verdicts; invoked via runMain(main). Child exit codes preserved via return.
- top-level-only scripts: imperative body extracted into main() so mid-flow
  aborts (throw ExitError) actually halt; pure consts/helpers stay at module scope.
- diff-touches-shipped-paths.cjs: stdin event handling restructured to an async
  read so the whole flow runs under runMain; uncaughtException/unhandledRejection
  nets replaced by an in-band catch that preserves EXIT_ERROR=2.

Exit codes verified unchanged for every converted script (success/error/help and
the 0/1/2 semantic codes in diff-touches). Rule stays warn here; flipped to error
in part 2 (#738) once gsd-core/bin/** is also clean.

Refs #739

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 16:13:13 -04:00
Tom Boucher
ba231ecbfc chore: clean up clear-cut ESLint warnings (#732) (#734)
Pay down pre-existing error→warn lint debt. Removes dead imports/vars, unused functions, redundant regex/string escapes, and stale eslint-disable directives; converts unused `catch (_e)` to optional catch binding (src/*.cts).

No behavior change. Lint 345→125 warnings (0 errors); deferred categories (n/no-process-exit, test-sleeps, control-regex) tracked in #732 for follow-up. Full test suite green (0 failures); code-review verified all removals unused and all escape fixes semantics-preserving.

Closes #732

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 11:24:48 -04:00
Tom Boucher
11afca2968 feat(#656): Research module — content-addressed cache + provider seam + registry-API legitimacy (#664)
* feat(#656): add Research Store module (content-addressed cache, TTL staleness)

Content-addressed research cache behind a clock seam: researchKey (sha256, deterministic), putResearch/getResearch ({hit,stale}, never throws), ttlForSource (curated HIGH 30d / MED 7d / web LOW 1d), two-tier resolveStorePath (curated -> ~/.gsd/research-cache, web/synthesis -> project .planning/research/.cache). 28 behavioral + property tests; boundary coverage at ttl-1/ttl/ttl+1.

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#656): add Research Provider module (waterfall + confidence + plan)

Single source of truth for the Balanced provider waterfall (docs Context7->Ref->Jina, web Exa+Tavily, fallback Perplexity/Brave, Firecrawl scrape-only). classifyConfidence stamps HIGH|MEDIUM|LOW by provider (never throws). providerAvailability maps config flags to usable providers. planResearch checks the Research Store (injected seam) and returns cache-hits + a per-question fetch plan, falling through the waterfall to the always-available websearch terminal. 22 behavioral + property tests.

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#656): add Package Legitimacy module (registry-API verdicts, slopcheck optional)

Replaces the pip-install-or-degrade slopcheck prose gate with code: classifyPackage (pure, never throws) computes OK|SUS|SLOP from tunable thresholds (minAgeDays 30, minWeeklyDownloads 1000, requireRepo). checkPackages queries injectable npm/PyPI/crates registry adapters (real https with 5s timeout, degraded-not-thrown on failure); slopcheck is one optional adapter that can only escalate severity, never degrade to [ASSUMED]. 34 behavioral + property tests; boundary coverage on age and downloads (limit-1/limit/limit+1).

Known follow-up: real npm adapter must add api.npmjs.org last-week downloads fetch (currently null -> unknown-downloads). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#656): detect Tavily/Ref/Perplexity/Jina provider keys; complete npm downloads adapter

config: add tavily_search/ref_search/perplexity/jina availability flags (env var or ~/.gsd/<x>_api_key), mirroring brave_search/exa_search/firecrawl, so the Research Provider waterfall can gate them. package-legitimacy: real npm adapter now fetches api.npmjs.org last-week downloads (bounded, degraded-not-thrown) so weeklyDownloads is populated. +12 config tests; 34 legitimacy tests unchanged.

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#656): expose Research seam via gsd-tools query (research-plan, research-store, package-legitimacy)

Routes the L2-hybrid surface so agents reach it as CLI: 'query research-store get/put' (cache, HOME-sandboxable), 'query research-plan --input' (cache-hits + fetch plan from planResearch), 'query package-legitimacy check --ecosystem' (async registry verdicts). Commands skip .planning root resolution and appear in top-level usage. 5 behavioral runGsdTools tests; command-contract unchanged (335).

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#656): document Research module (CONTEXT predicates, ADR-0656, architecture, changeset)

Adds GSD-RESEARCH.* + DEFECT.RESEARCH-PROVIDER-PROSE-DRIFT predicates to CONTEXT.md, ADR-0656 recording the L2-hybrid seam decision, a docs/ARCHITECTURE.md Research Module subsection, and an Added changeset fragment (pr:0, backfill on PR). Notes the #657 deferrals (agent collapse + install.js MCP mapping).

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#656): sync inventory for research modules

Regenerate INVENTORY-MANIFEST.json and bump docs/INVENTORY.md CLI Modules count 82->85 with rows for research-store/research-provider/package-legitimacy (DEFECT.INVENTORY-DRIFT).

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#656): eslint-ignore generated research .cjs artifacts (ADR-457)

research-store/research-provider/package-legitimacy .cjs are tsc-generated from src/*.cts, so they belong in the ESLint ignore block (lint the .cts source, not the emitted .cjs). Fixes tests/551-eslint-bin-lib-coverage.

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#656): backfill changeset pr number to #664

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#656): satisfy eslint lint-tests gate

Fix 20 eslint errors in the new research files: use helpers.cleanup() instead of raw fs.rmSync() in tests (local/no-raw-rmsync-in-tests, Windows-EBUSY retry budget); drop redundant '| string' union members and unnecessary type assertions; deterministic object normalization in researchKey (no-base-to-string). Logic unchanged; 6180 tests still green.

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): harden package legitimacy per review (W1/W2/I3/I4)

W1: httpsGet now reads statusCode; npm/PyPI/crates map 404 -> exists:false -> SLOP (registry-existence is the #1 slopsquatting defense; previously only npm caught it). Transport made injectable (_setHttpGet) for hermetic 404 tests. W2: suspicious-postinstall is now terminal SLOP independent of the optional slopcheck adapter, and the regex drops the bare https?:// arm (over-fired on esbuild/sharp/node-gyp) for shell-exec/download-exec signatures only. I3: checkPackages now threads version to registry.lookup and adapters verify that specific version exists. I4: moreServerVerdict -> moreSevereVerdict. +11 regression tests (all RED-first); 45 total green.

Addresses review by @davesienkowski on #664. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): research-store tier coherence + freshness + version TTL (W4/I1/I2/I4)

I1: tier now derives from source (curated -> user ~/.gsd, else -> project .planning), not kind, so put-tier and get-tier can't diverge; kind is a key component only. W4: getResearch searches both tiers and returns the freshest (non-stale preferred), never letting a stale curated entry shadow a fresh web one; blank version caps TTL at 1 day (no 30d on version-blind keys). I2: atomic platformWriteSync instead of raw fs.writeFileSync on the shared global path. I4: dropped the dead ttlForSource arm. CLI get now searches both tiers. +5 RED-first regression tests; 38 green.

Addresses review by @davesienkowski on #664. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): expose classifyConfidence as a CLI route, killing dead code (W3)

Adds 'gsd-tools query classify-confidence --provider X [--verified]' so research agents get the confidence tier FROM CODE (provider waterfall + verification lever) instead of asserting it in prose. classifyConfidence previously had no runtime caller. HIGH means 'trusted provider'; --verified raises web results to MEDIUM (verification semantics documented in ADR-0656). +4 behavioral tests.

Addresses review by @davesienkowski on #664 (W3). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): close Codex adversarial-review findings (path-traversal, version-age, malformed-cache)

HIGH: research key must be 64-hex sha256 (isValidResearchKey) + resolved-path containment check in put/get + CLI validation -> blocks '../../x' arbitrary-file-write. HIGH: package legitimacy now derives publishedAt from the REQUESTED version (npm time[version], PyPI releases[version] upload_time, crates versions[].created_at) so a new malicious version of an old package can't inherit old age and evade 'too-new'. MEDIUM: getResearch validates entry shape (finite fetched_at + positive ttl + required fields) -> malformed cache entry is a miss, not fresh-forever. +regression tests (RED-first); 111 green.

Codex adversarial review (required pre-PR gate). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): close code-review correctness findings

(1) package-legitimacy CLI now rejects unknown --flags instead of silently consuming the following package as a flag value; only --ecosystem takes a value. (2) crates recent_downloads (90-day) normalized to a weekly figure before the minWeeklyDownloads threshold (was ~13x too lenient). (3) research-plan --input validates parsed JSON is an object with an Array questions before destructuring -> clean usage error instead of an uncaught TypeError on null/bad input. (4) research-store put rejects a flag value that is itself a --flag (no more storing '--source' as content). (5) planResearch skips questions whose text is not a non-empty string instead of emitting question:undefined. +13 RED-first regression tests; 143 green.

Code-review gate. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(#657): extract researcher documentation_lookup to shared @-reference

6 researcher agents carried a near-duplicate <documentation_lookup> block; consolidate into gsd-core/references/research-documentation-lookup.md (@-included). Unifies the ctx7 CLI fallback to the safer 'command -v ctx7' guard (drops silent 'npx --yes ctx7@latest' execution in 5 agents). Behavior-preserving dedup; inventory 63->64 references. Phase A of the agent collapse.

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(#657): extract researcher philosophy + verification-protocol to shared @-references

philosophy and the pitfalls+pre-submission-checklist common-core were near-duplicated in project/phase researchers; consolidate into gsd-core/references/research-{philosophy,verification-protocol}.md (@-included). phase-researcher keeps its 3 extra checklist items inline. Pre-submission domains checklist made agent-agnostic so project-researcher doesn't lose features/architecture coverage. Write-contract intentionally left inline (bug-214 tests assert it verbatim). Inventory 64->66 refs. Behavior-preserving. Phase A.

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#657): wire gsd-phase-researcher to the Research seam (Phase B / S1)

The phase researcher now CALLS the code seam instead of carrying inline mechanics: provider waterfall -> 'gsd-tools query research-plan' (+ research-store put to cache digests); confidence-tier prose -> 'gsd-tools query classify-confidence'; slopcheck pip-install protocol -> 'gsd-tools query package-legitimacy check'. This makes the Research module a real runtime consumer (validates the seam end-to-end, addresses reviewer S1) and removes the duplicated waterfall/confidence/slopcheck prose. RESEARCH.md output contract, commit step, structured returns, and Phase-A @-includes unchanged. package-legitimacy-gate.test.cjs rewritten prose-grep -> behavioral (asserts the seam invocation).

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#657): wire gsd-project-researcher to the seam + add tavily/ref/jina MCP tools (Phase C.1)

project-researcher now calls gsd-tools query research-plan / classify-confidence (+ research-store put) instead of the inline provider waterfall + confidence-tier prose (mirrors the phase-researcher rewire; no package-legitimacy — phase-only). Output contract (STACK/FEATURES/ARCHITECTURE/PITFALLS/SUMMARY.md + sections, no-commit, structured returns, Phase-A @-includes) unchanged. Adds mcp__tavily/ref/jina__* to the project/phase/ui researcher tools frontmatter (Balanced provider set) so install.js MCP mapping (C.2) has a consumer.

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#657): cover tavily/ref/jina MCP install handling + frontmatter parity guard (Phase C.2)

Investigation: exa/firecrawl have no explicit per-runtime tool-mapping — every mcp__<server>__* except context7 rides the generic passthrough (Copilot lowercases; OpenCode/Cursor/Windsurf/Augment keep as-is; Gemini auto-discovers). tavily/ref/jina are handled identically, no install path broken. Added 12 copilot-install passthrough tests + a mcp-tool-inheritance parity guard (tavily co-declared with exa, jina with firecrawl, ref present across the 3 web researchers) so the MCP set can't drift. No io.github registry ids invented (none sourceable in-repo); documented as a follow-up. 488 tests green.

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#657): profiles as source of truth for researcher agents + drift-guard (Phase C.3)

scripts/research-profiles.cjs declares each of the 7 researcher agents' identity + contract (name, description, color, tools, required @-includes, required gsd-tools seam calls, output-contract markers). scripts/gen-research-agents.cjs --check validates every committed agent against its profile; --write regenerates ONLY the frontmatter from profiles (body untouched) and is a verified no-op against the current agents (zero diff = fidelity). tests/research-agent-profiles.test.cjs is the DEFECT.GENERATIVE-FIX drift guard. Design note: profiles govern the generatable/contract surface rather than destructively regenerating the disparate operational prose bodies (those were deduped via @-includes in Phase A). scripts/ is not inventoried (no inventory change).

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#657): complete agent provider-dispatch + parity guard; align legitimacy field; validate profiles

Adversarial-review findings: (HIGH) the seam-wired agents' Step-C dispatch only mapped 6 providers, so a planResearch result of jina/ref/perplexity/brave (reachable via the waterfall fallbacks) had no handling -> agent stall; completed both agents' dispatch to all 9 PROVIDER_WATERFALL ids + a catch-all, and added a parity test asserting agent dispatch stays in sync with research-provider PROVIDER_WATERFALL (DEFECT.GENERATIVE-FIX). (MEDIUM) phase-researcher package-legitimacy JSON example used 'package' but the module returns 'name' -> aligned. (LOW) gen-research-agents checkAgent now returns a clear failure for a malformed profile instead of throwing. +parity/validation tests (RED-first).

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): make classifyConfidence verification-evidence-driven (W3)

Confidence conflated provider authority with claim verification — context7/ref
stamped HIGH purely by provider identity, and the only verification lever was a
self-set --verified flag. Split into two axes: provider authority (static) +
verification evidence (code-computed). HIGH now requires ground-truth
corroboration (legitimacyVerdict OK), independent of provider; authority alone
caps at MEDIUM; SLOP caps at LOW; the self-reported --verified is demoted to a
MEDIUM-only web lever. HIGH = corroborated-against-authoritative-source, not a
correctness guarantee. Adds --legitimacy-verdict to the classify-confidence CLI;
updates CONTEXT.md predicate + ADR-0656 (tier set unchanged, ADR-consistent).

Addresses davesienkowski's W3 review on #664.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#656): bind classify-confidence verdict to code, closing CLI self-grading

Adversarial review found the new --legitimacy-verdict flag was caller-supplied,
so an agent could self-assert OK->HIGH without any real legitimacy check —
reintroducing the exact self-grading hole W3 closes. Remove the free flag; the
CLI now computes the verdict via checkPackages only when --package/--ecosystem
is given (code-computed, not agent-asserted). Update the stale CLI test
(context7 alone -> MEDIUM) and extend the property test to vary legitimacyVerdict.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 17:58:48 -04:00
Tom Boucher
afd60b809a chore(release): wire CHANGELOG render into release finalize job (#690 follow-up) (#715)
* feat(#690): wire CHANGELOG render into release finalize job

CHANGELOG promotion has always been a manual operator step, which is why
1.3.0/1.3.1 shipped unpromoted (#690). PR #694 added a `verify` latch that
fails a release lacking a dated heading, but nothing performed the promotion.

Wire `changeset render` into the finalize job, after build/test and before
the verify gate, committing the promoted CHANGELOG so it ships with the
release. Add a `--allow-empty` flag to cmdRender so a zero-fragment release
still emits a dated heading (with a '_No notable changes._' placeholder)
instead of writing nothing and tripping the verify gate.

Note: requires the changeset-archive cleanup (separate PR) to land first, so
the first render consumes only genuinely-unreleased fragments.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#713): set changeset pr number to 715

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 14:56:49 -04:00
Tom Boucher
3042b79178 fix(#690): promote CHANGELOG 1.3.x + gate release-notes promotion (#694)
* fix(#690): promote CHANGELOG 1.3.x + gate release-notes promotion

/gsd:update showed an empty "What's New" preview after updating to 1.3.1
because CHANGELOG.md's 1.3.x content was never promoted out of [Unreleased]
into dated sections, so `scripts/changeset/cli.cjs extract` returned exit 2
("no releases in range").

- CHANGELOG.md: split [Unreleased] into dated [1.3.0] and [1.3.1] sections
  (1.3.1 = hono advisory bump + installer-migration checksum self-heal, #670;
  1.3.0 = the feature release), restoring an empty [Unreleased].
- scripts/changeset/cli.cjs: new `verify` subcommand that exits non-zero when
  CHANGELOG has no dated `## [x.y.z]` heading for a version; hoist shared
  stripV/resolveChangelogPath helpers used by extract + verify.
- .github/workflows/release.yml: gate the finalize job on `verify` (after the
  build, before tag/publish) so an unpromoted CHANGELOG can never ship again.
- gsd-core/workflows/update.md: move `rm -f $CHANGELOG_TMP` after the
  human-readable extract re-run so the preview no longer degrades to
  "(changelog unavailable)".
- tests: regression guard for the 1.3.x headings + extract range + verify
  command coverage (present/absent/undated/v-prefixed/--json/prerelease).

Closes #690

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#690): add changeset fragment for #694

Fixed-type fragment for the user-facing /gsd:update preview fix and the
release-notes promotion gate.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 22:48:21 -04:00
Tom Boucher
cdd78bd2aa fix(#670): self-healing recovery for installer-migration checksum drift (#675)
Editing the body of an already-released installer migration drifts its computed
checksum (it hashes plan.toString()). The integrity guard then hard-aborted
every prior install on upgrade with "applied migration checksum changed" — a
100% reproducible blocker (v1.3.0, all platforms).

Already-applied migrations are filtered out of `pending` and never re-run, so
a drifted checksum is functionally inert. ADR-0008 anticipates checksum-mismatch
state as something the install-state layer must handle gracefully (plan -> apply
-> recover/report), not abort on.

This supersedes the published-checksum allowlist merged in #674 (per-release
maintenance debt — every historical checksum hand-pinned, still throws for any
unregistered value) with a general, self-healing recovery:

- Replace the throwing guard with non-fatal `collectAppliedChecksumDrift`,
  surfaced on `plan.checksumDrift`.
- Reconcile drifted stored checksums durably on the next state write
  (`reconcileDriftedChecksums`), idempotently (no perpetual writes).
- Relocate the "shipped migration bodies are immutable" rule to a CI baseline
  test that locks every shipped migration's checksum and fails on body drift —
  where #615 should have been caught, instead of blocking users.

Removes #674's legacyChecksums field, per-migration checksum pins,
published-checksums.json fixture, and compat test.

Fixes #670

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 13:44:05 -04:00