Commit Graph

3508 Commits

Author SHA1 Message Date
Tom Boucher
fbd62cd84f feat(#1035): phase 5a — author 16 role:runtime capability descriptors (registry-only) (#1039)
Author capabilities/<runtime>/capability.json for all 16 runtimes (role:runtime),
populating the registry's runtimes index ({} -> 16). Each carries the 6 ADR-1016
axes (configHome structured, configFormat, artifactLayout structured, commandStyle,
hooksSurface + hookEvents, sandboxTier, supportTier), extracted from the live
modules (runtime-homes/runtime-slash/runtime-artifact-layout/runtime-config-adapter
+ CODEX_AGENT_SANDBOX). validateRuntimeBody tightened to enforce the closed
vocabularies + structured configHome/artifactLayout (rejects old string configHome;
env required; skillsHome recursively validated).

Registry-only: bin/install.js + src/*.cts untouched, zero behavior change. The 5b-5f
drive steps consume these one axis at a time (ADR-1016 §8).

Codex accuracy review caught + fixed: opencode/kilo register ZERO lifecycle hooks
(install.js skips the whole block) -> hooksSurface:none (configFormat stays
settings-json); kimi probe selects on <root>/skills existence -> probeExists:"skills".
ArtifactKind required-field strictness + ConverterName closing deferred to 5d/5e.
ADR-1016 (Proposed) amended to match (Decisions 1 + 5).

Closes #1035

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 09:34:07 -04:00
github-actions[bot]
b2222fa769 docs(#860): ADR-3660 addendum — Nth-runtime-via-existing-layout is an addendum, not a new ADR
Maintainer governance decision (waiving a standalone ADR for Qoder, PR #1021):
registering a runtime that reuses the existing profile-marker-only install
surface + a single skills kind is an enhancement governed by ADR-3660 via
addendum. Records the addendum-vs-new-ADR qualifying criteria, the normative
agent-frontmatter contract (name+description only — the sibling converter
shape), and Qoder as the first runtime logged under this path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 08:39:22 -04:00
Tom Boucher
8bd5e07c58 feat(#1031): autonomous §3a.5 plan:pre ui-phase cutover — last inlined ui-phase site (step-only) (#1033)
Cut over autonomous.md §3a.5 (autonomous plan:pre ui-phase step) to the
loop.render-hooks plan:pre dispatch — completing the ui-phase migration begun
in #1026 (plan-phase.md §5.6). Step-only, non-blocking: autonomous is always
pipeline, so it fires active kind==step hooks and never runs the manual-only
plan:pre blocking gate.

Skip condition keys on "no active step hooks" (not empty activeHooks), so the
gate-only {ui_phase:false, ui_safety_gate:true} case skips silently with no
spurious warning — matching OLD §3a.5. Fires gsd-ui-phase under the identical
precondition (frontend + no UI-SPEC + workflow.ui_phase active), bare
${PHASE_NUM} args. Replaces the inline ui-safety-gate.cjs probe + config-get
with render-hooks + the ui.plan-gate check verb.

Codex caught the gate-only spurious-warning divergence on the first pass; fixed
+ re-confirmed equivalence-preserving. gsd-ui-phase skill, §5.6, §3d.5 untouched.

Closes #1031

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 08:31:35 -04:00
Tom Boucher
9ed8c7d574 feat(#1026): §5.6/ui-phase cutover — first gate dispatch (plan:pre step + blocking gate) (#1028)
Replace plan-phase.md §5.6 (a 6-branch inline UI gate) with a capability-driven
loop.render-hooks plan:pre dispatch — the FIRST gate dispatch in any workflow.
A step (ui-phase, when:workflow.ui_phase) + a new blocking gate
(when:workflow.ui_safety_gate). New ui.plan-gate check verb returns
{frontend, hasUiSpec, block}; the dispatch runs it unconditionally then fires
the active step (pipeline) or halts on the active blocking gate (manual). The
gate-handling (run check.query; halt if blocking+block) is the reusable
phase-6 template for blocking-gate cutovers.

Config semantics fixed per #1022 + maintainer call: ui_phase gates plan-time
UI-SPEC generation, ui_safety_gate gates the planning block. Common case + all
ui_phase=false cases are equivalence-preserving; the one intended change is
{ui_phase:true, ui_safety_gate:false} now auto-generating in pipelines.

Review found it broken twice (non-generic dispatch, phase-lookup divergence,
then the step-only check nested in a gate loop) — fixed; final Codex pass
verified all 8 (ui_phase,ui_safety_gate)x{pipeline,manual} cases correct.
gsd-ui-phase skill + autonomous §3a.5 untouched (§3a.5 deferred).
getRoadmapPhaseWithFallback mirrors cmdRoadmapGetPhase for lookup parity.

Closes #1026

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 00:58:21 -04:00
Tom Boucher
5f48a42514 docs(#1022): resolve step-vs-gate model question — steps additive, gates block, mode self-gates (#1025)
Record the resolution of #1022 (surfaced scoping the §5.6 ui-phase cutover):
a step is purely additive and never halts the host; host-blocking preconditions
are gates (blocking/onError:halt) — no hook-model change. Runtime/mode context
(auto/chain vs manual) self-gates in the skill, not via when (config-only).

§5.6 decomposes into the existing plan:pre step (ui-phase, self-gates on
frontend + pipeline) + a new plan:pre gate (frontend-and-no-UI-SPEC → halt,
when: workflow.ui_safety_gate) that blocks planning in manual mode — preserving
the "run /gsd:ui-phase first" UX (maintainer call: pipelines-only auto-fire).
The render-hooks dispatch template grows to handle gates, not just steps.

Recorded in ADR-894 (clarification) + CONTEXT.md
(RULESET.CAPABILITY.step-additive-gate-blocks). Unblocks the §5.6 cutover.

Closes #1022

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 22:51:31 -04:00
Tom Boucher
2ac6592096 feat(#1023): first phase-6 cutover — ui-review (verify:post) inline → loop.render-hooks dispatch (#1024)
Replace the inlined ui-review invocation in autonomous.md §3d.5 with a
loop.render-hooks verify:post dispatch — the first workflow to consume
render-hooks and fire a skill from it (closes the #1018 live-execution residual
as real wiring). Capability-driven, equivalence-preserving for the current
registry (only ui-review at verify:post, default on): fires gsd-ui-review under
the same precondition (UI-SPEC exists via consumes-gate + workflow.ui_review).

Gate findings (real pattern issues, fixed so every future cutover inherits them):
- bug-2643 static "Skill() references a real skill" check vs templated
  Skill(skill="gsd-${ref.skill}") dispatch → skip ${...}-templated names.
- Coverage moved, not lost: gen-capability-registry now validates
  steps[].ref.skill in skills + ref.agent in agents + rejects gsd- double-prefix.
- Tightened §3d.5 tests; markdown clarity (consumes rule, LLM-native JSON read,
  UI-REVIEW.md score hint).

gsd-ui-review skill + §3a.5/ui-phase untouched. §5.6/ui-phase cutover deferred
(#1022 step-can-halt-vs-gate model question).

Closes #1023

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 22:40:58 -04:00
Tom Boucher
7eea991884 test(#1018): hook-firing de-risk spike — structural off-means-off proof + cutover contract (#1020)
Spike disposition (a): prove the loop.render-hooks mechanism at the data level
and document the cutover-shape finding, without touching plan-phase.md §5.6.

- tests/loop-hook-firing-spike.test.cjs: makes "off means off" executable — a
  host-computed aggregate (hostConsume) that is a pure function of the active
  hook set; UI active-by-default → 1, off → 0 == independently-computed base
  (not tautological); synthetic 2-hook ordering; second point verify:post.
- CONTEXT.md RULESET.CAPABILITY.cutover-self-gating: the loop hook is coarse;
  phase-context detection + mode self-gate in the skill (ADR-894); §5.6 UI gate
  is the worked example of what must move into gsd-ui-phase before cutover.

Finding: the UI plan:pre hook is already inlined (§5.6) — cutover is move
host-logic-into-skill, not a 1:1 swap. Live LLM-execution of injected hook
markdown is the residual, carried into the first phase-6 cutover acceptance.

Closes #1018

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 21:40:30 -04:00
Tom Boucher
61b29d2af6 docs(#1016): ADR-1016 runtime capability descriptor (ADR-857 phase 5 design) (#1019)
Realize ADR-857 Branch 8 (host-CLI support as role:runtime Capabilities) +
materialize ADR-58's InstallPlan. Closes the 6-axis descriptor vocabulary,
absorbs the hard-case runtimes as data, and stages the install migration as a
5a→5g ladder (4/6 axes already modularized; InstallPlan is the 5g capstone,
reachable by collection, not a big-bang rewrite).

Amended after a plugin-side design grill (rubber-duck + grill-with-docs vs
#956 MemPalace / #999 Impeccable):
- New term Connected Capability (CONTEXT.md) — a Capability whose integration
  shape brings its own external process/service/state; orthogonal to authorship.
  Named, tracked gap (vehicle #956); current schema does not express it.
- Narrowed the dogfood claim: the descriptor dogfoods the runtime interface
  only, not the feature-plugin/Connected path.
- Structural "off means off" rule (CONTEXT.md RULESET.CAPABILITY.off-means-off):
  the host derives shared outputs from active hooks; a hook adds/is-counted,
  never mutates host source. Ratify in ADR-894; proven by spike #1018.
- Hand-waves resolved by code: sandboxTier real-but-thin; model-catalog
  orthogonal (not a 7th axis); converters closed into a ConverterName enum +
  added the kimi-agents artifact kind; configHome is pure read-only.
- Hook-firing path is unproven (render-hooks built, never consumed) → spike
  #1018 must prove render-hooks→live-workflow execution before phase-5 build.

Design-only; no code. Status: Proposed.

Closes #1016

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 21:00:37 -04:00
Colin Johnson
76f42ddb4b feat(#1014): add Claude Fable 5 model config (#1015) 2026-06-10 20:32:26 -04:00
Rezolv
28ac89d810 fix(#1008): tolerate EAGAIN + short writes in io output()/error() (#1009)
* fix(#1008): tolerate EAGAIN + short writes in io output()/error()

The I/O Module wrote stdout/stderr with a bare fs.writeSync(fd, data), assuming
it blocks until the kernel accepts every byte. That is false when the fd is a
non-blocking pipe (as under the parallel node:test runner on Linux CI): a full
pipe throws EAGAIN and a partially-drained pipe returns a short count. The former
caused spurious failures (e.g. bug-974 graphify property test threw EAGAIN); the
latter risked silently truncating output.

Add writeAllSync(fd, data): loop on short counts and retry EAGAIN/EINTR with a
bounded backoff. The backoff sleep buffer is allocated lazily on the first retry
(rare) and reused — keeping it out of module load avoids perturbing the
SharedArrayBuffer-allocation accounting in perf-316 and costs nothing on the
common no-retry path. Route output() and error() through it; non-transient
errors (EPIPE) still propagate. Mirrors the transient-errno handling already
applied to STATE.md lock acquisition (ACQUIRE_LOCK_RETRY_ERRNOS / #3776).

Regression cases live in tests/io.test.cjs (the owning module's file, per the
regression-test-name placement policy) and inject fs.writeSync via mock.method:
EAGAIN/EINTR retry, short-write no-truncation, EPIPE still surfaces, and error()
retries while still exit(1). Red against the pre-fix bare-writeSync io.cjs.

* chore(#1008): add Fixed changeset for io EAGAIN/short-write fix
2026-06-10 17:05:07 -04:00
Jeremy McSpadden
092340d18a fix(#711): wire autonomous convergence flag (#729)
* fix(#711): wire autonomous convergence flag

* Update wise-ibex-tumble.md

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 16:10:07 -04:00
Tom Boucher
19edab21da fix(#1006): rc CHANGELOG preview crash on malformed changeset fragment + validate fragment content at the gate (#1007)
* fix(#1006): harden render --preview against fragment parse failures

`render --preview` wrote `report.preview` unconditionally. When a `.changeset`
fragment fails to parse, `cmdRender` early-returns with `{exitCode:1, report:
{failures}}` and NO `preview` key, so `process.stdout.write(undefined)` threw
ERR_INVALID_ARG_TYPE and the rc release job's "Preview CHANGELOG" step died with
a cryptic TypeError that masked the real cause.

Guard the preview write on `typeof report.preview === 'string'` (ADR-227: shape,
not just type); when absent, fall through to the existing failure reporter that
names the offending fragment and exits non-zero — identical to a non-preview
render. Also backfills the stray placeholder `pr: 0` -> `pr: 939` in
.changeset/936-convergence-inline-plan-phase.md that triggered the live failure.

Regression test (red-then-green verified) added at the render --preview seam.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1006): validate changeset fragment content at the Changeset Required gate

The `Changeset Required` gate (scripts/changeset/lint.cjs) only checked that a
`.changeset/*.md` fragment EXISTS in the PR diff; it never validated the
fragment's contents. So a malformed fragment (e.g. an un-backfilled `pr: 0`
placeholder) silently merged to `next` and only detonated later in the rc
release job. This is the upstream prevention for #1006 — the crash hardening
turns the failure into a clear message, this stops the bad fragment ever
reaching the release path.

evaluateLint now accepts `fragmentFailures` and fails with the typed reason
`fail_invalid_fragment` (naming each offending file) before the existence/
opt-out checks — a malformed fragment beats `no-changelog`, since it will break
the render regardless. main() reads + parseFragment()s every changed fragment:
a deleted fragment (not on disk) is skipped, a present-but-unreadable one fails
closed. Tests assert on the typed LINT_REASON enum (no raw-text matching), a
precedence case over the opt-out label, and an end-to-end suite that drives the
real main() against a temp git repo (malformed -> fail, valid -> pass, deleted
-> skipped) so the wiring is regression-proof.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1006): assert the typed --json report in the preview regression test

Code review flagged the preview parse-failure regression test for positive
raw-text matching on CLI output (`combined.includes('bad-fragment.md')` /
`'invalid_pr'`), which this repo's testing standards forbid. Keep the non-json
`runRenderRaw` call for the negative crash proof (the ERR_INVALID_ARG_TYPE
crash lives only on the non-json stdout.write path), and add a `--json`
invocation that asserts the offending fragment + typed `invalid_pr` reason via
the structured `report.failures[]` surface instead of rendered prose.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 15:55:31 -04:00
Jeremy McSpadden
f61b97276e fix(#724): block convergence on actionable review findings (#728)
* fix(#724): block convergence on actionable review findings

* merge: integrate clean next (#936 inline) onto author tip + re-apply cursor fixes (Mode field, REVIEWS.md extraction) and review hardening (#724)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 14:54:15 -04:00
Jeremy McSpadden
fb37fa7dd5 fix(#725): route Codex gsd-tools calls through shim (#731)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 14:44:43 -04:00
Joe
5e8a723089 feat(templates): add optional Business Context section to PROJECT.md template (#756)
* feat(templates): add optional Business Context section to PROJECT.md template

Adds an optional `## Business Context` section (Customer, Revenue model,
Success metric, Strategy notes) between Core Value and Requirements, for
monetized or customer-facing projects. Optional by default — an HTML comment
tells non-business projects to delete it; capped at four one-line fields to
stay a constraint reference, not a business plan. The milestone evolution
review in complete-milestone.md checks it only when the section is present.

Refs #72

* chore(changeset): set pr number for #72 fragment

* test(#72): add source-text-is-the-product exemption marker

Addresses review Minor #1 on PR #756. The contract test reads the
PROJECT.md template and complete-milestone workflow .md files and
asserts on their content (the local/no-source-grep pattern). Those
.md files ARE the product surface, so this is a valid
source-text-is-the-product case. Add the explicit // allow-test-rule
marker per RULESET.TESTS.no-source-grep.exemption so intent is
audit-traceable before the rule promotes to error (#453).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-10 14:13:01 -04:00
Tom Boucher
4c10eb2253 fix(#991): inject configured agent_skills into code-review family subagents (#1005)
* fix(#991): inject configured agent_skills into code-review family subagents

code-review.md, code-review-fix.md, and eval-review.md spawned their
subagents (gsd-code-reviewer / gsd-code-fixer / gsd-eval-auditor) without
querying or injecting the project-configured agent_skills, while ~20 sibling
workflows do. Subagents don't inherit the orchestrator's auto-loaded context,
so this injection is the only channel — reviewers/fixers/auditors silently ran
without the configured rule/skill context.

Mirror the established sibling idiom: add
`VAR=$(gsd_run query agent-skills <agent-type>)` in each workflow's initialize
step and interpolate `${VAR}` into every Agent() spawn of that type. This
covers all spawn sites, including code-review-fix.md's --auto loop which
re-spawns gsd-code-reviewer in addition to the two gsd-code-fixer spawns.

Regression test reads the workflow text (source-text-is-the-product) and
asserts each file queries agent-skills for every agent type it spawns and
interpolates the result at least once per spawn.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#991): add changeset for code-review agent_skills injection fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 14:07:23 -04:00
github-actions[bot]
0976d849f2 chore: remove stray orchestration temp files (issue-996-regression.md, pr-1001-body.md) leaked via #1002 git add -A 2026-06-10 14:01:08 -04:00
Tom Boucher
adaf3e17d8 fix(#1001): make bug-969 hardening tests hermetic + move build tsbuildinfo out of shipped tree (regression from #996) (#1002)
* fix(#969): make bug-969 hardening tests hermetic and move build tsbuildinfo out of shipped tree (regression from #996)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#969): self-heal legacy bin-local tsbuildinfo and make sentinel test hermetic (adversarial-review follow-ups)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(changeset): set pr number to 1002

* docs(#1001): record DEFECT.SHARED-ARTIFACT-MUTATION-IN-CONCURRENT-TEST anti-pattern in CONTEXT.md

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 13:58:56 -04:00
Tom Boucher
88e30d5342 test(#969): fix stale-build flake (incremental + re-emit-on-missing) and make runGsdTools retry-once before surfacing subprocess kills (#996)
Closes #969

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 12:04:59 -04:00
Tom Boucher
1fd5c86a1e fix(#983): rewrite bare .claude paths in Trae/Windsurf converters (Codex/Cline parity) (#995)
* fix(#983): rewrite bare .claude paths in Trae/Windsurf converters (Codex/Cline parity)

Both convertClaudeToWindsurfMarkdown and convertClaudeToTraeMarkdown only
handled trailing-slash .claude/ forms; bare ~/.claude and $HOME/.claude
references (e.g. configDir = ~/.claude, RUNTIME_CONFIG_DIR=".../$HOME/.claude")
survived conversion and pointed users at the wrong config dir.

Fix: add bare-form replacements using negative lookahead (?![\w-]) to protect
.claude-plugin and .claudeignore, mirroring Cline (#782) and Codex (#570) precedent.
Also adds CLAUDE_CONFIG_DIR -> WINDSURF_CONFIG_DIR / TRAE_CONFIG_DIR rewrite.
_applyRuntimeRewrites windsurf case gets matching \b-anchored bare-form lines,
mirroring the existing trae case.

Closes #983

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#983): backfill changeset pr number (995)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 11:43:31 -04:00
radioflyer28
d617735bed fix(codex): avoid partial model effort pinning (#842)
Co-authored-by: Andrew Kriz <akriz@vt.edu>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-10 11:41:01 -04:00
Tom Boucher
36b68ac81d fix(#977): map ephemeral fnm multishell execPath to a stable fnm alias in normalizeNodePath (#992)
* fix(#977): map ephemeral fnm multishell execPath to a stable fnm alias in normalizeNodePath

Closes #977

* chore(#977): backfill changeset pr number (992)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 11:16:13 -04:00
Tom Boucher
972a41a528 fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:) (#990)
* fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:)

Closes #967

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#967): backfill changeset pr number (990)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 11:10:40 -04:00
Tom Boucher
921a7cd618 fix(#974): error on graphify --budget with missing/non-numeric value instead of silent NaN no-op (#986)
* fix(#974): error on graphify --budget with missing/non-numeric value instead of silent NaN no-op

When `--budget` was the last arg or followed by a non-numeric token,
parseInt(undefined/NaN-string, 10) produced NaN. NaN is falsy so both
the router check and applyBudget gate silently skipped budget trimming.
The query ran unbounded with no warning.

Fix: guard in graphify-command-router.cts — if args[budgetIdx+1] is
absent or parses to NaN, emit ERROR_REASON.USAGE and return early.
Defensive fix in graphify.cts: tighten `if (!budgetTokens)` →
`if (budgetTokens == null)` and `if (options.budget)` →
`if (options.budget != null)` so a real 0/NaN caller is handled
predictably by both independent guards.

Regression tests: 16 cases (unit/mock, subprocess, property-based)
covering boundary inputs: missing value, non-numeric, valid integers,
and fast-check properties over the budget parse contract.

Closes #974

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#974): backfill changeset pr number (986)

* fix(#974): convert static property (b) to real fc.assert property with constrained generator + deterministic seed

The test named "property: --budget as last arg always produces usage error"
was a static test with no fc.assert — it only checked a single hardcoded
term ("someterm") and could never flake or produce a fast-check path. This
is a generator/property bug (case b): the test was mislabeled as a property
test but lacked parameterization.

Root-cause of CI path "46:0:0" / seed 42 failure: a naive parameterized
version without the !startsWith('--') filter could feed term='--budget',
causing args.indexOf('--budget') to hit index 2 (the term slot) rather than
index 3 (the flag slot), placing the router in a different code path. The
property still holds — NaN detection fires on rawBudget='--budget' — but
the assertion text referenced the wrong invariant, making the failure appear
spurious. Fix: constrain the generator to non-flag terms (filter out
strings starting with '--') and pass explicit { seed: 42, numRuns: 200 } to
fc.assert so the test is fully deterministic in CI regardless of GSD_FC_SEED.

No change to src/graphify-command-router.cts (router is correct).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#974): constrain non-numeric budget property generator to genuinely-NaN values and pin fast-check seeds (CI determinism)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#974): use a strict valid-term generator and pin seeds so budget property tests are deterministic

Replace fc.string({minLength:1}).filter(!startsWith('--')) term generators in
properties (b) and (d) with a shared validTerm = fc.stringMatching(/^[A-Za-z0-9][A-Za-z0-9_.-]{0,29}$/)
that is alphanumeric-leading and contains no whitespace, flags, or sign-numerics.
This eliminates the class of CI failures where the old generator produced out-of-
contract inputs (" ", "+5", empty) that the router legitimately rejects for reasons
outside the budget-parse contract under test. Pin distinct seeds (1001–1004) on
every fc.assert for CI determinism. Stress-tested at numRuns=100000 per property
and across 10 seeds (1,2,7,13,42,43,44,99,12345,999999) at numRuns=2000 — all pass.
No src change: gsd-core/bin/lib/graphify-command-router.cjs is correct.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#974): replace flaky term-fuzzing properties with deterministic examples; keep value-fuzzing properties

The validTerm regex /^[A-Za-z0-9][A-Za-z0-9_.-]{0,29}$/ admitted single-digit
strings like "0" which are falsy; the router's `if (!term)` guard fires before
the budget-missing-value path, producing a spurious errFn call. Properties (b)
and (d), which test the BUDGET contract (not term handling), are replaced with
deterministic example loops over fixed valid terms. Properties (a) and (c),
which fuzz the BUDGET VALUE with a fixed term, are kept unchanged (seed+numRuns
pinned). The validTerm generator is fully removed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 10:58:42 -04:00
Tom Boucher
898d55788e fix(#976): detect command+args (wrapped) hook registrations in installer presence checks (#994)
* fix(#976): detect command+args (wrapped) hook registrations in installer presence checks

Add referencesHook() helper that inspects both h.command (standard form) and
h.args[] (args-form / wrapped-launcher form) when checking whether a managed
hook is already registered.  Rewrite all has*Hook predicates and the
alreadyHas* guards to use it so args-form registrations suppress the duplicate
stock string-command entry that was previously appended on every install/update.

Also add an explicit args-form skip to rewriteLegacyManagedNodeHookCommands so
entries with a non-empty args[] are left untouched (they are intentional user
wrappers, not legacy bare-node commands to migrate).

Extend isManagedHookCommand() in shell-command-projection.cts with an optional
args: unknown[] parameter that checks whether any arg's basename matches the
managed hook surface set — backward compatible; existing callers are unaffected.

Regression test added to tests/install-regressions.test.cjs:
- two-pass install with an args-form SessionStart entry pre-written to
  settings.local.json asserts exactly 1 hook entry remains after reinstall
  (previously 2 — the original args-form + a new stock string-command duplicate)
- rewriteLegacyManagedNodeHookCommands test asserts args-form entries unchanged

Closes #976

* chore(#976): backfill changeset pr number (994)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 10:44:24 -04:00
Tom Boucher
2981983bae fix(#973): add Edit to gsd-planner tools and forbid whole-file Write of ROADMAP.md (#989)
* fix(#973): add Edit to gsd-planner tools and forbid whole-file Write of ROADMAP.md

gsd-planner shipped Write but not Edit — the same writer-agent gap fixed for six
agents in #571/#581. Without Edit, an in-place ROADMAP update fell back to a
whole-file Write that truncated committed milestone history (292→16 lines in a
real incident).

Changes:
- agents/gsd-planner.md: add Edit to tools: frontmatter (adjacent to Write)
- agents/gsd-planner.md: update_roadmap step now directs Edit (scoped), with an
  explicit blocking prohibition on whole-file Write of ROADMAP.md or any existing
  curated .planning/ file
- agents/gsd-planner.md: Write contract section clarifies Write is authorized only
  for net-new PLAN.md creation; existing files must use Edit
- tests/agent-frontmatter.test.cjs: extend SECTION_WRITER_AGENTS list (#581 test)
  to cover gsd-planner — fails before fix, passes after
- .changeset/973-gsd-planner-edit-tool.md: Fixed changeset, pr:0

Closes #973

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#973): backfill changeset pr number (989)

* fix(#973): trim gsd-planner.md prose under agent size cap (keep Edit + scoped-Edit-for-ROADMAP rule)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 10:44:20 -04:00
Tom Boucher
22bb82e205 fix(#965): emit structured json error for unexpected handler throws under --json-errors (#987)
* fix(#965): emit structured json error for unexpected handler throws under --json-errors

When GSD_JSON_ERRORS=1 / --json-errors is active, an unexpected (non-ExitError) throw
in a handler now emits { ok: false, reason: "sdk_fail_fast", message } to stderr instead
of a raw stack trace. The plain-text behaviour (no json-error mode) is unchanged.

Closes #965

* chore(#965): backfill changeset pr number (987)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 10:40:12 -04:00
Tom Boucher
626575cbc5 fix(#978): parse --force in milestone complete dispatcher so the guard's documented override works (#982)
* fix(#978): parse --force in milestone complete dispatcher so the guard's documented override works

The dispatcher built `{ name, archivePhases }` but never parsed `--force`, so
`options.force` was always `undefined` and the guard inside `cmdMilestoneComplete`
(which tells users to "Re-run with --force to override") could never be bypassed.

Add `const force = args.includes('--force')` and pass it into the options object.
The guard already honors `options.force` — no changes to milestone.cts needed.

Closes #978

* chore(#978): backfill changeset pr number (982)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 10:40:06 -04:00
Tom Boucher
caca4d255c feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4) (#988)
* feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4)

Migrate the intel CLI command family from a hardcoded gsd-tools.cjs case arm to
a registry-dispatched Capability (commandFamilies mechanism, #961), mirroring
the graphify (#972) and audit (#984) cutovers. New src/intel-command-router.cts
exports routeIntelCommand reproducing all 9 subcommands verbatim (incl. the
status timeAgo non-raw post-processing), lazily requiring intel.cjs inside the
route fn. capabilities/intel/capability.json declares the intel command family;
commands-only (skills:[]), declares the existing intel.enabled gate (default
false — behavior unchanged).

Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing intel.test.cjs passes unchanged. Completes the first-party command-
family cutover sequence (graphify/audit/intel).

Closes #985

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#985): compute expected planningDir via path.join in intel cutover unit tests (Windows CI)

The intel-command-cutover unit-mock assertions hardcoded a POSIX `/.planning`
expectation while the router builds it with path.join(cwd, '.planning') →
backslashes on Windows, so the planningDir-arg assertions (query/status/diff/
snapshot/validate/update/api-surface) failed only on windows-latest CI. Compute
the expectation with path.join (cross-platform); production router unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 10:00:30 -04:00
Tom Boucher
9a03539c2d feat(#981): audit-uat + audit-open command cutover — commands-only capability (ADR-857 phase 4d-impl-3) (#984)
Migrate the audit-uat + audit-open CLI commands from hardcoded gsd-tools.cjs
case arms to a registry-dispatched Capability (commandFamilies mechanism, #961),
mirroring the graphify cutover (#972). New src/audit-command-router.cts exports
routeAuditUat/routeAuditOpen, each lazily requiring only its backing module
(uat.cjs/audit.cjs) inside the route fn — matching the old per-case lazy loads.
capabilities/audit/capability.json declares the two command families;
commands-only (skills:[]), no config gate, audit_review cluster untouched.

Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing audit regression tests pass unchanged.

Closes #981

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 08:34:49 -04:00
Tom Boucher
77bfd943dd feat(#972): graphify command cutover — first capability owning a command family (ADR-857 phase 4d-impl-2) (#975)
Migrate graphify into a Capability that owns its `graphify` command family,
dispatched via the registry (#961 mechanism) instead of a hardcoded case.
graphify is now an enable/disable plug-in.

- src/graphify-command-router.cts: routeGraphifyCommand (standard route*Command),
  reproduces the removed case EXACTLY (query +--budget, status, diff, build,
  hidden build snapshot, usage/unknown errors); injectable _graphify test seam.
- capabilities/graphify/capability.json: role feature, tier:full, skills:[graphify],
  config:{graphify.enabled default false}, commands:[{family:graphify, module,
  router:routeGraphifyCommand}].
- removed case 'graphify' from gsd-tools.cjs; graphify now flows
  default -> dispatchCapabilityCommand -> commandFamilies.graphify -> router.
- regenerated registry (commandFamilies/bySkill/configSchema/profileMembership/
  capabilityClusters for graphify); tier:full keeps 4c install/surface a no-op.

Equivalence-proven: 8 recording-mock unit tests assert the exact fn+args per
subcommand (budget, snapshot-vs-build); 9 subprocess tests assert distinguishing
output shapes; existing graphify tests pass unchanged through the new path.

Surfaced (not silently accepted): the pre-existing --budget-no-value NaN no-op
quirk, preserved for equivalence, filed separately.

Closes #972

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 07:12:57 -04:00
Colin Johnson
354e0e1b94 fix(ratchet): add --update drift repair + inherited-drift guidance to regression-name lint (#971) 2026-06-10 01:05:24 -04:00
Tom Boucher
0c567c17e4 feat(#961): capability command mechanism (commandFamilies index + default-case dispatch) — ADR-857 phase 4d-impl-1 (#964)
* feat(#961): capability command mechanism — commandFamilies index + default-case dispatch (ADR-857 phase 4d-impl-1)

Build the capability command mechanism per ADR-959: the `commands` declaration
field on the feature role, the registry commandFamilies index, and a real
dispatchCapabilityCommand consulted in runCommand's default case (replacing the
dead _dispatchNonFamily shim's role).

The registry DISCOVERS a standard route*Command (no rebuilt handler table). On a
default-case command, dispatch enforces a bare-.cjs-basename, resolves the module
under gsd-core/bin/lib/, asserts confinement, requires the resolved path, and
own-property-guards the router export before calling it. A require.main===module
guard makes gsd-tools.cjs importable for tests; the CLI path is unchanged.

Additive: commandFamilies is empty today, so the default case is behavior-
preserving for every command; the 10 dead _dispatchNonFamily sites are untouched
(future migration markers). The graphify cutover is the separate next step.

Closes #961

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#961): surface capability router failures structurally + enforce sync contract

Review follow-up: dispatchCapabilityCommand now wraps the router invocation so an
unexpected (non-ExitError) throw is converted to a structured, attributed
error(msg, SDK_FAIL_FAST) — honoring --json-errors — instead of escaping as a raw
stack trace; an intentional ExitError propagates unchanged. An async router
(returns a thenable) is rejected loudly with a structured error (the contract is
synchronous, like the 12 host routers). Not shipped as "consistent with existing
behavior": the host's pre-existing version of this gap is filed as #965.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 01:03:46 -04:00
Tom Boucher
e8cfb560b2 fix(#851): correct Codex quick adapter for generic multi_agent_v1 schema (#958)
* fix(#851): correct Codex adapter for generic multi_agent_v1 schema

The Codex skill adapter header in getCodexSkillAdapterHeader() documented
typed spawn_agent(agent_type=...) as a direct, unconditional mapping for all
Task()/Agent() calls. In sessions exposing only the generic multi_agent_v1
schema (message/items/fork_context — no agent_type field), this mapping is
silently invalid: the orchestrator cannot natively dispatch typed gsd-planner/
gsd-executor agents and may fall back to inline execution or produce errors.

Fix: Section C now requires schema detection before spawning. It documents the
typed mapping as conditional on the agent_type-capable schema (e.g. multi_agent_v2)
and introduces an explicitly-labeled generic-agent workaround for multi_agent_v1
sessions — read the agent TOML, inject its instructions as a role-preamble, and
call spawn_agent(message=...) — clearly marking the result as NOT equivalent to
typed gsd-planner/gsd-executor execution.

Regression test: tests/bug-851-codex-quick-adapter-agent-type-fallback.test.cjs
asserts schema-awareness language, the multi_agent_v1 fallback, the workaround
label, and backward compat with the existing bug-279 typed-spawn contract.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: add changeset for PR #958

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#851): keep Codex adapter block consistent with materialized skill surface

The `~/.codex/agents/<agent-name>.toml` literal introduced in the #851 prose
was being rewritten to the real install path by `_applyRuntimeRewrites` (the
`~/.codex/` → pathPrefix substitution) before the SKILL.md was written to
disk.  `getCodexSkillAdapterHeader()` still returned `~/.codex/agents/...` so
the test assertion (exact match between builder output and materialized file)
always failed.

Fix: replace the `~/.codex/agents/` literal with the runtime-neutral form
`agents/<agent-name>.toml` plus a parenthetical naming `$CODEX_HOME/` — which
is not matched by any rewrite pattern and survives the path-substitution step
unchanged.  The #851 schema-detection + generic-subagent-fallback intent is
fully preserved.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#851): resolve active Codex config root in fallback; strengthen tests (adversarial review)

- Rewrites the generic-agent workaround step 1 to explicitly describe
  active config root resolution (priority: $CODEX_HOME → --config-dir →
  --local .codex → default global dir) without the literal ~/.codex/
  substring that _applyRuntimeRewrites replaces, preventing bug-3582
  divergence.
- Replaces OR/loose-includes test assertions in bug-851 with AND-logic
  checks covering all four required elements: (a) schema-detection step,
  (b) active-config-root resolution for the TOML path including all three
  override mechanisms, (c) NOT-equivalent-to-typed-gsd-planner/gsd-executor
  label, and (d) fail-closed rule when typed dispatch is mandatory.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#851): register bug-851/947/948/950 in lint-regression-test-names allowlist

The ratchet (622e4be) bans NEW top-level bug-NNNN test files; the four
sibling PRs (#851, #947, #948, #950) landed AFTER the baseline was cut,
so their test files were not yet grandfathered. Add all four to the
identity allowlist so lint-regression-test-names passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 00:45:27 -04:00
Tom Boucher
1c86368785 fix(#947): restore gsd- prefix on Hermes skills for canonical dispatch (#955)
* test(#947): add regression tests and update stale Hermes assertions

- Add bug-947-hermes-gsd-prefix.test.cjs: 12 TDD tests covering fresh
  install canonical layout, bare-stem migration, manifest key format,
  and non-Hermes runtime isolation
- Update hermes-skills-migration.test.cjs: bare-stem → gsd-prefixed
  path and name assertions (#947 canonical layout)
- Update install-nested-layout.test.cjs: Hermes NEST matrix prefix ''
  → 'gsd-'
- Update install-regressions.test.cjs: Defect #1 now seeds bare-stem
  dirs (help/, quick/) and asserts gsd-help/ canonical output; use
  real GSD stems so readGsdCommandNames() migration finds them
- Update install-runtime-artifacts.test.cjs: Hermes nested layout and
  legacy migration assertions align with gsd- prefix
- Update install.test.cjs: Hermes install test uses gsd- prefixed paths
- Update runtime-artifact-layout.test.cjs: prefix '' → 'gsd-'

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#947): restore gsd- prefix on Hermes skills for canonical dispatch

Hermes skills were installing under bare-stem paths
(skills/gsd/<stem>/SKILL.md, name: <stem>) due to prefix: '' set in
ADR-3660 / #3664. This broke /gsd-<stem> dispatch and forced users to
invoke skills without the gsd- namespace prefix.

- src/runtime-artifact-layout.cts: change Hermes skillsKind prefix
  from '' to 'gsd-'; skills now land at skills/gsd/gsd-<stem>/SKILL.md
  with name: gsd-<stem>
- bin/install.js _runLegacyInstallMigrations: invert the #3664
  migration — remove stale bare-stem dirs (using readGsdCommandNames()
  to distinguish GSD-owned stems from user content), keep gsd-* dirs
  which are now canonical
- bin/install.js _runLegacyUninstallCleanup: also remove bare-stem
  dirs on uninstall for clean teardown
- bin/install.js uninstallRuntimeArtifacts: post-cleanup removes
  DESCRIPTION.md and empty skills/gsd/ category dir on Hermes
- bin/install.js: remove skillListPrefix Hermes exception (now uses
  shared 'gsd-' path)
- docs/adr/3660-runtime-artifact-layout-module.md: document #947
  reversal of the bare-stem sub-decision

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: add changeset for #947 fix (#955)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#947): remove ALL pre-migration bare-stem Hermes skills on reinstall (adversarial review)

Replace readGsdCommandNames()-based bare-stem cleanup (which missed skills
not in the commands source tree, e.g. dev-preferences) with
_removeHermesBareStemDirs(), called AFTER the install loop when the exact
set of installed gsd-<stem>/ dirs is authoritative. For every gsd-<stem>/
written this run, the corresponding bare skills/gsd/<stem>/ is removed.
User-owned bare dirs with no gsd-<stem> counterpart are preserved.

Add two adversarial-review regression tests that FAIL on old code:
- bare skills/gsd/dev-preferences/ removed when gsd-dev-preferences/ installed
- user-owned bare dir with no gsd-<stem> counterpart is preserved (no over-deletion)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 00:22:36 -04:00
Tom Boucher
a313a7e304 fix(#950): emit status: complete in quick-task SUMMARY frontmatter (#951)
* fix(#950): emit status: complete in quick-task SUMMARY frontmatter

Add `status: complete` to all four SUMMARY templates (summary.md,
summary-minimal.md, summary-standard.md, summary-complex.md), to the
executor agent's documented frontmatter field list, and to the quick.md
executor constraints block. The audit-open milestone-close scanner
(scanQuickTasks) reads this field to decide whether a quick task is done;
without it the scanner falls back to `[unknown]` and false-flags finished
tasks as open. Writer-side fix; the scanner is correct and unchanged.

Blast-radius: no other scanner reads `status:` from phase-plan SUMMARY
files. Phase disk_status is derived from file-count heuristics only.
Adding the field to the shared template is therefore safe and the value
`complete` is semantically accurate for a finished plan.

Regression test: tests/bug-950-quick-summary-status-complete.test.cjs
- RED: 4 template-contract tests fail before fix, behavioral tests pass
- GREEN: all 8 tests pass after fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: add changeset for fix/950-quick-summary-status-complete (#951)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#950): assert writer-path contract + scope template checks to YAML frontmatter (adversarial review)

- Add `// allow-test-rule: source-text-is-the-product` at file top (before block comment)
- Add `extractFrontmatter()` helper that handles both leading-frontmatter files
  (summary-minimal/standard/complex.md) and fenced-frontmatter files (summary.md,
  whose frontmatter is embedded inside a ```markdown fence) — assertions now
  target the actual YAML block, not the whole file
- Scope all four [TEMPLATE CONTRACT] tests through extractFrontmatter() so a stray
  `status: complete` in prose/examples cannot produce a false green; error messages
  now print the extracted block to aid diagnosis
- Add [WRITER-PATH] quick.md test: asserts the <constraints> block instructs the
  executor to write `status: complete` in SUMMARY frontmatter
- Add [WRITER-PATH] gsd-executor.md test: asserts the Frontmatter spec documents
  `status: complete` as a required field
- Sanity-checked: guards fail when `status: complete` is removed from a template
  or from quick.md, and pass once restored

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 00:22:33 -04:00
Tom Boucher
46967baae8 fix(#948): guard STATE.md no-op writes; preserve milestone_name/stopped_at (closes #944) (#952)
* test(#948): add regression tests for no-op write guard and record-session auto-create (#944)

Red before fix: 11/15 tests fail. Green after: 15/15.
Covers zero-match patch byte-identity, milestone_name preservation,
stopped_at frontmatter-wins, record-session auto-create fallback, and
adversarial fixtures (CRLF, empty body, non-canonical labels).

Also registers bug-948-state-noop-write-guard.test.cjs in the state
bucket of lint-test-file-count.allowlist.json.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#948): guard STATE.md no-op writes; preserve milestone_name/stopped_at (closes #944)

Shared root cause: `readModifyWriteStateMd` wrote STATE.md unconditionally
even when the transform produced no change, and `syncStateFrontmatter`
re-derived frontmatter from the possibly-stale body on every write.

Three coordinated fixes in src/state.cts:

1. readModifyWriteStateMd: add no-op guard — when transform result ===
   input content, skip the write entirely (no platformWriteSync, no
   last_updated bump, no frontmatter re-derive). Fixes #948 zero-match
   phantom write and the #944 phantom last_updated bump.

2. syncStateFrontmatter: extend existing-frontmatter preserve logic —
   fall back to existingFm['milestone_name'] / existingFm['milestone']
   when the derived value is the template placeholder 'milestone'
   (getMilestoneInfo returns this literal when it cannot match the
   version in ROADMAP.md); prefer existingFm['stopped_at'] /
   existingFm['paused_at'] over a body-derived value (the frontmatter
   value, written by the canonical record-session path, wins over stale
   historical body lines). Mirrors the fallback already in cmdStateJson.

3. cmdStateRecordSession: when --stopped-at / --resume-file are supplied
   but body labels are absent, DWIM auto-create a canonical ## Session
   section (mirroring how add-decision / add-blocker / record-metric
   auto-create their sections). Never return a silent recorded:false when
   the caller supplied values.

SDK check: no sdk/src/state.ts exists in this repo (the comment in
cmdStateSnapshot references a sibling concern in the TypeScript SDK
codebase, which is a separate repo not present here).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: add changeset for PR #952 (fix #948/#944)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#948): correct stopped_at preserve rule; adjust test for sync behaviour

The "always prefer frontmatter stopped_at" rule in syncStateFrontmatter
was too aggressive — it broke phase.complete which intentionally updates
stopped_at in the body and expects syncStateFrontmatter to pick it up.

The primary fix (no-op guard in readModifyWriteStateMd) already prevents
the stale-body-overwrites-frontmatter scenario from #948: the file is not
written when the transform produces no change, so syncStateFrontmatter
never runs on a zero-match patch. The body-derived value can only win when
an actual write occurs, which means the body was legitimately updated.

Reverted to the original #905 rule for stopped_at/paused_at: fall back to
existing frontmatter only when the derived value is absent (empty/null).

Also adjusted the sync-suite test to assert what state sync actually does
(milestone_name preservation) rather than a stopped_at-wins property that
state sync does not have by design.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#944): update existing session block in place (adversarial review)

HIGH finding: the DWIM auto-create in cmdStateRecordSession was appending
a second ## Session block unconditionally, even when one already existed
with non-canonical content (e.g. a markdown table). Both
buildStateFrontmatter and cmdStateSnapshot read only the FIRST ## Session
block via regex, so the newly-written Stopped at / Resume file values
landed in the second, invisible block — frontmatter stopped_at stayed
stale and state-snapshot returned nulls.

Fix: check for an existing ## Session heading. When one is present,
normalize that section in place by replacing its body with canonical
**Last session:** / **Stopped at:** / **Resume file:** bold-label lines.
Only append a brand-new section when NO ## Session heading exists.

LOW finding: the auto-create scaffold emits **Last session:** but
cmdStateSnapshot only matched **Last Date:**, so session.last_date was
null after auto-create despite a valid timestamp being written.

Fix: extend the lastDateMatch regex in cmdStateSnapshot to also accept
**Last session:** / Last session: (the form the scaffold writes).

Tests: 3 new tests added to bug-948-state-noop-write-guard.test.cjs that
confirmed failure against the previous HEAD and pass after this fix:
- exactly one ## Session block after record-session with non-canonical existing block
- state-snapshot sees correct stopped_at via first Session block (not a duplicate)
- state-snapshot session.last_date is non-null after auto-create on body-less file

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#944): improve in-place section replace to cleanly remove old body content

The previous regex `/(^## Session[ \t]*$)([\s\S]*?)(?=\n^## |\n*$)/im`
with a lazy match consumed nothing after the heading, so old non-canonical
body content (e.g. table rows) remained after the new canonical lines.

While functionally correct (parsers found the canonical lines first in the
FIRST ## Session block), it left stale content in the section. Replace with
a negative-lookahead per-line pattern that consumes all content from the
heading up to (but not including) the next ## heading, producing a clean
section with only the canonical bold-label lines.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 00:22:29 -04:00
Colin Johnson
baa169abe0 Merge pull request #963 from open-gsd/ci/audit-pipeline-fixes
ci(#962): pipeline audit fixes — scoped matrix, merged coverage gate, suite seeding, bug-* ratchet
2026-06-10 00:19:07 -04:00
Colin
533b518553 fix(security-scan): update scanner self-exemption allowlists for renamed suite files
The three shell scanners exempt their own adversarial test fixtures by exact
filename; the *.security.test.cjs renames broke those entries, so the PR diff
scan flagged the scanners' own test payloads. Verified locally with all three
scanners in --diff origin/next mode (0 findings) and the security suite
(207/207). The .sh files were missed in the original reference sweep because
the rename grep filtered to .cjs/.yml/.json/.md extensions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 00:15:02 -04:00
Colin
8bb6784009 chore(ratchet): regenerate bug-* allowlist after rebase onto next (244 -> 257)
13 bug-* files landed upstream between the audit baseline and this branch's
rebase; they predate the ratchet policy, so they are grandfathered.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:56:57 -04:00
Colin
9db7958a70 fix(review): single eslint home, threshold co-location, helper reuse, docs clarity
Review-pass fixes: lint:ci composes npm run lint (one eslint invocation
home); the scripts/ coverage floor moves to package.json
(test:coverage:scripts-floor) so both thresholds live together; the ratchet
test uses helpers.createTempDir; TESTING-SUITES.md clarifies what the
Windows scoped lane runs and why feat-*/enh-* files are exempt from the
bug-* ratchet.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:53:25 -04:00
Colin
a869df2acf ci(test.yml): fold coverage into ubuntu-24 lane, add scripts/ floor, single lint step
- Coverage gate (test:coverage:unit) now runs inside the ubuntu/24 full lane;
  the standalone coverage job duplicated that lane's entire unit run (~4 min
  of runner time per PR). required-tests gate updated accordingly.
- New second-tier floor: c8 check-coverage --lines 55 over scripts/**
  re-slices the same V8 data (measured 65.95%) — the CI/release/lint tooling
  was previously enforced at 0%.
- Coverage artifact now excludes coverage/tmp (>1 GB of raw V8 dumps).
- lint-tests runs npm run lint:ci — one orchestrated step, identical set
  locally and in CI; drops the no-op eslint --cache flag (CI never restored
  the cache directory).
- Delete unreferenced scripts/run-cross-platform-tests.cjs (+ its test);
  document the mutation UNMUTATED blind spot (~48% of lib lines) in
  stryker.config.mjs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:53:25 -04:00
Colin
622e4be6d8 test(ratchet): ban new top-level bug-NNNN test files via identity allowlist
244 one-off bug-* files (~38% of the suite) are grandfathered in
lint-regression-test-names.allowlist.json; new ones fail lint with
fold-into-module guidance, and deletions force allowlist pruning so the
baseline only shrinks. Wired into npm run lint:ci (new single entry point
for every CI lint). Policy documented in docs/TESTING-SUITES.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:50:41 -04:00
Colin
cd5db1f8db test(suites): seed security/slow/integration suites via measured retags
Renames (git mv) with all references updated (ci-test-scope RULES,
windows-parity allowlist, test-file-count allowlist, docs in 6 locales):

- 5 scanner tests -> *.security.test.cjs — the 'Run security tests' CI step
  ran zero files since the suite taxonomy landed; it is now honest.
- graphify-auto-update -> *.slow.test.cjs (36s, slowest file in the suite;
  e2e gsd-tools spawns) — runs on full-matrix lanes and push to next.
- installer-migration-install-integration -> *.integration.test.cjs
  (13s; an integration test by its own name).

Coverage gate measured after retags: 88.55% lines (gate 70%).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:50:41 -04:00
Colin
a647053dcf ci(scope): narrow #494 invariant — changed tests join windows lane, not full matrix
full_matrix fired on 15/15 sampled PRs because any tests/** change forced it,
costing ~25 runner-minutes each. A changed test file now always joins the
scoped windows lane (covering the #482 OS-specific failure class per-file)
and still runs on ubuntu 22/24 via targeted_tests; the residual macOS /
windows-node-22 cross-product is covered on every push to next.

Also narrows WINDOWS_HINTS from 6 substrings (102/633 files, a ~10-minute
scoped lane) to windows/win32/shell/path — the dropped hints (workflow,
install, hook) are either platform-independent lint tests or already covered
by fullMatrix rules.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:50:41 -04:00
Tom Boucher
19efcddfc7 fix(#941): track managed-hooks-registry.cjs in file manifest (#953)
* test(#941): regression test for managed-hooks-registry.cjs manifest omission

Adds bug-941-managed-hooks-registry-manifest.test.cjs which verifies:
- managed-hooks-registry.cjs appears in gsd-file-manifest.json after install
- manifest covers the full HOOKS_TO_COPY set (forward-proof)
- detect-custom-files reports 0 custom files after a clean install
- manifest hook keys use forward slashes (cross-platform)

All four assertions fail before the fix, confirming the bug is reproducible.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#941): track managed-hooks-registry.cjs in file manifest

The writeManifest() hooks loop in bin/install.js filtered hook filenames
with `file.startsWith('gsd-') && (file.endsWith('.js') || file.endsWith('.sh'))`.
managed-hooks-registry.cjs fails both predicates (wrong prefix, .cjs extension),
so it was never recorded in gsd-file-manifest.json even though it is shipped to
users as part of HOOKS_TO_COPY.

detect-custom-files scans the installed hooks/ dir and reports any file with no
manifest entry as a custom file, producing a perpetual false-positive
"Found 1 custom file(s)" warning on every /gsd-update for all users.

Fix: import HOOKS_TO_COPY from scripts/build-hooks.js and drive the manifest
hooks loop from that set (as a Set for O(1) lookup), so the manifest set is
structurally identical to the build set. Any future hook of any prefix or
extension added to HOOKS_TO_COPY is automatically covered. The new regression
test asserts full HOOKS_TO_COPY coverage and zero detect-custom-files
false-positives after a clean install.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#941): add changeset for managed-hooks-registry.cjs manifest fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#941): assert manifest hash matches installed hook contents (adversarial review)

Strengthen the regression test to not only verify that the manifest KEY
`hooks/managed-hooks-registry.cjs` is present after install, but also that
the stored hash equals the SHA256 of the actual installed file bytes — the
same algorithm used by the installer's fileHash() function.  A future
refactor that records the right key from the wrong path or content would
now fail this assertion immediately.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 23:28:02 -04:00
Tom Boucher
34435736e4 docs(#959): ADR-959 capability command contribution (ADR-857 phase 4d design) (#960)
Design how a Capability contributes a gsd-tools CLI command family and how the
hardcoded 73-case runCommand switch opens to registry-driven dispatch. Realizes
ADR-857 decision 7's reserved commands/module field (deferred by ADR-894).

Grilled to its leanest form: the registry DISCOVERS a standard route*Command
(no rebuilt handler table, no new arg convention); dispatch sits in the default
case (collision structurally impossible, no shadowing gate needed); graphify is
the first real cutover (lowest blast radius, has skill+cluster+config gate,
full-only so 4c stays no-op), proven equivalent and serving as the phase-6
template.

Design-only; CONTEXT.md gains a "Capability Command Family [Planned]" entry.

Closes #959

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 23:20:51 -04:00
Tom Boucher
dc139b38e0 feat(#949): install/surface consume derived profiles/clusters (ADR-857 phase 4c) (#954)
Make install + surface read the registry's derived profileMembership/
capabilityClusters so a capability's tier drives what installs + surfaces.
resolveProfile (when given the registry) unions capability skills for the
profiles its tier implies before the requires: closure; resolveSurface merges
capabilityClusters into the cluster map. bin/install.js, /gsd:surface, and the
capability-state resolver all thread the registry.

Shipped as a proven no-op: the UI capability is reconciled to tier:full (its
skills were full-only in the hand-authored profiles), so it contributes only to
the full profile (already the '*' sentinel) and core/standard are unchanged.
Equivalence tests prove resolveProfile/resolveSurface/listSurface/staging/
capability-state are identical with vs without the registry; the core-alias
staging path is verified equivalent (empty manifest → raw PROFILES.core).

Closes #949

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 22:35:42 -04:00
Tom Boucher
e006ff74fd docs(#956): MemPalace capability pre-proposal (PRD/ADR draft) (#957)
Combined PRD + ADR for wiring MemPalace (local-first AI memory) into the
GSD loop as an ADR-857 feature capability. Bidirectional sync, three
selectable memory-relationship modes (augment/kg_backend/replace),
loop-point recall+capture map, opt-in tier:full, MCP-primary/CLI-fallback.

Marked Pre-Proposal: the first-party-plugin proposal standard is not yet
established and ADR-857 phase-6 loop wiring is pending. First of a planned
series; PRD/ADR format is provisional pending PM-method evaluation.

Refs #956

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 22:30:48 -04:00
Tom Boucher
85cfa5dc13 feat(#945): unified capability-state resolver (ADR-857 phase 4b) (#946)
Add a read-side query composing the three toggle systems into one
per-capability view. resolveCapabilityState({registry, installedSkills,
surfacedSkills, config, cwd}) reports installed (skills ⊆ resolved install
profile), surfaced (skills ⊆ resolved surface), and per-hook active (no when →
active; non-empty-string when → resolved via _resolveActivationValue; empty/
non-string → inactive), with no forced composite verdict. cmdCapabilityState
does the I/O (resolveProfile + resolveSurface + loadConfig), resolves the
runtime config dir via the canonical getGlobalConfigDir (--config-dir override),
and surfaces resolution failures as warnings rather than a false installed='*'.
Routed as `gsd-tools capability state`.

Additive: install/surface/workflows untouched; consumed by nothing.

Closes #945

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 16:18:23 -04:00