Commit Graph

1584 Commits

Author SHA1 Message Date
Tom Boucher
2ae5fcdcf8 test(#1189): cover ADR-0006 planningPaths() consumption in init handlers (#1226)
Add a black-box regression guard (8 tests in tests/init.test.cjs) asserting the init handlers (execute-phase, plan-phase, phase-op, milestone-op) resolve workstream-scoped planning paths under GSD_WORKSTREAM via planningPaths()/planningDir(), never the flat .planning form.

Each positive case asserts both the scoped value and not-equal-to-flat (genuine guard, not coverage credit); fixtures seeded workstream-scoped; every CLI call pins GSD_WORKSTREAM and GSD_PROJECT for hermeticity. Verified by mutation (flat join fails exactly the 4 positive assertions) and re-verified under a polluted parent env. Test-only; src/ unmodified.

Closes #1189

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 14:20:47 -04:00
Tom Boucher
73b7f45140 feat(#1173): wire agent converters into descriptor-driven install path (#1227)
Extends `dispatchKindEntry` in `runtime-artifact-layout.cts` to route
agents-kind entries through a converter when the descriptor carries a
non-null `converter` field. Adds `stageAgentsForRuntimeWithConverter`
to `install-profiles.cts`, expands `VALID_CONVERTER_NAMES` with the 9
agent converter names, and adds a fail-first behavioral test suite
(9 tests) proving the new wiring end-to-end.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 14:20:21 -04:00
Tom Boucher
bf634b95c3 feat(#1213): Capability State Writer — write-side inverse of the resolver (#1225)
* feat(#1213): Capability State Writer — write-side inverse of the resolver

Adds src/capability-writer.cts (setCapabilityState + cmdCapabilitySet) and the
`gsd-tools capability set` subcommand: the write-side inverse of the capability
resolver (ADR-1213). One desired capability state projects onto the substrates —
`enabled` drives the runtime surface (canonical on/off), `gates` drive federated
config keys (hook granularity), install profile is a read-only floor — then
re-resolves and reports divergence (assert-and-report), so "off means off" holds
as a write-time invariant. Adds batched setConfigValues; routes gsd:settings
capability hook-gates through the writer. Docs: CLI-TOOLS reference, how-to,
ADR-1213, CONTEXT.md term.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1213): add changeset for Capability State Writer (#1225)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 12:51:17 -04:00
Tom Boucher
22f56f4431 ci(#1212): shard windows full-test lane to remove timeout cliff (#1222)
The `full test (windows-latest, *)` lane ran the entire unit suite (~740+
files) in one job whose wall-clock crept against the 20m cap and intermittently
CANCELLED (false-negative gate, observed on PR #1207). Prior tactical fixes
#869 (15→20m bump) and #1051 (handle-leak) deferred the cliff structurally.

Shard the unit suite across 3 parallel runners per OS/node leg so per-job
wall-clock is O(total/3) and stays under the cap as the suite grows.

- scripts/run-tests.cjs: add `--shard <i>/<n>` — a deterministic, balanced
  round-robin partition (fileIndex % n === i-1) over the SORTED selected file
  list. parseShardArg strictly validates i∈1..n, n≥1, integer-only; n=1 is a
  pure no-op. The 28K Windows argv chunking is preserved within each shard. A
  legitimately-empty shard (n > file count) exits 0; a selection empty BEFORE
  sharding still hits the discovery hard error. Composes with --suite and is
  order-independent (sorted before partition). Exports selectShard/parseShardArg.
- .github/workflows/test.yml: test-full becomes the 3 legs × 3 shards = 9-job
  cross-product (explicit include rows — a base shard dim does not cross-product
  with include legs, and a nested matrix.leg.os is unresolvable by the H1
  shell-policy linter). Unit suite runs sharded; integration/security run once
  per leg (shard 1). The Required tests fan-in is unchanged: it already needs
  test-full and checks the matrix-aggregate result, so a failed/cancelled shard
  fails the gate; the branch-protection check name is preserved.
- tests: partition/CLI + pure selectShard contract (completeness, disjointness,
  balance, determinism, boundaries, fast-check property) + parseShardArg
  validation, in run-tests-harness.test.cjs; a DEFECT.GENERATIVE-FIX parity
  guard (per-row shard values 1..N, every leg runs all shards, N == --shard /N
  denominator) + Required-tests name/needs pin, in ci-test-scope.test.cjs.

Closes #1212

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 12:25:29 -04:00
Tom Boucher
7edd18fd2b feat(#1165): async external_job_waiting half-state + resume/pause contract (#1221)
Core half of #1105: a legal external_job_waiting deferred state so an async-dispatched Execute step (committing a .planning/async-jobs/<job>.json manifest, deferring SUMMARY.md) is not an illegal partial. execute-phase safe-resume, resume-project, and pause-work reconcile against the versioned scheduler-agnostic manifest stability contract without re-dispatching; the producer is the capability half (#1164). Closes #1165.
2026-06-14 12:23:07 -04:00
Tom Boucher
82a0f561e1 fix(#1182): extract agent converters + tool-name tables into conversion module (#1220)
Extracts all 9 convertClaudeAgentTo* functions together with their full
dependency closure (claudeToCopilotTools, claudeToGeminiTools,
convertCopilotToolName, convertGeminiToolName) from bin/install.js into
src/runtime-artifact-conversion.cts and adds them to the module's export=
block.

Adds 6 regression tests in tests/copilot-install.test.cjs including a
DEFECT.GENERATIVE-FIX parity guard that asserts claudeToCopilotTools is
identical in module and bin/install.js.

Inline copies in bin/install.js are retained (#1175 will remove them).
This unblocks #1173 (descriptor-driven dispatch).

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 11:44:42 -04:00
Tom Boucher
8c6a0d65e6 fix(#1203): gate free-form ROADMAP deprecation warning on phase_id_convention (#1218)
* fix(#1203): gate free-form ROADMAP deprecation warning on phase_id_convention

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: backfill changeset PR number for #1218

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1203): use canonical gsd-tools roadmap upgrade command in warning

Match the migration command string to the canonical form used in
verify.cts and roadmap-command-router.cts (gsd-tools roadmap upgrade
--convention milestone-prefixed, dry-run by default) instead of the
non-canonical 'gsd roadmap upgrade --apply'.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 11:44:36 -04:00
Tom Boucher
2e8f4f6de1 fix(#1202): make verify key-links wave-aware for planned future files (#1219)
* fix(#1202): make verify key-links wave-aware for planned future files

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: backfill changeset PR number (#1219)

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 11:28:14 -04:00
Tom Boucher
1a186013a4 fix(#1205): roadmapper applies phase_id_convention to generated phase IDs (#1215)
* fix(#1205): roadmapper applies phase_id_convention to generated phase IDs

- Add Phase ID Convention section to <phase_identification> block:
  documents sequential (default) vs milestone-prefixed forms, and
  instructs the agent to read phase_id_convention from config.json
- Update <output_formats> to show both header and checklist forms for
  sequential and milestone-prefixed conventions with examples
  (e.g. ### Phase 1-01: Name, - [ ] **Phase 1-01: Name**)
- Add TDD regression test tests/bug-1205-roadmapper-convention.test.cjs
  (5 assertions, confirmed fail-first then pass after fix)
- Update tests/agent-size-baseline.json to reflect legitimate growth
- Add .changeset/brave-otters-leap.md (Fixed, pr:0 placeholder)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: backfill changeset pr: 1215 for fix/1205

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1205): move phase_id_convention regression into roadmapper-granularity.test.cjs

lint-regression-test-names rejects new standalone bug-NNNN-*.test.cjs files;
regression cases must live in the owning module's test file.

Move the 5 phase_id_convention assertions (#1205 regression) from the
removed tests/bug-1205-roadmapper-convention.test.cjs into
tests/roadmapper-granularity.test.cjs as a new describe block, alongside
the existing granularity calibration tests. Also update the allow-test-rule
comment to cover both #163 and #1205 surface contracts.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1205): fix lint-allow-test-rule-refs for roadmapper-granularity

- Add issue ref (see #1205) to allow-test-rule comment in
  tests/roadmapper-granularity.test.cjs so lint-allow-test-rule-refs
  passes (new exemptions require #NNN per ADR-456)
- Prune stale 'source-text-is-the-product' entry from
  scripts/lint-allow-test-rule-refs.allowlist.json (ratchet-down;
  comment now compliant and no longer needs grandfathering)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 10:46:22 -04:00
Tom Boucher
d444864bf8 fix(#1194): correct inverted statusline auto-compact buffer math (#1211)
* fix(#1194): correct inverted statusline auto-compact buffer math

The reserved-buffer percentage was computed as (acw/totalCtx)*100 — the
usable fraction — instead of (1 - acw/totalCtx)*100 — the reserved fraction.
When acw == totalCtx this produced buffer=100%, making the usable-range
denominator zero and pinning `used` at a constant 100% regardless of real
remaining context.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: backfill changeset PR number (#1211)

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 10:46:00 -04:00
Tom Boucher
e79e4e18b3 fix(#1186): guard hotfix version regex against leading zeros (ADR-218) (#1214)
Replace `^[0-9]+\.[0-9]+\.[1-9][0-9]*$` with
`^(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)\.[1-9][0-9]*$` in the hotfix
branch of the validate-version step, consistent with the two sibling
patterns already using the strict (0|[1-9][0-9]*) guard. Adds
regression assertions to tests/adr-218-release-version-validation.test.cjs
that prove `01.2.3` and `1.02.3` are rejected.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 10:45:42 -04:00
Tom Boucher
9e5d4b266b fix(#997): ensure canonical ~/.claude/gsd-core path for plugin installs (#1207)
* fix(#997): ensure canonical ~/.claude/gsd-core path for plugin installs via SessionStart hook

Claude Code marketplace plugin installs unpack the package into the
version-pinned plugin cache and never run bin/install.js, so
~/.claude/gsd-core/ is never created. Agents, commands, and templates
markdown-@-include the canonical ~/.claude/gsd-core/... path (which
expands ~ but NOT ${CLAUDE_PLUGIN_ROOT}), so every include resolved to
nothing and agents (e.g. the executor) failed.

Add a SessionStart hook (hooks/gsd-ensure-canonical-path.js) that, on a
plugin install, symlinks the canonical path's immutable subdirs (bin,
contexts, references, templates, workflows) to the plugin's bundled
gsd-core/ tree. It changes zero @-references, is a no-op in classic
installs, preserves user-generated files (USER-PROFILE.md, STATE.md),
prunes stale links so it self-heals after `claude plugin update`, uses
Windows junctions, and rejects bundled/canonical paths that escape the
resolved plugin root (no traversal, no clobber).

Registered in HOOKS_TO_COPY (build-hooks), MANAGED_HOOKS, hooks.json
SessionStart (runs first, timeout 5), and BUNDLED_GSD_HOOK_FILES.
Behavioral regression tests folded into issue-766-plugin-manifest.test.cjs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#997): backfill changeset PR number to #1207

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 09:37:41 -04:00
Tom Boucher
98866a0c69 feat(#1187): per-module Stryker mutation-score ratchet (ADR-456 80% floor) (#1200)
* feat(#1187): per-module mutation-score ratchet + graduate core-utils

ADR-456's 80% mutation floor was unenforceable as a single global break=50:
4 of 6 covered modules sit at 63-79% and forcing them to 80 would require
brittle exact-string assertions on equivalent string-literal mutants (a
Goodhart's-Law trap). Instead, each covered module declares a minScore floor
(locked at its measured score, TARGET 80) enforced per CI shard via
stryker --break, ratcheting up over time without brittle tests.
- mutation-matrix.cjs: minScore per module + TARGET_MUTATION_SCORE=80, emitted
  in the matrix; require.main guard + exports for testability.
- mutation.yml: per-shard --break <minScore>.
- stryker.config.mjs: global break 50->60 as a local backstop (CI uses minScore).
- Graduated core-utils (measured 77.5%, floor 75).
- context-utilization 79.5->92.3% via behavioral killers (state classification
  outputs + error-value contract, not exact-string matches) -> minScore 80 (TARGET).
- ratchet-integrity guard test (28 cases).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1187): pass mutation break via MUTATION_BREAK env (no stryker --break flag)

Adversarial review caught that Stryker 9.x has no --break CLI flag, so the
per-shard 'stryker run --break <minScore>' errored out every mutation shard.
Read the per-module floor from process.env.MUTATION_BREAK in stryker.config.mjs
and set it per shard via env in mutation.yml. Red-green verified: MUTATION_BREAK=99
exits 1, =80 exits 0. Also make the ratchet guard monotonic (RATCHET_BASELINE
floors; lowering a floor now fails the guard unless the baseline is edited).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1187): fail closed on bad MUTATION_BREAK + monotonic ratchet baseline

Code review: Number(env)||60 failed OPEN — an empty/invalid MUTATION_BREAK
(e.g. a future module missing minScore -> matrix expands to '') silently
degraded the shard to break 60, letting a high-floor module regress undetected.
resolveMutationBreak() now returns 60 only when the env is truly unset (local
backstop) and THROWS on present-but-empty/non-numeric/out-of-range (fail closed);
stryker.config.mjs imports it via createRequire. Also make RATCHET_BASELINE an
equality mirror (=== not >=) so any floor change is explicit in review and no
floor can be silently lowered. Tests: 46 (incl resolveMutationBreak cases).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1187): recalibrate config-schema/prompt-budget floors to CI scores

First CI mutation run failed two shards: the floors were set from local Stryker
runs whose TIMEOUTS were counted as kills (env-variable), inflating scores. CI
runs with timeout~0, so the real deterministic scores are lower:
- config-schema: local 69.7% -> CI 54.55% (5 local timeouts vanished) -> floor 52
- prompt-budget: local 99.6% -> CI 68.33% (239 local timeouts vanished) -> floor 66
Calibrate floors from CI (the documented source of truth) and record the lesson
in the comment so future floors aren't set from timeout-inflated local runs.
Baseline updated to match. The other 5 shards passed (deterministic CI scores
above their floors).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 08:25:41 -04:00
Tom Boucher
fdac556746 fix(#1160): resolve capability surface from installed skill layouts (#1206)
* fix(#1160): resolve capability surface from installed skill layouts

In a global skills-runtime install (e.g. Codex at ~/.codex), gsd-tools.cjs
runs from <configDir>/gsd-core/bin/ and the commands/gsd source tree is
absent — only <configDir>/skills/gsd-<stem>/SKILL.md files exist.
_resolveCommandsGsdDir() returned a path that does not exist there, so
loadSkillsManifest returned an empty Map. resolveSurface then materialised
the '*' (full) profile sentinel by enumerating that empty manifest → empty
surfaced Set → every capability reported surfaced=false/enabled=false
regardless of project config. As a result `loop render-hooks verify:post`
returned activeHooks:[] even with workflow.security_enforcement and
workflow.nyquist_validation enabled, silently disabling the security and
Nyquist gates.

Fix: add _loadInstalledSkillsManifest(configDir) that scans configDir/skills/
for gsd-<stem>/SKILL.md dirs and builds the same Map shape, and
_resolveManifest(commandsGsdDir, configDir) that prefers the source tree when
present (preserving repo-checkout behaviour) and falls back to the installed
skills layout otherwise. Both resolveCapabilityRuntimeState call sites use
_resolveManifest. Both helpers are exported for direct unit-testing.

Tests: capability-state.test.cjs gains a faithful installed-runtime e2e block
that copies gsd-core/bin + scripts + package.json into a temp install root
with no reachable commands/gsd, then runs the real gsd-tools.cjs against an
installed skills/ layout. It asserts capability state reports security &
nyquist enabled and verify:post includes security->secure-phase and
nyquist->validate-phase; a disabled-config negative confirms no
over-activation. This block FAILS before the fix (activeHooks:[]) and PASSES
after. Plus unit coverage for the two new helpers and the empty-surface
pre-fix scenario.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(changeset): backfill PR number

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 08:22:43 -04:00
Tom Boucher
e9f9ae49c8 fix(#1146): single base-branch resolver across forking workflows (#1198)
* fix(#1146): single base-branch resolver across forking workflows

Replaces duplicated per-workflow bash detection that silently fell through
to :-main on repos where origin/HEAD is unset (git init+remote add+fetch
without set-head, most CI checkouts, many worktrees).

New CJS module git-base-branch.cjs exposes `gsd_run query git.base-branch`
with full precedence ladder: git.base_branch config override → origin/HEAD
symref → git remote show origin (authoritative) → local branch presence →
"main". All git subprocesses bounded with timeouts; degrades gracefully.

Wires execute-phase, quick, ship, complete-milestone, and pr-branch to the
single resolver. Removes 14 lines of duplicated detection bash across the
five workflows.

Includes 7 behavioral tests covering the full precedence ladder including the
key regression case (master repo, origin/HEAD unset → must return "master",
NOT "main") and an anti-regression guard that fails if any workflow
re-introduces the :-main/:-master fallback pattern.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(changeset): backfill PR number #1198

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1146): drop stray PR-body file from branch

pr-1146-body.md was committed during changeset backfill but must not
be tracked in the repo. Content preserved externally for PR body use.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#1146): add tests for flat base_branch config key and both-branch tie-break

Closes two mutation gaps identified in adversarial review:
- A2: flat {base_branch: ...} at config root (legacy key form) was covered
  by code but unguarded against mutation of lines 74-75 in resolver
- H: tier-4 tie-break when both main+master exist locally (main wins,
  per tryLocalBranch JSDoc) was documented but untested

9/9 tests pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#1146): allowlist workflow-literal guard as runtime-contract exemption

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1146): degrade gracefully when gsd_run unavailable in handle_branching bash blocks

handle_branching (execute-phase.md) and step 2.5 (quick.md) are extracted
and run verbatim by behavioral tests that lack the gsd_run preamble.
Adding a || fallback ladder (git symbolic-ref then echo main) keeps the
unified resolver as primary in real workflows while letting the test harness
succeed without gsd_run defined.

Also propagates updated runtime-launcher preamble to pr-branch.md (added in
origin/next MemPalace PR) and regenerates workflow-size-baseline.json.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 08:22:40 -04:00
Tom Boucher
a375c4b354 feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) (#1201)
* feat(#956): add MemPalace memory capability (ADR-857 feature plug-in)

Adds an opt-in, default-resilient ADR-857 feature capability that wires
MemPalace (local-first memory: MCP server + CLI) into the GSD loop:
deliberate recall before discuss/plan and verbatim + temporal-KG capture
at phase boundaries. Three memory modes (augment default; kg_backend and
replace forward-declared). Master gate mempalace.enabled (default off);
every hook onError:skip, zero gates; absent/disabled MemPalace => loop
unchanged. Transport is rendered-markdown only — MemPalace runs
out-of-process, no third-party code in gsd-core (ADR-857 §7).

Capability: capabilities/mempalace/ (manifest + 2 fragments), skills
commands/gsd/mempalace-{recall,capture}.md, agent
agents/gsd-mempalace-curator.md. Registration: ns-context router,
utility cluster, KNOWN_SKILLS, help full.md, model-catalog, copilot
install list, size baselines; regenerated capability-registry +
inventory manifest. ship:post wired into ship.md (wire-on-demand).

HELD on #1196: this capability also declares hooks at discuss:pre and
discuss:post, which are structurally un-wireable until the host-loop
conformance model covers the discuss phase (discuss-phase.md is not in
HOST_LOOP_FILES). The phase6-capstone-conformance gate therefore fails
on exactly those two orphaned points by design — see #1196. Once #1196
lands, rebase onto next and the gate goes green with no further change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#956): backfill changeset PR number (#1201)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 02:09:10 -04:00
Tom Boucher
b1e8a74708 fix(#1196): wire discuss loop step for capability hooks (#1199)
* fix(#1196): wire discuss loop step for capability hooks

discuss was contract-declared (gsd:loop-host marker, in POINT_ORDER and
LOOP_HOST_CONTRACT) but structurally unwireable: discuss-phase.md had no
`loop render-hooks` dispatch and was absent from the conformance gate's
HOST_LOOP_FILES, so capabilities could never wire discuss:pre/discuss:post.

- discuss-phase.md: add minimal discuss:pre (before analyze_phase) and
  discuss:post (after write_context) render-hooks dispatch steps that
  delegate consumption to a new shared reference (kept under the 32KB
  #2551 budget; no inline subagent dispatch token).
- references/loop-hook-dispatch.md: new canonical, point-agnostic contract
  for consuming `loop render-hooks --raw` activeHooks (contribution/step/
  gate) — single source for hook consumption across host loops.
- gen-loop-host-contract.cjs: derive HOST_LOOP_FILES from STEP_WORKFLOWS and
  export scanWiredPoints()/getWiredLoopPoints() (throws on a missing host
  file) — one source of truth for the host-loop file + wired-point set.
- phase6-capstone-conformance.test.cjs: consume the derived HOST_LOOP_FILES
  and shared scanWiredPoints (was a hand-maintained duplicate omitting
  discuss-phase.md + a duplicated regex).
- gen-capability-registry.cjs: add validateHooksWired() gen-time guard that
  rejects a capability hook declared at a valid-but-unwired loop point, with
  a clear remediation message — failure now surfaces at gen --check/--write
  time instead of deep in the full conformance suite.
- tests (capability-registry.test.cjs): regression + anti-pattern parity
  guards (every loop-host marker is in STEP_WORKFLOWS/HOST_LOOP_FILES;
  POINT_ORDER === flattened LOOP_HOST_CONTRACT) so no step can drift into
  the discuss-class gap again.
- docs/INVENTORY*: register the new reference.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1196): backfill changeset PR number (#1199)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 01:39:54 -04:00
Tom Boucher
be132445aa fix(#1159): phase complete ignores historical/deferred requirement metadata (#1197)
* fix(#1159): phase complete ignores historical/deferred requirement metadata

Defect A: `phase complete` emitted a false "has unresolved gaps" warning
when a VERIFICATION.md file's frontmatter contained `status: passed` but
the body contained `previous_status: gaps_found`. The full-text regex
`/status: gaps_found/` matched the substring inside `previous_status:
gaps_found`, producing a spurious warning. Fixed by reading only the
frontmatter `status` key via `extractFrontmatter` in both `phase.cts`
(phase complete warning check) and `commands.cts` (determinePhaseStatus).

Defect B: requirement IDs under explicitly deferred/backlog/future/v2
section headings in REQUIREMENTS.md were flagged as missing from the
Traceability table. Fixed by splitting the body into markdown sections,
detecting deferred-intent headings via a keyword regex, and skipping
those sections when collecting IDs to check against the table.

Both fixes include boundary tests (false positive suppressed; genuine
gap/missing warnings still fire) in 4-phase-complete-cjs-regression.test.cjs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1159): address adversarial-review findings (subheading + case sensitivity)

- Defect B: Replace section-split approach with line-by-line depth-tracking
  to correctly propagate deferred status to sub-headings and ignore headings
  inside fenced code blocks. Previously `## Future Backlog` / `### Sub` would
  leak sub-heading IDs (split created a new non-deferred section per heading).
- Defect A/commands: restore case-insensitive status matching via toLowerCase()
  to match the prior /status:\s*passed/i regex semantics.
- Add test #1159-B-4 covering the nested-subheading case.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(changeset): backfill PR number for #1159 fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 00:00:37 -04:00
Tom Boucher
5fa4dcd78c fix: recover silently-excluded test dirs + test-architecture audit hardening (#1195)
* fix: recurse test discovery so subdir test suites actually run

scripts/run-tests.cjs discovered tests with a flat readdirSync(testDir),
silently excluding tests/observability/ (4 files), tests/dispatch/ (1) and
tests/installer-migrations/ (1) — 94 passing tests — from `npm test` and all
CI lanes. Walk the tree recursively (relative subpaths preserved), classify
suites by basename, and add a fail-on-zero-executed guard for suite/default
runs (escape hatch GSD_ALLOW_EMPTY_SUITE=1) while preserving the empty
--files/--files-from path the CI inert lane relies on.

Unit suite 735 -> 741 files; surfaces ADR-227's observability/dispatch seam.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: retire 5 verified-worthless tests

Adversarial verification confirmed these 5 prove nothing — their coverage is
provided more strictly elsewhere:
- enh-2790 'has a name: field' spot-checks (command-contract enforces /^gsd[:-]/)
- command-routing-hub duplicate construct + duplicate ERROR_KINDS assertions
- no-cjs-sdk-handsync-tooling (guarded files that never existed on main; bug-190
  covers the real retired SDK artifacts)
- runtime-artifact-layout cline edge case (subsumed by the explicit-global test
  and bug-782-cline-skills-emission)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: add ADR-218 release version-validation coverage

ADR-218 (reject leading-zero versions like 1.01.0; npm duplicate pre-check) had
zero tests — the logic lived only in release.yml bash. Add a test that extracts
the actual rejection regexes from the workflow and exercises them against a
boundary table (leading-zero/malformed rejected, valid accepted) plus structural
wiring assertions. Goes red if the regex is reverted to [0-9]+.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: redesign weak tests into behavioral, deterministic assertions

Per the ADR test audit, rewrite 27 weak test files (test-only, no source
changes) so each can go red for the defect it guards:
- kill pass-always assert.ok(true) placeholders (research-cli, worktree-baseref,
  bug-260 security guard, eslint-rules x24, clusters '|| true')
- replace source-text grep with behavioral calls (install Kilo, sh-hook-paths,
  plan-review-convergence) and add a repo-layout governance test
- de-flake real-clock/Math.random coupling (phase last_updated, bug-3707 mtime,
  context-utilization property, feat-3594)
- fix independence/shared-state violations (bug-492 singleton, issue-844 tmpRoot,
  core reapStaleTempFiles, active-workstream TTY, feat-488 GSD_HOME)
- strengthen property/shape-only tests (research-provider/store classification +
  collision) and unconditional plugin.json schema validation (issue-766)

Verified: all 28 files run together 1220 pass / 0 fail / 1 skip.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: add no-tautological-assert lint rule, error in test suite

New custom ESLint rule (eslint-rules/no-tautological-assert.cjs) bans asserts
that can never fail: assert(true)/assert.ok(<always-truthy literal>),
'cond || true' inside an assert, and equality asserts comparing two identical
literals. Wired as error on tests/**; full sweep confirmed zero existing
violations so the suite stays green. Prevents the placeholder-assert regressions
the audit redesigns just removed. RuleTester coverage added (6 valid, 8 invalid).

Note: no-only-tests was already enforced via eslint-plugin-no-only-tests, so no
duplicate rule was added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: gate new allow-test-rule exemptions to require an issue ref

ADR-456 requires any allow-test-rule exemption added after the ADR to carry a
tracking issue number, but nothing enforced it. New ratchet gate
(scripts/lint-allow-test-rule-refs.cjs, wired into lint:ci) fails when a NEW
allow-test-rule comment lacks a #NNN/URL reference; the 323 existing untracked
exemptions are grandfathered in an allowlist that ratchets down as they gain
refs. Red-green verified (novel untracked offender fails; compliant passes).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: add ADR test-audit evidence report (#1192)

Full risk-first qa-test-architect audit of the ADR portfolio (37 ADRs + 4
platform lenses, adversarial verification of retire verdicts) that drove the
P0 discovery fix, ADR-218 coverage, 5 retires, 27 redesigns, and the two new
lint gates. Filed as point-in-time evidence under docs/issueevidence/, named
for tracking issue #1192.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: replace pre-existing raw NUL byte with escape in feat-3594 fixture

feat-3594's null-byte parser fixture contained a literal NUL byte (pre-existing
on next at b10e5681 — confirmed: base blob has 1 NUL, this fix has 0), which
made git treat the file as binary and would break grep/editors. Switch to the
\x00 escape; the runtime string value (a real NUL in the parser input) is
unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: address adversarial-review findings

Codex adversarial pass over the branch:
- capability-registry drift test no longer mutates the committed generated
  capability-registry.cjs in place (concurrency hazard) — uses in-memory
  checkPipeline comparison instead.
- allow-test-rule ratchet now detects exemptions in ALL comment forms (block
  /* */ too, matching no-source-grep) so a block comment can't bypass it;
  one newly-surfaced pre-existing offender grandfathered (323->324).
- install.test Kilo case asserts on what install(false,'kilo') actually writes
  rather than manually calling configureKiloPermissions (masked the call site).
- issue-766 drops the undeclared transitive ajv dep for explicit structural
  assertions from the schema fixture.
- adr-218 test notes the hotfix leading-zero gap is tracked in #1186.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: address code-review findings (subdir discovery, rule + test gaps)

xhigh code review surfaced 15 confirmed issues, all fixed:
- run-tests.cjs --files now resolves subdir tests by bare basename + handles
  Windows backslash paths (ambiguous basenames error clearly).
- affected-tests-lib.cjs listTestFiles made recursive — the targeted CI lane was
  silently dropping changed subdir tests (same false-green class the audit fixed).
- no-tautological-assert now catches 'true || cond' and empty []/{}  equality.
- verify-test-quality: restore provenance-classification coverage, tighten the
  writeFile circular-detection check, guard the module-level file read.
- sh-hook-paths: cover the global-install .sh delegation branch (#2045 guard).
- active-workstream null-guard runs deterministically (no longer skipped on TTY).
- adr-218 structural guards tightened (major/minor leading-zero; needs: membership).
- repo-layout AGENTS.md guard no longer false-alarms on equivalent refactors.
- cross-ai ordering guard fails red when the step is missing.
- issue-766 parses required fields from the schema fixture (auto-enforced).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: stub USERPROFILE alongside HOME in feat-488 (Windows parity)

The feat-488 redesign stubbed process.env.HOME but not USERPROFILE; os.homedir()
resolves from USERPROFILE on Windows, so the home stub was not hermetic there —
caught by windows-test-parity-guard (stubsHomeNoUserProfile). Save/set/restore
USERPROFILE symmetrically with HOME (delete-if-originally-undefined).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: reconcile allow-test-rule allowlist after rebase onto next

Rebasing onto current next pulled in merged PR #1170, which added
inventory-headings-countfree.test.cjs (a baseline allow-test-rule exemption) and
deleted inventory-counts.test.cjs. Grandfather the former and prune the latter so
the ratchet matches the merged tree. No new debt from this PR.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 23:35:08 -04:00
Tom Boucher
3ebd13c45d fix(#1145): implement query user-story.validate handler (#1193)
* fix(#1145): implement query user-story.validate handler

`query user-story.validate` was a phantom command invoked by mvp-phase.md
(line 102) and verify-work.md (line 170) but had no CJS handler. Every
call exited with "Unknown command: user-story". The dotted-form dispatcher
strips the `query` prefix, splits on `.`, yielding command='user-story'
which fell through to the `default:` case with no registered capability
handler.

Adds a `case 'user-story':` handler inline in gsd-tools.cjs (same pattern
as the #1140 fix in PR #1148). The handler:

- Validates "As a [role], I want to [capability], so that [outcome]."
- Uses \S anchors to require non-whitespace content in each slot
  (whitespace-only slots like "As a  , I want to ..." now correctly
  return valid:false — found by adversarial Codex review)
- Returns { valid: boolean, errors: string[], slots: {role, capability,
  outcome} | null }
- Supports --pick valid for bare boolean output (verify-work.md usage)
- Added 'user-story' to SKIP_ROOT_RESOLUTION (pure string validation)
- Added 'user-story' to TOP_LEVEL_USAGE command list

Regression tests added to tests/commands.test.cjs (per lint-regression-
test-names policy; new bug-NNNN standalone files are banned).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(changeset): backfill PR number for #1145 fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-13 23:34:44 -04:00
Tom Boucher
3827054954 fix(#1156): support table-format STATE.md and insert missing roadmap plan rows (#1172)
* fix(#1162): support table-format STATE.md in state field read/replace

Extend stateExtractField and stateReplaceField in state-document.cts to
detect and operate on pipe-table rows (| Field | value |).  The separator
row | --- | --- | is excluded from matching.  updateCurrentPositionFields
in state.cts now falls through to stateReplaceField for table-format
Current Position sections when the inline Status:/Last activity: patterns
do not match.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1163): insert missing plan checklist rows in roadmap update-plan-progress

cmdRoadmapUpdatePlanProgress now inserts `- [ ] NN-XX-PLAN.md` checkbox
rows under the phase Plans: line when no per-plan checkbox rows exist yet
(fresh template).  Rows are sorted ascending and any already-summarised
plans are immediately marked [x].  The planCountPattern is extended to
also match plain `Plans:` (in addition to bold `**Plans:**`) so plan
counts are updated in both template variants.  The existing-rows check
covers both top-level and indented checkbox forms to preserve idempotency.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1163): fill partial plan-row gaps and scope insertion to active milestone

- Finding 1 (HIGH): replace all-or-nothing rowsAlreadyPresent guard with per-file
  set-difference so missing rows are inserted even when SOME plan rows already exist
- Finding 2 (MEDIUM): extend planCountPattern to recognise **Plans**: (canonical
  template form — bold word + outer colon) alongside **Plans:** and plain Plans:;
  use two-pattern fallback for row insertion to anchor under Plans: checklist header
  rather than the **Plans**: summary line
- Finding 3 (MEDIUM): scope row insertion to the active (post-</details>) milestone
  region so duplicate phase headings in archived sections never receive new rows
- Finding 4 (LOW): rename misleading "pipe-like content" test to honestly describe
  what it tests (multi-row table isolation); add out-of-scope escaped-pipe comment
- Remove now-unused anyCheckboxMatched variable (lint clean)
- Add 5 adversarial regression tests (pre-fix failures confirmed)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1163): scope missing-plan detection to active milestone; preserve authored state fields in table format

- roadmap.cts: compute activeRegion (post-</details> slice) once and use it
  for BOTH missingPlans detection and row insertion, so archived <details>
  rows no longer suppress active-section inserts (Finding 1 code-review)
- roadmap.cts: change (Plans:) inner capture to non-capturing (?:Plans:)
  in insertRowsPatternA to prevent group-numbering shift (Finding 3)
- state.cts: mirror inline-branch preserve-authored guards onto table-format
  branches in updateCurrentPositionFields — Status table branch checks
  isInList/matchesPattern before replacing; Last Activity table branch
  checks isDateShape/inList, preserving executor-authored narrative prose
  (Finding 2 code-review)
- tests: add three failing-first regression tests confirming each finding

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1162,#1163): fold table-format + plan-row regressions into owning module tests

Move bug-1162 cases into tests/state.test.cjs and bug-1163 cases into
tests/roadmap.test.cjs under named regression describe blocks; delete the
standalone bug-NNNN files and prune their allowlist entries.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1156): add changeset fragment for table-format state + roadmap insert fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 21:56:18 -04:00
Tom Boucher
4db185da74 fix(#1161): make phase complete idempotent for roadmap completion dates (#1177)
* fix(#1161): preserve existing roadmap completion date on repeat phase complete

Guard the Completed cell write in cmdRoadmapUpdatePlanProgress so that a
non-empty, non-placeholder date is never overwritten on repeat invocations.
Also routes the date source through realClock.today() so GSD_NOW_MS pins
the written date deterministically in tests (clock seam parity with the
rest of the codebase). Regression tests added to
4-phase-complete-cjs-regression.test.cjs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1161): preserve completion date in cmdPhaseComplete (real handler) + drive regression via phase complete CLI

- Import realClock in phase.cts; switch cmdPhaseComplete from bare `new Date()` to `realClock.today()` so the clock seam is honoured in tests
- Apply preserve-or-stamp guard to both 5-col (cells[4]) and 4-col (cells[3]) paths in cmdPhaseComplete: keep existing non-empty, non-dash date; only stamp today on first completion
- Rewrite the #1161 describe block in 4-phase-complete-cjs-regression.test.cjs to drive the real handler via runGsdTools('phase complete 1') end-to-end; cases (a)/(b)/(c)/(d) all FAIL against unfixed build and pass post-fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1161): preserve only date-shaped completion cells (self-heal garbage); correct test helper for 5-col

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1161): add changeset fragment for completion-date idempotence

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 21:46:10 -04:00
Tom Boucher
ae8bb707bc refactor(#1170): remove hand-maintained INVENTORY count scalars (#1179)
* refactor(#1170): remove hand-maintained INVENTORY count scalars

The `(N shipped)` heading counts in docs/INVENTORY.md were absolute
scalars that collided silently on merge: two branches each bumping the
same integer to N+1 produced a clean git merge whose value the merged
filesystem (N+2) contradicted, hard-failing inventory-counts.test.cjs on
the CI merge commit across all platforms (DEFECT.INVENTORY-MERGE-UNDERCOUNT).

- Strip the six `(N shipped)` heading counts + the two prose footnote
  counts; repoint the intro to INVENTORY-MANIFEST.json as the registry.
- Drop the decorative `generated` date from the manifest + its
  strip-before-compare branch in gen-inventory-manifest.cjs (it conflicted
  on cross-day merges and is read by nothing).
- Delete inventory-counts.test.cjs (scalar-vs-disk gate, the collision
  source); its drift protection is subsumed by the merge-safe set-membership
  test inventory-manifest-sync.test.cjs, which stays as the sole gate.
- Add inventory-headings-countfree.test.cjs guard (fails if a count is
  re-added to a heading).
- Fix already-broken count-bearing cross-doc anchors to stable count-free
  slugs in ARCHITECTURE.md + multi-agent-orchestration.md.
- Retire the now-impossible DEFECT.INVENTORY-MERGE-UNDERCOUNT + obsolete
  RULESET.DOC-CONSISTENCY in CONTEXT.md; de-count DEFECT.INVENTORY-DRIFT;
  correct stale MANIFEST-CANONICAL-KEY (all six families canonical).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1170): backfill changeset PR number (#1179)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 21:30:03 -04:00
Tom Boucher
b10e56818b feat(#1169): complete ADR-857 phase 6 — migrate features to Capabilities, revive dead gates, harden conformance gate (#1183)
* test(#1168): make phase-6 gate un-gameable — reject empty stubs + require loop shrink

The migration assertion previously checked only role==feature, so a registration-only stub (empty hooks, logic left inline) would turn the gate green while phase 6 stayed incomplete — the exact false-completion pattern this gate exists to prevent. Strengthen it: each ADR-named feature must OWN its behavior (>=1 hook, or a command family); and plan-phase.md/execute-phase.md must shrink strictly below their frozen pre-phase-6 sizes (94519/93166 LF bytes), which also defeats double-run gaming (declare a hook but keep the inline block -> file does not shrink -> red).

Gate now 5 pass / 4 fail (orphaned execute:wave:post, empty/unregistered features, config-key leaks, no shrink). Green is now reachable only by REAL migration. Refs #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate gap-analysis to a Capability (plan:post gate)

First real ADR-857 phase-6 migration (pattern-defining tracer). gap-analysis moves from an inline post_planning_gaps branch in plan-phase.md to a real plan:post gate Capability:

- capabilities/gap-analysis/capability.json: role:feature, plan:post gate (when=workflow.post_planning_gaps, blocking:false advisory), OWNS workflow.post_planning_gaps (federated out of central schema). - plan-phase.md: inline config-get + gsd_run gap-analysis block replaced with a plan:post render-hooks call site dispatching the gate; file shrinks 94519->93279. - src/check-command-router.cts: cmdGapAnalysisPlanPost runs the real gap analysis via gap-checker. - post_planning_gaps removed from central manifest; resolves via federated config (default true preserved). - tests/post-planning-gaps-2493: re-pointed to assert capability ownership.

Verified: gate 5 pass / 4 fail (gap-analysis cleared from migration, plan:post-orphan, config-leak, and plan-phase shrink checks); loadConfig still returns post_planning_gaps=true; check command runs real analysis; 392/392 in the config/registry/federation/router net. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate profile-pipeline to a command-family Capability

ADR-857 Decision 7: profile-pipeline becomes a command-family Capability (like audit/intel/graphify). capabilities/profile-pipeline/capability.json declares an 8-command family (scan-sessions, extract-messages, profile-sample, write-profile, profile-questionnaire, generate-dev-preferences, generate-claude-profile, generate-claude-md) backed by a new gsd-core/bin/lib/profile-pipeline-command-router.cjs; the inline case arms are removed from gsd-tools.cjs. Owns profile-pipeline.enabled (federated).

Verified: registry shows role:feature with commands.length=8; scan-sessions/profile-sample run live via the family; gate cleared profile-pipeline from the empty-stub failure (only tdd/schema-gate/drift remain); 296/296 registry+inventory+gsd-tools tests; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1167): wire execute:wave:post + implement ui.safety-gate check

Revives the second dead gate from #1167: ui.gates@execute:wave:post was declared but never dispatched AND its check.query (ui.safety-gate) was unimplemented. Adds the per-wave execute:wave:post render-hooks call site in execute-phase.md (fires after each wave's merge/cleanup, before the next forks) and implements cmdUiSafetyGate (frontend + UI-SPEC aware, mirrors cmdUiPlanGate) in check-command-router. +17 regression tests.

Verified: phase-6 orphaned-points conformance test now PASSES (gate 6 pass / 3 fail); ui-safety-gate routable in dot+hyphen forms; check-ui-safety-gate 17/17, check-ui-plan-gate 18/18; lint 0 errors. Refs #1167, #1168.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate drift (schema + codebase) to execute:wave:post gates

Removes the inline schema_drift_gate + codebase_drift_gate steps (77 lines) from execute-phase.md; drift becomes a Capability with two execute:wave:post gates (verify.schema-drift blocking, verify.codebase-drift advisory) dispatched via the per-wave render-hooks call site. check-command-router routes verify.schema-drift / verify.codebase-drift to the real detectors. Federates workflow.drift_threshold / drift_action / schema_drift_gate out of central.

Also fixes the execute:wave:post dispatch prose to run NON-blocking (advisory) gates too — the prior version only ran blocking gates, which would have silently dropped the codebase-drift advisory after its inline step was removed. Behavior preserved.

Verified: gate 7 pass / 2 fail (drift cleared from stub + config-leak; execute-phase.md 92297 < 93166 frozen -> shrink passes); both drift checks run real detection; loadConfig defaults preserved (threshold=3, action=warn, gate=true); drift-detection 56/56 + schema-drift 34/34; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate tdd to a Capability (plan:pre contribution + execute:post gate)

tdd becomes a real Capability: a plan:pre contribution injects the <tdd_mode_active> planner guidance (rendered from PLAN_PRE_HOOKS_JSON like security's contribution), and an execute:post gate (tdd.review-checkpoint, advisory) runs the real end-of-phase RED/GREEN review via a new check-command handler. Inline tdd_mode reads + the inline planner block + the tdd_review_checkpoint step are removed; workflow.tdd_mode is federated out of central. The MVP+TDD per-task RED-commit gate is preserved — TDD_MODE is now derived from the execute:post hooks (capId==tdd active), not an inline config-get.

BEHAVIOR CHANGE (documented, not silent): the --tdd CLI flag now persists workflow.tdd_mode=true via config-set instead of being per-invocation. Rationale: tdd is now a config-toggled Capability, and env vars do not persist across the workflow's separate bash blocks (config does), so an ephemeral override isn't cleanly achievable; --tdd therefore enables the tdd capability, consistent with how all capabilities are toggled.

Verified: gate 7 pass / 2 fail (tdd cleared from stub + config-leak; plan-phase + execute-phase both < frozen sizes); contribution injection + execute:post gate dispatch wired; MVP+TDD gate preserved; tdd.review-checkpoint runs real review; full unit suite 556/0; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate schema-gate to a plan:pre contribution Capability

The plan-time schema-push detection (former plan-phase.md §5.7) becomes a schema-gate Capability: a plan:pre contribution (into:planner, when:workflow.schema_push_detection) whose fragment carries the full ORM-detection + [BLOCKING] schema-push-task injection logic, rendered into the planner via the existing plan:pre render-hooks dispatch. The inline §5.7 block is removed (plan-phase.md 94519->90445). workflow.schema_push_detection is a new capability-owned (federated) key, default true. (The execute-side schema-drift gate was migrated separately into the drift capability.)

Verified: registry inlines the fragment (len 2704) so it is actually delivered at plan:pre; gate 8 pass / 1 fail — all 5 ADR-named features now real Capabilities, only the config-leak test remains (intel/security, next unit). Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): close the 3 capability config-key leaks — phase-6 gate now GREEN

Removes the last inline config-get reads of capability-owned keys from plan-phase.md. security_asvs_level/security_block_on now flow through the security plan:pre contribution via a new loop-resolver configValues mechanism (resolves declared config keys with the same 4-level precedence as activation and attaches them to the rendered hook); the §5.55 banner reads them from PLAN_PRE_HOOKS_JSON. intel.enabled becomes a real intel plan:pre step (ref.command: intel api-surface) dispatched via render-hooks; the inline intel branch is gone. gen-capability-registry now validates ref.command as a third dispatch shape.

Verified: phase-6 capstone conformance gate is FULLY GREEN (9/0); 3 leaks gone (grep=0); security configValues resolve to {2,medium}/default {1,high}; intel step present only when enabled; loop-render-hooks 62/0, capability-registry 287/0, capability-state/federated-config 113/0; lint 0 errors. Closes the migration half of #1169. Refs #1139, #1167, #1168.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): address adversarial review — restore schema-drift block, generic planner injection, uniform gate contract

Adversarial review caught 2 real regressions the green gate missed: (1) schema-drift no longer blocked — the execute:wave:post dispatch read GATE_RESULT.block but verify.schema-drift emitted drift_detected/blocking, and onError:skip wrongly bypassed positive blocks; (2) only tdd's plan:pre contribution was injected into the planner, dropping schema-gate's schema-push detection and security's threat-model guidance.

Fixes: (A) every gate check returns a uniform boolean 'block' under --raw (the dispatch form), with advisory gates (tdd/gap) carrying their report in 'message'; (B) gate-dispatch contract corrected at all sites — onError governs command errors only, a blocking gate's positive block always halts; (C) generic planner injection of all plan:pre contributions where into=='planner' (tdd + schema-gate + security incl configValues); (D) two new conformance assertions: planner contributions injected generically + every gate check.query returns boolean block under --raw.

Verified: gate 11/11; all 6 gate checks return boolean block under --raw; full suite 595/0; lint 0 errors. Refs #1167, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): restore MVP+TDD end-of-phase blocking escalation (2nd adversarial pass)

The migrated tdd execute:post gate is statically blocking:false, but the contract (references/execute-mvp-tdd.md + CONTEXT.md) requires the end-of-phase TDD review to ESCALATE from advisory to blocking when MVP_MODE && TDD_MODE && a TDD plan misses a RED/GREEN commit. The migration prose had downgraded this to a 'strong advisory recommendation' — silent loss of the blocking escalation. Restore it: the tdd-gate dispatch now refuses to mark the phase complete (Phase blocked message) under MVP+TDD when GATE_RESULT.block is true; advisory otherwise.

Also strengthen tests/execute-mvp-tdd-gate.test.cjs: hasBlockingEscalation previously matched any line with 'blocking'+'mvp+tdd' (so 'advisory (blocking: false) ... under MVP+TDD' was a false green); now it requires the real refusal semantics ('refuse to mark the phase complete' / 'phase blocked'). Caught by 2nd adversarial review pass.

Verified: execute-phase.md 92702 < 93166 frozen; mvp-tdd-gate + phase-6 gate 19/0; full suite green; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): restore MVP+TDD proceed-block, codebase auto-remap, schema skip-flag (3rd adversarial pass)

3rd adversarial pass found 4 more silent regressions: (1) the tdd MVP+TDD 'refuse to mark complete' was nullified by a downstream 'ALWAYS proceed regardless of gate results' line — proceed is now conditional (stops on an active MVP+TDD block); (2) the test now asserts the proceed is NOT an unconditional override; (3) codebase-drift auto-remap (spawn gsd-codebase-mapper when drift_action=auto-remap) was dropped — the execute:wave:post advisory dispatch now consumes spawn_mapper/directive; (4) GSD_SKIP_SCHEMA_CHECK bypass was lost from the gate path — cmdVerifySchemaDrift now honors the env var (block:false when set).

Verified: no unconditional proceed; GSD_SKIP_SCHEMA_CHECK=true -> block:false; gate 11/11 + mvp-tdd 9/9; full suite 569/0; lint 0; execute-phase.md 93109 < 93166. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): init.cts reads federated config keys from nested path (4th adversarial pass)

Config federation moved tdd_mode/research/nyquist_validation from flat config.<key> to nested config.workflow.<key>, but src/init.cts still read them flat — so init.plan-phase/init.execute-phase emitted tdd_mode:false / research_enabled:undefined / nyquist:undefined regardless of config (a public command-contract regression; the migrated loops use render-hooks so enforcement was unaffected). Read via config.workflow (type-safe Record cast). Now init reflects the same resolved values + federated defaults (research/nyquist default true) as the render-hooks path.

Verified: build clean; init.plan-phase emits tdd_mode:true/research:false/nyquist:false for set config, defaults true for empty; full suite 591/0; lint 0. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1169): add changeset for ADR-857 phase-6 completion (PR #1183)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): complete phase-6 migration fallout — restore TEXT_MODE, fix registry .claude leak, re-point stale workflow-contract tests

The capability migration left real regressions and stale consumer tests that
the per-module unit suite missed but the full cross-platform suite caught (27
failing tests):

Real source regressions (fixed):
- execute-phase.md lost its AskUserQuestion TEXT_MODE plain-text fallback when
  the inline schema_drift_gate step was removed — non-Claude runtimes would
  stall. Restored, and the execute:post gate-dispatch prose de-duplicated to
  cite the execute:wave:post contract (loop body shrinks below the frozen
  pre-phase-6 ceiling while keeping every onError/blocking nuance).
- capabilities/tdd inline fragment hardcoded `@~/.claude/gsd-core/references/tdd.md`,
  baked verbatim into the committed capability-registry.cjs and leaked the
  install path on 11 non-Claude runtimes (registry .cjs is copied, not
  path-converted). Made the fragment path-free; regenerated the registry. The
  phase-6 conformance gate now guards this (no ~/.claude install path in any
  capability source or the generated registry).
- plan-phase.md: removed a §5.7 stub re-added in error and routed Branch 2 to
  step 6 (schema-gate is a plan:pre capability, §5.7 is gone).

Stale workflow-contract tests re-pointed to the capability dispatch they now
must assert (behavior verified preserved in source first, assertions kept
equal-or-stronger): bug-621 + bug-2851 (gap-analysis via gsd_run render-hooks
plan:post + registry binding), feat-2527 (tdd_mode federated out of central),
phase6-planning + plan-phase-ui-redirect (§5.6 bounded by ## 6.),
plan-phase-drift-guard (intel when:intel.enabled skip branch).

profile-pipeline-command-router.cjs un-ignored from eslint (hand-written, no
TS source) + stale disable comments removed. Size baseline regenerated.

Verified: full suite 15140 tests / 0 fail; lint 0 errors; conformance gate green
legitimately. Refs #1139, #1167, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1169): add ADR-857 E2E content-test coverage for the 12 loop points + capability deliverables

Grounds the capability engine in behavioral E2E tests (drive the real
render-hooks/check CLI + the real registry, assert typed result content — no
source-grep), structured around what ADR-857 says to deliver. 207 tests; each
genuineness-checked (flip the expectation, confirm it fails).

Per-loop-point dispatch (7 files): empty-point negative-space across the 6
no-hook points; verify:post 3-step resolution+ordering+onError; plan:pre
contribution/configValues + ui.plan-gate + intel; plan:post gap-analysis;
execute:wave:post drift+ui gates via the check route (schema-drift block/skip,
codebase-drift threshold BVA, auto-remap); execute:post tdd.review-checkpoint
RED/GREEN; ship:pre security gate resolution + frontmatter-get predicate pieces.

ADR-deliverable coverage (4 files): predicate boundary held (edge/prohibition
probes stay core, not off-by-default Feature Capabilities — phase-6 exception);
core loop runs with zero capabilities (all 12 points empty, init bundles
resolve); contribution merge (multiple ordered <contribution from=> blocks);
federated-config key removal on uninstall.

federated-config allowlisted for its 3-file split (unit + integration +
lifecycle). Refs #1139, #1167, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): remove dead drifted converter dups + address adversarial review

Lint cleanup (root-caused, not waved off): src/runtime-artifact-conversion.cts
carried 11 agent-converter functions (+5 orphaned consts/helpers) that were
never exported, never called, and had silently DRIFTED from the live
hand-authored copies in bin/install.js (one even referenced an undefined
`claudeToCopilotTools`). Deleted the dead duplicates; install.js's live copies
are untouched (it never imported these). Lint now 0 errors / 0 warnings.

Adversarial-review (Codex) findings fixed:
- HIGH: execute-phase.md TDD_MODE used `jq ... || echo false`, silently
  disabling the MVP+TDD blocking gate on jq-less runtimes. Reverted to the
  `node -e` form (node is guaranteed; matches the file's other node-e usages) so
  a missing optional tool can no longer fail-open a blocking safety path.
- MEDIUM: federated-config-key-removal orphan-key test was vacuous (it skipped
  the orphan assertion). Now asserts the removed capability's key is genuinely
  not surfaced/validated after uninstall.
- LOW: phase-6 conformance leak regex broadened to catch absolute-home and
  Windows-backslash `.claude/(gsd-core|commands|agents|hooks)` paths, not only
  `~`/`$HOME` forward-slash forms.
- LOW: bug-2851 plan:post dispatch assertion now requires `--raw` (matched its
  stated contract).
- nit: plan-pre intel-step test duplicate assertion replaced with a distinct
  structured-output check.

Size baseline regenerated (execute-phase.md 93089 < 93166 frozen). Refs #1167, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1169): make runtime-homes-descriptor-drive titles environment-independent

The descriptor-equivalence test embedded the absolute golden config path
(`os.homedir()`-derived) directly in each `test(...)` title, so titles differed
between macOS (`/Users/x/.claude`) and Docker (`/home/gsdtest/.claude`). Every
test PASSES on both platforms (15885/0 leaf tests each), but gsd-test-summary
compares results by title and reported 29+29 false "only in Mac / only in
Docker" discrepancies for tests that actually pass everywhere.

Move the golden path out of the title and into the assertion message (still
shown on failure); titles are now byte-identical across platforms so the
cross-platform comparator matches them. No assertion logic or golden values
changed. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): derive TDD_MODE via gsd_run --active-cap, not node -e (fix prompt-injection CI gate)

The prior fix reverted execute-phase.md:181 from jq to `node -e` to close a
Codex HIGH (jq||echo-false silently disabling the MVP+TDD blocking gate on
jq-less runtimes) — but the CI prompt-injection scanner BLOCKS new `node -e` in
workflow markdown (inline code-exec = injection vector), turning the security
gate red. Both forms were wrong: node -e fails the scanner; jq fail-opens a
blocking safety gate; `config-get workflow.tdd_mode` is forbidden by the
conformance leak gate (tdd_mode is capability-owned).

Correct fix (what Codex recommended): a gsd_run-native boolean. Add an
`--active-cap <capId>` flag to `loop render-hooks <point>` that resolves hooks
the normal way and prints exactly `true`/`false` for whether a capId is active
— scanner-safe (canonical launcher, no inline code), node-reliable (no optional
jq to fail-open), and leak-free (render-hooks resolution, not config-get).
execute-phase.md:181 now `TDD_MODE=$(gsd_run loop render-hooks execute:post
--active-cap tdd)`. +5 behavioral tests for the flag.

Verified: prompt-injection-scan --diff origin/next → 0 findings; conformance
gate 13/13 (execute-phase.md 92934 < 93166); execute-mvp-tdd + tdd-mode +
loop-render-hooks 87/0; lint 0/0. Refs #1167, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 21:07:55 -04:00
Tom Boucher
83e87e3ec0 feat(#1180): hard-gate the GSD Version requirement on bug reports (#1181)
Auto-close bug reports opened without a valid GSD Version (Issue Forms enforce required only in the web UI). Bug reports only; version-shaped validation; version-exempt opt-out. Closes #1180
2026-06-13 16:13:00 -04:00
Tom Boucher
a19a709e62 test(#1171): add agent-classification parity guard (#1176)
Makes docs/AGENTS.md section structure the single source of truth for the
primary-vs-advanced agent classification and fails when docs/INVENTORY.md's
"Primary doc" column or the AGENTS.md prose counts drift from it.

The classification ("primary" = full role card, "advanced stub" = concise
stub) was hand-duplicated across three doc surfaces with no enforcement.
This drift guard derives the expected class from AGENTS.md section placement
and cross-checks the INVENTORY.md column, the prose counts, the parenthetical
advanced-agent list, and full roster completeness (21 primary / 12 advanced
/ 0 inventory-only today).

No agent frontmatter field is added (avoids the capability-registry `tier`
collision and the research-profiles.cjs ripple); the classification stays
documentary, enforced from where it is defined.

Closes #1171

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 11:28:05 -04:00
Tom Boucher
8b037fe1e3 test(#1168): make phase-6 conformance gate fail on real incompleteness
The phase-6 capstone gate (#1139/#1158) passed green while ADR-857's acceptance criteria were unmet — a false green that let phase 6 read as complete. Drop the paper-over (the call-site allowlist) and assert the real criteria with no exemptions, so green means 'phase 6 conformant', not 'no new regression'.

RED by design until the work lands: (1) every declared hook point must have a render-hooks call site → execute:wave:post (ui_safety_gate) is dead; (2) every ADR-857-named optional feature must be a Capability → tdd/schema-gate/drift/gap-analysis/profile-pipeline un-migrated; (3) the host loop must read no capability-owned config key inline → intel.enabled/security_asvs_level/security_block_on leak in plan-phase.md.

Refs #1139, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 11:09:21 -04:00
Tom Boucher
70f35ebc7a feat(#1169): wire ship:pre security gate via render-hooks (ADR-857 phase 6)
Part of the ADR-857 phase-6 migration (#1169, epic #857) — not a shipped bug fix. The capability system (registry, render-hooks, gates, capability-state) lives only on next / 1.5.0-rc; npm latest is 1.4.5 and contains none of it, so no released user can hit this.

The security capability declares a blocking gates@ship:pre hook (when=workflow.security_enforcement, predicate SECURITY.md.threats_open==0), but ship.md never called `loop render-hooks ship:pre` and had no inline fallback, so the declared gate never fired. Wire it into ship.md preflight_checks via the render-hooks idiom — fail-closed: the ship blocks unless threats_open is exactly 0.

Add a phase-6 conformance assertion that every declared hook point has a render-hooks call site; execute:wave:post (ui_safety_gate) stays in KNOWN_UNWIRED pending its unimplemented ui.safety-gate check.

Part of #1169. Refs #857.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 11:09:21 -04:00
Tom Boucher
f116b76128 test(#1139): add phase 6 capstone conformance gate (#1158) 2026-06-13 02:12:16 -04:00
Tom Boucher
aec3374bc2 feat(#1138): make runtime descriptors authoritative (#1157) 2026-06-13 01:49:25 -04:00
Tom Boucher
0c0a8966ba fix(#1151): drive codex sandbox_mode emission from runtime descriptor sandboxTier axis (#1152)
* fix(#1151): drive codex sandbox_mode emission from runtime descriptor sandboxTier axis

The sandboxTier runtime-capability axis was cosmetic: declared and validated
on all 16 descriptors but read by nothing. The codex per-agent sandbox_mode
line was emitted unconditionally from the hardcoded CODEX_AGENT_SANDBOX map, so
the descriptor field drove no behaviour (ADR-857 audit finding F10; the
"rides along in 5e/5g" promise in ADR-1016 §8 never landed).

Make the axis load-bearing:
- resolveInstallPlan projects sandboxTier as a 7th InstallPlan axis and fails
  loud (throws) on a missing/invalid value rather than coercing to 'none'.
- installCodexConfig / generateCodexAgentToml gate sandbox_mode emission on
  sandboxTier !== 'none'.
- The per-agent CODEX_AGENT_SANDBOX map is kept: it is GSD agent policy, not a
  runtime-descriptor property (different layer). Full removal of that map is
  tracked under #1138 (phase-6 descriptor-residue removal).

For codex (sandboxTier === 'codex-agent-sandbox') the emitted TOML is
byte-identical to before; for 'none' runtimes sandbox_mode is omitted.
Adds leaf, projection, and installCodexConfig threading-seam regression tests;
updates the enh-1082 InstallPlan golden master with sandboxTier for all 16
runtimes. Confirmed hypothesis: schema-first vocabulary closure outran consumer
wiring, with no conformance gate to catch the orphaned axis.

Closes #1151
Refs #857, #1138

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1151): stamp changeset with PR number 1152

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 00:06:52 -04:00
Tom Boucher
607813f5d0 feat(#1136): consume resolved capability state (#1153)
* feat(#1136): consume resolved capability state

* chore(#1136): add capability state changeset
2026-06-12 23:43:07 -04:00
Tom Boucher
b431b1fab4 fix(#1140): implement state add-roadmap-evolution CJS handler (#1148)
* fix(#1140): implement state add-roadmap-evolution CJS handler

`query state.add-roadmap-evolution` was unreachable: the CJS state router
listed it in the `unsupported` map with a circular message ("...is SDK-only.
Use: gsd-tools query state.add-roadmap-evolution ...") and no CJS handler
existed after the SDK retirement (ADR-0174). Every `/gsd:phase insert` and
`/gsd:phase --edit` run hit a dead end recording Roadmap Evolution.

Re-implement `cmdStateAddRoadmapEvolution` in CJS (src/state.cts) and wire it
into the state router; remove the now-stale `unsupported` entry. The handler
appends a single-line bullet under `## Accumulated Context` → `### Roadmap
Evolution` (creating the subsection/section if missing, deduping identical
entries), scoping every lookup to the Accumulated Context body so a decoy
heading in an unrelated section is never targeted, and flattening multiline
notes to a single bullet. Section-boundary regexes mirror the sibling
add-decision/add-blocker handlers and preserve following sections on CRLF input.

Regression cases live in tests/state.test.cjs (per the no-new-bug-NNNN-files
policy) and cover the literal issue repro plus the CLI/parser QA matrix
(missing/empty/whitespace note, flag-shaped value, duplicate flags, hostile
shell metacharacters, Unicode, decoy section, CRLF, missing STATE.md).

Closes #1140

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1140): backfill changeset PR number (1148)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 22:27:02 -04:00
Tom Boucher
1b880fefd7 feat(#1137): migrate review verification hooks to capabilities (#1147)
* feat(#1137): migrate review verification hooks to capabilities

* chore(#1147): add changeset
2026-06-12 21:24:09 -04:00
Tom Boucher
44024aa535 fix(#1133): honor model_policy on the claude runtime (forward-port to next) (#1144)
Forward-port of the #1133 hotfix (commit 22f237a5, 1.4.5 hotfix line) onto
next. The hotfix was authored against src/core.cts (v1.4.4); on next the
resolver logic lives in src/model-resolver.cts (ADR-457 extraction, #888),
so the patch is re-applied there rather than cherry-picked.

resolveModelInternal step 2.5 now honors model_policy on the claude runtime:
the policy-resolved full model ID is mapped back to a Claude Code agent alias
via CLAUDE_POLICY_ID_TO_ALIAS (reverse of MODEL_ALIAS_MAP + claude-fable-5 ->
fable). Bare aliases (opus/sonnet/haiku/fable) pass through; an ID with no
Claude alias warns once to stderr (deduped by agentType::policyModel::tier)
and falls back to the configured tier alias. Non-claude runtimes return full
IDs verbatim (unchanged). resolveModelForTier is intentionally unchanged.

The warn-dedupe cache lives in model-resolver.cts; core.cts composes the
exported _resetRuntimeWarningCacheForTests to clear both that cache and the
config-loader warning cache (config-loader cannot import model-resolver --
circular dependency).

Ports the 6 #1133 tests (rewriting the old claude-no-op test that asserted
the bug) plus one added test covering the MODEL_ALIAS_MAP reverse-map path.

Forward-port of #1133

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 19:56:54 -04:00
Tom Boucher
4ab5c7b3f2 feat(#1135): migrate planning hooks to capabilities (#1141)
* feat(#1135): migrate planning hooks to capabilities

* chore(#1135): add phase 6 planning capabilities changeset

* fix(#1135): satisfy lint for agent hook rendering
2026-06-12 19:51:54 -04:00
Tom Boucher
fd01e7a12e feat(#1132): complete contribution hook prerequisite
Closes #1132
2026-06-12 18:39:57 -04:00
Tom Boucher
1d90ad3c30 fix(commands): route nested sub_repos by longest prefix, not array order (#1130)
groupFilesBySubrepo selected the first sub_repos entry in array order
whose prefix matched a file, so a file under a more-specific nested
sub-repo (e.g. packages/core/widget.js with sub_repos
["packages", "packages/core"]) was mis-routed to the less-specific
parent ("packages").

Select the longest (most-specific) matching prefix within each
first-segment bucket instead, making routing independent of sub_repos
array order. String()-guard the length comparison so non-string entries
still never throw (preserves the #311 tolerance). Update the stale doc
comment that claimed first-match semantics.

Closes #391.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 17:33:26 -04:00
Tom Boucher
eb051ea696 feat(#1123,#1124): enforce duplicate-producer invariant + fail-loud loadCentralConfigKeys in gen-capability-registry (#1131)
Closes #1123
Closes #1124
Refs #857
2026-06-12 16:51:02 -04:00
Tom Boucher
827011b865 fix(#1098): guard generate-claude-md against clobbering hand-crafted files; redirect default to .claude/CLAUDE.md (#1118)
/gsd-new-project wrote a repo-root CLAUDE.md full of broad project docs,
overwriting/diluting a hand-crafted instruction file. --force was parsed but
silently dropped, and nothing guarded an existing non-GSD file.

- Guard: an existing instruction file with no `<!-- GSD:<section>-start` markers
  (hand-crafted) is left untouched; report action:"skipped". --force (now wired
  through CmdGenerateClaudeMdOptions) overwrites intentionally. The marker check
  uses /<!-- GSD:[a-z]+-start/ so a file merely documenting GSD syntax is safe.
- Redirect: the Claude-family default output is now ./.claude/CLAUDE.md (a valid
  auto-loaded project-memory location) instead of repo-root ./CLAUDE.md. Aligned
  across the handler default, config-defaults.manifest.json, buildNewProjectConfig,
  the config template, new-project.md, and cmdGenerateClaudeProfile; advisory
  read-CLAUDE.md hints in plan-phase/quick/profile-user updated. Codex still
  writes AGENTS.md.

Closes #1098

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 16:03:51 -04:00
Tom Boucher
fc37ae0c4f fix(#1115): capability-probe codex hook-trust bypass flag in /gsd:review; fail loud on empty output (#1122)
On codex-cli < 0.137.0 the review.md `codex exec` invocation passed
--dangerously-bypass-hook-trust (added in 0.137) unconditionally and discarded
stderr, so codex exited "unexpected argument" before reading the prompt and the
empty output file was treated as a completed review — a silent degraded review.

- Capability-probe the flag (`codex exec --help | grep`) and apply it via
  $CODEX_BYPASS_FLAG only when supported (works fine without it on older CLIs).
- Capture codex stderr to a .err file instead of /dev/null, and replace an empty
  output with a diagnostic so a broken reviewer is surfaced (mirrors Cursor/OpenCode).
- Update enh-773 enforcement test to require the capability gate + fail-loud guard
  instead of the unconditional flag.

Closes #1115

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 15:55:17 -04:00
Tom Boucher
1c073c1c81 fix(#1107): progress consults verification.status before reporting a phase complete (#1116)
/gsd-progress derived phase completeness from plan/summary counts only and never
consulted the verification.status query (the #651 seam), so a phase whose
VERIFICATION.md ended human_needed or gaps_found was reported complete and
routing skipped to the next phase. Add Step 1.7 (consult verification.status for
the current phase) and routing rows that send gaps_found to plan-phase --gaps
(Route V.gaps) and human_needed to verify-work (Route V.human) before the generic
complete row. passed/missing/unknown still route as complete so unverified phases
are not falsely blocked.

Closes #1107

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 15:55:01 -04:00
Tom Boucher
fb8e7a3a65 fix(#1114): write-profile resolves the active runtime config home for USER-PROFILE.md (#1119)
Under Codex, `write-profile` wrote ~/.claude/gsd-core/USER-PROFILE.md while Codex
discuss-phase advisor-mode (installed under ~/.codex) checked the Codex home and
never found the profile, so advisor-mode silently stayed disabled. Resolve the
default output via the runtime-aware getGlobalConfigDir (GSD_RUNTIME / config.runtime),
mirroring cmdGenerateDevPreferences. Claude unchanged; --output still wins.

Closes #1114

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 15:54:48 -04:00
Tom Boucher
e837bfc1de fix(#1101): record-session updates ## Session Continuity in place, no duplicate block (#1113)
The reported symptom (recorded:false yet STATE.md frontmatter still mutated) was
already resolved on next by #944/#948. This fixes the residual: the DWIM
auto-create recognised only the canonical `## Session` heading, so a bootstrap
`## Session Continuity` section fell through to the append branch and produced a
second `## Session` block. Insert only the missing canonical fields after the
`## Session Continuity` heading — preserving the heading and any prose (e.g.
"Next recommended action") — and teach the snapshot / frontmatter readers to
recognise that heading (the optional ` Continuity` group still excludes
`## Session Continuity Archive`, preserving #2444 scoping).

Closes #1101

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 15:54:28 -04:00
Tom Boucher
4beceb02fe fix(#1103): preserve newline before Plans: header in roadmap.annotate-dependencies (#1111)
The match regex's `(?:^|\n)` anchor consumes a leading newline on mid-string
matches (the common case where a `**Plans:** N plans` summary line precedes a
bare `Plans:` block). The replacement dropped that newline, fusing the summary
line onto the header — e.g. `**Plans:** 3 plansPlans:`. Re-emit the consumed
newline so markdown line boundaries are preserved.

Closes #1103

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 15:54:25 -04:00
Tom Boucher
7f1d49935c ci(#1104): keep next package.json in sync with the last published release (#1109)
* ci(#1104): sync next package.json version to the last published release

next rested on a -dev stream per ADR-660 (1.3.1-dev.0) — a never-published
placeholder that leaked to source/dev installs. Make every release type write
its exact published version back to next:

- finalize/hotfix (push main): auto-backmerge sets next's version to main's
  released version, folded into the existing back-merge PR (+ pinned setup-node).
- rc (no main push): the rc job opens + admin-merges a sync PR after publish.

Shared, fail-closed scripts/sync-next-version.cjs stamps package.json + the
runtime manifests via the npm version hook and refuses any non-release version.
Amends ADR-660 (supersedes the -dev stream decision).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* ci(#1104): harden next-version sync against post-publish failure modes

Review hardening (Codex + code-review gates) on the #1104 sync helper and
its workflow callers:

- release.yml rc Sync step: continue-on-error so a post-publish sync hiccup
  cannot fail an already-published release (npm immutability would block re-run).
- auto-backmerge.yml inline sync: set -euo pipefail + validate VERSION before
  any shell use (closes a ${VERSION}-in-commit-message injection vector); git
  add -u instead of -A.
- sync-next-version.cjs: reuse an existing open PR instead of failing gh pr
  create on rc re-runs; regex-parse the PR number and fail loud; discriminate
  the git diff --cached --quiet exit code (only status 1 == has-diff, else
  rethrow); git add -u to avoid sweeping runner artifacts into next; tolerate
  already-merged on admin merge.
- tests: +2 (existing-PR reuse, non-diff rethrow); 14/14 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 13:29:45 -04:00
Rezolv
3e836fef0d feat(spec-phase): spec-completeness edge-probe (#550) (#584)
* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/

Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.

Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.

Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
  planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
  source-checkout-gated build fallback) instead of LLM re-derivation; the engine
  capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
  resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
  per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
  validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.

Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.

* test(#550): RED — status×verification re-cut + probe-core engine specs

Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):

  status: resolved | dismissed | unresolved   (lifecycle, shared)
  verification: explicit | backstop | null     (only when resolved)

- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
  to be extracted — validateResolution(r, validators), validateRequirement,
  analyzeCoverage(items, resolutions?, validators), byVerification rollup,
  runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
  {resolved,backstop}; coverage gains byVerification.{explicit,backstop};
  proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
  COUNT preserved on every fixture (closed set = resolved+dismissed; doc
  line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
  rewritten to the two-axis model.

Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).

* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)

Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.

probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
  verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
  items[] (core never assumes propose is deterministic — edge resolves via
  LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
  count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
  {categories, verification, requiredFieldsByVerification} (ADR-550 #5)

edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.

* chore(#550): register probe-core.cjs artifact in ledgers + inventory

New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:

- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
  never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
  is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
  edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).

probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.

* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]

trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.

Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.

* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard

Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:

- validateResolution now enforces the 'verification is null unless resolved'
  invariant for EVERY status (not just resolved): a dismissed/unresolved
  resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
  (was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
  it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
  contract; a new test locks that an all-dismissed run is NOT affirmatively covered
  (byVerification is the honest gate).

Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.

* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics

Re-review #5 (trek-e) clarity edits:

- Decision 5: annotate that only contract item (a) ships on #584 (the edge
  adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
  parenthetical — the blessed/implemented semantics are count-preserved = the
  CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
  carrying the per-tier resolved-status breakdown. The old parenthetical
  contradicted the shipped count.

* test(#550): cover runProbeCli structural-guard numeric-count branch

Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.

* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)

The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.

* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)

A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.

* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)

templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.

* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)

The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.

* docs(#550): add how-to for resolving edge-coverage findings (B1)

Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.

* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)

trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.

* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)

trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.

* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)

trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.

* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe

Rebased onto next (e4f0910d), which replaced the line-based tier-max
size ratchet (#597) with the byte-based per-file baseline guard (#1074).
The edge-probe feature legitimately grows two workflows:

  - spec-phase.md  15131 -> 23094 (+7963): Step 5.5 Edge-Completeness Probe
  - plan-phase.md  93135 -> 94253 (+1118): covered/backstop edge lift into
    must_haves.truths (the live <downstream_consumer> block)

Both remain under their tier hard caps (plan-phase XL 98304, ~4KB
headroom; spec-phase DEFAULT 40960). Growth is real inline workflow
content the feature requires at that step — not eager @-import proxy
gaming. Drops the obsolete line-based XL_BUDGET 93000->94000 bump
(superseded by the byte baseline) via rebase.

* chore(#550): reconcile INVENTORY headline counts after rebase onto next

Rebase onto current next dropped the prior reconcile commit (stale counts
refs 68 / CLI 107). Current next + the edge-probe additions yield:
  - References (68 -> 69 shipped): + gsd-core/references/edge-probe.md
  - CLI Modules (107 -> 109 shipped): + edge-probe.cjs + probe-core.cjs

Rows for all three already present; only the headline counts were stale.
Caught by tests/inventory-counts.test.cjs (CI ubuntu-24 leg).
2026-06-12 11:05:31 -04:00
Rezolv
e4f0910d62 test(#1074): agent-size-budget per-file baseline + line→byte rebase (PR 3/3) (#1097)
Extends the #1074 scheme to tests/agent-size-budget.test.cjs, which used the
same assertTightCeiling tier ratchet but was still line-based (never rebased in
#717). Completes the migration — the last part of the #1074 epic.

- Rebase agent sizing from lines to LF-normalized bytes (#717/#683).
- Delete the 'SIZE: tier anti-creep' describe (3 assertTightCeiling tests);
  add a per-agent baseline (tests/agent-size-baseline.json) as the primary
  anti-creep, and byte hard caps (XL 56 KiB / LARGE 48 KiB / DEFAULT 24 KiB),
  each above its tier high-water with real headroom. No separate new-file cap:
  a net-new agent is DEFAULT-tier, already bounded by the DEFAULT cap.
- Keep the agent-classification tests verbatim.
- scripts/workflow-size.cjs: add generic measureMdFiles(dir, predicate)
  (workflows + agents share one byte-measurement path); measureWorkflows now
  delegates to it.
- scripts/update-size-baseline.cjs: one 'npm run size:baseline' now regenerates
  BOTH the workflow and agent baselines (gsd-* filter for agents).

Rebased onto next after PR 2/3 (#1096) merged: replicate the
scripts/lib/workflow-size.cjs -> scripts/workflow-size.cjs move (PR 1/3) across
the generator and the agent test's require; regenerate the agent baseline
against current agents (a uniform +170 B preamble drift on all 33 since
authoring).

Addresses the #1097 review (trek-e):
- BLOCKER (acceptance criterion 5): document the agent contract in CONTEXT.md.
  Adds RULESET.AGENT_SIZE_BUDGET (caps 57344/49152/24576, per-agent baseline,
  dual size:baseline, shared measureMdFiles seam) and disambiguates it from the
  separate DEFECT.AGENT-FILE-SIZE-CAP-BREACH 45K-CHAR guard (two units, two
  purposes).
- Docs: now that #1096's docs/TESTING-SUITES.md "Workflow size budget" section
  is in next, fold in the agent coverage here (renamed to "Workflow & agent
  size budget"): agent caps + per-agent baseline + the how-to + reference rows,
  and the disambiguation from the 45K-char guard.
- Minor (negative proof): add a boundary-fixture test exercising the hard-cap
  comparison at cap-1/cap/cap+1 through the real lfByteCount path, so a future
  threshold/operator edit can't silently neuter a cap.
- Nit: align the tier test name wording ("stays within") with the <= operator.

Negative proof on a real tracked agent (gsd-planner): baseline catches +10 B;
XL hard cap catches 57,516 > 57,344 with the baseline current.

Closes #1095 (PR 3/3 child); landing this completes the #1074 epic.
2026-06-12 09:58:44 -04:00
Rezolv
b055ca14e2 test(#1074): swap workflow size enforcement to baseline + loose hard caps (PR 2/3) (#1096)
Completes the #1074 migration for workflows. The per-file baseline (PR 1) is
now the primary anti-creep guard, so the tier-max tighten-only ceilings are
retired here.

- Delete the 'SIZE: tier anti-creep' describe (3 assertTightCeiling tests) and
  the GRACE constant — per-file baseline already guards every file by name,
  strictly stronger than max(tier).
- Convert the per-file tier test into 'SIZE: workflow tier hard caps': absolute
  red lines (XL 96 KiB / LARGE 60 KiB / DEFAULT 40 KiB) that mean 'extract, do
  not raise', each with real headroom above its high-water file.
- Add a 32 KiB (Codex project_doc_max_bytes) cap for net-new workflow files not
  yet in the baseline and not explicitly tiered.
- Rewrite the header doc comment for the new two-guard model; drop the
  assertTightCeiling import (now unused here; still exported + used by the agent
  test until PR 3).

Addresses the three PR-2 items from the #1089 review:
- Finish the enumeration consolidation (Minor #1): the tier hard-cap and
  new-file guards now read both their file list and byte sizes from
  measureWorkflows()/listWorkflowStems(), removing the inline readdirSync +
  per-file byteCount split-brain. Enumeration and measurement share one source.
- Rewrite CONTEXT.md RULESET.WORKFLOW_SIZE_BUDGET (Minor #2): stale caps
  (XL<=90000/LARGE<=54000/DEFAULT<=38000) replaced with the baseline-first
  model, current hard caps, the new-file anchor, and the size:baseline
  remediation.
- Ship docs (contract-change requirement): a Diataxis how-to + reference for
  the size guard in docs/TESTING-SUITES.md, beside the sibling regression-name
  ratchet (the 'file grew, CI red -> npm run size:baseline, commit the one-line
  diff, justify or extract lazily' workflow). CONTRIBUTING.md's workflow-tree
  note is updated to the baseline model and points at the new guide.

Negative proof: each guard fails independently when violated (new-file cap at
33 KB; hard cap with baseline current; baseline on any per-file growth).

Refs #1074. Part 2 of 3.
2026-06-12 09:30:44 -04:00
Tom Boucher
9e2ef2c94d fix(#1091): thread install scope into skill converters so local Antigravity/Copilot installs use workspace paths (#1092)
The skills layout wrapper (skillsKind) invoked every per-runtime skill
converter as realConverter(content, skillName, runtime, cmdNames). The
3rd positional arg is overloaded: claude/kimi/cline converters read
`runtime` there, but the copilot/antigravity converters read `isGlobal`
there — so they received the truthy runtime string and always took the
global path branch, leaking ~/.gemini/antigravity/ and ~/.copilot/ into
local/workspace installs instead of .agent/ and .github/.

Thread `scope` from resolveRuntimeArtifactLayout -> dispatchKindEntry ->
skillsKind, derive isGlobal = scope === 'global', and pass it as a
non-colliding 5th positional arg. Move isGlobal out of the colliding 3rd
slot in the two converter signatures (3rd/4th become ignored
_runtime/_cmdNames, matching the kimi convention). The fix flows through
the shared ArtifactKind.stage closure, so applySurface re-apply inherits
it via the same seam.

Regression test exercises the wrapper seam (installRuntimeArtifacts at
local scope) for both runtimes and asserts workspace paths, not global.

Closes #1091

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 00:20:28 -04:00