Commit Graph

1629 Commits

Author SHA1 Message Date
Tom Boucher
a13101ee5c fix(#1329): existence-filter scoped-CI fallback so a deleted test can't crash the lane
ci-prepare-test-scope.cjs's empty-detection FALLBACK hardcoded
tests/core.test.cjs, deleted in #1291. Every scoped lane (scope=targeted|
windows) that hit the fallback wrote the stale path into .ci-selected-tests.txt
and crashed run-tests with "requested test file(s) not found: core.test.cjs".
The full/sharded lanes glob the suite and were immune, so only the scoped
lanes went red (e.g. run 27599149212 on #1308).

Existence-filter the FALLBACK at write time and fall back to the 'unit' suite
sentinel (the #408/#641 path, resolved live by run-tests) when nothing
survives, so a stale reference degrades instead of crashing the lane. Detected
lists still pass through verbatim (they may carry a suite sentinel and are
already filtered by affected-tests-lib). Refactor to an exported, testable
resolveSelection().

Add a generative parity guard (DEFECT.GENERATIVE-FIX) asserting every FALLBACK
entry resolves on disk or is a known suite sentinel — it fails the instant a
refactor deletes a listed file, which #1291 did and CI did not catch — plus
resolveSelection unit tests and an end-to-end subprocess test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 08:38:55 -04:00
Tom Boucher
3a53255244 refactor(#1308): consolidate config-key precedence engine into single owner (#1322)
Make src/capability-activation.cts the sole owner of the four-level config-key
precedence walk via a new raw-value primitive resolveConfigKey(dotKey,{config,
cwd,registry}); the boolean wrapper _resolveActivationValue and loop-resolver's
resolveConfigValues are both rebuilt on it. loop-resolver deletes its
byte-identical copy and inline re-walk and imports the engine.

resolveCapabilityRuntimeState no longer returns registry/config (leaked internal
detail); the caller loads one fail-closed config snapshot and threads it in via a
new optional configOverride param, so capability `active` and hook when/configValues
resolve against the same object. capability-writer requires the registry module
directly.

Adds a DEFECT.GENERATIVE-FIX parity gate (identity + behavioral matrix +
end-to-end resolveLoopHooks configValues) that fails if the two precedence
surfaces ever diverge. Pure internal refactor, no user-facing change.

Closes #1308

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 02:38:47 -04:00
Tom Boucher
2d264da661 enhance(#968): region-scoped negative-grep idiom + cross-task conflict warning (#1320)
Adds a region/function-scoped negative-grep idiom to the gsd-planner verification guidance plus a warn-only `validate_plan` check (`scanFileWideNegativeGateConflict`) that flags when a task's file-wide negative grep bans a construct a sibling task legitimately requires elsewhere in the same file. ReDoS-safe (linear, no RegExp on author patterns); region-scoped gates are exempt. Warn-only — never errors, never flips `valid`.

Closes #968

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-16 02:20:34 -04:00
Tom Boucher
14e0709eed refactor(#1307): gate intel on isCapabilityActive + loop-resolver honors capability active (#1315)
* refactor(#1307): gate intel on isCapabilityActive + loop-resolver honors capability active

Part A: intel's command gate moves from config-only isIntelEnabled to the
shared isCapabilityActive('intel', cwd) — a consistency cutover (intel has
skills:[] so its tri-state collapses to the intel.enabled config leg; the gate
now flows through the resolver's precedence + runtime-aware resolution).
Part B: loop-resolver hook rendering now gates on capability state.active
(=== true, fail-closed) instead of state.enabled, so the capability config
gate is honored by the hook consumer, not just per-hook 'when'. active is now
required in the loop-resolver input types. Regression test proves a config-
disabled (active=false) capability's unconditional hook is not rendered.
Part of #1302.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1307): add changeset for intel + loop-resolver active gate

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 00:12:54 -04:00
Rezolv
e50ead7ad2 enhance(verify-phase): deterministic auto-locate of the prohibition check descriptor (#1278) (#1301)
* test(1278): RED-first descriptor parity + fail-closed guards + CHK-07 byte-stability (wave 1)

- CHK-03 (RED): extend PROB-14 parity in prohibition-probe.schema.test.cjs to carry the flat
  check_kind/check_target/check_rule scalars through project->write->parseMustHavesBlock; the
  non-droppable check_kind-presence assertion is the load-bearing RED trigger (fails because
  projectProhibitions strips check_* on the current build).
- CHK-07 (GREEN forward-guard): probe-core.test.cjs pins descriptor-less byte-stability +
  dispositionForProhibition fail-closed policy, with a t.todo marker forward-locking plan 01-02.
- CHK-06 (RED): prohibition-enforcement.test.cjs asserts descriptorFromProjection export +
  fail-closed on absent/partial/unknown descriptors via the projection adapter (RED until 01-03).
- No src/*.cts or .cjs edits; no new test files; lint-test-file-count clean.

* feat(1278): add optional flat-scalar check descriptor fields to Prohibition interface (wave 2)

- check_kind?/check_target?/check_rule? mirror CheckDescriptor.kind/target/rule (minus caller-attested failFirst, #1279)
- optional so existing Prohibition consumers compile unchanged

* feat(1278): project check descriptor as flat scalars in projectProhibitions (wave 2)

- emit check_kind/check_target (+ check_rule only for lint-rule with a rule) when descriptor well-formed
- under-specified/descriptor-less items project byte-identically (CHK-07); flat scalars ride existing parseMustHavesBlock continuation-KV path (no parser rewrite)
- add CHK-02 probe-core unit cases pinning the projection
- turns CHK-03 parity test GREEN; dispositionForProhibition untouched

* feat(1278): descriptorFromProjection read-back adapter feeds fail-closed locate (wave 3)

- Add descriptorFromProjection(projected) -> CheckDescriptor | null to
  src/prohibition-enforcement.cts: renames the projected flat scalars
  check_kind/check_target/check_rule -> {kind,target,rule?}, or null when
  the descriptor is absent/non-object (no check_kind key).
- failFirst is NEVER sourced from the projection (stays caller-attested; #1279).
- rule is set only when check_rule is a non-empty string; the adapter does NOT
  re-validate kind/target/rule — an under-specified descriptor reconstructs to
  one the EXISTING runProhibitionEnforcement LOCATE guard rejects (located:false,
  never green). The merged #1259 guard stays the single source of fail-closed truth.
- Turns the RED CHK-06 fail-closed tests (plan 01-01) GREEN end-to-end; CHK-03 /
  CHK-07 stay green. CheckDescriptor type, locate guard, dispositionForProhibition,
  and parseMustHavesBlock are unchanged (additive +36/-0).

* feat(1278): verify-phase locates prohibition check from projected descriptor (wave 3)

- request.check kind/target/rule sourced from projected check_kind/check_target/check_rule via descriptorFromProjection, not verifier invention (CHK-05)
- replaces the #1278 author-supplied / tracked-follow-up note with the delivered deterministic-locate behavior
- preserves fail-closed routing: absent/partial descriptor -> never green, hard-gate in both modes
- failFirst stays a verify-time caller attestation; #1279 bounds the remaining fail-first proof

* feat(1278): spec-phase captures wired-check descriptor on test-tier resolution (wave 3)

- Step 5.6 'Keep it' / verification: test path captures check_kind/check_target/check_rule, projected onto must_haves.prohibitions for verify-phase deterministic locate (CHK-04)
- SOFT capture: a test-tier prohibition without a descriptor is still allowed (no hard authoring block); stays fail-closed/flagged downstream
- --auto captures only an unambiguous descriptor, never fabricates a check path
- failFirst NOT captured at spec-phase (verify-time attestation; #1279)
- PROB-06 soft-gate + text-mode (PROB-09) behavior unchanged

* chore(1278): re-baseline workflow size for grown verify-phase + spec-phase prose (wave 3)

- spec-phase.md 28438 -> 30343 (+1905), verify-phase.md 35362 -> 36498 (+1136)
- regenerated via npm run size:baseline (no hand-picked numbers); growth is the #1278 deterministic-locate + descriptor-capture prose
- workflow-size-budget guard green (122/122)

* docs(1278): ratify optional check descriptor in dated ADR-550 addendum + type:Changed changeset

- Append dated 2026-06-15 ADR-550 addendum ratifying the D3 prohibition-item
  shape extension (optional flat-scalar check_kind/check_target/check_rule)
- Document flat-scalar rationale, deterministic projection/read-back,
  fail-closed on partial/invalid/absent, #1279/policy out-of-scope
- Add .changeset/1278-prohibition-check-descriptor.md (type: Changed)

* docs(1278): document optional check descriptor in prohibition-probe reference + FEATURES

- Add 'Optional wired-check descriptor (deterministic locate, #1278)' section
  to the prohibition-probe reference (flat-scalar keys, projection/read-back,
  fail-closed + backward-compat, failFirst stays attested)
- Add deterministic prohibition-check descriptor source entry to FEATURES.md
- No CONTEXT.md glossary change: descriptor reuses existing wired-check /
  verification:test vocabulary, no new glossary term introduced

* fix(1278): pass packaging gates — changeset pr field + retired slash-form fix

- Add required pr: 1278 to changeset (lint:changeset MISSING_PR hard requirement;
  plan's 'omit if unknown' was inaccurate — issue number per #1259 convention,
  updated to real PR number when opened) [Rule 3 - blocking]
- Fix retired /gsd-spec-phase -> /gsd:spec-phase at verify-phase.md:83 (wave-3
  prose; caught by slash-namespace invariant #3443/bug-2543, blocked CHK-09
  full-suite-green) [Rule 1 - bug]
- size:baseline + INVENTORY manifest verified in-sync post-build (no diff)

* docs(1278): add check descriptor + descriptorFromProjection to CONTEXT.md prohibition glossary

* fix(1278): harden descriptorFromProjection round-trip (numeric-coercion + stray-rule) per review

- MD-01/LW-01: narrow projected scalars to primitives + String()-coerce, so a
  numeric-looking check_target (parseMustHavesBlock coerces ^\d+$ to number)
  reconstructs as a string and locates instead of silently un-locating; no
  as-string type-lie, satisfies no-base-to-string.
- LW-02: attach rule only for the lint-rule kind (drop a stray node-test rule).
- LW-03: document the optional check_* keys in the reference Output schema.
RED->GREEN tests added in prohibition-enforcement.test.cjs.

* chore(1278): set changeset pr to 1301

* test(1278): add fast-check property for the check-descriptor round-trip + fail-closed (trek-e review)

RULESET.TESTS.property-based-testing: the projectProhibitions -> render ->
parseMustHavesBlock -> descriptorFromProjection chain is a bijective/transformation
contract. Adds 2 fc properties to tests/probe-core.property.test.cjs (no new file;
ratchet stays at 2 for probe-core):
- well-formed descriptors survive the round-trip across the full string domain
  incl. the numeric-coercion case (target/rule reconstruct as strings);
- under-specified/invalid descriptors (absent / target-less / rule-less /
  unknown-kind) are always fail-closed (never green, flagged, unlocated).
Stability is asserted at the descriptorFromProjection layer (the raw parse step is
intentionally lossy for numeric scalars; the shared parser is unchanged).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-16 00:01:03 -04:00
Tom Boucher
e3b829e765 refactor(#1306): gate graphify on isCapabilityActive (tri-state, runtime-aware) (#1313)
* refactor(#1306): gate graphify on isCapabilityActive (tri-state), not config-only

graphify's command gate moves from the config-only isGraphifyEnabled to the
shared isCapabilityActive('graphify', cwd) — so graphify is off unless installed
AND surfaced AND graphify.enabled. Fixes a latent claude-hardcoding in the
resolver: resolveCapabilityRuntimeState now detects the active runtime via
resolveRuntime(cwd) (GSD_RUNTIME -> config.runtime -> 'claude') so non-Claude
runtimes (Codex/Cursor) read their own surface, not ~/.claude. Hermetic
regression test proves config-on+unsurfaced -> disabled; cross-runtime test
proves GSD_RUNTIME=codex honors CODEX_HOME. Gate fails closed on every error
path. Part of #1302.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1306): add changeset for graphify tri-state gate

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 23:19:21 -04:00
Tom Boucher
e9ee7e9ba6 feat(#1305): per-capability active tri-state + isCapabilityActive in Capability State Resolver (#1311)
* feat(#1305): add per-capability active tri-state + isCapabilityActive to the Capability State Resolver

CapabilityStateEntry gains active = enabled && configActivation, where
configActivation resolves the capability's optional activationKey via the
shared _resolveActivationValue (absent activationKey -> true). enabled stays
installed && surfaced (unchanged). Each hook's active now also cascades the
capability config gate (active && configured). Adds isCapabilityActive(capId,
cwd) — a thin convenience over resolveCapabilityRuntimeState. cmdCapabilityState
emits active per capability. No consumer cutover yet (graphify/intel: #1306/#1307;
loop-resolver: #1310). Part of #1302.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1305): add changeset for capability active tri-state

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 22:13:20 -04:00
Tom Boucher
1f41a0ce9a feat(#1304): add optional activationKey capability manifest field (#1309)
Add an optional activationKey to the feature role of capability.json — the
dotted config key that gates the whole capability (e.g. graphify.enabled).
gen-capability-registry validates it (non-empty string, reserved-name guard,
must be declared in the capability's own config slice, feature-only) and emits
it per-capability in the generated registry. Declared on graphify + intel.
No runtime consumption yet (resolver wiring lands in #1305). Part of #1302.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 21:33:06 -04:00
Tom Boucher
91bd82f9a6 enhance(execute): isolated-executor rejected/over-reaching run fails safe — never default recovery to main (#1292) (#1303)
* enhance(execute): isolated-executor rejected/over-reaching run fails safe (#1292)

When an isolated (worktree) executor run is rejected — the user declines to
merge it, the orchestrator surfaces recovery for a blocked/halted plan, or the
run over-reached the requested scope — the orchestrator must no longer
default/propose recovery by editing the primary checkout (`main`). Absent an
explicit guardrail, the LLM orchestrator could improvise "continue on main",
inverting the isolation contract at the moment it matters most.

Added an ISOLATED-RUN RECOVERY — FAIL SAFE policy: default to a safe halt that
offers a fresh, narrowly-scoped worktree or inspect/discard; editing the primary
checkout requires explicit, clearly-labeled confirmation and is never the
default/proposed option.

To respect the ADR-857 phase-6 host-loop size cap on execute-phase.md (it sits
just under the pre-phase-6 baseline), the policy is delivered as an extracted
reference fragment rather than inline:
- New `execute-phase/steps/worktree-recovery-policy.md` holds the recovery policy
  (the existing FAIL-CLOSED rule #48 for base/HEAD mismatches + the #1292
  fail-safe guardrail). No #48 behavior change — moved verbatim.
- execute-phase.md references the fragment at the worktree-spawn recovery point,
  the step-5.5 merge decision, and the stalled-agent "switch to inline execution"
  menu (which for an isolated run now follows the fail-safe policy). Net effect:
  execute-phase.md shrinks below its cap.
- quick.md carries the fail-safe guardrail inline at its post-return merge/discard
  decision (quick.md is not size-capped).

Scoped to the recovery offer only — no automatic scope-overreach detection
(explicitly out of scope per the issue) and no new config key. Adds content
regression tests, a USER-GUIDE note, and a workflow size-baseline update.

Closes #1292

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(changeset): Changed fragment for #1292 isolated-executor fail-safe recovery

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 20:53:38 -04:00
Tom Boucher
137760a655 fix(#1296): align config docs/prompts/schema with consumers (#1299)
* fix(#1296): align config docs/prompts/schema with consumers

The user-facing config surface disagreed with what the consumers actually do
(subset of the #1216 audit). No runtime consumption behavior changes.

- workflow.subagent_timeout: settings-advanced.md prompt + docs/CONFIGURATION.md
  said "seconds (default 600)" but the consumer (map-codebase.md) uses
  milliseconds (default 300000). Relabeled all four spots in settings-advanced.md
  (prompt, parse-default list, example, confirmation table) + the CONFIGURATION.md
  row.
- review.models.<cli>: settings-integrations.md, docs/CONFIGURATION.md (Integration
  Settings), and docs/CLI-TOOLS.md documented a shell command, but review.md injects
  the value into a --model/-m flag. Relabeled to a bare model id and reconciled the
  contradictory CONFIGURATION.md sections.
- workflow.test_command + workflow.build_command: consumed via config-get
  (test_command in verify-phase/execute-phase/audit-fix/post-merge-gate;
  build_command in post-merge-gate) and documented, but absent from validKeys so
  `config set` rejected them. Registered both in config-schema.manifest.json and
  documented them in references/planning-config.md (overview + complete reference).

Regression tests: behavioral config-set tests (tests/config.test.cjs) + doc-parity
content guards (tests/config-field-docs.test.cjs).

Deferred to other #1216 clusters: security-gate wiring, search_gitignored wiring,
mvp_mode, source_grounding_authority labeling, and config-set enum enforcement.

Closes #1296
Refs #1216

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(changeset): Fixed fragment for #1296 config-surface alignment

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 19:57:11 -04:00
Tom Boucher
8c3d934a90 refactor(#1294): T-final — delete the core.cjs re-export spine (epic #1267 complete) (#1295)
* refactor(#1294): T-final — delete the core.cjs re-export spine (epic #1267 complete)

After T0–T6 nothing imports core, so retire the spine and its scaffolding:
- delete src/core.cts (and the gitignored gsd-core/bin/lib/core.cjs artifact;
  remove its .gitignore + eslint-ignore entries)
- delete scripts/lint-core-spine-imports.cjs + its allowlist; drop it from the
  package.json lint:ci chain
- regenerate docs/INVENTORY-MANIFEST.json (drops the core.cjs surface)
- sweep stale references: CONTEXT.md glossary back-compat clauses (spine retired,
  callers import the leaf directly), planning-config.md CONFIG_DEFAULTS owner,
  and false present-tense core.cjs claims in leaf-module docstrings

The ADR-857 decomposition is complete: the former Core god-module is fully
dissolved into its leaf modules; no re-export spine remains. No behaviour change.

Closes #1294

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1294): migrate the computed-path core.cjs importers the literal grep missed

bin/install.js used require(path.join(_gsdLibDir, 'core.cjs')) (a computed
path, and bin/install.js was never in the convergence lint's scan roots), and
~8 test files referenced core.cjs via path.join/readFileSync/existsSync/FILE_ARG
forms the literal-string migration grep missed. Route install.js's symbols to
their leaves (RUNTIME_PROFILE_MAP->model-catalog, resolveTierEntry/EFFORT_SET->
model-resolver) and repoint/adjust the test references to the leaves. Recovers
the 161 'Cannot find module core.cjs' failures from the spine deletion.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 18:50:46 -04:00
Tom Boucher
c76827afbc refactor(#1291): T6 — migrate test files off the core spine ahead of deletion (#1293)
The convergence lint only scanned src/ + gsd-core/bin, so ~35 test files
still imported core.cjs. Repoint all 33 behaviour importers to the leaf
modules directly (same symbol->leaf map as the src migration; leaves are the
objects core re-exported by reference), delete the now-meaningless
shim-identity describe blocks in the 8 leaf tests, and delete tests/core.test.cjs
(forwarded-behaviour coverage now lives at the leaves; resolveWorktreeRoot
test relocated to worktree-safety in T0) and tests/lint-core-spine-imports.test.cjs
(the lint is removed in T-final). Dropped the stale core.test.cjs entries from
the allow-test-rule-refs allowlist; eslint-rules RuleTester fixture path
pointed at io.cjs.

After T6: ZERO test imports core.cjs. core.cts still builds (now fully unused);
T-final deletes it. No behaviour change.

Closes #1291

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 18:11:20 -04:00
Tom Boucher
b108f101b0 fix(#1284): grant mcp__perplexity__* to researcher agents + dispatch-table parity guard (#1288)
Adds mcp__perplexity__* to both researcher profiles (generated source-of-truth) and regenerates the agents; adds a generative dispatch-table↔tools parity guard so future provider drift fails CI. Regenerates the agent-size baseline for the +20-byte frontmatter growth.

Fixes #1284

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 17:20:08 -04:00
Tom Boucher
13414cb168 Merge pull request #1273 from davesienkowski/feat/1259-test-tier-enforcement
enhance(verify-phase): enforce test-tier prohibitions — wire the deferred negative-test gate (#1259, ADR-550 D5d)
2026-06-15 15:54:16 -04:00
Tom Boucher
54420bae9e refactor(#1277): T1 — decouple agent-install-check + git-base-branch from the core spine (#1280)
First leaf-migration tranche of epic #1267 (after T0 #1268). Migrate the
via-core callers of the two leaves T0 created to import from the leaf
modules directly, and stop core re-exporting their symbols:
- checkAgentsInstalled: docs.cts, verify.cts, init.cts -> agent-install-check.cjs
- gitWorktreeInfoInternal: init.cts -> git-base-branch.cjs
- getAgentsDir had no external via-core caller (internal to the leaf)

core no longer re-exports getAgentsDir / checkAgentsInstalled /
gitWorktreeInfoInternal; the now-unused agent-install-check + git-base-branch
requires are dropped from core; the shim-identity assertions for these are
deleted (behaviour tests retained). Convergence lint stays green (0 new).
No behaviour change.

Closes #1277

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 15:46:41 -04:00
Dave
0d783fa57d Merge remote-tracking branch 'origin/next' into feat/1259-test-tier-enforcement
# Conflicts:
#	tests/workflow-size-baseline.json
2026-06-15 15:42:23 -04:00
Tom Boucher
0a856f06cd enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier (#1271)
* enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier

Introduce a per-truth PRESENT_BEHAVIOR_UNVERIFIED state for must-haves that assert a state transition or a cancellation/cleanup/ordering invariant whose only evidence is symbol presence + wiring. Such truths are excluded from the verified_truths score, reported as a behavior_unverified count, recorded in an always-on behavior_unverified_items frontmatter list, and routed to the existing human_needed sink — so a clean N/N can no longer be reached on symbol presence alone.

The overall-status vocabulary and the src/verification.cts seam are unchanged (the new state is per-truth only); gaps_found keeps decision-tree precedence; override-passed truths still count toward verified_truths. Mirrors the calibration into the shipped verify-phase.md workflow (with an infra/foundation carve-out), the VERIFICATION.md templates, and docs (planning-artifacts.md, AGENTS.md). gsd-verifier.md kept under its 48KB LARGE cap; size baselines regenerated.

Closes #966

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#966): add changeset fragment for gsd-verifier behavior-unverified calibration

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 15:24:56 -04:00
Dave
95e5abfb66 Merge remote-tracking branch 'origin/next' into feat/1259-test-tier-enforcement 2026-06-15 15:20:54 -04:00
Dave
6a9e6cb0ae fix(1259-01): address trek-e review — B1 (fatal/suppressed), B2 (bounded subprocess), M1/M2, minors
Maintainer CHANGES_REQUESTED (reviewed e08667e5, pre-portability-fix):
- B1: lint-rule no longer greens an unparseable target (eslintHasFatalError -> fail
  closed on any fatal/parse error) or an inline-suppressed violation (eslintJsonHasRule
  now scans suppressedMessages too). RED-first + real-runner repros.
- B2: both child spawns get a bounded timeout (30s node / 60s eslint) + 16MiB maxBuffer;
  timeout fails closed. Injectable timeoutMs (positive-only — 0/negative can't disable
  the bound) enables a fast 1.5s hang test.
- M1: scoped verify-phase.md — the check descriptor is author-supplied for now; filed
  #1278 for deterministic auto-locate of the descriptor (the locate half).
- M2: added tests/prohibition-enforcement.property.test.cjs (fast-check fail-closed
  invariants; within the <=2-file budget).
- m1: tapTestNames excludes # SKIP/# TODO; parseNodeTestSummary tracks # cancelled;
  isNonVacuousNodeTestPass requires cancelled===0.
- m2: scoped the determinism claim to the decision/parse layer (real runner is env-dependent).
- m3: filed #1279 for machine-proven fail-first (violation-fixture probe).
- n1: -- before target in both arg builders (option-injection). n2: dropped dead token.
- B3 (Windows npx) was already fixed in 2af76306 (pushed ~65s after the review).

Verified node 22 + 24; size baseline regenerated for the verify-phase note.
2026-06-15 15:19:04 -04:00
Tom Boucher
48d9cec6fe refactor(#1268): re-home core re-export-spine squatters + migration-convergence lint (#1272)
Re-home the 6 implementation functions squatting in the core.cjs re-export
spine (ADR-857) into the modules whose interface they belong to, with core
re-exporting them BY REFERENCE so all 32 callers + the shim-identity tests
keep resolving unchanged:
- worktree-safety: resolveWorktreeRoot, pruneOrphanedWorktrees
- git-base-branch (broadened to the Git Query Module): gitWorktreeInfoInternal
- agent-install-check (new leaf): getAgentsDir, checkAgentsInstalled
- delete the _resetRuntimeWarningCacheForTests wrapper; consumers use a
  shared resetRuntimeWarningCaches() helper in tests/helpers.cjs

Add scripts/lint-core-spine-imports.cjs (migration-convergence lint with a
30-importer allowlist, wired into lint:ci) so the staged spine retirement
provably converges: CI fails on any new ./core import. Register the new
generated agent-install-check.cjs in eslint-ignore + .gitignore +
INVENTORY-MANIFEST.json.

No behaviour change. First tranche (T0) of epic #1267.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 15:08:58 -04:00
Dave
564f1da3ad test(1259-01): pin target-side basename normalization (round-4 WR-R4-01)
A mutant dropping baseOf() from the TARGET side of the node-test vacuity compare
survived — add the mirror case (relative TAP name vs absolute descriptor target,
both basename-normalized -> vacuous).
2026-06-15 14:27:29 -04:00
Dave
2af7630642 fix(1259-01): make real runner cross-platform/version portable (CI ubuntu-24 + windows-24)
CI failed on node 24 + Windows in the real-runner E2E tests — real portability
bugs in the shipping runner, not test flakiness:

- node 24 names a zero-test file's TAP entry by ABSOLUTE path (node 22 used the
  basename), so the exact-string vacuity discriminator misfired -> empty file
  falsely greened. Now compares BASENAMES (separator-agnostic) -> robust across
  OS + node version. Verified: empty file is non-green under node 24 locally.
- the lint-rule runner spawned 'npx', which execFileSync can't launch on Windows
  (and failed fast on ubuntu-24). Now resolves the project's eslint CLI via
  eslint/package.json and runs 'node <cli>' through process.execPath (portable;
  no shell -> no injection). eslint absent -> fail-closed.
- node-test also spawns via process.execPath, not bare 'node'.

Also folds in round-3 WR-01: the probe-core inline comments still said
enforcement was 'deferred to a follow-up PR' — corrected (landed in #1259).
Strengthened the isNonVacuousNodeTestPass unit test to pin the basename clause
(absolute-path target). Verified 71/71 green under BOTH node 22 and node 24.
2026-06-15 14:19:35 -04:00
Dave
e08667e5aa test(1259-01): regenerate workflow-size baseline for verify-phase.md growth (+1459 B)
verify-phase.md grew 30875 -> 32334 from the test-tier enforcement consumer step
+ descriptor docs. Well under the 40960 DEFAULT hard cap; regenerate the per-file
baseline ratchet (the growth is the load-bearing consumer wiring, not bloat).
2026-06-15 14:00:03 -04:00
Dave
97931249bc fix(1259-01): close re-review NEW-BL-01 (ignored-target vacuous green) + no-throw hardening
Round-2 adversarial review found the lint-rule runner falsely greened an
eslint-IGNORED target: eslint returns a length-1 "File ignored" result (ruleId
null) that passed the >=1-file vacuity guard while nothing was linted — reopening
the vacuous-green class. Fix: buildLintArgs now passes --no-warn-ignored so an
ignored path returns [] -> fails closed (verified + E2E test on an ignored
bin/lib artifact).

Also: wrap runCheck() so even a (test-injected) throwing runner fails closed
(NEW-WR-01, full no-throw contract); document the benign basename-naming
constraint on wired node-test names (NEW-WR-02, fail-closed).
2026-06-15 13:45:03 -04:00
Dave
31b822b312 fix(1259-01): close adversarial review findings — genuine enforcement, honest fail-first scope
Adversarial pre-submission review found the injected-runCheck tests masked a
non-functional real runner. Fixes:

- BL-01 (false green on vacuous test): the node-test runner now parses the TAP
  summary and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail) AND a
  reported test named distinctly from the file — node --test counts an empty file
  as one passing test, so counts alone could not catch it.
- SF-01 (lint anchor never greened): the lint-rule runner now runs the project
  eslint as --format json and filters by ruleId, so plugin rules (local/*) load
  via the flat config — bare --rule cannot load a plugin. local/no-source-grep
  now genuinely greens (covered by a real, non-injected test).
- BL-02 (tautological fail-first): the runner no longer echoes the caller's
  failFirst as if confirmed. failFirst is documented as caller-ATTESTED; the
  producer requires attestation + a genuine non-vacuous pass. Machine-proven
  fail-first (needs a violation fixture) is flagged as a tracked follow-up in
  ADR-550, the changeset, FEATURES, the reference doc, and verify-phase.
- SF-02: added real-runner end-to-end tests (no injected runCheck) + pure,
  exported parse/filter helpers (parseNodeTestSummary, tapTestNames,
  eslintJsonHasRule, eslintFileResultCount) so the shipping branches are
  mutation-pinned.
- NIT-01/02: LOCATE guard rejects empty-string rule and unknown kinds.
- Hardening: spawn checks with NODE_TEST_CONTEXT/NODE_OPTIONS scrubbed so an
  ambient test-runner context cannot corrupt a verify-time result.
- Docs reconciled to the shipped behavior (no 'confirms fail-first' overclaim).
2026-06-15 13:38:15 -04:00
Dave
676436258a fix(1259-01): lint-rule real runner — keep rule id distinct from lint target
The default lint-rule runner passed check.target as BOTH the --rule id and the
eslint path, so it could never pass (eslint tried to lint a file named after the
rule). Add a distinct check.rule field (rule id) vs check.target (path to lint),
extract a pure exported buildLintArgs() so the mapping is mutation-testable
without spawning eslint, fail-closed on a lint-rule missing its rule id, and carry
the rule into enforcement evidence. Updates verify-phase descriptor docs.
2026-06-15 13:10:49 -04:00
Dave
1919c2fd8f test(1259-01): RED — lint-rule real runner must keep rule id distinct from lint target 2026-06-15 13:08:33 -04:00
Dave
8f428bdd41 feat(1259-01): deterministic prohibition-enforcement producer + check route
- Author src/prohibition-enforcement.cts: the test-tier prohibition PRODUCER/gate (ADR-550 D5d heavy half)
  locate wired check (node-test|lint-rule) -> confirm fail-first -> run -> build enforcementEvidence -> dispositionForProhibition
  pure/deterministic with injectable runCheck; missing/failing/non-fail-first -> hard-gate (both modes); passing -> green
- Route check prohibition-enforcement in src/check-command-router.cts (same family as ui-plan-gate / tdd-review-checkpoint)
- Add tests/prohibition-enforcement.test.cjs (behavioral, typed-field, injected runner) — new module within <=2 budget
- Register the built bin/lib surface: docs/INVENTORY.md row + regenerated docs/INVENTORY-MANIFEST.json
- .gitignore: add the emitted gsd-core/bin/lib/prohibition-enforcement.cjs (build artifact, ADR-457)
- No src/probe-core.cts edit — the green/fail-closed policy seam already exists
2026-06-15 12:44:29 -04:00
Dave
6115ab216d test(1259-01): RED — test-tier enforcement (both check kinds) before producer exists
- Extend prohibition-probe.verify-tier.test.cjs with the ENFORCEMENT half (#1259, ADR-550 D5d)
- Require the not-yet-built gsd-core/bin/lib/prohibition-enforcement.cjs (RED)
- Cover both wired-check kinds (node-test + no-source-grep lint-rule) and miss/fail hard-gate
- Typed-field assertions only; injected runCheck (no real subprocess)
- Keep the 2 original fail-closed tests verbatim
2026-06-15 12:38:48 -04:00
Tom Boucher
cf68841220 enh(#1243): consume Claude plugin-provided skills in agent_skills (epic #1258 Phase B) (#1261)
* feat(#1243): consume Claude plugin-provided skills via native Skill-tool directive + grant Skill to agent_skills-consumer agents

- Relax global skill name validation to accept namespaced form `^[A-Za-z0-9_-]+(:[A-Za-z0-9_-]+)*$`
- Namespaced names (containing colon) on claude runtime emit a Skill-tool load directive instead of a @-include line
- Namespaced names on non-claude runtimes are skipped with a warning
- Bare unresolved names retain existing warn-and-skip behavior (no promotion to directive)
- Grant `Skill` tool to all 22 agent_skills consumer agents; 5 generated agents updated via research-profiles.cjs + regen, 17 hand-authored agents edited directly
- Add 16 TDD tests in describe('bug #1243') covering happy/mixed/precedence/negative/cross-runtime/regression/grant cases

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#1243): document plugin-provided skills in agent_skills

Update the Agent Skills Injection reference in CONFIGURATION.md with
the three entry forms (project-relative, global:<name>,
global:<plugin>:<skill>), the Claude-only runtime behaviour of the
namespaced form and the warn-skip on other runtimes, the plugin
pre-install prerequisite, and the consumer-agent Skill tool grant.

Add docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md with a
step-by-step guide for installing the plugin, locating the namespaced
skill name, wiring it into agent_skills, and verifying injection.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1243): align agent_skills docs with emitted block format + mixed-block regression test (code-review)

- Replace two-section mixed-block example (bogus "Load these plugin-provided skills using the Skill tool:" header) with the actual single-section inline format in CONFIGURATION.md and docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md
- Fix quoted warning text in how-to doc to exactly match the emitted string: [agent-skills] WARNING: Plugin-namespaced skill "global:<name>" requires a Skill-tool-capable runtime (claude) — skipping on runtime "<runtime>"
- Replace phantom agent slugs (gsd-checker, gsd-researcher, gsd-advisor, gsd-synthesizer) in CONFIGURATION.md Supported Agent Types with real agents/gsd-*.md examples (gsd-plan-checker, gsd-phase-researcher, gsd-code-reviewer, gsd-ui-auditor, gsd-research-synthesizer)
- Add byte-identical mixed-block regression test: one path-resolvable global skill + one plugin-namespaced skill on claude runtime → asserts r.ir.block === single-section interleaved block, no secondary header

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1243): regenerate agent-size baseline for the Skill-tool grant

The 22 agent_skills-consumer agents each grew +7 bytes from adding `Skill`
to their tools list; refresh the committed per-agent size baseline (#1074 guard).

* chore(#1243): add Added changeset fragment

* fix(#1243): traceable allow-test-rule ref + separator-agnostic byte-identical tests (CI)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-15 00:49:12 -04:00
Tom Boucher
00c447701d fix(#1257): update pipe-table Status/Phase/Plan cells in planned-phase + begin-phase (#1260)
* fix(#1257): update pipe-table Status/Phase/Plan cells in planned-phase + begin-phase

cmdStatePlannedPhase ran its body-field replacements on the full file content,
so the case-insensitive ^Status: pattern matched the YAML frontmatter `status:`
line before the body `| Status | … |` cell — the cell never advanced to
'Ready to execute' and syncStateFrontmatter re-derived the stale 'planning'
status (the #1230 delta heuristic preserved it). cmdStateBeginPhase had
pipe-table else-branches only for Status/Last activity (#1256), so the Current
Position `| Phase |` / `| Plan |` cells were left stale while a spurious inline
`Phase: N — EXECUTING` line was prepended.

Both handlers now strip frontmatter before body-field replacement and update the
pipe-table cells in place via stateReplaceField, matching inline-format
behaviour. Systemic residual of #1255 / #1256.

Adds 4 regression tests (#1257 block in tests/state.test.cjs) covering both
findings — RED before the fix, GREEN after; full state-area suite (437) and
local unit suite stay green.

Closes #1257

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1257): backfill changeset pr number (#1260)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 00:29:18 -04:00
Tom Boucher
2668fbbeb0 fix(#1255): advance frontmatter status for pipe-table STATE.md on begin/complete-phase (#1256)
* fix(#1255): advance frontmatter status for pipe-table STATE.md on begin/complete-phase

state begin-phase/complete-phase called stateReplaceField on the FULL file content, so its case-insensitive ^Status: plain-pattern matched the YAML frontmatter status: line first (no g flag) and never updated the body pipe-table Status cell; syncStateFrontmatter then re-derived the stale status from the unchanged body, freezing the frontmatter status. Fix: strip frontmatter before the body-field replacements (operate on body only), reassemble with frontmatter preserved, so the body Status cell updates and the frontmatter derives correctly — for inline AND pipe-table body formats. Also corrects the Current Position pipe-table else-branches (Status/Phase/Last-activity) to write bare, consistent cell values. Pipe-table Status is a supported body format (not rewritten to inline).

Closes #1255

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1255): add changeset for pipe-table state status fix (#1256)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1255): fold pipe-table regression into state.test.cjs + Windows-portable frontmatter regex

Per the 2026-06 audit, new tests/bug-NNNN-*.test.cjs files are banned (lint-regression-test-names) — folded the 7 #1255 regressions into tests/state.test.cjs and removed the standalone file + its lint-test-file-count allowlist entry. Also fixed the frontmatter assertions' /^---\n/ anchors to /^---\r?\n/ (windows-test-parity-guard frontmatterAnchorLiteralNewline).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 22:42:28 -04:00
Rezolv
3556450b0d feat(spec-phase): surface zero-classification edge-probe requirements as unclassified candidates (#1110) (#1117)
Edge-probe now surfaces a zero-classification requirement (non-empty prose, no
shape cue matched, no `shapes` override) as a single soft `unclassified — review
manually` candidate instead of silently dropping it — the exact blind spot the
probe exists to catch. Dismissible like any edge; the `shapes: []` opt-out stays
silent; `TAXONOMY` (the closed 8 categories) is unchanged. Under `--auto` the
candidate is left `unresolved`, never auto-`backstop` (a missing shape is not
evidence an edge exists).

Closes #1110
2026-06-14 21:44:21 -04:00
Rezolv
395fb519e7 feat(spec-phase): prohibition probe — surface "must-NOT" constraints (#644) (#1149)
Adds the spec-time prohibition probe (spec-phase Step 5.6) — the second adapter of the probe-core resolution model. Surfaces unwritten must-NOT constraints as negative SPEC acceptance criteria with test/judgment verification tiers; fail-closed at verify time. Per ADR-550. Closes #644.
2026-06-14 21:29:11 -04:00
Tom Boucher
b783410815 refactor(#1191): inject clock/reset testability seams + handle valid-null settings (#1233)
* refactor(#1191): inject clock/reset testability seams + handle valid-null settings

- worktree-safety reapOrphanWorktrees: injectable deps.nowMs clock for deterministic stale-lock boundary tests (mirrors snapshotWorktreeInventory's options.nowMs).

- active-workstream-store: _resetControllingTtyCacheForTests() seam clears the memoized controlling-TTY probe cache; test replaces require.cache busting.

- gen-capability-registry: export stripGeneratedComment (additive); test imports the real helper + equivalence assertion, keeping the deliberate drift oracle.

- install.js readSettings: a successfully-parsed JSON null is treated as empty settings ({}) instead of being mis-reported as malformed; genuine parse failures still warn. readSettings/stripJsonComments exported (GSD_TEST_MODE-guarded require) for real behavioral tests.

Closes #1191

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1191): add changeset for valid-null settings fix (#1233)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1191): replace Stryker-incompatible structural reset test with behavioral isTTY-spy

The seam-2 reset test read the BUILT active-workstream-store.cjs and grepped for 'didProbeControllingTtyToken = false' — Stryker instruments that file so the literal is absent, failing the mutation DRY RUN. Replaced with a behavioral test that spies on process.stdin.isTTY access count to prove a post-reset probe re-runs (kills the didProbe-reset mutant) without reading source text. Local stryker: dry run passes, score 85.21% >= 80.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 20:52:56 -04:00
Tom Boucher
55eb1fd72e refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam (#1242)
* refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam

ADR-22's severity mapping, authority auto-upgrade, and rung>=3 hard-block lived only as prose in plan-review-convergence.md — untestable. Extracted into src/plan-drift-guard.cts (pure: AUTHORITY_RUNGS, getEffectiveAuthority, classifyDriftSeverity) + a gsd-tools drift-guard CLI seam (authority/severity), and rewired the workflow to call the seam deterministically instead of reasoning the decision in prose. 47 unit/e2e/structural tests cover the full severity table, the grep->intel auto-upgrade, and the rung>=3 HIGH hard-block.

Closes #1190

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1190): add changeset for ADR-22 drift-guard seam (#1242)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1190): register ADR-22 module in eslint-ignore + inventory manifest + docs-exempt changeset

Full-matrix CI surfaced new-module/command governance ripples beyond the lint-tests chain: (1) tsc-generated plan-drift-guard.cjs must be in the eslint ignore list (551-eslint-bin-lib-coverage); (2) docs/INVENTORY-MANIFEST.json must include the new module/command (regen via gen-inventory-manifest.cjs --write); (3) a type:Added changeset triggers docs-required — added a docs-exempt marker (internal seam, no user-facing surface).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 20:52:36 -04:00
Tom Boucher
6ef78fe417 feat(#1188): branch-coverage floors + promote no-source-grep to error on tests/** (#1250)
* feat(#1188): promote no-source-grep to error on tests/**

Apply local/no-source-grep as an error to tests/**/*.test.cjs (was warn on bin/scripts only). Only 2 real violations surfaced (repo-layout.test.cjs structural guard-placement checks on bin/install.js) — marked allow-test-rule with #1188 reason. PART A (branch-coverage floor) follows separately.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1188): add c8 branch-coverage floors (60 global, 70 per-file on UNMUTATED modules)

c8 enforced line floors only. Add a --branches 60 global floor (current ~82.5%, margin mirrors the lines-70-vs-91% gap) to test:coverage + test:coverage:unit, plus a chained 'c8 check-coverage --per-file --branches 70' for the high-risk UNMUTATED modules state/phase/verify/init (current min 78.3% on verify). The per-file check reuses the coverage data the suite run just produced (the proven scripts-floor pattern) so it adds no second suite run.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1188): per-module branch check must pin lines/funcs/stmts to 0

Standalone 'c8 check-coverage' defaults unspecified metrics to 90 and enforces them, so --branches 70 alone also failed on lines (verify.cjs 85.29% < default 90). Pin --lines 0 --functions 0 --statements 0 so only the branch floor (70) is enforced on the UNMUTATED modules. Data-reuse confirmed working (it computed verify's real %).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 20:52:13 -04:00
Tom Boucher
6724581285 fix(#1230): preserve frontmatter status/stopped_at when a state write doesn't change the body source (#1252)
* fix(#1230): preserve frontmatter status/stopped_at when a state write doesn't change the body source

Adds a delta heuristic in `readModifyWriteStateMd`: snapshots the body
Status and Stopped At fields before the transform, then after
syncStateFrontmatter runs, restores the existing frontmatter values for
any field whose body source was not changed by this write. Integrates
cleanly with the existing !resync progress-restore block (computes postFm
once, applies both restorations, reconstructs frontmatter once).

Regression tests added to bug-905 test file (same theme, no new file).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1230): add Fixed changeset fragment

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-14 20:51:53 -04:00
Tom Boucher
8da48e22ae fix(#1229): count bullet-only phases so phase.add stops reusing an existing phase number (#1249)
* fix(#1229): count bullet-only phases + guard against number collision in phase.add

Before this fix, the phase.add number scan only checked ### Phase N: section
headers and on-disk phases/N-* directories. A phase that existed only as a
roadmap bullet (e.g. "- [ ] **Phase 11: ...**") was invisible to both scans,
causing phase.add to silently assign a duplicate number.

Fix: add a bullet-entry regex scan (all checkbox variants: [ ], [x], [~], with
or without ** bold markers) to the set-based phase-number collection in
cmdPhaseAdd. Also added a post-compute collision guard that advances the
candidate past any already-used number.

Regression tests added to tests/phase.test.cjs (bug #1229 describe block):
bullet-only collision, [x]/[~] variants, plain-bullet, and baseline preservation.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1229): add Fixed changeset fragment

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-14 17:52:15 -04:00
Tom Boucher
10ae85cbbf fix(#1238): keep node resolvable in bug-891 PATH isolation so home-fallback subtests don't false-fail when node co-locates with a global gsd-tools shim (#1247) 2026-06-14 17:18:17 -04:00
Tom Boucher
1fa7bc594c refactor(#1190): extract ADR-230 PR-target branch policy into a tested, fork-safe seam (#1246)
ADR-230's branching-model gate (pr-target-validator.yml) decided allowed/blocked PR targets via inline regex in github-script — untestable. Extracted the decision into committed scripts/pr-target-policy.cjs (classifyPrTarget(base,head)->{decision}), and rewired the workflow to checkout the BASE ref (trusted; fork-tamper-safe) + require the module. Behavior-identical (Codex-verified char-by-char regex equivalence + all side-effects preserved). 70 tests incl. an equivalence oracle battery + hyphen-boundary negatives. Added contents:read for the checkout. Re-attribution: no ADR-230 test references exist (issue's '2 misattributed files' claim not borne out).

Closes #1190

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 17:16:19 -04:00
Tom Boucher
00acbc8868 fix(#1223): install scripts/fix-slash-commands.cjs so gsd-tools loads (#1240)
* fix(#1223): install scripts/fix-slash-commands.cjs so gsd-tools loads

Before this fix, bin/install.js copied scripts/changeset/ and scripts/lib/
into the runtime config dir but omitted scripts/fix-slash-commands.cjs.
gsd-core/bin/lib/command-roster.cjs requires this file at module load via
require('../../../scripts/fix-slash-commands.cjs'), so every gsd-tools
command crashed with MODULE_NOT_FOUND on every installed runtime.

Four changes:
- bin/install.js copy step: copy fix-slash-commands.cjs into <configDir>/scripts/
  with source-missing hard-fail and verifyFileInstalled smoke check
- bin/install.js writeManifest: track scripts/fix-slash-commands.cjs (not
  covered by the changeset/lib subdir loops)
- bin/install.js uninstall: best-effort unlinkSync before scripts/ rmdir
- scripts/fix-slash-commands.cjs readCmdNames(): wrap readdirSync in
  try/catch returning [] so skill-based/global installs without a local
  commands/gsd/ directory do not throw ENOENT

Tests added to tests/install.test.cjs (6 new tests):
- smoke: install() copies fix-slash-commands.cjs
- e2e: spawned gsd-tools.cjs does not crash with MODULE_NOT_FOUND
- manifest: writeManifest() tracks the file
- uninstall: uninstall() removes the file
- readCmdNames unit: export returns an array
- readCmdNames spawn: absent COMMANDS_DIR returns exit 0 (no throw)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1223): backfill changeset PR number (#1240)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 16:03:26 -04:00
Tom Boucher
0b3a2e5f9c feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15) (#1237)
* feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15)

ADR-15 designates /gsd-progress --next --auto --converge as the PRIMARY plan-convergence surface, but only the secondary surface (autonomous.md) was wired. next.md now parses --converge/--cross-ai into a plan strategy, gates on workflow.plan_review_convergence, forwards reviewer flags + --max-cycles, and routes Route-3 planning through /gsd:plan-review-convergence (mirroring autonomous.md); --auto chaining preserves converge mode. Adds argument-hint + help/full.md + COMMANDS.md + how-to parity and a structural regression test.

Closes #1190

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1190): add changeset for progress --converge surface (#1237)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1190): keep progress --converge docs skill-dep-clean + regen workflow size baseline

CI surfaced two ripples from the ADR-15 workflow edits: (1) lint-skill-deps + profile-closure flagged /gsd:plan-phase and /gsd:plan-review-convergence SlashCommand tokens in progress.md's flag docs as undeclared deps — reworded to plain prose since progress.md only advertises the flag (the real invocation lives in next.md); (2) the per-file workflow size baseline needed regenerating after the next.md/help edits.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 15:52:22 -04:00
Tom Boucher
e97b5d6b7b fix(#1217): bound acquireStateLock retries on recoverable errno (no busy-spin) (#1236)
* fix(#1217): bound acquireStateLock retries on recoverable errno (no busy-spin)

The recoverable-errno branch (`ACQUIRE_LOCK_RETRY_ERRNOS`) previously called
`continue` directly, skipping both `clock.sleep()` and the 30 000 ms budget
check. A permanently-failing ENOENT (e.g. parent dir removed) would spin at
100% CPU forever with the event loop fully blocked, making the OS-level
`timeout` the only escape.

Fix: extract a `checkBudgetAndSleep(context)` helper and call it from BOTH
the recoverable-errno path and the EEXIST contention path so every retry is
bounded and backed off identically.

Regression tests added to tests/clock-seam.test.cjs:
- persistent ENOENT throws budget-exceeded error (clock must advance via sleep)
- transient ENOENT (2 retries then success) acquires lock normally

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1217): backfill changeset PR number (#1236)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 15:10:59 -04:00
Tom Boucher
cafb874c4a fix(#1224): accept --pr 0 placeholder at changeset creation (#1231)
* fix(#1224): accept --pr 0 placeholder at changeset creation

The required-field guard `!opts.pr` treated the integer 0 as falsy,
rejecting the documented `pr: 0` two-push placeholder with a usage
error (exit 2). Non-numeric `--pr abc` (NaN) was also silently
accepted before (passes `!NaN === true`... actually `!NaN` is true, so
NaN would trigger the guard already). The new explicit checks use
`opts.pr === null` for missing flag and `Number.isNaN` for non-numeric
input, accepting all finite integer values including 0.

The merge-time safety net in parse.cjs (`pr <= 0` → INVALID_PR) is
unchanged — a pr:0 fragment is still rejected at lint/render time.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1224): backfill changeset PR number (#1231)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 15:02:54 -04:00
Tom Boucher
2ae5fcdcf8 test(#1189): cover ADR-0006 planningPaths() consumption in init handlers (#1226)
Add a black-box regression guard (8 tests in tests/init.test.cjs) asserting the init handlers (execute-phase, plan-phase, phase-op, milestone-op) resolve workstream-scoped planning paths under GSD_WORKSTREAM via planningPaths()/planningDir(), never the flat .planning form.

Each positive case asserts both the scoped value and not-equal-to-flat (genuine guard, not coverage credit); fixtures seeded workstream-scoped; every CLI call pins GSD_WORKSTREAM and GSD_PROJECT for hermeticity. Verified by mutation (flat join fails exactly the 4 positive assertions) and re-verified under a polluted parent env. Test-only; src/ unmodified.

Closes #1189

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 14:20:47 -04:00
Tom Boucher
73b7f45140 feat(#1173): wire agent converters into descriptor-driven install path (#1227)
Extends `dispatchKindEntry` in `runtime-artifact-layout.cts` to route
agents-kind entries through a converter when the descriptor carries a
non-null `converter` field. Adds `stageAgentsForRuntimeWithConverter`
to `install-profiles.cts`, expands `VALID_CONVERTER_NAMES` with the 9
agent converter names, and adds a fail-first behavioral test suite
(9 tests) proving the new wiring end-to-end.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 14:20:21 -04:00
Tom Boucher
bf634b95c3 feat(#1213): Capability State Writer — write-side inverse of the resolver (#1225)
* feat(#1213): Capability State Writer — write-side inverse of the resolver

Adds src/capability-writer.cts (setCapabilityState + cmdCapabilitySet) and the
`gsd-tools capability set` subcommand: the write-side inverse of the capability
resolver (ADR-1213). One desired capability state projects onto the substrates —
`enabled` drives the runtime surface (canonical on/off), `gates` drive federated
config keys (hook granularity), install profile is a read-only floor — then
re-resolves and reports divergence (assert-and-report), so "off means off" holds
as a write-time invariant. Adds batched setConfigValues; routes gsd:settings
capability hook-gates through the writer. Docs: CLI-TOOLS reference, how-to,
ADR-1213, CONTEXT.md term.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1213): add changeset for Capability State Writer (#1225)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 12:51:17 -04:00
Tom Boucher
22f56f4431 ci(#1212): shard windows full-test lane to remove timeout cliff (#1222)
The `full test (windows-latest, *)` lane ran the entire unit suite (~740+
files) in one job whose wall-clock crept against the 20m cap and intermittently
CANCELLED (false-negative gate, observed on PR #1207). Prior tactical fixes
#869 (15→20m bump) and #1051 (handle-leak) deferred the cliff structurally.

Shard the unit suite across 3 parallel runners per OS/node leg so per-job
wall-clock is O(total/3) and stays under the cap as the suite grows.

- scripts/run-tests.cjs: add `--shard <i>/<n>` — a deterministic, balanced
  round-robin partition (fileIndex % n === i-1) over the SORTED selected file
  list. parseShardArg strictly validates i∈1..n, n≥1, integer-only; n=1 is a
  pure no-op. The 28K Windows argv chunking is preserved within each shard. A
  legitimately-empty shard (n > file count) exits 0; a selection empty BEFORE
  sharding still hits the discovery hard error. Composes with --suite and is
  order-independent (sorted before partition). Exports selectShard/parseShardArg.
- .github/workflows/test.yml: test-full becomes the 3 legs × 3 shards = 9-job
  cross-product (explicit include rows — a base shard dim does not cross-product
  with include legs, and a nested matrix.leg.os is unresolvable by the H1
  shell-policy linter). Unit suite runs sharded; integration/security run once
  per leg (shard 1). The Required tests fan-in is unchanged: it already needs
  test-full and checks the matrix-aggregate result, so a failed/cancelled shard
  fails the gate; the branch-protection check name is preserved.
- tests: partition/CLI + pure selectShard contract (completeness, disjointness,
  balance, determinism, boundaries, fast-check property) + parseShardArg
  validation, in run-tests-harness.test.cjs; a DEFECT.GENERATIVE-FIX parity
  guard (per-row shard values 1..N, every leg runs all shards, N == --shard /N
  denominator) + Required-tests name/needs pin, in ci-test-scope.test.cjs.

Closes #1212

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 12:25:29 -04:00
Tom Boucher
7edd18fd2b feat(#1165): async external_job_waiting half-state + resume/pause contract (#1221)
Core half of #1105: a legal external_job_waiting deferred state so an async-dispatched Execute step (committing a .planning/async-jobs/<job>.json manifest, deferring SUMMARY.md) is not an illegal partial. execute-phase safe-resume, resume-project, and pause-work reconcile against the versioned scheduler-agnostic manifest stability contract without re-dispatching; the producer is the capability half (#1164). Closes #1165.
2026-06-14 12:23:07 -04:00