Commit Graph

3659 Commits

Author SHA1 Message Date
Tom Boucher
2d264da661 enhance(#968): region-scoped negative-grep idiom + cross-task conflict warning (#1320)
Adds a region/function-scoped negative-grep idiom to the gsd-planner verification guidance plus a warn-only `validate_plan` check (`scanFileWideNegativeGateConflict`) that flags when a task's file-wide negative grep bans a construct a sibling task legitimately requires elsewhere in the same file. ReDoS-safe (linear, no RegExp on author patterns); region-scoped gates are exempt. Warn-only — never errors, never flips `valid`.

Closes #968

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-16 02:20:34 -04:00
Tom Boucher
14e0709eed refactor(#1307): gate intel on isCapabilityActive + loop-resolver honors capability active (#1315)
* refactor(#1307): gate intel on isCapabilityActive + loop-resolver honors capability active

Part A: intel's command gate moves from config-only isIntelEnabled to the
shared isCapabilityActive('intel', cwd) — a consistency cutover (intel has
skills:[] so its tri-state collapses to the intel.enabled config leg; the gate
now flows through the resolver's precedence + runtime-aware resolution).
Part B: loop-resolver hook rendering now gates on capability state.active
(=== true, fail-closed) instead of state.enabled, so the capability config
gate is honored by the hook consumer, not just per-hook 'when'. active is now
required in the loop-resolver input types. Regression test proves a config-
disabled (active=false) capability's unconditional hook is not rendered.
Part of #1302.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1307): add changeset for intel + loop-resolver active gate

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 00:12:54 -04:00
Rezolv
e50ead7ad2 enhance(verify-phase): deterministic auto-locate of the prohibition check descriptor (#1278) (#1301)
* test(1278): RED-first descriptor parity + fail-closed guards + CHK-07 byte-stability (wave 1)

- CHK-03 (RED): extend PROB-14 parity in prohibition-probe.schema.test.cjs to carry the flat
  check_kind/check_target/check_rule scalars through project->write->parseMustHavesBlock; the
  non-droppable check_kind-presence assertion is the load-bearing RED trigger (fails because
  projectProhibitions strips check_* on the current build).
- CHK-07 (GREEN forward-guard): probe-core.test.cjs pins descriptor-less byte-stability +
  dispositionForProhibition fail-closed policy, with a t.todo marker forward-locking plan 01-02.
- CHK-06 (RED): prohibition-enforcement.test.cjs asserts descriptorFromProjection export +
  fail-closed on absent/partial/unknown descriptors via the projection adapter (RED until 01-03).
- No src/*.cts or .cjs edits; no new test files; lint-test-file-count clean.

* feat(1278): add optional flat-scalar check descriptor fields to Prohibition interface (wave 2)

- check_kind?/check_target?/check_rule? mirror CheckDescriptor.kind/target/rule (minus caller-attested failFirst, #1279)
- optional so existing Prohibition consumers compile unchanged

* feat(1278): project check descriptor as flat scalars in projectProhibitions (wave 2)

- emit check_kind/check_target (+ check_rule only for lint-rule with a rule) when descriptor well-formed
- under-specified/descriptor-less items project byte-identically (CHK-07); flat scalars ride existing parseMustHavesBlock continuation-KV path (no parser rewrite)
- add CHK-02 probe-core unit cases pinning the projection
- turns CHK-03 parity test GREEN; dispositionForProhibition untouched

* feat(1278): descriptorFromProjection read-back adapter feeds fail-closed locate (wave 3)

- Add descriptorFromProjection(projected) -> CheckDescriptor | null to
  src/prohibition-enforcement.cts: renames the projected flat scalars
  check_kind/check_target/check_rule -> {kind,target,rule?}, or null when
  the descriptor is absent/non-object (no check_kind key).
- failFirst is NEVER sourced from the projection (stays caller-attested; #1279).
- rule is set only when check_rule is a non-empty string; the adapter does NOT
  re-validate kind/target/rule — an under-specified descriptor reconstructs to
  one the EXISTING runProhibitionEnforcement LOCATE guard rejects (located:false,
  never green). The merged #1259 guard stays the single source of fail-closed truth.
- Turns the RED CHK-06 fail-closed tests (plan 01-01) GREEN end-to-end; CHK-03 /
  CHK-07 stay green. CheckDescriptor type, locate guard, dispositionForProhibition,
  and parseMustHavesBlock are unchanged (additive +36/-0).

* feat(1278): verify-phase locates prohibition check from projected descriptor (wave 3)

- request.check kind/target/rule sourced from projected check_kind/check_target/check_rule via descriptorFromProjection, not verifier invention (CHK-05)
- replaces the #1278 author-supplied / tracked-follow-up note with the delivered deterministic-locate behavior
- preserves fail-closed routing: absent/partial descriptor -> never green, hard-gate in both modes
- failFirst stays a verify-time caller attestation; #1279 bounds the remaining fail-first proof

* feat(1278): spec-phase captures wired-check descriptor on test-tier resolution (wave 3)

- Step 5.6 'Keep it' / verification: test path captures check_kind/check_target/check_rule, projected onto must_haves.prohibitions for verify-phase deterministic locate (CHK-04)
- SOFT capture: a test-tier prohibition without a descriptor is still allowed (no hard authoring block); stays fail-closed/flagged downstream
- --auto captures only an unambiguous descriptor, never fabricates a check path
- failFirst NOT captured at spec-phase (verify-time attestation; #1279)
- PROB-06 soft-gate + text-mode (PROB-09) behavior unchanged

* chore(1278): re-baseline workflow size for grown verify-phase + spec-phase prose (wave 3)

- spec-phase.md 28438 -> 30343 (+1905), verify-phase.md 35362 -> 36498 (+1136)
- regenerated via npm run size:baseline (no hand-picked numbers); growth is the #1278 deterministic-locate + descriptor-capture prose
- workflow-size-budget guard green (122/122)

* docs(1278): ratify optional check descriptor in dated ADR-550 addendum + type:Changed changeset

- Append dated 2026-06-15 ADR-550 addendum ratifying the D3 prohibition-item
  shape extension (optional flat-scalar check_kind/check_target/check_rule)
- Document flat-scalar rationale, deterministic projection/read-back,
  fail-closed on partial/invalid/absent, #1279/policy out-of-scope
- Add .changeset/1278-prohibition-check-descriptor.md (type: Changed)

* docs(1278): document optional check descriptor in prohibition-probe reference + FEATURES

- Add 'Optional wired-check descriptor (deterministic locate, #1278)' section
  to the prohibition-probe reference (flat-scalar keys, projection/read-back,
  fail-closed + backward-compat, failFirst stays attested)
- Add deterministic prohibition-check descriptor source entry to FEATURES.md
- No CONTEXT.md glossary change: descriptor reuses existing wired-check /
  verification:test vocabulary, no new glossary term introduced

* fix(1278): pass packaging gates — changeset pr field + retired slash-form fix

- Add required pr: 1278 to changeset (lint:changeset MISSING_PR hard requirement;
  plan's 'omit if unknown' was inaccurate — issue number per #1259 convention,
  updated to real PR number when opened) [Rule 3 - blocking]
- Fix retired /gsd-spec-phase -> /gsd:spec-phase at verify-phase.md:83 (wave-3
  prose; caught by slash-namespace invariant #3443/bug-2543, blocked CHK-09
  full-suite-green) [Rule 1 - bug]
- size:baseline + INVENTORY manifest verified in-sync post-build (no diff)

* docs(1278): add check descriptor + descriptorFromProjection to CONTEXT.md prohibition glossary

* fix(1278): harden descriptorFromProjection round-trip (numeric-coercion + stray-rule) per review

- MD-01/LW-01: narrow projected scalars to primitives + String()-coerce, so a
  numeric-looking check_target (parseMustHavesBlock coerces ^\d+$ to number)
  reconstructs as a string and locates instead of silently un-locating; no
  as-string type-lie, satisfies no-base-to-string.
- LW-02: attach rule only for the lint-rule kind (drop a stray node-test rule).
- LW-03: document the optional check_* keys in the reference Output schema.
RED->GREEN tests added in prohibition-enforcement.test.cjs.

* chore(1278): set changeset pr to 1301

* test(1278): add fast-check property for the check-descriptor round-trip + fail-closed (trek-e review)

RULESET.TESTS.property-based-testing: the projectProhibitions -> render ->
parseMustHavesBlock -> descriptorFromProjection chain is a bijective/transformation
contract. Adds 2 fc properties to tests/probe-core.property.test.cjs (no new file;
ratchet stays at 2 for probe-core):
- well-formed descriptors survive the round-trip across the full string domain
  incl. the numeric-coercion case (target/rule reconstruct as strings);
- under-specified/invalid descriptors (absent / target-less / rule-less /
  unknown-kind) are always fail-closed (never green, flagged, unlocated).
Stability is asserted at the descriptorFromProjection layer (the raw parse step is
intentionally lossy for numeric scalars; the shared parser is unchanged).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-16 00:01:03 -04:00
Tom Boucher
e3b829e765 refactor(#1306): gate graphify on isCapabilityActive (tri-state, runtime-aware) (#1313)
* refactor(#1306): gate graphify on isCapabilityActive (tri-state), not config-only

graphify's command gate moves from the config-only isGraphifyEnabled to the
shared isCapabilityActive('graphify', cwd) — so graphify is off unless installed
AND surfaced AND graphify.enabled. Fixes a latent claude-hardcoding in the
resolver: resolveCapabilityRuntimeState now detects the active runtime via
resolveRuntime(cwd) (GSD_RUNTIME -> config.runtime -> 'claude') so non-Claude
runtimes (Codex/Cursor) read their own surface, not ~/.claude. Hermetic
regression test proves config-on+unsurfaced -> disabled; cross-runtime test
proves GSD_RUNTIME=codex honors CODEX_HOME. Gate fails closed on every error
path. Part of #1302.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1306): add changeset for graphify tri-state gate

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 23:19:21 -04:00
Tom Boucher
e9ee7e9ba6 feat(#1305): per-capability active tri-state + isCapabilityActive in Capability State Resolver (#1311)
* feat(#1305): add per-capability active tri-state + isCapabilityActive to the Capability State Resolver

CapabilityStateEntry gains active = enabled && configActivation, where
configActivation resolves the capability's optional activationKey via the
shared _resolveActivationValue (absent activationKey -> true). enabled stays
installed && surfaced (unchanged). Each hook's active now also cascades the
capability config gate (active && configured). Adds isCapabilityActive(capId,
cwd) — a thin convenience over resolveCapabilityRuntimeState. cmdCapabilityState
emits active per capability. No consumer cutover yet (graphify/intel: #1306/#1307;
loop-resolver: #1310). Part of #1302.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1305): add changeset for capability active tri-state

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 22:13:20 -04:00
Tom Boucher
1f41a0ce9a feat(#1304): add optional activationKey capability manifest field (#1309)
Add an optional activationKey to the feature role of capability.json — the
dotted config key that gates the whole capability (e.g. graphify.enabled).
gen-capability-registry validates it (non-empty string, reserved-name guard,
must be declared in the capability's own config slice, feature-only) and emits
it per-capability in the generated registry. Declared on graphify + intel.
No runtime consumption yet (resolver wiring lands in #1305). Part of #1302.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 21:33:06 -04:00
Tom Boucher
91bd82f9a6 enhance(execute): isolated-executor rejected/over-reaching run fails safe — never default recovery to main (#1292) (#1303)
* enhance(execute): isolated-executor rejected/over-reaching run fails safe (#1292)

When an isolated (worktree) executor run is rejected — the user declines to
merge it, the orchestrator surfaces recovery for a blocked/halted plan, or the
run over-reached the requested scope — the orchestrator must no longer
default/propose recovery by editing the primary checkout (`main`). Absent an
explicit guardrail, the LLM orchestrator could improvise "continue on main",
inverting the isolation contract at the moment it matters most.

Added an ISOLATED-RUN RECOVERY — FAIL SAFE policy: default to a safe halt that
offers a fresh, narrowly-scoped worktree or inspect/discard; editing the primary
checkout requires explicit, clearly-labeled confirmation and is never the
default/proposed option.

To respect the ADR-857 phase-6 host-loop size cap on execute-phase.md (it sits
just under the pre-phase-6 baseline), the policy is delivered as an extracted
reference fragment rather than inline:
- New `execute-phase/steps/worktree-recovery-policy.md` holds the recovery policy
  (the existing FAIL-CLOSED rule #48 for base/HEAD mismatches + the #1292
  fail-safe guardrail). No #48 behavior change — moved verbatim.
- execute-phase.md references the fragment at the worktree-spawn recovery point,
  the step-5.5 merge decision, and the stalled-agent "switch to inline execution"
  menu (which for an isolated run now follows the fail-safe policy). Net effect:
  execute-phase.md shrinks below its cap.
- quick.md carries the fail-safe guardrail inline at its post-return merge/discard
  decision (quick.md is not size-capped).

Scoped to the recovery offer only — no automatic scope-overreach detection
(explicitly out of scope per the issue) and no new config key. Adds content
regression tests, a USER-GUIDE note, and a workflow size-baseline update.

Closes #1292

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(changeset): Changed fragment for #1292 isolated-executor fail-safe recovery

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 20:53:38 -04:00
Tom Boucher
137760a655 fix(#1296): align config docs/prompts/schema with consumers (#1299)
* fix(#1296): align config docs/prompts/schema with consumers

The user-facing config surface disagreed with what the consumers actually do
(subset of the #1216 audit). No runtime consumption behavior changes.

- workflow.subagent_timeout: settings-advanced.md prompt + docs/CONFIGURATION.md
  said "seconds (default 600)" but the consumer (map-codebase.md) uses
  milliseconds (default 300000). Relabeled all four spots in settings-advanced.md
  (prompt, parse-default list, example, confirmation table) + the CONFIGURATION.md
  row.
- review.models.<cli>: settings-integrations.md, docs/CONFIGURATION.md (Integration
  Settings), and docs/CLI-TOOLS.md documented a shell command, but review.md injects
  the value into a --model/-m flag. Relabeled to a bare model id and reconciled the
  contradictory CONFIGURATION.md sections.
- workflow.test_command + workflow.build_command: consumed via config-get
  (test_command in verify-phase/execute-phase/audit-fix/post-merge-gate;
  build_command in post-merge-gate) and documented, but absent from validKeys so
  `config set` rejected them. Registered both in config-schema.manifest.json and
  documented them in references/planning-config.md (overview + complete reference).

Regression tests: behavioral config-set tests (tests/config.test.cjs) + doc-parity
content guards (tests/config-field-docs.test.cjs).

Deferred to other #1216 clusters: security-gate wiring, search_gitignored wiring,
mvp_mode, source_grounding_authority labeling, and config-set enum enforcement.

Closes #1296
Refs #1216

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(changeset): Fixed fragment for #1296 config-surface alignment

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 19:57:11 -04:00
Tom Boucher
8c3d934a90 refactor(#1294): T-final — delete the core.cjs re-export spine (epic #1267 complete) (#1295)
* refactor(#1294): T-final — delete the core.cjs re-export spine (epic #1267 complete)

After T0–T6 nothing imports core, so retire the spine and its scaffolding:
- delete src/core.cts (and the gitignored gsd-core/bin/lib/core.cjs artifact;
  remove its .gitignore + eslint-ignore entries)
- delete scripts/lint-core-spine-imports.cjs + its allowlist; drop it from the
  package.json lint:ci chain
- regenerate docs/INVENTORY-MANIFEST.json (drops the core.cjs surface)
- sweep stale references: CONTEXT.md glossary back-compat clauses (spine retired,
  callers import the leaf directly), planning-config.md CONFIG_DEFAULTS owner,
  and false present-tense core.cjs claims in leaf-module docstrings

The ADR-857 decomposition is complete: the former Core god-module is fully
dissolved into its leaf modules; no re-export spine remains. No behaviour change.

Closes #1294

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1294): migrate the computed-path core.cjs importers the literal grep missed

bin/install.js used require(path.join(_gsdLibDir, 'core.cjs')) (a computed
path, and bin/install.js was never in the convergence lint's scan roots), and
~8 test files referenced core.cjs via path.join/readFileSync/existsSync/FILE_ARG
forms the literal-string migration grep missed. Route install.js's symbols to
their leaves (RUNTIME_PROFILE_MAP->model-catalog, resolveTierEntry/EFFORT_SET->
model-resolver) and repoint/adjust the test references to the leaves. Recovers
the 161 'Cannot find module core.cjs' failures from the spine deletion.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 18:50:46 -04:00
Tom Boucher
c76827afbc refactor(#1291): T6 — migrate test files off the core spine ahead of deletion (#1293)
The convergence lint only scanned src/ + gsd-core/bin, so ~35 test files
still imported core.cjs. Repoint all 33 behaviour importers to the leaf
modules directly (same symbol->leaf map as the src migration; leaves are the
objects core re-exported by reference), delete the now-meaningless
shim-identity describe blocks in the 8 leaf tests, and delete tests/core.test.cjs
(forwarded-behaviour coverage now lives at the leaves; resolveWorktreeRoot
test relocated to worktree-safety in T0) and tests/lint-core-spine-imports.test.cjs
(the lint is removed in T-final). Dropped the stale core.test.cjs entries from
the allow-test-rule-refs allowlist; eslint-rules RuleTester fixture path
pointed at io.cjs.

After T6: ZERO test imports core.cjs. core.cts still builds (now fully unused);
T-final deletes it. No behaviour change.

Closes #1291

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 18:11:20 -04:00
Tom Boucher
76765bc24d refactor(#1289): T5 — migrate the final idiom-hard callers off the core spine (#1290)
The last 4 core importers, migrated off non-destructure idioms:
- gsd-tools.cjs: core.{error,ERROR_REASON,setJsonErrorMode,output} -> io.cjs;
  core.findProjectRoot -> project-root.cjs (lazy wrapper preserved); inline
  resolveWorktreeRoot require -> worktree-safety.cjs
- audit-command-router: DI default `_core ?? core` -> `_core ?? io` (seam preserved)
- intel-command-router: DI default -> `{ output: io.output, timeAgo: coreUtils.timeAgo }` (seam preserved)
- check-command-router: io destructure -> io.cjs; dynamic core['planningDir']
  -> planning-workspace, core['findPhaseInternal'] -> phase-locator (typed
  imports, dropped the Record-cast bracket hack)

Allowlist is now EMPTY — NO file imports the core spine. core.cts re-exports
are dead weight; T-final deletes core.cts + remaining shim tests + the lint.
No behaviour change.

Closes #1289

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 17:44:25 -04:00
Tom Boucher
b108f101b0 fix(#1284): grant mcp__perplexity__* to researcher agents + dispatch-table parity guard (#1288)
Adds mcp__perplexity__* to both researcher profiles (generated source-of-truth) and regenerates the agents; adds a generative dispatch-table↔tools parity guard so future provider drift fails CI. Regenerates the agent-size baseline for the +20-byte frontmatter growth.

Fixes #1284

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 17:20:08 -04:00
Tom Boucher
645601a10d refactor(#1286): T4 — migrate 5 large destructure callers off the core spine (batch 3) (#1287)
Migrate the entire core surface of commands (~23 symbols), phase (~17),
roadmap, state, template to the leaf modules directly (behaviour-identical —
leaves are the objects core re-exports by reference). All 5 now import zero
core symbols and are removed from the allowlist (9 -> 4). Dropped a dead
`void replaceInCurrentMilestone` from phase.cts; stale core.cjs docstrings
fixed. core.cts re-exports untouched (serve the remaining 4 idiom-hard
files); teardown is T-final. No behaviour change.

Closes #1286

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 17:03:22 -04:00
Tom Boucher
ec2ecdf28b refactor(#1283): T3 — migrate 9 multi-leaf callers off the core spine (batch 2) (#1285)
Migrate 9 files' entire core surface to the leaf modules directly
(behaviour-identical — leaves are the objects core re-exports by reference):
config, docs, gap-checker, graphify-command-router (namespace core.output ->
io.output), init (17 core symbols -> 8 leaves), profile-output, uat,
verification, workstream.

All 9 now import zero core symbols and are removed from the allowlist
(18 -> 9). Stale core.* docstrings corrected. core.cts re-exports untouched
(still serve the remaining 9 files); teardown is T-final. No behaviour change.

Closes #1283

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 16:40:48 -04:00
Tom Boucher
a5f213e73e refactor(#1281): T2 — migrate 12 single-leaf callers off the core spine (batch 1) (#1282)
Per the T1 design rubber-duck, batch by FILE so each tranche drops
convergence-lint allowlist entries. Migrate 12 files' core imports to the
leaf modules directly (behaviour-identical — leaves are the objects core
re-exports by reference):
- io (output/error/ERROR_REASON): agent-command-router, capability-state,
  capability-writer, frontmatter, gsd2-import, learnings, loop-resolver,
  task-command-router
- roadmap-command-router -> config-loader; workstream-inventory -> core-utils
- milestone, verify -> their full leaf sets (both were multi-leaf, not
  single-leaf as first scoped; migrated completely)

All 12 files now import zero core symbols and are removed from the
allowlist (30 -> 18). core.cts re-exports untouched (still serve the
remaining 18 files); teardown is T-final. Stale core.cjs docstrings in the
migrated files corrected to reference io.cjs. No behaviour change.

Closes #1281

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 16:18:54 -04:00
Tom Boucher
13414cb168 Merge pull request #1273 from davesienkowski/feat/1259-test-tier-enforcement
enhance(verify-phase): enforce test-tier prohibitions — wire the deferred negative-test gate (#1259, ADR-550 D5d)
2026-06-15 15:54:16 -04:00
Tom Boucher
54420bae9e refactor(#1277): T1 — decouple agent-install-check + git-base-branch from the core spine (#1280)
First leaf-migration tranche of epic #1267 (after T0 #1268). Migrate the
via-core callers of the two leaves T0 created to import from the leaf
modules directly, and stop core re-exporting their symbols:
- checkAgentsInstalled: docs.cts, verify.cts, init.cts -> agent-install-check.cjs
- gitWorktreeInfoInternal: init.cts -> git-base-branch.cjs
- getAgentsDir had no external via-core caller (internal to the leaf)

core no longer re-exports getAgentsDir / checkAgentsInstalled /
gitWorktreeInfoInternal; the now-unused agent-install-check + git-base-branch
requires are dropped from core; the shim-identity assertions for these are
deleted (behaviour tests retained). Convergence lint stays green (0 new).
No behaviour change.

Closes #1277

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 15:46:41 -04:00
Dave
0d783fa57d Merge remote-tracking branch 'origin/next' into feat/1259-test-tier-enforcement
# Conflicts:
#	tests/workflow-size-baseline.json
2026-06-15 15:42:23 -04:00
Tom Boucher
0a856f06cd enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier (#1271)
* enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier

Introduce a per-truth PRESENT_BEHAVIOR_UNVERIFIED state for must-haves that assert a state transition or a cancellation/cleanup/ordering invariant whose only evidence is symbol presence + wiring. Such truths are excluded from the verified_truths score, reported as a behavior_unverified count, recorded in an always-on behavior_unverified_items frontmatter list, and routed to the existing human_needed sink — so a clean N/N can no longer be reached on symbol presence alone.

The overall-status vocabulary and the src/verification.cts seam are unchanged (the new state is per-truth only); gaps_found keeps decision-tree precedence; override-passed truths still count toward verified_truths. Mirrors the calibration into the shipped verify-phase.md workflow (with an infra/foundation carve-out), the VERIFICATION.md templates, and docs (planning-artifacts.md, AGENTS.md). gsd-verifier.md kept under its 48KB LARGE cap; size baselines regenerated.

Closes #966

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#966): add changeset fragment for gsd-verifier behavior-unverified calibration

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 15:24:56 -04:00
Dave
05c0c65b41 fix(1259-01): import io.cjs leaf directly, not the core.cjs re-export spine (#1268 lint)
Merged latest origin/next, which added lint-core-spine-imports (#1268, core.cjs
spine retirement). The new prohibition-enforcement module imported output/error/
ERROR_REASON from the core.cjs re-export spine — banned for new files. Re-homed to
the leaf './io.cjs' (matches graphify-/intel-command-router). lint:ci now clean.
2026-06-15 15:22:27 -04:00
Dave
95e5abfb66 Merge remote-tracking branch 'origin/next' into feat/1259-test-tier-enforcement 2026-06-15 15:20:54 -04:00
Dave
6a9e6cb0ae fix(1259-01): address trek-e review — B1 (fatal/suppressed), B2 (bounded subprocess), M1/M2, minors
Maintainer CHANGES_REQUESTED (reviewed e08667e5, pre-portability-fix):
- B1: lint-rule no longer greens an unparseable target (eslintHasFatalError -> fail
  closed on any fatal/parse error) or an inline-suppressed violation (eslintJsonHasRule
  now scans suppressedMessages too). RED-first + real-runner repros.
- B2: both child spawns get a bounded timeout (30s node / 60s eslint) + 16MiB maxBuffer;
  timeout fails closed. Injectable timeoutMs (positive-only — 0/negative can't disable
  the bound) enables a fast 1.5s hang test.
- M1: scoped verify-phase.md — the check descriptor is author-supplied for now; filed
  #1278 for deterministic auto-locate of the descriptor (the locate half).
- M2: added tests/prohibition-enforcement.property.test.cjs (fast-check fail-closed
  invariants; within the <=2-file budget).
- m1: tapTestNames excludes # SKIP/# TODO; parseNodeTestSummary tracks # cancelled;
  isNonVacuousNodeTestPass requires cancelled===0.
- m2: scoped the determinism claim to the decision/parse layer (real runner is env-dependent).
- m3: filed #1279 for machine-proven fail-first (violation-fixture probe).
- n1: -- before target in both arg builders (option-injection). n2: dropped dead token.
- B3 (Windows npx) was already fixed in 2af76306 (pushed ~65s after the review).

Verified node 22 + 24; size baseline regenerated for the verify-phase note.
2026-06-15 15:19:04 -04:00
Tom Boucher
48d9cec6fe refactor(#1268): re-home core re-export-spine squatters + migration-convergence lint (#1272)
Re-home the 6 implementation functions squatting in the core.cjs re-export
spine (ADR-857) into the modules whose interface they belong to, with core
re-exporting them BY REFERENCE so all 32 callers + the shim-identity tests
keep resolving unchanged:
- worktree-safety: resolveWorktreeRoot, pruneOrphanedWorktrees
- git-base-branch (broadened to the Git Query Module): gitWorktreeInfoInternal
- agent-install-check (new leaf): getAgentsDir, checkAgentsInstalled
- delete the _resetRuntimeWarningCacheForTests wrapper; consumers use a
  shared resetRuntimeWarningCaches() helper in tests/helpers.cjs

Add scripts/lint-core-spine-imports.cjs (migration-convergence lint with a
30-importer allowlist, wired into lint:ci) so the staged spine retirement
provably converges: CI fails on any new ./core import. Register the new
generated agent-install-check.cjs in eslint-ignore + .gitignore +
INVENTORY-MANIFEST.json.

No behaviour change. First tranche (T0) of epic #1267.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 15:08:58 -04:00
Dave
564f1da3ad test(1259-01): pin target-side basename normalization (round-4 WR-R4-01)
A mutant dropping baseOf() from the TARGET side of the node-test vacuity compare
survived — add the mirror case (relative TAP name vs absolute descriptor target,
both basename-normalized -> vacuous).
2026-06-15 14:27:29 -04:00
Dave
2af7630642 fix(1259-01): make real runner cross-platform/version portable (CI ubuntu-24 + windows-24)
CI failed on node 24 + Windows in the real-runner E2E tests — real portability
bugs in the shipping runner, not test flakiness:

- node 24 names a zero-test file's TAP entry by ABSOLUTE path (node 22 used the
  basename), so the exact-string vacuity discriminator misfired -> empty file
  falsely greened. Now compares BASENAMES (separator-agnostic) -> robust across
  OS + node version. Verified: empty file is non-green under node 24 locally.
- the lint-rule runner spawned 'npx', which execFileSync can't launch on Windows
  (and failed fast on ubuntu-24). Now resolves the project's eslint CLI via
  eslint/package.json and runs 'node <cli>' through process.execPath (portable;
  no shell -> no injection). eslint absent -> fail-closed.
- node-test also spawns via process.execPath, not bare 'node'.

Also folds in round-3 WR-01: the probe-core inline comments still said
enforcement was 'deferred to a follow-up PR' — corrected (landed in #1259).
Strengthened the isNonVacuousNodeTestPass unit test to pin the basename clause
(absolute-path target). Verified 71/71 green under BOTH node 22 and node 24.
2026-06-15 14:19:35 -04:00
Tom Boucher
51b23c248a fix(#1274): bump ws to ^8.21.0 to clear advisory GHSA-96hv-2xvq-fx4p (#1275)
The direct production dependency ws was pinned to 8.20.1, in the vulnerable
range of GHSA-96hv-2xvq-fx4p (high — memory-exhaustion DoS). The freshly
disclosed advisory turned the #3588 `npm audit --omit=dev reports zero
advisories` CI gate red repo-wide (next + every open PR). Bump to the
patched ^8.21.0 (backward-compatible). npm audit --omit=dev now reports 0.

Closes #1274

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 14:07:52 -04:00
Dave
e08667e5aa test(1259-01): regenerate workflow-size baseline for verify-phase.md growth (+1459 B)
verify-phase.md grew 30875 -> 32334 from the test-tier enforcement consumer step
+ descriptor docs. Well under the 40960 DEFAULT hard cap; regenerate the per-file
baseline ratchet (the growth is the load-bearing consumer wiring, not bloat).
2026-06-15 14:00:03 -04:00
Dave
566f99089b chore(1259-01): backfill changeset pr number (#1273) 2026-06-15 13:48:43 -04:00
Dave
1977b8097c docs(1259-01): correct probe-core fail-closed reason — enforcement landed, not deferred
The test-tier fail-closed reason string (surfaced to users via the producer) and
the dispositionForProhibition docstring still said the negative-test enforcement
was 'deferred to a follow-up PR'. This PR IS that follow-up, so the claim is now
false. Updated the user-facing reason and the comment to say the producer landed
in #1259. Policy/branching logic is byte-unchanged (only message strings).
2026-06-15 13:47:02 -04:00
Dave
97931249bc fix(1259-01): close re-review NEW-BL-01 (ignored-target vacuous green) + no-throw hardening
Round-2 adversarial review found the lint-rule runner falsely greened an
eslint-IGNORED target: eslint returns a length-1 "File ignored" result (ruleId
null) that passed the >=1-file vacuity guard while nothing was linted — reopening
the vacuous-green class. Fix: buildLintArgs now passes --no-warn-ignored so an
ignored path returns [] -> fails closed (verified + E2E test on an ignored
bin/lib artifact).

Also: wrap runCheck() so even a (test-injected) throwing runner fails closed
(NEW-WR-01, full no-throw contract); document the benign basename-naming
constraint on wired node-test names (NEW-WR-02, fail-closed).
2026-06-15 13:45:03 -04:00
Dave
31b822b312 fix(1259-01): close adversarial review findings — genuine enforcement, honest fail-first scope
Adversarial pre-submission review found the injected-runCheck tests masked a
non-functional real runner. Fixes:

- BL-01 (false green on vacuous test): the node-test runner now parses the TAP
  summary and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail) AND a
  reported test named distinctly from the file — node --test counts an empty file
  as one passing test, so counts alone could not catch it.
- SF-01 (lint anchor never greened): the lint-rule runner now runs the project
  eslint as --format json and filters by ruleId, so plugin rules (local/*) load
  via the flat config — bare --rule cannot load a plugin. local/no-source-grep
  now genuinely greens (covered by a real, non-injected test).
- BL-02 (tautological fail-first): the runner no longer echoes the caller's
  failFirst as if confirmed. failFirst is documented as caller-ATTESTED; the
  producer requires attestation + a genuine non-vacuous pass. Machine-proven
  fail-first (needs a violation fixture) is flagged as a tracked follow-up in
  ADR-550, the changeset, FEATURES, the reference doc, and verify-phase.
- SF-02: added real-runner end-to-end tests (no injected runCheck) + pure,
  exported parse/filter helpers (parseNodeTestSummary, tapTestNames,
  eslintJsonHasRule, eslintFileResultCount) so the shipping branches are
  mutation-pinned.
- NIT-01/02: LOCATE guard rejects empty-string rule and unknown kinds.
- Hardening: spawn checks with NODE_TEST_CONTEXT/NODE_OPTIONS scrubbed so an
  ambient test-runner context cannot corrupt a verify-time result.
- Docs reconciled to the shipped behavior (no 'confirms fail-first' overclaim).
2026-06-15 13:38:15 -04:00
Dave
676436258a fix(1259-01): lint-rule real runner — keep rule id distinct from lint target
The default lint-rule runner passed check.target as BOTH the --rule id and the
eslint path, so it could never pass (eslint tried to lint a file named after the
rule). Add a distinct check.rule field (rule id) vs check.target (path to lint),
extract a pure exported buildLintArgs() so the mapping is mutation-testable
without spawning eslint, fail-closed on a lint-rule missing its rule id, and carry
the rule into enforcement evidence. Updates verify-phase descriptor docs.
2026-06-15 13:10:49 -04:00
Dave
1919c2fd8f test(1259-01): RED — lint-rule real runner must keep rule id distinct from lint target 2026-06-15 13:08:33 -04:00
Dave
ac28ef3a55 chore(1259-01): ignore the emitted prohibition-enforcement.cjs in eslint
- ADR-457: lint the src/*.cts source, not the tsc-emitted gitignored .cjs artifact
- Mirrors the existing per-file ignore entries (probe-core.cjs, check-command-router.cjs)
2026-06-15 12:48:59 -04:00
Dave
ce01e1376b docs(1259-01): wire verify-phase consumer + ADR-550/FEATURES/reference + changeset
- verify-phase.md: replace test-tier 'fail-closed/deferred' bullet with the check prohibition-enforcement
  enforcement step (locate -> fail-first -> run -> evidence -> green-or-hard-gate); update determine_status tree
- ADR-550 addendum: mark D5d enforcement half LANDED (#1259); cite ADR-857 open-question §147 + D6 (core verify rail)
- FEATURES §146: enforcement wording + add REQ-PROHIB-07; keep REQ-PROHIB-06 intact
- references/prohibition-probe.md: test-tier enforced + hard-gates via check prohibition-enforcement
- changeset (type: Changed) with the D5 '2 no-source-grep invalid cases, not 96' correction
2026-06-15 12:47:39 -04:00
Dave
8f428bdd41 feat(1259-01): deterministic prohibition-enforcement producer + check route
- Author src/prohibition-enforcement.cts: the test-tier prohibition PRODUCER/gate (ADR-550 D5d heavy half)
  locate wired check (node-test|lint-rule) -> confirm fail-first -> run -> build enforcementEvidence -> dispositionForProhibition
  pure/deterministic with injectable runCheck; missing/failing/non-fail-first -> hard-gate (both modes); passing -> green
- Route check prohibition-enforcement in src/check-command-router.cts (same family as ui-plan-gate / tdd-review-checkpoint)
- Add tests/prohibition-enforcement.test.cjs (behavioral, typed-field, injected runner) — new module within <=2 budget
- Register the built bin/lib surface: docs/INVENTORY.md row + regenerated docs/INVENTORY-MANIFEST.json
- .gitignore: add the emitted gsd-core/bin/lib/prohibition-enforcement.cjs (build artifact, ADR-457)
- No src/probe-core.cts edit — the green/fail-closed policy seam already exists
2026-06-15 12:44:29 -04:00
Dave
6115ab216d test(1259-01): RED — test-tier enforcement (both check kinds) before producer exists
- Extend prohibition-probe.verify-tier.test.cjs with the ENFORCEMENT half (#1259, ADR-550 D5d)
- Require the not-yet-built gsd-core/bin/lib/prohibition-enforcement.cjs (RED)
- Cover both wired-check kinds (node-test + no-source-grep lint-rule) and miss/fail hard-gate
- Typed-field assertions only; injected runCheck (no real subprocess)
- Keep the 2 original fail-closed tests verbatim
2026-06-15 12:38:48 -04:00
Tom Boucher
b370787155 Merge pull request #1262 from open-gsd/chore/sync-next-version-1.5.0-rc.4
chore: sync next package version to 1.5.0-rc.4
2026-06-15 01:07:49 -04:00
github-actions[bot]
04e4d631c4 chore: sync next package version to 1.5.0-rc.4 2026-06-15 05:07:43 +00:00
Tom Boucher
cf68841220 enh(#1243): consume Claude plugin-provided skills in agent_skills (epic #1258 Phase B) (#1261)
* feat(#1243): consume Claude plugin-provided skills via native Skill-tool directive + grant Skill to agent_skills-consumer agents

- Relax global skill name validation to accept namespaced form `^[A-Za-z0-9_-]+(:[A-Za-z0-9_-]+)*$`
- Namespaced names (containing colon) on claude runtime emit a Skill-tool load directive instead of a @-include line
- Namespaced names on non-claude runtimes are skipped with a warning
- Bare unresolved names retain existing warn-and-skip behavior (no promotion to directive)
- Grant `Skill` tool to all 22 agent_skills consumer agents; 5 generated agents updated via research-profiles.cjs + regen, 17 hand-authored agents edited directly
- Add 16 TDD tests in describe('bug #1243') covering happy/mixed/precedence/negative/cross-runtime/regression/grant cases

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#1243): document plugin-provided skills in agent_skills

Update the Agent Skills Injection reference in CONFIGURATION.md with
the three entry forms (project-relative, global:<name>,
global:<plugin>:<skill>), the Claude-only runtime behaviour of the
namespaced form and the warn-skip on other runtimes, the plugin
pre-install prerequisite, and the consumer-agent Skill tool grant.

Add docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md with a
step-by-step guide for installing the plugin, locating the namespaced
skill name, wiring it into agent_skills, and verifying injection.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1243): align agent_skills docs with emitted block format + mixed-block regression test (code-review)

- Replace two-section mixed-block example (bogus "Load these plugin-provided skills using the Skill tool:" header) with the actual single-section inline format in CONFIGURATION.md and docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md
- Fix quoted warning text in how-to doc to exactly match the emitted string: [agent-skills] WARNING: Plugin-namespaced skill "global:<name>" requires a Skill-tool-capable runtime (claude) — skipping on runtime "<runtime>"
- Replace phantom agent slugs (gsd-checker, gsd-researcher, gsd-advisor, gsd-synthesizer) in CONFIGURATION.md Supported Agent Types with real agents/gsd-*.md examples (gsd-plan-checker, gsd-phase-researcher, gsd-code-reviewer, gsd-ui-auditor, gsd-research-synthesizer)
- Add byte-identical mixed-block regression test: one path-resolvable global skill + one plugin-namespaced skill on claude runtime → asserts r.ir.block === single-section interleaved block, no secondary header

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1243): regenerate agent-size baseline for the Skill-tool grant

The 22 agent_skills-consumer agents each grew +7 bytes from adding `Skill`
to their tools list; refresh the committed per-agent size baseline (#1074 guard).

* chore(#1243): add Added changeset fragment

* fix(#1243): traceable allow-test-rule ref + separator-agnostic byte-identical tests (CI)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-15 00:49:12 -04:00
Tom Boucher
00c447701d fix(#1257): update pipe-table Status/Phase/Plan cells in planned-phase + begin-phase (#1260)
* fix(#1257): update pipe-table Status/Phase/Plan cells in planned-phase + begin-phase

cmdStatePlannedPhase ran its body-field replacements on the full file content,
so the case-insensitive ^Status: pattern matched the YAML frontmatter `status:`
line before the body `| Status | … |` cell — the cell never advanced to
'Ready to execute' and syncStateFrontmatter re-derived the stale 'planning'
status (the #1230 delta heuristic preserved it). cmdStateBeginPhase had
pipe-table else-branches only for Status/Last activity (#1256), so the Current
Position `| Phase |` / `| Plan |` cells were left stale while a spurious inline
`Phase: N — EXECUTING` line was prepended.

Both handlers now strip frontmatter before body-field replacement and update the
pipe-table cells in place via stateReplaceField, matching inline-format
behaviour. Systemic residual of #1255 / #1256.

Adds 4 regression tests (#1257 block in tests/state.test.cjs) covering both
findings — RED before the fix, GREEN after; full state-area suite (437) and
local unit suite stay green.

Closes #1257

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1257): backfill changeset pr number (#1260)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 00:29:18 -04:00
Tom Boucher
2668fbbeb0 fix(#1255): advance frontmatter status for pipe-table STATE.md on begin/complete-phase (#1256)
* fix(#1255): advance frontmatter status for pipe-table STATE.md on begin/complete-phase

state begin-phase/complete-phase called stateReplaceField on the FULL file content, so its case-insensitive ^Status: plain-pattern matched the YAML frontmatter status: line first (no g flag) and never updated the body pipe-table Status cell; syncStateFrontmatter then re-derived the stale status from the unchanged body, freezing the frontmatter status. Fix: strip frontmatter before the body-field replacements (operate on body only), reassemble with frontmatter preserved, so the body Status cell updates and the frontmatter derives correctly — for inline AND pipe-table body formats. Also corrects the Current Position pipe-table else-branches (Status/Phase/Last-activity) to write bare, consistent cell values. Pipe-table Status is a supported body format (not rewritten to inline).

Closes #1255

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1255): add changeset for pipe-table state status fix (#1256)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1255): fold pipe-table regression into state.test.cjs + Windows-portable frontmatter regex

Per the 2026-06 audit, new tests/bug-NNNN-*.test.cjs files are banned (lint-regression-test-names) — folded the 7 #1255 regressions into tests/state.test.cjs and removed the standalone file + its lint-test-file-count allowlist entry. Also fixed the frontmatter assertions' /^---\n/ anchors to /^---\r?\n/ (windows-test-parity-guard frontmatterAnchorLiteralNewline).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 22:42:28 -04:00
Rezolv
3556450b0d feat(spec-phase): surface zero-classification edge-probe requirements as unclassified candidates (#1110) (#1117)
Edge-probe now surfaces a zero-classification requirement (non-empty prose, no
shape cue matched, no `shapes` override) as a single soft `unclassified — review
manually` candidate instead of silently dropping it — the exact blind spot the
probe exists to catch. Dismissible like any edge; the `shapes: []` opt-out stays
silent; `TAXONOMY` (the closed 8 categories) is unchanged. Under `--auto` the
candidate is left `unresolved`, never auto-`backstop` (a missing shape is not
evidence an edge exists).

Closes #1110
2026-06-14 21:44:21 -04:00
Rezolv
395fb519e7 feat(spec-phase): prohibition probe — surface "must-NOT" constraints (#644) (#1149)
Adds the spec-time prohibition probe (spec-phase Step 5.6) — the second adapter of the probe-core resolution model. Surfaces unwritten must-NOT constraints as negative SPEC acceptance criteria with test/judgment verification tiers; fail-closed at verify time. Per ADR-550. Closes #644.
2026-06-14 21:29:11 -04:00
Tom Boucher
9a4a448c27 docs(#1235): ADR — migrate agent conversion to the descriptor-driven install path (#1253)
Records the design gate for moving agent conversion off the inline bin/install.js loop onto the descriptor path (ADR-3660): the two-(really three-)path problem, the 10 parity behaviors the descriptor agents path must gain (verified against code — the issue's 7 plus Qwen/Hermes branding, the Codex TOML sidecar, and stale-agent cleanup), an AgentConverterContext contract to fix the leaky (content)=>string converter signature, and an incremental per-runtime cutover gated on byte-for-byte golden parity (full + minimal mode). Status: Proposed. Codex-reviewed for technical accuracy.

Closes #1235

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 20:54:20 -04:00
Tom Boucher
b783410815 refactor(#1191): inject clock/reset testability seams + handle valid-null settings (#1233)
* refactor(#1191): inject clock/reset testability seams + handle valid-null settings

- worktree-safety reapOrphanWorktrees: injectable deps.nowMs clock for deterministic stale-lock boundary tests (mirrors snapshotWorktreeInventory's options.nowMs).

- active-workstream-store: _resetControllingTtyCacheForTests() seam clears the memoized controlling-TTY probe cache; test replaces require.cache busting.

- gen-capability-registry: export stripGeneratedComment (additive); test imports the real helper + equivalence assertion, keeping the deliberate drift oracle.

- install.js readSettings: a successfully-parsed JSON null is treated as empty settings ({}) instead of being mis-reported as malformed; genuine parse failures still warn. readSettings/stripJsonComments exported (GSD_TEST_MODE-guarded require) for real behavioral tests.

Closes #1191

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1191): add changeset for valid-null settings fix (#1233)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1191): replace Stryker-incompatible structural reset test with behavioral isTTY-spy

The seam-2 reset test read the BUILT active-workstream-store.cjs and grepped for 'didProbeControllingTtyToken = false' — Stryker instruments that file so the literal is absent, failing the mutation DRY RUN. Replaced with a behavioral test that spies on process.stdin.isTTY access count to prove a post-reset probe re-runs (kills the didProbe-reset mutant) without reading source text. Local stryker: dry run passes, score 85.21% >= 80.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 20:52:56 -04:00
Tom Boucher
55eb1fd72e refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam (#1242)
* refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam

ADR-22's severity mapping, authority auto-upgrade, and rung>=3 hard-block lived only as prose in plan-review-convergence.md — untestable. Extracted into src/plan-drift-guard.cts (pure: AUTHORITY_RUNGS, getEffectiveAuthority, classifyDriftSeverity) + a gsd-tools drift-guard CLI seam (authority/severity), and rewired the workflow to call the seam deterministically instead of reasoning the decision in prose. 47 unit/e2e/structural tests cover the full severity table, the grep->intel auto-upgrade, and the rung>=3 HIGH hard-block.

Closes #1190

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1190): add changeset for ADR-22 drift-guard seam (#1242)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1190): register ADR-22 module in eslint-ignore + inventory manifest + docs-exempt changeset

Full-matrix CI surfaced new-module/command governance ripples beyond the lint-tests chain: (1) tsc-generated plan-drift-guard.cjs must be in the eslint ignore list (551-eslint-bin-lib-coverage); (2) docs/INVENTORY-MANIFEST.json must include the new module/command (regen via gen-inventory-manifest.cjs --write); (3) a type:Added changeset triggers docs-required — added a docs-exempt marker (internal seam, no user-facing surface).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 20:52:36 -04:00
Tom Boucher
6ef78fe417 feat(#1188): branch-coverage floors + promote no-source-grep to error on tests/** (#1250)
* feat(#1188): promote no-source-grep to error on tests/**

Apply local/no-source-grep as an error to tests/**/*.test.cjs (was warn on bin/scripts only). Only 2 real violations surfaced (repo-layout.test.cjs structural guard-placement checks on bin/install.js) — marked allow-test-rule with #1188 reason. PART A (branch-coverage floor) follows separately.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1188): add c8 branch-coverage floors (60 global, 70 per-file on UNMUTATED modules)

c8 enforced line floors only. Add a --branches 60 global floor (current ~82.5%, margin mirrors the lines-70-vs-91% gap) to test:coverage + test:coverage:unit, plus a chained 'c8 check-coverage --per-file --branches 70' for the high-risk UNMUTATED modules state/phase/verify/init (current min 78.3% on verify). The per-file check reuses the coverage data the suite run just produced (the proven scripts-floor pattern) so it adds no second suite run.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1188): per-module branch check must pin lines/funcs/stmts to 0

Standalone 'c8 check-coverage' defaults unspecified metrics to 90 and enforces them, so --branches 70 alone also failed on lines (verify.cjs 85.29% < default 90). Pin --lines 0 --functions 0 --statements 0 so only the branch floor (70) is enforced on the UNMUTATED modules. Data-reuse confirmed working (it computed verify's real %).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 20:52:13 -04:00
Tom Boucher
6724581285 fix(#1230): preserve frontmatter status/stopped_at when a state write doesn't change the body source (#1252)
* fix(#1230): preserve frontmatter status/stopped_at when a state write doesn't change the body source

Adds a delta heuristic in `readModifyWriteStateMd`: snapshots the body
Status and Stopped At fields before the transform, then after
syncStateFrontmatter runs, restores the existing frontmatter values for
any field whose body source was not changed by this write. Integrates
cleanly with the existing !resync progress-restore block (computes postFm
once, applies both restorations, reconstructs frontmatter once).

Regression tests added to bug-905 test file (same theme, no new file).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1230): add Fixed changeset fragment

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-14 20:51:53 -04:00
Tom Boucher
8da48e22ae fix(#1229): count bullet-only phases so phase.add stops reusing an existing phase number (#1249)
* fix(#1229): count bullet-only phases + guard against number collision in phase.add

Before this fix, the phase.add number scan only checked ### Phase N: section
headers and on-disk phases/N-* directories. A phase that existed only as a
roadmap bullet (e.g. "- [ ] **Phase 11: ...**") was invisible to both scans,
causing phase.add to silently assign a duplicate number.

Fix: add a bullet-entry regex scan (all checkbox variants: [ ], [x], [~], with
or without ** bold markers) to the set-based phase-number collection in
cmdPhaseAdd. Also added a post-compute collision guard that advances the
candidate past any already-used number.

Regression tests added to tests/phase.test.cjs (bug #1229 describe block):
bullet-only collision, [x]/[~] variants, plain-bullet, and baseline preservation.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1229): add Fixed changeset fragment

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-14 17:52:15 -04:00