Commit Graph

222 Commits

Author SHA1 Message Date
Tom Boucher
bb88a78faa fix(#1472,#1454): workstream-aware health paths; exclude active worktree from W017 (#1483)
* fix(#1472,#1454): validate health workstream-aware paths; exclude active worktree from W017

#1472: cmdValidateHealth now uses planningRoot(cwd) for shared-root files
(PROJECT.md, config.json, MILESTONES.md) and planningDir(cwd) for
workstream-scoped files (ROADMAP.md, STATE.md, phases/). Previously a
single planningDir() call was used for all paths, causing false
E002/E003/E004/W003 when GSD_WORKSTREAM is set.

#1454: W017 no longer fires for a stale worktree whose path equals or is
an ancestor of process.cwd(), preventing advice to remove the active
session's own worktree.

Regression tests added for both bugs; all 40 existing health tests pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: correct changeset format

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 13:37:09 -04:00
Tom Boucher
7c93d9e222 feat(#1463): add capability outdated (per-source update check); drop phantom slash-command docs
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 10:50:12 -04:00
Tom Boucher
08d1c57d6e fix(#1460): verify-or-reject capability --integrity per source; confine hook commands to the bundle
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 09:10:04 -04:00
Tom Boucher
f7c6ce7a1f fix(#1461): loader never crashes on a malformed overlay (skip-with-warning); bound the source fetch
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 03:24:47 -04:00
Tom Boucher
e7855bc217 fix(#1459): user-owned consent store gates third-party capability activation; env/cwd in disclosure; loader validator parity (#1473) 2026-06-20 01:59:03 -04:00
Tom Boucher
c866ac1b24 fix(#1462): fail closed without data loss on a corrupt capability ledger; atomic ledger write (#1469) 2026-06-19 19:55:10 -04:00
Tom Boucher
0d56f544d2 feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation (#1458)
* feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation

ADR-1244 Phase 6. Adds the capability matrix reference, generated FROM the committed registry so it
can never drift from the actual capability set:
- scripts/gen-capability-matrix.cjs (--write / --check); --check is a CI drift guard.
- tests/capability-matrix-sync.test.cjs (4 tests): drift guard, buildMatrix==committed, every cap
  present, no placeholders.
- docs/reference/capability-matrix.md regenerated from the registry (release-stable: shows engines.gsd,
  omits the lockstep per-cap version that would churn the file every release).
- Consolidated the duplicate trust-model doc: deleted docs/explanation/the-capability-trust-model.md,
  merged its content into capability-trust-model.md, redirected ~10 references; no stale links remain.
- Diataxis verification (now that gsd capability is a real command): corrected the matrix's third-party
  section — the matrix is the first-party catalogue; the overlay-aware view of installed third-party
  capabilities is 'gsd capability list', not this generated file.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1435): Added changeset for the capability matrix reference

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1435): address code-review — non-vacuous matrix test + generator polish

- capability-matrix-sync.test.cjs: assert the 'security registers a ship:pre gate' precondition
  unconditionally so the extension-point check can never degrade to a vacuous pass on registry drift.
- gen-capability-matrix.cjs: warn (stderr) on an unknown loop point at generation time; rename
  enginesOf -> fmtEngines for consistency with the other fmt* helpers (output unchanged).
- capability-trust-model.md: point the two how-to links at the real files
  (import-a-capability-from-a-url.md, version-a-capability.md) instead of the bare directory.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1435): backfill changeset PR number → #1458

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 12:45:26 -04:00
Tom Boucher
34bc096ec2 feat(#1451): wire gsd capability install/update/remove/list/disable/enable management CLI (#1457)
* feat(#1451): wire gsd capability install/update/remove/list/disable/enable CLI

ADR-1244 D5/D6: the management command was built as a library (capability-lifecycle.cjs
install/upgrade/remove + capability-ledger.cjs) across Phases 3-5 but never wired to a
user-facing command — gsd-tools.cjs 'capability' only handled state/set. This adds the
six subcommands, dispatching to the existing lifecycle/ledger:

- install <spec> [--integrity] [--scope global|project] [--yes] [--shared-file <rel>]…
- update [<id>|--all] [--scope] [--yes] [--shared-file]  (re-resolves recorded source)
- remove <id> [--purge-data] [--scope]  (first-party rejected)
- list [--json]  (first-party + overlay, both scopes, JSON array)
- disable|enable <id>  (activation-state alias of capability set --off/--on)

Scope→runtimeDir mapping matches capability-loader exactly (global=$GSD_HOME||home,
project=project root; caps at <root>/.gsd/capabilities/<id>, ledger at <root>/.gsd-capabilities.json).
Consent is non-interactive: --yes grants; without it an executable install aborts after
printing the disclosure and writes nothing. Best-effort reconcile before each mutation.

Tests: tests/capability-cli.test.cjs (20 behavioral, real resolver via local specs,
GSD_HOME-sandboxed) — install consent/block/usage matrix, list, update round-trip,
remove round-trip + first-party guard, disable/enable, unknown subcommand.
Docs: docs/reference/gsd-capability-command.md reconciled to the real surface
(ledger paths, --shared-file, consent model, disable mechanism, outdated marked planned);
docs/COMMANDS.md gains the gsd capability entry.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1451): resolve adversarial-review findings + root-cause the --raw silent-output bug

Adversarial-review (Codex) fixes:
- capReadStrict passes a malformed strict_known_registries value THROUGH so the trust gate
  fail-closes on it (was silently downgrading to permissive)
- installCapability/upgradeCapability gain an expectedId guard + first-party-id rejection
  (capability-lifecycle.cts): an overlay can't shadow a first-party id, and 'update <id>' can't
  act on a different id if the recorded source was retargeted
- capability update: prints the consent disclosure, exits non-zero on --all partial failure,
  no longer masks the resolved id
- capability remove: ledger-first ordering so an overlay is removable even if it shadows a
  first-party name; first-party guard only fires for ids not in the ledger
- gsd-capability-command.md: disable/enable doc corrected (registry-known ids; overlay toggle
  not yet wired through this path)

Silent-output bug (root cause, not waved off as pre-existing):
- captureStdoutSyncWrites buffered fd-1 output and DISCARDED it on the throw path — any --raw
  command that emitted a result/error envelope then threw (to set a non-zero exit) lost ALL of
  stdout. Now it flushes the captured buffer before re-throwing (exit code preserved).
- cmdCapabilitySet threw via process.exit() (bypassing the capture wrapper entirely); now throws
  ExitError so the wrapper flushes — matches the repo's no-process-exit architecture.
- Regression test: capability disable <unknown> --raw must emit the JSON error envelope on stdout.

Verified: capability suite 165/165, @file/json-errors/phase 183/183, lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1451): address adversarial-review R2 — shared-file confinement, MCP no-clobber, config fail-closed

- confinedSharedFile(): realpath-confine every shared-config write/strip to the scope root (mirrors
  safeRmUnder), so a --shared-file whose parent is a symlink escaping the scope can't write outside it.
- mcpServers shared edits: never overwrite an UNOWNED entry — a name collision with the user's (or
  another capability's) server is skipped, so install/remove can't silently clobber user MCP config
  (hooks already append; the map-keyed mcpServers path was the gap).
- capReadStrict: a PRESENT-but-unparseable .planning/config.json now fails CLOSED (lockdown) instead
  of silently downgrading the strict_known_registries policy to permissive.
- Tests: symlink-escape shared-file writes nothing outside scope; colliding user mcpServers entry
  preserved; unparseable config blocks an external install. capability suite 83/83, lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1451): address code-review — aborted-status robustness + coverage + project-scoped strict doc

- install/update: handle an 'aborted' result independently of the requiresConsent flag so it can never
  fall through to the generic 'blocked: unknown reason' arm (aborted always means consent-needed per
  the lifecycle contract; latent today, hardened for future status additions).
- Clarify capResolveScope comment (project scope === already-resolved cwd) and document that
  strict_known_registries is a PROJECT-scoped policy (read regardless of --scope; no machine-wide
  allowlist) in gsd-capability-command.md.
- Tests: update --all over an empty ledger returns an empty result set (exit 0); a flag value that
  looks like another flag (--integrity --scope) is rejected, not swallowed. CLI suite 33/33, lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1451): FEATURES.md entry #147 + Added/Fixed changesets for the capability CLI

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1451): backfill changeset PR number → #1457

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 12:02:36 -04:00
Tom Boucher
1abebbf4fd feat(#1434): registry-driven dispatch for third-party capabilities (ADR-1244 Phase 5) (#1450)
ADR-1244 Phase 5 (D7). dispatchOverlayCapabilityCommand in gsd-tools.cjs dispatches an installed third-party capability command family via loadRegistry({includeInstalled}), gated on a committed ledger entry (consent) and confined to the capability's install root (defaultRequireFromInstallRoot: bare-.cjs basename + realpath containment, rejects ../ traversal + symlink escape); same own-property/function/sync/ExitError guards as the first-party path. capability-loader records _overlay.commandRoots only for accepted overlay caps with a committed, structurally-valid ledger entry (fail closed). First-party graphify/intel/audit unchanged (already on the registry seam). 3 Codex rounds converged + /security-review (no HIGH) + /code-review (Approve); gsd-test green both platforms; CI green.

Closes #1434.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 23:37:43 -04:00
Behruz Nassre Esfahani
c26947808a feat(#1269): expand same-prefix numeric ID ranges in --phase-req-ids (#1419)
* feat(#1269): expand same-prefix numeric ID ranges in --phase-req-ids

normalizePhaseReqIds treated a range token like "SEL-01..SEL-03" as a single
literal ID, so gap-analysis reported the range string as a missing requirement
even when SEL-01/02/03 existed individually.

Add a per-token range expander (run AFTER the existing split, preserving the
string[] | null | undefined contract): a token matching <PREFIX>-<NN>..<PREFIX>-<MM>
with identical prefixes, ascending bounds, and EQUAL digit width expands to the
individual IDs preserving that width; anything ambiguous stays literal
(fail-closed). Differing-width bounds stay literal so the expander never invents
a zero-padding the author didn't type, and ranges beyond MAX_PHASE_REQ_RANGE
(1000) stay literal as a DoS guard.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1269): add changeset for --phase-req-ids range expansion

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1269): mark changeset docs-exempt (internal flag, no user docs surface)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1269): isolate the DoS-cap branch with a same-width range; doc nits

Review fixes: the AC4 DoS test used REQ-1..REQ-100000, whose differing digit
widths trip the width guard before the cap is reached. Use a same-width
REQ-0001..REQ-1001 (span 1001 > 1000) so the test actually exercises the cap.
Clarify the PHASE_REQ_RANGE_RE capture-group JSDoc and the property-test width
assertion comment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-18 21:59:10 -04:00
Rezolv
dcceb1a004 fix(#1441): antigravity resolver prefers GSD-owned dir over first-existing (#1442)
* fix(#1441): antigravity resolver prefers GSD-owned dir over first-existing

resolveConfigHomeFromDescriptor's dot-home-nested probe returned the first
bare-existing candidate, so a CLI user (~/.gemini/antigravity-cli) who also
had the IDE's ~/.gemini/antigravity dir was silently shadowed to the legacy
dir (probed first). Regression from #217 — pre-#217 returned
~/.gemini/antigravity unconditionally.

Add an optional probeMarker (gsd-core/VERSION) to the dot-home-nested
descriptor: a two-pass probe prefers the candidate GSD installed into, then
bare existence, then probe[0]. Behavior is byte-identical when probeMarker is
absent (windsurf etc. unaffected). Adds detectAntigravityDirAmbiguity() for
installer/operator guidance on already-misinstalled users (auto-relocation
ruled out per ADR-0008's single-configDir migration bound).

Regression tests fail before / pass after: coexistence + marker-priority
cases, end-to-end through the registry descriptor.

Claude-Session: https://claude.ai/code/session_01JU2WB23JLE3QTysPXuFJjB

* chore(#1441): add changeset for antigravity resolver fix

Claude-Session: https://claude.ai/code/session_01JU2WB23JLE3QTysPXuFJjB
2026-06-18 21:47:18 -04:00
Tom Boucher
9219af3360 feat(#1433): capability trust gate + upgrade/compat (ADR-1244 Phase 4) (#1449)
ADR-1244 Phase 4 (D5 trust + D6 upgrade/compat). capability-trust.cjs (disclosure/consent, strict_known_registries, engines+compatVersions, reserved namespace) + capability-lifecycle.cjs (install/upgrade/remove/reconcile; ledger-as-commit-point _pending intent; atomic stage-then-swap; surgical marker-isolated shared-edit strip; owner-token lock) + capability-source promote/skipEnginesGate seams + loader pending-skip + config keys. No sandbox re-derived (consent+integrity+reversibility). 6 Codex adversarial rounds + /security-review (no HIGH) + /code-review; gsd-test green both platforms; CI green. Phase 5 (#1434) wires the CLI dispatch.

Closes #1433.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 21:40:37 -04:00
Tom Boucher
1abb0d3427 feat(#1432): capability source resolver + ledger (ADR-1244 Phase 3) (#1443)
* feat(#1432): capability source resolver + ledger (ADR-1244 Phase 3)

ADR-1244 D3 + D4 — additive, testable-in-isolation modules (the /gsd:capability
install command + consent gate are Phase 4). NEW modules only; not yet wired into
bin/install.js.

- src/capability-source.cts → resolveCapabilitySource(spec, opts): one seam, one
  adapter per source kind — local (fs copy), git (execGit clone+checkout, https/ssh/
  git transports only), npm (execNpm pack --ignore-scripts + tar, NEVER npm install),
  tarball (https download + sha512 integrity-before-extraction + tar), registry (stub).
  SECURITY: install never executes capability code (copy/extract only, --ignore-scripts);
  integrity verified before staging; symlink members rejected at interior, source-root,
  AND tar-member (verbose-listing) layers; tar-slip member paths rejected pre-extraction;
  npm specs with shell metacharacters (incl %) rejected (execNpm uses a Windows shell);
  git ext::/file:// transports + leading-dash/metachar refs rejected; atomic staging
  (.staging→rename, restore-on-failure); full Phase 1/2 validator suite + engines.gsd
  pre-check on the fetched manifest. Test seam _setCapabilitySourceHttpGet.
- src/capability-ledger.cts → per-runtime .gsd-capabilities.json install manifest:
  readLedger/writeLedger (atomic via platformWriteSync)/recordInstall (idempotent,
  prototype-pollution-guarded)/removeEntry/reconcile (orphan report, hardened against
  hostile/non-string/'..' files[] members — never throws, never oracles outside runtimeDir).
- 48 new tests (34 source incl. the full security matrix, 14 ledger). Two Codex
  adversarial rounds; all 12 findings fixed + regression-tested.
- .gitignore + eslint.config.mjs: built capability-source/ledger.cjs git+eslint-ignored
  (ADR-457 #551 migration coverage); CONTEXT glossary + INVENTORY rows + manifest regen.

Closes #1432

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1432): add changeset for capability source resolver + ledger

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 16:23:20 -04:00
Tom Boucher
353f63d170 feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) (#1440)
* feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2)

Promote the registry from a frozen data file to loadRegistry({includeInstalled}),
composing the first-party registry with a validated installed overlay (ADR-1244 D2):

- Extract the conformance validator to a shared runtime-callable module
  (gsd-core/bin/lib/capability-validator.cjs); the generator re-exports it
  verbatim, guarded by a generative-parity test (no build-time/runtime drift).
- capability-loader.cts: loadRegistry({includeInstalled}) composes first-party
  ∪ validated overlay from $GSD_HOME/.gsd/capabilities (global) and
  <root>/.gsd/capabilities (project) via the canonical buildRegistry. First-party
  always wins (id/skill/agent/config/command-family + reserved gsd-/anthropic-
  prefixes); full merged-set cross-capability validation; engines.gsd load-time
  re-gate (skip-with-warning); gate-kind capabilities FAIL CLOSED; fragment-path
  escapes rejected.
- semverSatisfies (hand-written, no dep) for the engines.gsd gate, fail-closed.
- Wire surface/state + loop to the overlay; loop injects a blocking gate for each
  skipped gate-kind overlay (fail-closed).
- cwd-aware overlay config-key federation: config-loader _federatedConfigSchema(cwd)
  + config-schema isValidConfigKey(key, cwd) compose the overlay per loadConfig/
  config-set call (never eager at module load, never wrong-cwd); first-party path
  unchanged with no cwd.
- run-tests.cjs sandboxes GSD_HOME (idempotent — nested spawns reuse it) for test
  hermeticity; capability-loader.cjs git+eslint-ignored (tsc artifact);
  capability-validator.cjs stays linted (#551 migration coverage).

Closes #1431

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1431): add changeset for runtime capability registry overlay

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1431): kill config-schema cwd-aware federation mutants (Stryker ≥52)

The cwd-aware overlay config-key federation added to config-schema.cts
(_capabilityConfigSchema(cwd) + isCapabilityConfigKey/isValidConfigKey cwd
threading) introduced mutable surface uncovered by config-schema's mutation
test set, dropping its score to 39.58% (below the 52 break threshold). Add a
real-overlay-fixture describe block exercising every branch (cwd guard, overlay
loadRegistry, found-branch, first-party fallback, cwd threading); local Stryker
score 39.58% -> 77.08%.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 14:41:20 -04:00
Tom Boucher
b0c774c2e3 feat(#1416): formalize Resolution convention + agent-skills value envelope (Resolution Provenance P3) (#1425)
Narrows P3 of ADR-1411 (Resolution Provenance, epic #1411) based on an
adversarial fit-analysis that showed a single Resolution<T> envelope adopted
by agent-skills, capability-state, and capability-writer fails the deletion
test: configured/reason are meaningless for capability verbs, and
capability-writer's errors[] (operation-not-applied) cannot fold into
warnings[]. The only genuinely shared seam is warnings: string[].

Changes:

- src/resolution.cts: new pure types+builder leaf — exports Resolution<T>
  {value, configured, reason, warnings}, makeResolution<T>() builder, and
  AgentSkillsValue {block, skills_count}. No other src/ imports.

- src/init.cts: cmdAgentSkills --json IR gains additive value:{block,
  skills_count} field (built via makeResolution). All existing flat fields
  (agent_type, block, skills_count, warnings, configured, reason, source,
  degraded) are retained unchanged for back-compat.

- src/capability-state.cts: doc comment on ResolveCapabilityRuntimeStateResult
  naming it the canonical read-verb envelope. No JSON change.

- src/capability-writer.cts: doc comment on SetCapabilityStateResult naming it
  the canonical mutation-verb result (warnings=advisory, errors=operation-
  not-applied). No JSON change.

- CONTEXT.md: new ### Resolution Convention glossary entry after
  ### Resolution Provenance.

- docs/adr/1411-resolution-provenance.md: P3 narrowing amendment appended.

- tests/resolution.test.cjs: 9 unit tests for makeResolution (new).
- tests/agent-skills.test.cjs: 2 P3 tests for value.block/value.skills_count
  and back-compat of all flat fields.

All 277 tests pass (5 suites). npm run lint clean. All lint checks pass.

Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 07:44:52 -04:00
Tom Boucher
484a5b7b86 fix(#1415): loadConfig provenance + agent-skills diagnostic (Resolution Provenance P2) (#1424)
Implements ADR-1411 P2 / #1415, closing #1366.

Part 1 — config-loader.cts:
- Adds `loadConfigResolved(cwd, options) → ConfigResolution { config, source, degraded }`
  with six tagged branch return paths:
  A1 ws+wsconfig → source:'workstream', degraded:false
  A2 no-ws+config → source:'root', degraded:false
  B  ws requested, wsconfig absent → source:'root', degraded:true (intercepts recursive call)
  C  .planning/ exists, no config → source:'builtin-defaults', degraded:false
  D  no .planning/, global defaults readable → source:'global-defaults', degraded:false
  E  no .planning/, no global → source:'builtin-defaults', degraded:false
- `loadConfigResolved` calls `findProjectRoot` at entry to anchor resolution to the
  nearest .planning/ ancestor (cwd-drift fix, heuristic 4 from P1)
- `loadConfig` becomes a one-line delegation: `return loadConfigResolved(cwd, options).config`
- Exports `loadConfigResolved` in the `export =` block

Part 2 — init.cts cmdAgentSkills:
- Imports `findProjectRoot` from `./project-root.cjs`
- Anchors to project root before loading config (fixes #1366 cwd-drift)
- Uses `loadConfigResolved` for provenance; passes projectRoot to buildAgentSkillsBlock
- Computes `configured` + `reason` (AgentSkillsReason enum): 'resolved' |
  'not_configured' | 'configured_empty' | 'configured_unresolved'
- configured_empty and configured_unresolved emit stderr WARNING; not_configured is silent
- --json IR gains: configured, reason, source, degraded (in addition to existing
  agent_type, block, skills_count, warnings)

Tests (TDD):
- tests/config-loader.test.cjs: 8 new provenance tests (RED before impl, GREEN after)
- tests/agent-skills.test.cjs: 7 new diagnostic tests (RED before impl, GREEN after)
- All 306 tests across 4 suites pass (config-loader:33, agent-skills:72, init:105, workstream:96)

Docs:
- CONTEXT.md: Config Loader Module entry updated with loadConfigResolved interface;
  Resolution Provenance entry notes P2 is now implemented
- docs/CLI-TOOLS.md: --json field reference table added for agent-skills

Closes #1366
Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 07:05:15 -04:00
Tom Boucher
d220ba6fd6 fix(#1414): resolve project root from descendant via nearest ancestor .planning/ (Resolution Provenance P1) (#1423)
Add heuristic (4) to findProjectRoot: after heuristics (1)-(3) (sub_repos,
multiRepo, .git+parent-.planning/) are exhausted without a match, perform a
second bounded walk-up within FIND_PROJECT_ROOT_MAX_DEPTH to locate the
nearest ancestor directory containing a .planning/ subdirectory. Returns that
ancestor as the project root, so loadConfig finds the correct config instead
of falling through to defaults when gsd-tools is invoked from a plain
descendant subdirectory of a single-repo project (#1366 cwd-drift gap).

Ordering is load-bearing: the new walk runs AFTER the existing loop so
sub_repos workspaces (where a child sub-repo may have its own .planning/)
still resolve correctly to the parent workspace. The existing own-.planning/
guard (heuristic 0, #1362) and the depth bound (FIND_PROJECT_ROOT_MAX_DEPTH=10)
are both preserved unchanged.

Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 00:50:30 -04:00
Tom Boucher
666d933e16 ci(#1401): add no-adhoc-markdown-parsing rule (fence-strip + section-collect) + grandfather burn-down (#1402)
Tightens the over-broad heading-walk detection: removes heading-walk
entirely and narrows fence-regex to require a multiline body ([\s\S]),
so single-line tests like /^```/ and /^###\s+/ are no longer flagged.
Grandfathers the 10 genuine section-collect sites across state.cts,
milestone.cts, audit.cts, and phase-lifecycle.cts with concrete reasons.
Adds 12 RuleTester tests (3 positive, 9 negative) to eslint-rules.test.cjs.

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 20:14:12 -04:00
Tom Boucher
ce8bcb1b95 refactor(#1398): migrate state.cts section-collects onto markdown-sectionizer seam, byte-identical STATE.md (epic #1372 T6) (#1399)
* refactor(#1398): migrate state.cts section-collect regexes onto markdown-sectionizer seam (epic #1372 T6)

Replace ≈18 hand-rolled `/(heading)([\s\S]*?)(?=stop)/` regex splices in
state.cts with direct `tokenizeHeadings` calls that compute the exact
[bodyStart, stopOffset) span, preserving byte-identical STATE.md output.

Non-migratable site left in place: `cmdStateRecordMetric`'s metricsPattern
captures table-header rows in group 1 — not a standard heading+body shape.
Write orchestration (readModifyWriteStateMd / syncStateFrontmatter /
shouldPreserveExistingProgress / #952 no-op guard) is UNTOUCHED.

Verification: t6-headtohead.cjs head-to-head harness runs 25 ops across 6
STATE.md fixture variants (inline, trailing-blanks, CRLF, no-frontmatter,
nested-acc, post-milestone) against origin/next and reports 0 diffs.
No-op guard confirmed: record-session on recorded:false leaves STATE.md
byte-identical. All 150 state tests and 62 milestone/forensics tests pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#1398): capture T6 state-section-splice characterization; drop throwaway harness

Remove scripts/t6-headtohead.cjs (committed throwaway HEAD-vs-origin/next
byte-compare harness). Capture its coverage as 23 behavioral characterization
tests appended to tests/state.test.cjs, exercising all migrated cmdState*
write-ops across 7 fixture variants (inline, trailing-blanks, CRLF,
no-frontmatter, nested-acc, no-current-pos, post-milestone). Includes the
#952 no-op guard (recorded:false + byte-unchanged assertion) and
CRLF/trailing-blanks edge-case coverage. Also removes the dead
spliceStateSection helper (defined but never called) that was surfacing as
an @typescript-eslint/no-unused-vars warning.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 17:56:17 -04:00
Tom Boucher
308c7be481 refactor(#1396): T5 — migrate uat-predicate + uat onto markdown-sectionizer seam (#1397)
- uat-predicate.cts: replace local _stripFencedBlocks (and its private
  FenceState/StripFencedResult types) with a call to stripFencedCode from
  markdown-sectionizer.cjs (ADR-1372 T5). Both stripFalsePositiveContexts
  step (c) and analyzeMarkdown now route through the seam. The three other
  passes in stripFalsePositiveContexts — frontmatter strip, HTML-comment
  strip, blockquote-line filter — remain caller-side (seam does not do these).
  The unterminatedFence signal consumed by analyzeMarkdown is preserved; it
  is now returned by stripFencedCode (same machine, same contract).

- uat.cts: migrate the ## Current Test, ## Tests, and ## Human Verification
  section-collect patterns onto collectSection/tokenizeHeadings from the seam.
  UAT-specific item parsing (### N. Name blocks, expected/result fields,
  categorization logic) stays caller-side. The HTML-comment strip within the
  Current Test body remains caller-side (UAT document structure, not seam scope).

- tests/markdown-sectionizer.test.cjs: remove the 18-case tautological parity
  guard (DEFECT.GENERATIVE-FIX). Once uat-predicate imports the seam the guard
  compares the seam to itself — removing it is the T5 commitment per ADR-1372.

4-space-indent behavior change (CommonMark correctness improvement): the seam
uses /^( {0,3})/ (CommonMark §4.5 ≤3-space indent); the retired
_stripFencedBlocks used /^(\s*)/ (any indent). A 4-space-indented ``` is no
longer treated as a fence opener (it is an indented code block per CommonMark).
Head-to-head over 9 corpus inputs: 0 diffs on all standard cases; 2 diffs only
on the synthetic 4-space-indent edge cases. No UAT fixture or test in the suite
exercises 4-space-indented fences. The change is a correctness improvement.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 15:44:11 -04:00
Tom Boucher
6e9f8bf50b refactor(#1393): migrate roadmap-parser onto markdown-sectionizer seam (epic #1372 T4) (#1395)
Removes all three inline copies of the fenced-code state machine from
src/roadmap-parser.cts and replaces them with calls to
tokenizeHeadings() from the canonical markdown-sectionizer seam
(ADR-1372 T4).

Changes:
- Drop stripFencedLines() function (copy 1 of 3 — standalone helper)
- Rewrite computeSectionEnd() using tokenizeHeadings() offsets into
  the original content (copy 2 — inline fence loop); the returned
  character offset is preserved exactly for all inputs
- Rewrite getMilestonePhaseFilter() versionOverride fence loop using
  tokenizeHeadings() (copy 3 — inline); sectionEnd offset preserved
- Replace stripFencedLines(roadmap) + unanchored phasePattern.exec()
  with tokenizeHeadings(roadmap) filtered by level and phase-heading
  pattern; headings in inline HTML comments (<!-- ## Phase N: -->)
  are no longer mis-counted (anchored ATX detection is correct)
- Add import { tokenizeHeadings } from ./markdown-sectionizer.cjs

Offset preservation: tokenizeHeadings() records h.offset as the
character index of '#' in the ORIGINAL content; computeSectionEnd()
and the versionOverride path both use h.offset directly as the
section-end character offset — no stripping, no shift.

Corpus head-to-head: 29/30 slots are byte-identical to origin/next.
The 1 diff (getMilestonePhaseFilter phaseCount for the HTML-comment
fixture: 3→2) is an improvement: the old unanchored regex counted
"## Phase 998:" embedded in "<!-- ## Phase 998: ... -->" mid-line;
tokenizeHeadings() correctly requires '#' at line-start (ATX rule).
No existing test asserts on that count; feat-3594 passes unchanged.

All 122 roadmap/milestone/phase tests green.

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 15:13:14 -04:00
Tom Boucher
3a1961ebae refactor(#1390): migrate check-command-router + gap-checker onto markdown-sectionizer seam (epic #1372 T3) (#1392)
- check-command-router: stripCommentsAndFences delegates fenced-code
  stripping to seam's stripFencedCode; HTML-comment stripping stays
  caller-side. extractPlanDesignatedSections replaces hand-rolled
  split(/\r?\n/) + /^#{1,6}\s+/ heading walk with collectSections
  driven by DESIGNATED_HEADINGS_RE.

- gap-checker: parseRequirements checkbox-bullet detection migrates
  to iterateBullets (checkbox markers); **ID** extracted caller-side.
  Table-row path and separator-row skip stay caller-side.

- Adds refactor-1390-t3-characterization.test.cjs (43 behavioral
  tests) that were green before and remain green after.

- T1 fail-loud / could-not-parse gate semantics unchanged (verified
  by decisions.test.cjs 59/59 and direct gate invocation).

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 14:32:43 -04:00
Tom Boucher
75c2e259d0 refactor(#1387): migrate adr-parser onto markdown-sectionizer seam (epic #1372 T2) (#1388)
* refactor(#1387): migrate adr-parser onto markdown-sectionizer seam (T2)

Replace hand-rolled parseSections (split/heading-regex walk/body accumulation)
with collectSections(content, () => true) from the canonical seam. Replace
hand-rolled splitEntries bullet-strip regex with iterateBullets from the seam,
preserving the plain-text-line fallback for byte-identical output. Removes the
last inline heading/bullet scanning from adr-parser.cts; normalizeAdrHeader and
all ADR-specific classification logic are unchanged. 218/218 tests pass before
and after.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#1387): keep splitEntries flat (iterateBullets changed its contract); seam adoption stays in parseSections

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#1387): drop dead preamble reconstruction; add targeted adr-parser mutation tests

parseSections' preamble reconstruction block (heading: null entry) was dead code:
both consumers (parseAdrMarkdown and parseStatusFromSections) skip heading: null
sections immediately on entry. Confirmed via analytical trace and zero-diff corpus
head-to-head across all 44 docs/adr/*.md files.

Adds 11 targeted behavioral tests to kill cheap surviving mutants:
- pushUnique intra-values dedup (kills seen.add removal mutant)
- body split/join round-trip with multi-line prose and entries
- parseStatusFromSections [0] indexing (only first line determines status)
- classifyHeader equality vs prefix-match boundary (exact match, prefix match,
  synonym+letter non-match)
- goal section prose vs entries distinction (bullet markers preserved in context)
- normalizeAdrHeader non-word char removal (parens and slash behavior)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 13:45:05 -04:00
Tom Boucher
e58b5e1721 fix(#1364): decisions adopt markdown-sectionizer seam + fail-loud coverage gate (epic #1372 T1) (#1386)
* test(#1364,#1365): add decisions regression tests (fail-first proof)

Adds tests/decisions.test.cjs with:
- #1364 recall tests: parseDecisions from markdown-header + em-dash bullets
  (these FAIL on pre-T1 code, proving the bug is present before the fix)
- #1365 fail-loud tests: check.decision-coverage-plan must return passed:false
  for decision-shaped but 0-extracted content (FAIL pre-T1, gate silently passed)
- extractDecisions outcome enum tests (could-not-parse/none-present/parsed)
- Parser QA matrix: CRLF, unicode headings, fenced-code suppression, both bullet forms
- Boundary/threshold tests at limit-1 (0), limit (1)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1364,#1365): adopt markdown-sectionizer seam in decisions.cts; add fail-loud gate

#1364 — Recall: decisions.cts now uses the seam's extractTaggedBlocks and
collectSection for the markdown-header fallback path. Em-dash bullet form
(- **D-NN — title** body) is now recognised alongside the existing colon form.

#1365 — Fail-loud: adds extractDecisions() returning a typed DecisionExtraction
{ decisions, outcome } where outcome is 'parsed' | 'none-present' | 'could-not-parse'.
The blocking gate (cmdDecisionCoveragePlan) now treats could-not-parse as
passed:false with a format-mismatch reason instead of the prior silent passed:true/skip.
gap-checker runGapAnalysis surfaces 'extracted 0 of N — possible format mismatch'
for could-not-parse instead of 'No requirements or decisions to check'.

parseDecisions remains a thin delegate over extractDecisions, so all existing
callers are unaffected.

Seam adoption: stripFencedCode (seam), extractTaggedBlocks(content,'decisions') (seam),
collectSection(content, /decisions?/i, {levelBounded,stripFences}) (seam).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1364,#1365): tighten could-not-parse, parse-miss fail-loud, curly-quote discretion, gap-checker FIX D

FIX A: empty <decisions> scaffolds and all-prose sections no longer return
could-not-parse; outcome is none-present unless the block/section contains
a \bD- token or a parse-miss, preventing false blocks on legitimate phases.

FIX B: parseDecisionLines now tracks parse-misses (D-NN-shaped bullets that
fail both regexes); extractDecisions returns could-not-parse when parseMisses>0
even if some decisions parsed — silent drops no longer mask format errors.

FIX C: curly-quote normalization regex now includes actual U+2018/U+2019
characters so '### Claude's Discretion' (curly apostrophe) correctly yields
trackable:false (regression vs pre-T1 behavior).

FIX D: gap-checker runGapAnalysis surfaces the decision could-not-parse
format-mismatch signal independently of whether requirements items exist —
previously masked inside `if (items.length === 0)`.

Adds 14 behavioral regression tests (fail-first verified manually before fixes).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1365): fail-loud gate on parse-miss regardless of covered decisions

Change the `could-not-parse` guard in `cmdDecisionCoveragePlan` and
`cmdDecisionCoverageVerify` from `decisions.length === 0 && outcome ===
'could-not-parse'` to fire on `outcome === 'could-not-parse'` alone.

Previously a CONTEXT.md with a valid D-01 (covered by the plan) plus a
malformed D-02 (parse-miss) would skip the guard (length === 1), proceed
to coverage, find D-01 covered, and silently return passed:true — hiding
the D-02 parse-miss entirely.

Adds a gate-level fail-first test that places D-01 into a ## Must Haves
section (DESIGNATED_HEADINGS_RE match) so coverage of D-01 would pass on
its own, proving the only path to passed:false is the parse-miss fix.
Also adds the matching verify-side advisory assertion.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1364,#1365): add Fixed changeset (pr:0 placeholder)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1364): backfill changeset PR number (1386)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 12:35:38 -04:00
Tom Boucher
7ca8011cd9 refactor(#1373): add canonical markdown-sectionizer seam (epic #1372 T0) (#1381)
* refactor(#1373): add markdown-sectionizer seam (ADR-1372 T0)

Establishes the canonical markdown-structure parsing seam per ADR-1372.
No existing parsers are modified; this is the foundational T0 tier only.

- docs/adr/1372-markdown-sectionizer-seam.md: Accepted ADR defining the
  seam interface, the tiered migration plan (T0-T7), and the prohibition
  enforcement approach (no-adhoc-markdown-parsing ESLint rule in T7).
- src/markdown-sectionizer.cts: Pure module, Node built-ins only.
  Exports: stripFencedCode (CommonMark-correct state machine ported from
  uat-predicate.cts _stripFencedBlocks, CRLF-safe, unterminatedFence
  signal), tokenizeHeadings (ATX headings outside fenced blocks),
  collectSections (line-by-line predicate-driven section collection),
  collectSection (single named section, levelBounded stop, optional
  stripFences), iterateBullets (dash/checkbox/numbered + continuation).
- tests/markdown-sectionizer.test.cjs: 54-test behavioral suite covering
  the parser QA matrix (LF/CRLF, Unicode headings, headings-inside-fences,
  unterminated fences, nested levels, all bullet markers, continuation
  lines, empty/non-string input) plus 4 fast-check property tests
  (idempotence, output shape, never-throws, length monotonicity).
- CONTEXT.md: Markdown Sectionizer glossary entry added (PR review gate).

Tests: 54 pass, 0 fail. Existing adr-parser + uat-passed tests: 22 pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#1373): add extractTaggedBlocks + replaceSection to seam; register inventory

- src/markdown-sectionizer.cts: extend Section type with bodyStart/bodyEnd offsets;
  add extractTaggedBlocks(content, tagName) (inner text of <tag>…</tag> blocks,
  tagName regex-escaped, caller decides fence-stripping) and replaceSection(content,
  section, newBody) (pure character-offset splice for read-modify-write callers);
  update collectSections/collectSection to populate bodyStart/bodyEnd.
- tests/markdown-sectionizer.test.cjs: add 33 new behavioral tests for
  extractTaggedBlocks, replaceSection, and a DEFECT.GENERATIVE-FIX parity guard
  that asserts stripFencedCode and uat-predicate's _stripFencedBlocks agree on a
  shared 9-item corpus; documents the known 4-space-indent divergence.
- docs/adr/1372-markdown-sectionizer-seam.md: list extractTaggedBlocks and
  replaceSection in §"The seam".
- CONTEXT.md: update ### Markdown Sectionizer glossary entry with the two new exports.
- docs/INVENTORY.md: add markdown-sectionizer.cjs row (alphabetically between
  loop-resolver and milestone).
- docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest --write.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1373): clear no-unsafe-assignment + unused-var lint in markdown-sectionizer

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1373): correct section offset/round-trip + CommonMark heading/fence edges; register eslint coverage

FIX 1 (CRITICAL): Enforce content.slice(bodyStart,bodyEnd) === body invariant in both
collectSection and collectSections. bodyEnd is now bodyStart + body.length instead of
the raw stop-line offset, eliminating the trailing-newline overcounting that caused
replaceSection to drop separator newlines (## A\nbody## B gluing bug).

FIX 2 (MED): tokenizeHeadings now accepts ≤3-space indent (CommonMark §4.5) and empty
ATX headings (## / ##   ), text=''. 4-space indent correctly excluded.

FIX 3 (MED): collectSection gains stopAtLevel option — stops at the next heading whose
level ≤ stopAtLevel, independent of the opener's level. Enables state.cts ## sections
that also stop at ### without abusing levelBounded.

FIX 4 (MED): Backtick fence opener info string must not contain a backtick (CommonMark).
Applied in both stripFencedCode and tokenizeHeadings fence state machines. Tilde fences
unaffected.

FIX 5 (LOW): "byte offset" → "character (string-index) offset" in HeadingToken / Section
doc comments.

FIX 6 (LOW): extractTaggedBlocks doc comment documents nested-tag non-support; test locks
the non-greedy close-at-first-</tag> behavior.

FIX 7: Add gsd-core/bin/lib/markdown-sectionizer.cjs to eslint.config.mjs ignores so
tests/551-eslint-bin-lib-coverage.test.cjs passes (3/3).

Tests: 107 pass / 0 fail (was 87; +20 new tests for FIX 1–4, 6).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1373): gitignore tsc-built markdown-sectionizer.cjs (ADR-457 build-at-publish)

The seam's compiled artifact must be a build-at-publish output like every other
src/*.cts->bin/lib/*.cjs module (decisions, core, state, ...), not a committed
file. Add it to the ADR-457 ignore list and untrack it; build:lib/CI regenerate it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 11:14:17 -04:00
Tom Boucher
66085d0080 fix(#1374): surface diagnostic when configured agent skills all fail to resolve (#1376)
* fix(#1374): surface diagnostic when configured agent skills all fail to resolve

buildAgentSkillsBlock returned '' (only ad-hoc per-path stderr warnings) when an agent configured via agent_skills had paths that all failed to resolve — missing SKILL.md, unsafe path, invalid global name, OR a malformed (non-string/non-array) value. query agent-skills --json reported skills_count>0 with an empty block and no machine-readable signal, so a fully-dropped configuration was indistinguishable from a resolved one.

Thread an optional diagnostics collector through buildAgentSkillsBlock: route every skip warning through a warn() helper (stderr + collector), flag truthy-but-malformed config values, emit an aggregate warning when configured paths resolve to zero skills, and surface the collected reasons in a new warnings[] field on the query agent-skills --json IR. Empty arrays and falsy values stay silent (skills_count is honestly 0). skills_count semantics unchanged. Docs updated for the new IR field.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1374): backfill changeset PR number (#1376)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 10:16:51 -04:00
Tom Boucher
120f85164b feat(#1355): detect-and-warn guard for claude-code agent-teams (#1371)
* feat(#1355): detect-and-warn guard for claude-code agent-teams

GSD's multi-agent orchestration can stall under claude-code's experimental
agent-teams (a subagent's completion fails to route to the orchestrator). Per
the maintainer decision, the accepted scope is a read-only detector + one
non-fatal warning — NOT the declined run_in_background/TaskOutput conversion.

- New Teams Status Module (src/teams-status.cts → gsd-core/bin/lib/teams-status.cjs):
  pure resolveTeamsStatus({runtime, env}) + thin CLI cmdTeamsStatus reusing
  resolveRuntime. active = strictly-truthy env flag AND runtime === 'claude'.
- Wire `gsd-tools query teams-status [--active]` (read-only; no capability
  registration needed — conformance gates govern features, not query commands).
- One non-fatal warning in plan-phase.md before the first Agent spawn, gated on
  `query teams-status --active`; zero behavior change on non-claude/teams-off.
- Hermeticity: clear CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS in run-tests.cjs +
  SESSION_ENV_KEYS. Docs reference + CONTEXT.md glossary. Built lib gitignored.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1355): add changeset for teams-detect guard

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1355): bump plan-phase.md workflow size baseline (+407B for teams warning)

The non-fatal agent-teams warning block added to plan-phase.md grew it
92759 → 93166 bytes, past its committed per-file baseline ratchet. The growth
is small, deliberate, and still well under the workflow tier hard cap. Regenerate
the baseline via `npm run size:baseline` (only plan-phase.md changed).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1355): register teams-status.cjs in the inventory manifest

The new teams-status CLI module is a tracked surface; regenerate
docs/INVENTORY-MANIFEST.json (cli_modules family) via
gen-inventory-manifest.cjs --write so the inventory-manifest-sync gate passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 08:52:23 -04:00
Tom Boucher
c03f3cc6af fix(#1348): canonicalize Codex hooks.json writes to the nested { hooks } shape (#1363)
* fix(#1348): canonicalize Codex hooks.json writes to the nested { hooks } shape

reconcileCodexHooksJsonEvent preserved whatever shape it read, so on an empty,
absent, or legacy top-level hooks.json it wrote top-level event keys
(`{ "SessionStart": [...] }`) that current Codex (deny_unknown_fields) rejects,
instead of the canonical `{ "hooks": { "SessionStart": [...] } }`.

- Lift any top-level event arrays (legacy, empty, or mixed nested+top-level)
  into the nested `hooks` table, merging same-named events so no user/legacy
  entry is dropped and no stray top-level event key survives. Mirrors
  reconcileCursorHooksJson.
- Collapse an empty hook table back to `{}` so removal on an absent file does
  not write a spurious `{ "hooks": {} }`.
- Read path still tolerates both shapes; dedup/removal unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1348): add changeset for Codex hooks.json canonicalization

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 08:51:11 -04:00
Tom Boucher
c53fd1f654 fix(#1324): resolve glued phase tokens (#1353) 2026-06-16 21:55:58 -04:00
Tom Boucher
aab26c7bf4 fix(#1343): parse decision bullets with text before the colon (#1358)
* fix(#1343): parse decision bullets with text before the colon

parseDecisions() silently dropped any `- **D-NN ...:**` decision bullet
whose header had freeform text (a parenthetical, em-dash, or prose) before
the `:**`, so the blocking check.decision-coverage-plan gate computed
coverage over a narrowed set and reported a false pass.

- Broaden bulletRe to tolerate a freeform run before the colon while
  preserving the optional [bracket] tag capture (drives `trackable`).
- Add a parse-miss guard: a line that looks like a D-NN bullet but still
  fails the regex flushes the current decision and warns instead of
  vanishing — the gate-integrity floor.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1343): add changeset for decision-coverage false-pass fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1343): relocate decision-parser regression into owning module test file

CI's lint-regression-test-names bans new bug-NNNN-*.test.cjs files. Move the
9 regression cases from tests/bug-1343-parsedecisions-drop.test.cjs into the
owning parser test file tests/post-planning-gaps-2493.test.cjs (which already
exercises parseDecisions) and delete the banned file.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-16 21:55:51 -04:00
Rezolv
00c05eb717 Merge branch 'next' into feat/1279-fail-first-prover 2026-06-16 17:35:55 -04:00
Tom Boucher
a0dbf8bbdf fix(#1319): use portable Claude skill effort (#1352) 2026-06-16 15:30:17 -04:00
Tom Boucher
c20d741dc9 fix(#1316): preserve prose STATE phase names (#1351) 2026-06-16 15:11:23 -04:00
Tom Boucher
284dc7bc44 fix: resume UAT checkpoint from paused placeholder (#1350) 2026-06-16 14:47:17 -04:00
Dave
56d4a1bf39 enhance(#1279): project check_violation_fixture scalar — #1278 locate + #1279 proof compose end-to-end (#1346)
Delivers option (a) from the #1314 maintainer review: thread a fourth flat
scalar check_violation_fixture through the projection so a prohibition authored
at spec-phase machine-proves fail-first and greens through the deterministic
path alone — zero hand-authoring at verify time.

- src/probe-core.cts: Prohibition gains check_violation_fixture?; projectProhibitions
  emits it (both kinds) ONLY for a well-formed descriptor and ONLY when non-empty
  (blank/absent -> projects absent -> producer hard-gates, never a partial green).
- src/prohibition-enforcement.cts: descriptorFromProjection reads it back into
  violationFixture via the same numeric-coercion-safe scalar() normalizer.
- Tests (RED-first, proven non-vacuous by reverting both src edits): CHK-02(#1346)
  projection emit, CHK-08(#1346) read-back, CHK-03(D) example round-trip, the
  fast-check round-trip property extended to the 4th scalar (the contract trek-e
  blocked #1301 on), and a real-subprocess COMPOSE capstone greening end-to-end
  through project -> descriptorFromProjection -> default prover+runner.
- Docs flipped from 'hard-gates until #1346' to 'composes end-to-end': verify-phase.md,
  prohibition-probe.md, spec-phase.md authoring, ADR-550 addendum, changeset.
  #1346 now tracks only the node-test causation residual.

190 affected-suite tests green; eslint + tsc clean; size baseline regenerated.
2026-06-16 14:00:49 -04:00
Tom Boucher
6e242bd76a fix: allow quick worktree parent plan base (#1347) 2026-06-16 14:00:06 -04:00
Dave
3cbbcdab42 Merge branch 'next' into feat/1279-fail-first-prover 2026-06-16 13:06:37 -04:00
Dave
33b6ee6f0c fix(#1279): node-test prover fail-closes on missing violationFixture + honest projection docs
Addresses the #1314 maintainer review (trek-e):

- Major 1 (fail-OPEN): defaultProveFailFirst's node-test branch only guarded
  `if (!fixture)`. A missing/typo'd/stale violationFixture made GSD_PROHIB_SUBJECT
  point at a missing file; an honest negative test threw ENOENT *inside its
  callback* (a failing test named distinctly from the file), which
  isNonVacuousNodeTestRed accepted as proof -> a green forged from a setup crash.
  Now requires fs.existsSync(path.resolve(cwd, fixture)) before spawning, symmetric
  with the lint-rule path's file-result guard. Regression test pins it (RED without
  the guard); a second test pins cwd-relative fixture resolution.

- Major 2 (misleading prose): the #1278 projection carries no violationFixture, so
  the deterministic-locate path always hard-gates (fail-closed) until a
  check_violation_fixture scalar is threaded through. verify-phase.md and
  prohibition-probe.md no longer read as if the projected path produces greens;
  the ADR-550 addendum records both items. Tracked as follow-up #1346.

- Documented residual: existence is necessary but not sufficient (a red caused by
  the env being set vs the subject's content); recorded as a constraint, in #1346.

- Nit: stale 'NOT attested fail-first' comment -> 'NOT machine-proven fail-first'.

64 tests pass; eslint + tsc clean; changeset valid.
2026-06-16 13:05:20 -04:00
Tom Boucher
80014109c1 fix: preserve state patch progress counters (#1345) 2026-06-16 11:52:09 -04:00
Tom Boucher
ee9ef7a8e4 Merge branch 'next' into feat/1279-fail-first-prover 2026-06-16 11:32:57 -04:00
Tom Boucher
704d7bc2a6 fix(#1263): restore init phase requirements from flat Phase Details (#1344)
* fix: resolve init phase details fallback

* chore: add changeset for phase details fallback
2026-06-16 11:03:47 -04:00
Tom Boucher
3a53255244 refactor(#1308): consolidate config-key precedence engine into single owner (#1322)
Make src/capability-activation.cts the sole owner of the four-level config-key
precedence walk via a new raw-value primitive resolveConfigKey(dotKey,{config,
cwd,registry}); the boolean wrapper _resolveActivationValue and loop-resolver's
resolveConfigValues are both rebuilt on it. loop-resolver deletes its
byte-identical copy and inline re-walk and imports the engine.

resolveCapabilityRuntimeState no longer returns registry/config (leaked internal
detail); the caller loads one fail-closed config snapshot and threads it in via a
new optional configOverride param, so capability `active` and hook when/configValues
resolve against the same object. capability-writer requires the registry module
directly.

Adds a DEFECT.GENERATIVE-FIX parity gate (identity + behavioral matrix +
end-to-end resolveLoopHooks configValues) that fails if the two precedence
surfaces ever diverge. Pure internal refactor, no user-facing change.

Closes #1308

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 02:38:47 -04:00
Tom Boucher
2d264da661 enhance(#968): region-scoped negative-grep idiom + cross-task conflict warning (#1320)
Adds a region/function-scoped negative-grep idiom to the gsd-planner verification guidance plus a warn-only `validate_plan` check (`scanFileWideNegativeGateConflict`) that flags when a task's file-wide negative grep bans a construct a sibling task legitimately requires elsewhere in the same file. ReDoS-safe (linear, no RegExp on author patterns); region-scoped gates are exempt. Warn-only — never errors, never flips `valid`.

Closes #968

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-16 02:20:34 -04:00
Dave
49bef1927e Merge remote-tracking branch 'origin/next' into feat/1279-fail-first-prover
# Conflicts:
#	docs/adr/550-spec-phase-probe-contract.md
#	gsd-core/workflows/verify-phase.md
#	tests/workflow-size-baseline.json
2026-06-16 00:32:22 -04:00
Dave
da52c53f34 enhance(#1279): harden node-test fail-first proof to require a NON-VACUOUS red
A violation fixture that crashes the negative test at load emits a file-named
# fail 1 — a crash, not the assertion firing red. Require a failing test named
distinctly from the file (isNonVacuousNodeTestRed), symmetric with the clean-pass
non-vacuity guard and the lint-rule specific-rule-id requirement. Closes the one
soundness asymmetry surfaced by adversarial review (code-review IN-01).
2026-06-16 00:18:38 -04:00
Tom Boucher
14e0709eed refactor(#1307): gate intel on isCapabilityActive + loop-resolver honors capability active (#1315)
* refactor(#1307): gate intel on isCapabilityActive + loop-resolver honors capability active

Part A: intel's command gate moves from config-only isIntelEnabled to the
shared isCapabilityActive('intel', cwd) — a consistency cutover (intel has
skills:[] so its tri-state collapses to the intel.enabled config leg; the gate
now flows through the resolver's precedence + runtime-aware resolution).
Part B: loop-resolver hook rendering now gates on capability state.active
(=== true, fail-closed) instead of state.enabled, so the capability config
gate is honored by the hook consumer, not just per-hook 'when'. active is now
required in the loop-resolver input types. Regression test proves a config-
disabled (active=false) capability's unconditional hook is not rendered.
Part of #1302.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1307): add changeset for intel + loop-resolver active gate

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 00:12:54 -04:00
Rezolv
e50ead7ad2 enhance(verify-phase): deterministic auto-locate of the prohibition check descriptor (#1278) (#1301)
* test(1278): RED-first descriptor parity + fail-closed guards + CHK-07 byte-stability (wave 1)

- CHK-03 (RED): extend PROB-14 parity in prohibition-probe.schema.test.cjs to carry the flat
  check_kind/check_target/check_rule scalars through project->write->parseMustHavesBlock; the
  non-droppable check_kind-presence assertion is the load-bearing RED trigger (fails because
  projectProhibitions strips check_* on the current build).
- CHK-07 (GREEN forward-guard): probe-core.test.cjs pins descriptor-less byte-stability +
  dispositionForProhibition fail-closed policy, with a t.todo marker forward-locking plan 01-02.
- CHK-06 (RED): prohibition-enforcement.test.cjs asserts descriptorFromProjection export +
  fail-closed on absent/partial/unknown descriptors via the projection adapter (RED until 01-03).
- No src/*.cts or .cjs edits; no new test files; lint-test-file-count clean.

* feat(1278): add optional flat-scalar check descriptor fields to Prohibition interface (wave 2)

- check_kind?/check_target?/check_rule? mirror CheckDescriptor.kind/target/rule (minus caller-attested failFirst, #1279)
- optional so existing Prohibition consumers compile unchanged

* feat(1278): project check descriptor as flat scalars in projectProhibitions (wave 2)

- emit check_kind/check_target (+ check_rule only for lint-rule with a rule) when descriptor well-formed
- under-specified/descriptor-less items project byte-identically (CHK-07); flat scalars ride existing parseMustHavesBlock continuation-KV path (no parser rewrite)
- add CHK-02 probe-core unit cases pinning the projection
- turns CHK-03 parity test GREEN; dispositionForProhibition untouched

* feat(1278): descriptorFromProjection read-back adapter feeds fail-closed locate (wave 3)

- Add descriptorFromProjection(projected) -> CheckDescriptor | null to
  src/prohibition-enforcement.cts: renames the projected flat scalars
  check_kind/check_target/check_rule -> {kind,target,rule?}, or null when
  the descriptor is absent/non-object (no check_kind key).
- failFirst is NEVER sourced from the projection (stays caller-attested; #1279).
- rule is set only when check_rule is a non-empty string; the adapter does NOT
  re-validate kind/target/rule — an under-specified descriptor reconstructs to
  one the EXISTING runProhibitionEnforcement LOCATE guard rejects (located:false,
  never green). The merged #1259 guard stays the single source of fail-closed truth.
- Turns the RED CHK-06 fail-closed tests (plan 01-01) GREEN end-to-end; CHK-03 /
  CHK-07 stay green. CheckDescriptor type, locate guard, dispositionForProhibition,
  and parseMustHavesBlock are unchanged (additive +36/-0).

* feat(1278): verify-phase locates prohibition check from projected descriptor (wave 3)

- request.check kind/target/rule sourced from projected check_kind/check_target/check_rule via descriptorFromProjection, not verifier invention (CHK-05)
- replaces the #1278 author-supplied / tracked-follow-up note with the delivered deterministic-locate behavior
- preserves fail-closed routing: absent/partial descriptor -> never green, hard-gate in both modes
- failFirst stays a verify-time caller attestation; #1279 bounds the remaining fail-first proof

* feat(1278): spec-phase captures wired-check descriptor on test-tier resolution (wave 3)

- Step 5.6 'Keep it' / verification: test path captures check_kind/check_target/check_rule, projected onto must_haves.prohibitions for verify-phase deterministic locate (CHK-04)
- SOFT capture: a test-tier prohibition without a descriptor is still allowed (no hard authoring block); stays fail-closed/flagged downstream
- --auto captures only an unambiguous descriptor, never fabricates a check path
- failFirst NOT captured at spec-phase (verify-time attestation; #1279)
- PROB-06 soft-gate + text-mode (PROB-09) behavior unchanged

* chore(1278): re-baseline workflow size for grown verify-phase + spec-phase prose (wave 3)

- spec-phase.md 28438 -> 30343 (+1905), verify-phase.md 35362 -> 36498 (+1136)
- regenerated via npm run size:baseline (no hand-picked numbers); growth is the #1278 deterministic-locate + descriptor-capture prose
- workflow-size-budget guard green (122/122)

* docs(1278): ratify optional check descriptor in dated ADR-550 addendum + type:Changed changeset

- Append dated 2026-06-15 ADR-550 addendum ratifying the D3 prohibition-item
  shape extension (optional flat-scalar check_kind/check_target/check_rule)
- Document flat-scalar rationale, deterministic projection/read-back,
  fail-closed on partial/invalid/absent, #1279/policy out-of-scope
- Add .changeset/1278-prohibition-check-descriptor.md (type: Changed)

* docs(1278): document optional check descriptor in prohibition-probe reference + FEATURES

- Add 'Optional wired-check descriptor (deterministic locate, #1278)' section
  to the prohibition-probe reference (flat-scalar keys, projection/read-back,
  fail-closed + backward-compat, failFirst stays attested)
- Add deterministic prohibition-check descriptor source entry to FEATURES.md
- No CONTEXT.md glossary change: descriptor reuses existing wired-check /
  verification:test vocabulary, no new glossary term introduced

* fix(1278): pass packaging gates — changeset pr field + retired slash-form fix

- Add required pr: 1278 to changeset (lint:changeset MISSING_PR hard requirement;
  plan's 'omit if unknown' was inaccurate — issue number per #1259 convention,
  updated to real PR number when opened) [Rule 3 - blocking]
- Fix retired /gsd-spec-phase -> /gsd:spec-phase at verify-phase.md:83 (wave-3
  prose; caught by slash-namespace invariant #3443/bug-2543, blocked CHK-09
  full-suite-green) [Rule 1 - bug]
- size:baseline + INVENTORY manifest verified in-sync post-build (no diff)

* docs(1278): add check descriptor + descriptorFromProjection to CONTEXT.md prohibition glossary

* fix(1278): harden descriptorFromProjection round-trip (numeric-coercion + stray-rule) per review

- MD-01/LW-01: narrow projected scalars to primitives + String()-coerce, so a
  numeric-looking check_target (parseMustHavesBlock coerces ^\d+$ to number)
  reconstructs as a string and locates instead of silently un-locating; no
  as-string type-lie, satisfies no-base-to-string.
- LW-02: attach rule only for the lint-rule kind (drop a stray node-test rule).
- LW-03: document the optional check_* keys in the reference Output schema.
RED->GREEN tests added in prohibition-enforcement.test.cjs.

* chore(1278): set changeset pr to 1301

* test(1278): add fast-check property for the check-descriptor round-trip + fail-closed (trek-e review)

RULESET.TESTS.property-based-testing: the projectProhibitions -> render ->
parseMustHavesBlock -> descriptorFromProjection chain is a bijective/transformation
contract. Adds 2 fc properties to tests/probe-core.property.test.cjs (no new file;
ratchet stays at 2 for probe-core):
- well-formed descriptors survive the round-trip across the full string domain
  incl. the numeric-coercion case (target/rule reconstruct as strings);
- under-specified/invalid descriptors (absent / target-less / rule-less /
  unknown-kind) are always fail-closed (never green, flagged, unlocated).
Stability is asserted at the descriptorFromProjection layer (the raw parse step is
intentionally lossy for numeric scalars; the shared parser is unchanged).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-16 00:01:03 -04:00
Dave
9aa5e88186 docs(#1279): correct two stale CALLER-ATTESTED docstrings to machine-proven 2026-06-15 23:24:15 -04:00
Tom Boucher
e3b829e765 refactor(#1306): gate graphify on isCapabilityActive (tri-state, runtime-aware) (#1313)
* refactor(#1306): gate graphify on isCapabilityActive (tri-state), not config-only

graphify's command gate moves from the config-only isGraphifyEnabled to the
shared isCapabilityActive('graphify', cwd) — so graphify is off unless installed
AND surfaced AND graphify.enabled. Fixes a latent claude-hardcoding in the
resolver: resolveCapabilityRuntimeState now detects the active runtime via
resolveRuntime(cwd) (GSD_RUNTIME -> config.runtime -> 'claude') so non-Claude
runtimes (Codex/Cursor) read their own surface, not ~/.claude. Hermetic
regression test proves config-on+unsurfaced -> disabled; cross-runtime test
proves GSD_RUNTIME=codex honors CODEX_HOME. Gate fails closed on every error
path. Part of #1302.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1306): add changeset for graphify tri-state gate

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 23:19:21 -04:00
Dave
711ef092d7 enhance(#1279): wire machine-proven fail-first into the green verdict + evidence method
- passed = proof.provenFailFirst === true && run.passed === true (FF-01 flip)
- caller failFirst attestation removed from the green AND (FF-08 demote)
- proveFailFirst seam defaults to defaultProveFailFirst, no-throw-wrapped (FF-04/FF-05)
- evidence carries failFirstProof: proof.method on a proven green (FF-07)
- CheckDescriptor.failFirst demoted to non-authoritative hint; kept for route-JSON shape
2026-06-15 22:47:16 -04:00