Commit Graph

98 Commits

Author SHA1 Message Date
Dave
31b822b312 fix(1259-01): close adversarial review findings — genuine enforcement, honest fail-first scope
Adversarial pre-submission review found the injected-runCheck tests masked a
non-functional real runner. Fixes:

- BL-01 (false green on vacuous test): the node-test runner now parses the TAP
  summary and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail) AND a
  reported test named distinctly from the file — node --test counts an empty file
  as one passing test, so counts alone could not catch it.
- SF-01 (lint anchor never greened): the lint-rule runner now runs the project
  eslint as --format json and filters by ruleId, so plugin rules (local/*) load
  via the flat config — bare --rule cannot load a plugin. local/no-source-grep
  now genuinely greens (covered by a real, non-injected test).
- BL-02 (tautological fail-first): the runner no longer echoes the caller's
  failFirst as if confirmed. failFirst is documented as caller-ATTESTED; the
  producer requires attestation + a genuine non-vacuous pass. Machine-proven
  fail-first (needs a violation fixture) is flagged as a tracked follow-up in
  ADR-550, the changeset, FEATURES, the reference doc, and verify-phase.
- SF-02: added real-runner end-to-end tests (no injected runCheck) + pure,
  exported parse/filter helpers (parseNodeTestSummary, tapTestNames,
  eslintJsonHasRule, eslintFileResultCount) so the shipping branches are
  mutation-pinned.
- NIT-01/02: LOCATE guard rejects empty-string rule and unknown kinds.
- Hardening: spawn checks with NODE_TEST_CONTEXT/NODE_OPTIONS scrubbed so an
  ambient test-runner context cannot corrupt a verify-time result.
- Docs reconciled to the shipped behavior (no 'confirms fail-first' overclaim).
2026-06-15 13:38:15 -04:00
Dave
676436258a fix(1259-01): lint-rule real runner — keep rule id distinct from lint target
The default lint-rule runner passed check.target as BOTH the --rule id and the
eslint path, so it could never pass (eslint tried to lint a file named after the
rule). Add a distinct check.rule field (rule id) vs check.target (path to lint),
extract a pure exported buildLintArgs() so the mapping is mutation-testable
without spawning eslint, fail-closed on a lint-rule missing its rule id, and carry
the rule into enforcement evidence. Updates verify-phase descriptor docs.
2026-06-15 13:10:49 -04:00
Dave
ce01e1376b docs(1259-01): wire verify-phase consumer + ADR-550/FEATURES/reference + changeset
- verify-phase.md: replace test-tier 'fail-closed/deferred' bullet with the check prohibition-enforcement
  enforcement step (locate -> fail-first -> run -> evidence -> green-or-hard-gate); update determine_status tree
- ADR-550 addendum: mark D5d enforcement half LANDED (#1259); cite ADR-857 open-question §147 + D6 (core verify rail)
- FEATURES §146: enforcement wording + add REQ-PROHIB-07; keep REQ-PROHIB-06 intact
- references/prohibition-probe.md: test-tier enforced + hard-gates via check prohibition-enforcement
- changeset (type: Changed) with the D5 '2 no-source-grep invalid cases, not 96' correction
2026-06-15 12:47:39 -04:00
Rezolv
3556450b0d feat(spec-phase): surface zero-classification edge-probe requirements as unclassified candidates (#1110) (#1117)
Edge-probe now surfaces a zero-classification requirement (non-empty prose, no
shape cue matched, no `shapes` override) as a single soft `unclassified — review
manually` candidate instead of silently dropping it — the exact blind spot the
probe exists to catch. Dismissible like any edge; the `shapes: []` opt-out stays
silent; `TAXONOMY` (the closed 8 categories) is unchanged. Under `--auto` the
candidate is left `unresolved`, never auto-`backstop` (a missing shape is not
evidence an edge exists).

Closes #1110
2026-06-14 21:44:21 -04:00
Rezolv
395fb519e7 feat(spec-phase): prohibition probe — surface "must-NOT" constraints (#644) (#1149)
Adds the spec-time prohibition probe (spec-phase Step 5.6) — the second adapter of the probe-core resolution model. Surfaces unwritten must-NOT constraints as negative SPEC acceptance criteria with test/judgment verification tiers; fail-closed at verify time. Per ADR-550. Closes #644.
2026-06-14 21:29:11 -04:00
Tom Boucher
55eb1fd72e refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam (#1242)
* refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam

ADR-22's severity mapping, authority auto-upgrade, and rung>=3 hard-block lived only as prose in plan-review-convergence.md — untestable. Extracted into src/plan-drift-guard.cts (pure: AUTHORITY_RUNGS, getEffectiveAuthority, classifyDriftSeverity) + a gsd-tools drift-guard CLI seam (authority/severity), and rewired the workflow to call the seam deterministically instead of reasoning the decision in prose. 47 unit/e2e/structural tests cover the full severity table, the grep->intel auto-upgrade, and the rung>=3 HIGH hard-block.

Closes #1190

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1190): add changeset for ADR-22 drift-guard seam (#1242)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1190): register ADR-22 module in eslint-ignore + inventory manifest + docs-exempt changeset

Full-matrix CI surfaced new-module/command governance ripples beyond the lint-tests chain: (1) tsc-generated plan-drift-guard.cjs must be in the eslint ignore list (551-eslint-bin-lib-coverage); (2) docs/INVENTORY-MANIFEST.json must include the new module/command (regen via gen-inventory-manifest.cjs --write); (3) a type:Added changeset triggers docs-required — added a docs-exempt marker (internal seam, no user-facing surface).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 20:52:36 -04:00
Tom Boucher
0b3a2e5f9c feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15) (#1237)
* feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15)

ADR-15 designates /gsd-progress --next --auto --converge as the PRIMARY plan-convergence surface, but only the secondary surface (autonomous.md) was wired. next.md now parses --converge/--cross-ai into a plan strategy, gates on workflow.plan_review_convergence, forwards reviewer flags + --max-cycles, and routes Route-3 planning through /gsd:plan-review-convergence (mirroring autonomous.md); --auto chaining preserves converge mode. Adds argument-hint + help/full.md + COMMANDS.md + how-to parity and a structural regression test.

Closes #1190

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1190): add changeset for progress --converge surface (#1237)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1190): keep progress --converge docs skill-dep-clean + regen workflow size baseline

CI surfaced two ripples from the ADR-15 workflow edits: (1) lint-skill-deps + profile-closure flagged /gsd:plan-phase and /gsd:plan-review-convergence SlashCommand tokens in progress.md's flag docs as undeclared deps — reworded to plain prose since progress.md only advertises the flag (the real invocation lives in next.md); (2) the per-file workflow size baseline needed regenerating after the next.md/help edits.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 15:52:22 -04:00
Tom Boucher
bf634b95c3 feat(#1213): Capability State Writer — write-side inverse of the resolver (#1225)
* feat(#1213): Capability State Writer — write-side inverse of the resolver

Adds src/capability-writer.cts (setCapabilityState + cmdCapabilitySet) and the
`gsd-tools capability set` subcommand: the write-side inverse of the capability
resolver (ADR-1213). One desired capability state projects onto the substrates —
`enabled` drives the runtime surface (canonical on/off), `gates` drive federated
config keys (hook granularity), install profile is a read-only floor — then
re-resolves and reports divergence (assert-and-report), so "off means off" holds
as a write-time invariant. Adds batched setConfigValues; routes gsd:settings
capability hook-gates through the writer. Docs: CLI-TOOLS reference, how-to,
ADR-1213, CONTEXT.md term.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1213): add changeset for Capability State Writer (#1225)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 12:51:17 -04:00
Tom Boucher
7edd18fd2b feat(#1165): async external_job_waiting half-state + resume/pause contract (#1221)
Core half of #1105: a legal external_job_waiting deferred state so an async-dispatched Execute step (committing a .planning/async-jobs/<job>.json manifest, deferring SUMMARY.md) is not an illegal partial. execute-phase safe-resume, resume-project, and pause-work reconcile against the versioned scheduler-agnostic manifest stability contract without re-dispatching; the producer is the capability half (#1164). Closes #1165.
2026-06-14 12:23:07 -04:00
Tom Boucher
e9f9ae49c8 fix(#1146): single base-branch resolver across forking workflows (#1198)
* fix(#1146): single base-branch resolver across forking workflows

Replaces duplicated per-workflow bash detection that silently fell through
to :-main on repos where origin/HEAD is unset (git init+remote add+fetch
without set-head, most CI checkouts, many worktrees).

New CJS module git-base-branch.cjs exposes `gsd_run query git.base-branch`
with full precedence ladder: git.base_branch config override → origin/HEAD
symref → git remote show origin (authoritative) → local branch presence →
"main". All git subprocesses bounded with timeouts; degrades gracefully.

Wires execute-phase, quick, ship, complete-milestone, and pr-branch to the
single resolver. Removes 14 lines of duplicated detection bash across the
five workflows.

Includes 7 behavioral tests covering the full precedence ladder including the
key regression case (master repo, origin/HEAD unset → must return "master",
NOT "main") and an anti-regression guard that fails if any workflow
re-introduces the :-main/:-master fallback pattern.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(changeset): backfill PR number #1198

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1146): drop stray PR-body file from branch

pr-1146-body.md was committed during changeset backfill but must not
be tracked in the repo. Content preserved externally for PR body use.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#1146): add tests for flat base_branch config key and both-branch tie-break

Closes two mutation gaps identified in adversarial review:
- A2: flat {base_branch: ...} at config root (legacy key form) was covered
  by code but unguarded against mutation of lines 74-75 in resolver
- H: tier-4 tie-break when both main+master exist locally (main wins,
  per tryLocalBranch JSDoc) was documented but untested

9/9 tests pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#1146): allowlist workflow-literal guard as runtime-contract exemption

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1146): degrade gracefully when gsd_run unavailable in handle_branching bash blocks

handle_branching (execute-phase.md) and step 2.5 (quick.md) are extracted
and run verbatim by behavioral tests that lack the gsd_run preamble.
Adding a || fallback ladder (git symbolic-ref then echo main) keeps the
unified resolver as primary in real workflows while letting the test harness
succeed without gsd_run defined.

Also propagates updated runtime-launcher preamble to pr-branch.md (added in
origin/next MemPalace PR) and regenerates workflow-size-baseline.json.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 08:22:40 -04:00
Tom Boucher
a375c4b354 feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) (#1201)
* feat(#956): add MemPalace memory capability (ADR-857 feature plug-in)

Adds an opt-in, default-resilient ADR-857 feature capability that wires
MemPalace (local-first memory: MCP server + CLI) into the GSD loop:
deliberate recall before discuss/plan and verbatim + temporal-KG capture
at phase boundaries. Three memory modes (augment default; kg_backend and
replace forward-declared). Master gate mempalace.enabled (default off);
every hook onError:skip, zero gates; absent/disabled MemPalace => loop
unchanged. Transport is rendered-markdown only — MemPalace runs
out-of-process, no third-party code in gsd-core (ADR-857 §7).

Capability: capabilities/mempalace/ (manifest + 2 fragments), skills
commands/gsd/mempalace-{recall,capture}.md, agent
agents/gsd-mempalace-curator.md. Registration: ns-context router,
utility cluster, KNOWN_SKILLS, help full.md, model-catalog, copilot
install list, size baselines; regenerated capability-registry +
inventory manifest. ship:post wired into ship.md (wire-on-demand).

HELD on #1196: this capability also declares hooks at discuss:pre and
discuss:post, which are structurally un-wireable until the host-loop
conformance model covers the discuss phase (discuss-phase.md is not in
HOST_LOOP_FILES). The phase6-capstone-conformance gate therefore fails
on exactly those two orphaned points by design — see #1196. Once #1196
lands, rebase onto next and the gate goes green with no further change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#956): backfill changeset PR number (#1201)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 02:09:10 -04:00
Tom Boucher
b1e8a74708 fix(#1196): wire discuss loop step for capability hooks (#1199)
* fix(#1196): wire discuss loop step for capability hooks

discuss was contract-declared (gsd:loop-host marker, in POINT_ORDER and
LOOP_HOST_CONTRACT) but structurally unwireable: discuss-phase.md had no
`loop render-hooks` dispatch and was absent from the conformance gate's
HOST_LOOP_FILES, so capabilities could never wire discuss:pre/discuss:post.

- discuss-phase.md: add minimal discuss:pre (before analyze_phase) and
  discuss:post (after write_context) render-hooks dispatch steps that
  delegate consumption to a new shared reference (kept under the 32KB
  #2551 budget; no inline subagent dispatch token).
- references/loop-hook-dispatch.md: new canonical, point-agnostic contract
  for consuming `loop render-hooks --raw` activeHooks (contribution/step/
  gate) — single source for hook consumption across host loops.
- gen-loop-host-contract.cjs: derive HOST_LOOP_FILES from STEP_WORKFLOWS and
  export scanWiredPoints()/getWiredLoopPoints() (throws on a missing host
  file) — one source of truth for the host-loop file + wired-point set.
- phase6-capstone-conformance.test.cjs: consume the derived HOST_LOOP_FILES
  and shared scanWiredPoints (was a hand-maintained duplicate omitting
  discuss-phase.md + a duplicated regex).
- gen-capability-registry.cjs: add validateHooksWired() gen-time guard that
  rejects a capability hook declared at a valid-but-unwired loop point, with
  a clear remediation message — failure now surfaces at gen --check/--write
  time instead of deep in the full conformance suite.
- tests (capability-registry.test.cjs): regression + anti-pattern parity
  guards (every loop-host marker is in STEP_WORKFLOWS/HOST_LOOP_FILES;
  POINT_ORDER === flattened LOOP_HOST_CONTRACT) so no step can drift into
  the discuss-class gap again.
- docs/INVENTORY*: register the new reference.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1196): backfill changeset PR number (#1199)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 01:39:54 -04:00
Tom Boucher
3ebd13c45d fix(#1145): implement query user-story.validate handler (#1193)
* fix(#1145): implement query user-story.validate handler

`query user-story.validate` was a phantom command invoked by mvp-phase.md
(line 102) and verify-work.md (line 170) but had no CJS handler. Every
call exited with "Unknown command: user-story". The dotted-form dispatcher
strips the `query` prefix, splits on `.`, yielding command='user-story'
which fell through to the `default:` case with no registered capability
handler.

Adds a `case 'user-story':` handler inline in gsd-tools.cjs (same pattern
as the #1140 fix in PR #1148). The handler:

- Validates "As a [role], I want to [capability], so that [outcome]."
- Uses \S anchors to require non-whitespace content in each slot
  (whitespace-only slots like "As a  , I want to ..." now correctly
  return valid:false — found by adversarial Codex review)
- Returns { valid: boolean, errors: string[], slots: {role, capability,
  outcome} | null }
- Supports --pick valid for bare boolean output (verify-work.md usage)
- Added 'user-story' to SKIP_ROOT_RESOLUTION (pure string validation)
- Added 'user-story' to TOP_LEVEL_USAGE command list

Regression tests added to tests/commands.test.cjs (per lint-regression-
test-names policy; new bug-NNNN standalone files are banned).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(changeset): backfill PR number for #1145 fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-13 23:34:44 -04:00
Tom Boucher
b10e56818b feat(#1169): complete ADR-857 phase 6 — migrate features to Capabilities, revive dead gates, harden conformance gate (#1183)
* test(#1168): make phase-6 gate un-gameable — reject empty stubs + require loop shrink

The migration assertion previously checked only role==feature, so a registration-only stub (empty hooks, logic left inline) would turn the gate green while phase 6 stayed incomplete — the exact false-completion pattern this gate exists to prevent. Strengthen it: each ADR-named feature must OWN its behavior (>=1 hook, or a command family); and plan-phase.md/execute-phase.md must shrink strictly below their frozen pre-phase-6 sizes (94519/93166 LF bytes), which also defeats double-run gaming (declare a hook but keep the inline block -> file does not shrink -> red).

Gate now 5 pass / 4 fail (orphaned execute:wave:post, empty/unregistered features, config-key leaks, no shrink). Green is now reachable only by REAL migration. Refs #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate gap-analysis to a Capability (plan:post gate)

First real ADR-857 phase-6 migration (pattern-defining tracer). gap-analysis moves from an inline post_planning_gaps branch in plan-phase.md to a real plan:post gate Capability:

- capabilities/gap-analysis/capability.json: role:feature, plan:post gate (when=workflow.post_planning_gaps, blocking:false advisory), OWNS workflow.post_planning_gaps (federated out of central schema). - plan-phase.md: inline config-get + gsd_run gap-analysis block replaced with a plan:post render-hooks call site dispatching the gate; file shrinks 94519->93279. - src/check-command-router.cts: cmdGapAnalysisPlanPost runs the real gap analysis via gap-checker. - post_planning_gaps removed from central manifest; resolves via federated config (default true preserved). - tests/post-planning-gaps-2493: re-pointed to assert capability ownership.

Verified: gate 5 pass / 4 fail (gap-analysis cleared from migration, plan:post-orphan, config-leak, and plan-phase shrink checks); loadConfig still returns post_planning_gaps=true; check command runs real analysis; 392/392 in the config/registry/federation/router net. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate profile-pipeline to a command-family Capability

ADR-857 Decision 7: profile-pipeline becomes a command-family Capability (like audit/intel/graphify). capabilities/profile-pipeline/capability.json declares an 8-command family (scan-sessions, extract-messages, profile-sample, write-profile, profile-questionnaire, generate-dev-preferences, generate-claude-profile, generate-claude-md) backed by a new gsd-core/bin/lib/profile-pipeline-command-router.cjs; the inline case arms are removed from gsd-tools.cjs. Owns profile-pipeline.enabled (federated).

Verified: registry shows role:feature with commands.length=8; scan-sessions/profile-sample run live via the family; gate cleared profile-pipeline from the empty-stub failure (only tdd/schema-gate/drift remain); 296/296 registry+inventory+gsd-tools tests; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1167): wire execute:wave:post + implement ui.safety-gate check

Revives the second dead gate from #1167: ui.gates@execute:wave:post was declared but never dispatched AND its check.query (ui.safety-gate) was unimplemented. Adds the per-wave execute:wave:post render-hooks call site in execute-phase.md (fires after each wave's merge/cleanup, before the next forks) and implements cmdUiSafetyGate (frontend + UI-SPEC aware, mirrors cmdUiPlanGate) in check-command-router. +17 regression tests.

Verified: phase-6 orphaned-points conformance test now PASSES (gate 6 pass / 3 fail); ui-safety-gate routable in dot+hyphen forms; check-ui-safety-gate 17/17, check-ui-plan-gate 18/18; lint 0 errors. Refs #1167, #1168.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate drift (schema + codebase) to execute:wave:post gates

Removes the inline schema_drift_gate + codebase_drift_gate steps (77 lines) from execute-phase.md; drift becomes a Capability with two execute:wave:post gates (verify.schema-drift blocking, verify.codebase-drift advisory) dispatched via the per-wave render-hooks call site. check-command-router routes verify.schema-drift / verify.codebase-drift to the real detectors. Federates workflow.drift_threshold / drift_action / schema_drift_gate out of central.

Also fixes the execute:wave:post dispatch prose to run NON-blocking (advisory) gates too — the prior version only ran blocking gates, which would have silently dropped the codebase-drift advisory after its inline step was removed. Behavior preserved.

Verified: gate 7 pass / 2 fail (drift cleared from stub + config-leak; execute-phase.md 92297 < 93166 frozen -> shrink passes); both drift checks run real detection; loadConfig defaults preserved (threshold=3, action=warn, gate=true); drift-detection 56/56 + schema-drift 34/34; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate tdd to a Capability (plan:pre contribution + execute:post gate)

tdd becomes a real Capability: a plan:pre contribution injects the <tdd_mode_active> planner guidance (rendered from PLAN_PRE_HOOKS_JSON like security's contribution), and an execute:post gate (tdd.review-checkpoint, advisory) runs the real end-of-phase RED/GREEN review via a new check-command handler. Inline tdd_mode reads + the inline planner block + the tdd_review_checkpoint step are removed; workflow.tdd_mode is federated out of central. The MVP+TDD per-task RED-commit gate is preserved — TDD_MODE is now derived from the execute:post hooks (capId==tdd active), not an inline config-get.

BEHAVIOR CHANGE (documented, not silent): the --tdd CLI flag now persists workflow.tdd_mode=true via config-set instead of being per-invocation. Rationale: tdd is now a config-toggled Capability, and env vars do not persist across the workflow's separate bash blocks (config does), so an ephemeral override isn't cleanly achievable; --tdd therefore enables the tdd capability, consistent with how all capabilities are toggled.

Verified: gate 7 pass / 2 fail (tdd cleared from stub + config-leak; plan-phase + execute-phase both < frozen sizes); contribution injection + execute:post gate dispatch wired; MVP+TDD gate preserved; tdd.review-checkpoint runs real review; full unit suite 556/0; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate schema-gate to a plan:pre contribution Capability

The plan-time schema-push detection (former plan-phase.md §5.7) becomes a schema-gate Capability: a plan:pre contribution (into:planner, when:workflow.schema_push_detection) whose fragment carries the full ORM-detection + [BLOCKING] schema-push-task injection logic, rendered into the planner via the existing plan:pre render-hooks dispatch. The inline §5.7 block is removed (plan-phase.md 94519->90445). workflow.schema_push_detection is a new capability-owned (federated) key, default true. (The execute-side schema-drift gate was migrated separately into the drift capability.)

Verified: registry inlines the fragment (len 2704) so it is actually delivered at plan:pre; gate 8 pass / 1 fail — all 5 ADR-named features now real Capabilities, only the config-leak test remains (intel/security, next unit). Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): close the 3 capability config-key leaks — phase-6 gate now GREEN

Removes the last inline config-get reads of capability-owned keys from plan-phase.md. security_asvs_level/security_block_on now flow through the security plan:pre contribution via a new loop-resolver configValues mechanism (resolves declared config keys with the same 4-level precedence as activation and attaches them to the rendered hook); the §5.55 banner reads them from PLAN_PRE_HOOKS_JSON. intel.enabled becomes a real intel plan:pre step (ref.command: intel api-surface) dispatched via render-hooks; the inline intel branch is gone. gen-capability-registry now validates ref.command as a third dispatch shape.

Verified: phase-6 capstone conformance gate is FULLY GREEN (9/0); 3 leaks gone (grep=0); security configValues resolve to {2,medium}/default {1,high}; intel step present only when enabled; loop-render-hooks 62/0, capability-registry 287/0, capability-state/federated-config 113/0; lint 0 errors. Closes the migration half of #1169. Refs #1139, #1167, #1168.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): address adversarial review — restore schema-drift block, generic planner injection, uniform gate contract

Adversarial review caught 2 real regressions the green gate missed: (1) schema-drift no longer blocked — the execute:wave:post dispatch read GATE_RESULT.block but verify.schema-drift emitted drift_detected/blocking, and onError:skip wrongly bypassed positive blocks; (2) only tdd's plan:pre contribution was injected into the planner, dropping schema-gate's schema-push detection and security's threat-model guidance.

Fixes: (A) every gate check returns a uniform boolean 'block' under --raw (the dispatch form), with advisory gates (tdd/gap) carrying their report in 'message'; (B) gate-dispatch contract corrected at all sites — onError governs command errors only, a blocking gate's positive block always halts; (C) generic planner injection of all plan:pre contributions where into=='planner' (tdd + schema-gate + security incl configValues); (D) two new conformance assertions: planner contributions injected generically + every gate check.query returns boolean block under --raw.

Verified: gate 11/11; all 6 gate checks return boolean block under --raw; full suite 595/0; lint 0 errors. Refs #1167, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): restore MVP+TDD end-of-phase blocking escalation (2nd adversarial pass)

The migrated tdd execute:post gate is statically blocking:false, but the contract (references/execute-mvp-tdd.md + CONTEXT.md) requires the end-of-phase TDD review to ESCALATE from advisory to blocking when MVP_MODE && TDD_MODE && a TDD plan misses a RED/GREEN commit. The migration prose had downgraded this to a 'strong advisory recommendation' — silent loss of the blocking escalation. Restore it: the tdd-gate dispatch now refuses to mark the phase complete (Phase blocked message) under MVP+TDD when GATE_RESULT.block is true; advisory otherwise.

Also strengthen tests/execute-mvp-tdd-gate.test.cjs: hasBlockingEscalation previously matched any line with 'blocking'+'mvp+tdd' (so 'advisory (blocking: false) ... under MVP+TDD' was a false green); now it requires the real refusal semantics ('refuse to mark the phase complete' / 'phase blocked'). Caught by 2nd adversarial review pass.

Verified: execute-phase.md 92702 < 93166 frozen; mvp-tdd-gate + phase-6 gate 19/0; full suite green; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): restore MVP+TDD proceed-block, codebase auto-remap, schema skip-flag (3rd adversarial pass)

3rd adversarial pass found 4 more silent regressions: (1) the tdd MVP+TDD 'refuse to mark complete' was nullified by a downstream 'ALWAYS proceed regardless of gate results' line — proceed is now conditional (stops on an active MVP+TDD block); (2) the test now asserts the proceed is NOT an unconditional override; (3) codebase-drift auto-remap (spawn gsd-codebase-mapper when drift_action=auto-remap) was dropped — the execute:wave:post advisory dispatch now consumes spawn_mapper/directive; (4) GSD_SKIP_SCHEMA_CHECK bypass was lost from the gate path — cmdVerifySchemaDrift now honors the env var (block:false when set).

Verified: no unconditional proceed; GSD_SKIP_SCHEMA_CHECK=true -> block:false; gate 11/11 + mvp-tdd 9/9; full suite 569/0; lint 0; execute-phase.md 93109 < 93166. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): init.cts reads federated config keys from nested path (4th adversarial pass)

Config federation moved tdd_mode/research/nyquist_validation from flat config.<key> to nested config.workflow.<key>, but src/init.cts still read them flat — so init.plan-phase/init.execute-phase emitted tdd_mode:false / research_enabled:undefined / nyquist:undefined regardless of config (a public command-contract regression; the migrated loops use render-hooks so enforcement was unaffected). Read via config.workflow (type-safe Record cast). Now init reflects the same resolved values + federated defaults (research/nyquist default true) as the render-hooks path.

Verified: build clean; init.plan-phase emits tdd_mode:true/research:false/nyquist:false for set config, defaults true for empty; full suite 591/0; lint 0. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1169): add changeset for ADR-857 phase-6 completion (PR #1183)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): complete phase-6 migration fallout — restore TEXT_MODE, fix registry .claude leak, re-point stale workflow-contract tests

The capability migration left real regressions and stale consumer tests that
the per-module unit suite missed but the full cross-platform suite caught (27
failing tests):

Real source regressions (fixed):
- execute-phase.md lost its AskUserQuestion TEXT_MODE plain-text fallback when
  the inline schema_drift_gate step was removed — non-Claude runtimes would
  stall. Restored, and the execute:post gate-dispatch prose de-duplicated to
  cite the execute:wave:post contract (loop body shrinks below the frozen
  pre-phase-6 ceiling while keeping every onError/blocking nuance).
- capabilities/tdd inline fragment hardcoded `@~/.claude/gsd-core/references/tdd.md`,
  baked verbatim into the committed capability-registry.cjs and leaked the
  install path on 11 non-Claude runtimes (registry .cjs is copied, not
  path-converted). Made the fragment path-free; regenerated the registry. The
  phase-6 conformance gate now guards this (no ~/.claude install path in any
  capability source or the generated registry).
- plan-phase.md: removed a §5.7 stub re-added in error and routed Branch 2 to
  step 6 (schema-gate is a plan:pre capability, §5.7 is gone).

Stale workflow-contract tests re-pointed to the capability dispatch they now
must assert (behavior verified preserved in source first, assertions kept
equal-or-stronger): bug-621 + bug-2851 (gap-analysis via gsd_run render-hooks
plan:post + registry binding), feat-2527 (tdd_mode federated out of central),
phase6-planning + plan-phase-ui-redirect (§5.6 bounded by ## 6.),
plan-phase-drift-guard (intel when:intel.enabled skip branch).

profile-pipeline-command-router.cjs un-ignored from eslint (hand-written, no
TS source) + stale disable comments removed. Size baseline regenerated.

Verified: full suite 15140 tests / 0 fail; lint 0 errors; conformance gate green
legitimately. Refs #1139, #1167, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1169): add ADR-857 E2E content-test coverage for the 12 loop points + capability deliverables

Grounds the capability engine in behavioral E2E tests (drive the real
render-hooks/check CLI + the real registry, assert typed result content — no
source-grep), structured around what ADR-857 says to deliver. 207 tests; each
genuineness-checked (flip the expectation, confirm it fails).

Per-loop-point dispatch (7 files): empty-point negative-space across the 6
no-hook points; verify:post 3-step resolution+ordering+onError; plan:pre
contribution/configValues + ui.plan-gate + intel; plan:post gap-analysis;
execute:wave:post drift+ui gates via the check route (schema-drift block/skip,
codebase-drift threshold BVA, auto-remap); execute:post tdd.review-checkpoint
RED/GREEN; ship:pre security gate resolution + frontmatter-get predicate pieces.

ADR-deliverable coverage (4 files): predicate boundary held (edge/prohibition
probes stay core, not off-by-default Feature Capabilities — phase-6 exception);
core loop runs with zero capabilities (all 12 points empty, init bundles
resolve); contribution merge (multiple ordered <contribution from=> blocks);
federated-config key removal on uninstall.

federated-config allowlisted for its 3-file split (unit + integration +
lifecycle). Refs #1139, #1167, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): remove dead drifted converter dups + address adversarial review

Lint cleanup (root-caused, not waved off): src/runtime-artifact-conversion.cts
carried 11 agent-converter functions (+5 orphaned consts/helpers) that were
never exported, never called, and had silently DRIFTED from the live
hand-authored copies in bin/install.js (one even referenced an undefined
`claudeToCopilotTools`). Deleted the dead duplicates; install.js's live copies
are untouched (it never imported these). Lint now 0 errors / 0 warnings.

Adversarial-review (Codex) findings fixed:
- HIGH: execute-phase.md TDD_MODE used `jq ... || echo false`, silently
  disabling the MVP+TDD blocking gate on jq-less runtimes. Reverted to the
  `node -e` form (node is guaranteed; matches the file's other node-e usages) so
  a missing optional tool can no longer fail-open a blocking safety path.
- MEDIUM: federated-config-key-removal orphan-key test was vacuous (it skipped
  the orphan assertion). Now asserts the removed capability's key is genuinely
  not surfaced/validated after uninstall.
- LOW: phase-6 conformance leak regex broadened to catch absolute-home and
  Windows-backslash `.claude/(gsd-core|commands|agents|hooks)` paths, not only
  `~`/`$HOME` forward-slash forms.
- LOW: bug-2851 plan:post dispatch assertion now requires `--raw` (matched its
  stated contract).
- nit: plan-pre intel-step test duplicate assertion replaced with a distinct
  structured-output check.

Size baseline regenerated (execute-phase.md 93089 < 93166 frozen). Refs #1167, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1169): make runtime-homes-descriptor-drive titles environment-independent

The descriptor-equivalence test embedded the absolute golden config path
(`os.homedir()`-derived) directly in each `test(...)` title, so titles differed
between macOS (`/Users/x/.claude`) and Docker (`/home/gsdtest/.claude`). Every
test PASSES on both platforms (15885/0 leaf tests each), but gsd-test-summary
compares results by title and reported 29+29 false "only in Mac / only in
Docker" discrepancies for tests that actually pass everywhere.

Move the golden path out of the title and into the assertion message (still
shown on failure); titles are now byte-identical across platforms so the
cross-platform comparator matches them. No assertion logic or golden values
changed. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): derive TDD_MODE via gsd_run --active-cap, not node -e (fix prompt-injection CI gate)

The prior fix reverted execute-phase.md:181 from jq to `node -e` to close a
Codex HIGH (jq||echo-false silently disabling the MVP+TDD blocking gate on
jq-less runtimes) — but the CI prompt-injection scanner BLOCKS new `node -e` in
workflow markdown (inline code-exec = injection vector), turning the security
gate red. Both forms were wrong: node -e fails the scanner; jq fail-opens a
blocking safety gate; `config-get workflow.tdd_mode` is forbidden by the
conformance leak gate (tdd_mode is capability-owned).

Correct fix (what Codex recommended): a gsd_run-native boolean. Add an
`--active-cap <capId>` flag to `loop render-hooks <point>` that resolves hooks
the normal way and prints exactly `true`/`false` for whether a capId is active
— scanner-safe (canonical launcher, no inline code), node-reliable (no optional
jq to fail-open), and leak-free (render-hooks resolution, not config-get).
execute-phase.md:181 now `TDD_MODE=$(gsd_run loop render-hooks execute:post
--active-cap tdd)`. +5 behavioral tests for the flag.

Verified: prompt-injection-scan --diff origin/next → 0 findings; conformance
gate 13/13 (execute-phase.md 92934 < 93166); execute-mvp-tdd + tdd-mode +
loop-render-hooks 87/0; lint 0/0. Refs #1167, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 21:07:55 -04:00
Tom Boucher
70f35ebc7a feat(#1169): wire ship:pre security gate via render-hooks (ADR-857 phase 6)
Part of the ADR-857 phase-6 migration (#1169, epic #857) — not a shipped bug fix. The capability system (registry, render-hooks, gates, capability-state) lives only on next / 1.5.0-rc; npm latest is 1.4.5 and contains none of it, so no released user can hit this.

The security capability declares a blocking gates@ship:pre hook (when=workflow.security_enforcement, predicate SECURITY.md.threats_open==0), but ship.md never called `loop render-hooks ship:pre` and had no inline fallback, so the declared gate never fired. Wire it into ship.md preflight_checks via the render-hooks idiom — fail-closed: the ship blocks unless threats_open is exactly 0.

Add a phase-6 conformance assertion that every declared hook point has a render-hooks call site; execute:wave:post (ui_safety_gate) stays in KNOWN_UNWIRED pending its unimplemented ui.safety-gate check.

Part of #1169. Refs #857.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 11:09:21 -04:00
Tom Boucher
f116b76128 test(#1139): add phase 6 capstone conformance gate (#1158) 2026-06-13 02:12:16 -04:00
Tom Boucher
aec3374bc2 feat(#1138): make runtime descriptors authoritative (#1157) 2026-06-13 01:49:25 -04:00
Tom Boucher
607813f5d0 feat(#1136): consume resolved capability state (#1153)
* feat(#1136): consume resolved capability state

* chore(#1136): add capability state changeset
2026-06-12 23:43:07 -04:00
Tom Boucher
1b880fefd7 feat(#1137): migrate review verification hooks to capabilities (#1147)
* feat(#1137): migrate review verification hooks to capabilities

* chore(#1147): add changeset
2026-06-12 21:24:09 -04:00
Tom Boucher
44024aa535 fix(#1133): honor model_policy on the claude runtime (forward-port to next) (#1144)
Forward-port of the #1133 hotfix (commit 22f237a5, 1.4.5 hotfix line) onto
next. The hotfix was authored against src/core.cts (v1.4.4); on next the
resolver logic lives in src/model-resolver.cts (ADR-457 extraction, #888),
so the patch is re-applied there rather than cherry-picked.

resolveModelInternal step 2.5 now honors model_policy on the claude runtime:
the policy-resolved full model ID is mapped back to a Claude Code agent alias
via CLAUDE_POLICY_ID_TO_ALIAS (reverse of MODEL_ALIAS_MAP + claude-fable-5 ->
fable). Bare aliases (opus/sonnet/haiku/fable) pass through; an ID with no
Claude alias warns once to stderr (deduped by agentType::policyModel::tier)
and falls back to the configured tier alias. Non-claude runtimes return full
IDs verbatim (unchanged). resolveModelForTier is intentionally unchanged.

The warn-dedupe cache lives in model-resolver.cts; core.cts composes the
exported _resetRuntimeWarningCacheForTests to clear both that cache and the
config-loader warning cache (config-loader cannot import model-resolver --
circular dependency).

Ports the 6 #1133 tests (rewriting the old claude-no-op test that asserted
the bug) plus one added test covering the MODEL_ALIAS_MAP reverse-map path.

Forward-port of #1133

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 19:56:54 -04:00
Tom Boucher
4ab5c7b3f2 feat(#1135): migrate planning hooks to capabilities (#1141)
* feat(#1135): migrate planning hooks to capabilities

* chore(#1135): add phase 6 planning capabilities changeset

* fix(#1135): satisfy lint for agent hook rendering
2026-06-12 19:51:54 -04:00
Tom Boucher
827011b865 fix(#1098): guard generate-claude-md against clobbering hand-crafted files; redirect default to .claude/CLAUDE.md (#1118)
/gsd-new-project wrote a repo-root CLAUDE.md full of broad project docs,
overwriting/diluting a hand-crafted instruction file. --force was parsed but
silently dropped, and nothing guarded an existing non-GSD file.

- Guard: an existing instruction file with no `<!-- GSD:<section>-start` markers
  (hand-crafted) is left untouched; report action:"skipped". --force (now wired
  through CmdGenerateClaudeMdOptions) overwrites intentionally. The marker check
  uses /<!-- GSD:[a-z]+-start/ so a file merely documenting GSD syntax is safe.
- Redirect: the Claude-family default output is now ./.claude/CLAUDE.md (a valid
  auto-loaded project-memory location) instead of repo-root ./CLAUDE.md. Aligned
  across the handler default, config-defaults.manifest.json, buildNewProjectConfig,
  the config template, new-project.md, and cmdGenerateClaudeProfile; advisory
  read-CLAUDE.md hints in plan-phase/quick/profile-user updated. Codex still
  writes AGENTS.md.

Closes #1098

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 16:03:51 -04:00
Tom Boucher
fc37ae0c4f fix(#1115): capability-probe codex hook-trust bypass flag in /gsd:review; fail loud on empty output (#1122)
On codex-cli < 0.137.0 the review.md `codex exec` invocation passed
--dangerously-bypass-hook-trust (added in 0.137) unconditionally and discarded
stderr, so codex exited "unexpected argument" before reading the prompt and the
empty output file was treated as a completed review — a silent degraded review.

- Capability-probe the flag (`codex exec --help | grep`) and apply it via
  $CODEX_BYPASS_FLAG only when supported (works fine without it on older CLIs).
- Capture codex stderr to a .err file instead of /dev/null, and replace an empty
  output with a diagnostic so a broken reviewer is surfaced (mirrors Cursor/OpenCode).
- Update enh-773 enforcement test to require the capability gate + fail-loud guard
  instead of the unconditional flag.

Closes #1115

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 15:55:17 -04:00
Tom Boucher
1c073c1c81 fix(#1107): progress consults verification.status before reporting a phase complete (#1116)
/gsd-progress derived phase completeness from plan/summary counts only and never
consulted the verification.status query (the #651 seam), so a phase whose
VERIFICATION.md ended human_needed or gaps_found was reported complete and
routing skipped to the next phase. Add Step 1.7 (consult verification.status for
the current phase) and routing rows that send gaps_found to plan-phase --gaps
(Route V.gaps) and human_needed to verify-work (Route V.human) before the generic
complete row. passed/missing/unknown still route as complete so unverified phases
are not falsely blocked.

Closes #1107

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 15:55:01 -04:00
Rezolv
c70842b6df enhance(spec-phase): edge-probe precision probe should name tie-breaking / rounding-mode (#1108)
* feat(spec-phase): name tie-breaking/rounding mode in the edge-probe precision probe (#1102)

The precision edge probe fired on numeric-range requirements but its text never
named the most common rounding failure mode — tie-breaking / rounding mode
(half-up vs half-to-even, ceil/floor/truncate). In an A/B experiment the weak
tier (haiku) false-passed 50% of tie-class defects against the surfaced-
unresolved spec; naming the rule in the probe text drove honest abstention
83%→100% (false-pass 17%→0%, haiku 50%→0%) with no category regression
(precision still fires only on numeric-range).

Sharpen the single TAXONOMY probe string and update every rendering together so
the doc↔fixture↔machine contract (edge-probe-docs-fixtures.test.cjs) stays green:
src/edge-probe.cts, the edge-probe.md taxonomy table + 2 worked examples, the
01-round-half-even and 04-money-rounding fixtures, and the resolve-edge-coverage
how-to. Prose-only — no consumer keys on the probe text; SHAPE_CUES firing
unchanged; fully backward compatible.

* chore(changeset): Changed fragment for edge-probe precision probe text (#1108)
2026-06-12 13:04:31 -04:00
Rezolv
3e836fef0d feat(spec-phase): spec-completeness edge-probe (#550) (#584)
* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/

Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.

Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.

Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
  planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
  source-checkout-gated build fallback) instead of LLM re-derivation; the engine
  capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
  resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
  per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
  validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.

Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.

* test(#550): RED — status×verification re-cut + probe-core engine specs

Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):

  status: resolved | dismissed | unresolved   (lifecycle, shared)
  verification: explicit | backstop | null     (only when resolved)

- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
  to be extracted — validateResolution(r, validators), validateRequirement,
  analyzeCoverage(items, resolutions?, validators), byVerification rollup,
  runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
  {resolved,backstop}; coverage gains byVerification.{explicit,backstop};
  proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
  COUNT preserved on every fixture (closed set = resolved+dismissed; doc
  line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
  rewritten to the two-axis model.

Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).

* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)

Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.

probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
  verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
  items[] (core never assumes propose is deterministic — edge resolves via
  LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
  count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
  {categories, verification, requiredFieldsByVerification} (ADR-550 #5)

edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.

* chore(#550): register probe-core.cjs artifact in ledgers + inventory

New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:

- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
  never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
  is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
  edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).

probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.

* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]

trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.

Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.

* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard

Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:

- validateResolution now enforces the 'verification is null unless resolved'
  invariant for EVERY status (not just resolved): a dismissed/unresolved
  resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
  (was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
  it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
  contract; a new test locks that an all-dismissed run is NOT affirmatively covered
  (byVerification is the honest gate).

Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.

* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics

Re-review #5 (trek-e) clarity edits:

- Decision 5: annotate that only contract item (a) ships on #584 (the edge
  adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
  parenthetical — the blessed/implemented semantics are count-preserved = the
  CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
  carrying the per-tier resolved-status breakdown. The old parenthetical
  contradicted the shipped count.

* test(#550): cover runProbeCli structural-guard numeric-count branch

Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.

* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)

The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.

* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)

A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.

* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)

templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.

* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)

The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.

* docs(#550): add how-to for resolving edge-coverage findings (B1)

Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.

* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)

trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.

* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)

trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.

* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)

trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.

* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe

Rebased onto next (e4f0910d), which replaced the line-based tier-max
size ratchet (#597) with the byte-based per-file baseline guard (#1074).
The edge-probe feature legitimately grows two workflows:

  - spec-phase.md  15131 -> 23094 (+7963): Step 5.5 Edge-Completeness Probe
  - plan-phase.md  93135 -> 94253 (+1118): covered/backstop edge lift into
    must_haves.truths (the live <downstream_consumer> block)

Both remain under their tier hard caps (plan-phase XL 98304, ~4KB
headroom; spec-phase DEFAULT 40960). Growth is real inline workflow
content the feature requires at that step — not eager @-import proxy
gaming. Drops the obsolete line-based XL_BUDGET 93000->94000 bump
(superseded by the byte baseline) via rebase.

* chore(#550): reconcile INVENTORY headline counts after rebase onto next

Rebase onto current next dropped the prior reconcile commit (stale counts
refs 68 / CLI 107). Current next + the edge-probe additions yield:
  - References (68 -> 69 shipped): + gsd-core/references/edge-probe.md
  - CLI Modules (107 -> 109 shipped): + edge-probe.cjs + probe-core.cjs

Rows for all three already present; only the headline counts were stale.
Caught by tests/inventory-counts.test.cjs (CI ubuntu-24 leg).
2026-06-12 11:05:31 -04:00
Tom Boucher
77a671ec53 feat(#791): migrate antigravity workspace base dir .agent → .agents (#1090)
Fresh antigravity workspace installs write under the canonical .agents/ (plural) base; legacy .agent/ stays recognized (dual-read). Global ~/.gemini/antigravity/ path unchanged. Closes #791.
2026-06-11 22:37:55 -04:00
Tom Boucher
94872662e9 feat(#1082): complete phase 5 — descriptor-drive all install surfaces + materialize the InstallPlan — ADR-857/1016/58 (#1080)
* feat(#1077): phase 5f-2 — drive the hookEvents dialect (PostToolUse/AfterTool) from the descriptor

postToolEvent (bin/install.js) and preToolEvent (applySettingsJsonHooks in
runtime-hooks-surface.cts) now select the event-name dialect from
registry.runtimes[id].runtime.hookEvents instead of the hardcoded
(runtime === 'gemini' || runtime === 'antigravity') check: hookEvents === 'gemini'
→ AfterTool/BeforeTool; else → PostToolUse/PreToolUse. hookEvents threaded into the
applySettingsJsonHooks opts bag. Equivalence-preserving (Codex-verified): hookEvents
'gemini' is exactly {gemini, antigravity}, 'claude' the rest; undefined → claude
dialect (matches the old else).

The per-event SET guards (isQwen||claude → SubagentStop/Stop/PreCompact;
runtime==='claude' → FileChanged; isGemini → Gemini agent-events) stay HARDCODED —
hookEvents (2-value) is too coarse to drive them (the event set differs within
hookEvents='claude'); per-event-set drive tracked in #1076.

Registry-parity test (enh-1077): asserts BOTH post-tool (AfterTool/PostToolUse) AND
pre-tool (BeforeTool/PreToolUse) dialects are a pure function of hookEvents, for
gemini/antigravity/claude/augment — non-vacuous (catches a broken hookEvents thread).

Closes #1077

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1077): build hooks/dist in before() so dialect-drive test passes in scoped CI

hooks/dist is gitignored and absent in scoped/windows CI jobs that do not
pre-run build:hooks. Without it, install() finds no hook files and all
AfterTool/BeforeTool/PostToolUse/PreToolUse event arrays come back empty,
failing every hook-presence assertion. Added an idempotent ensureHooksDist()
called in a top-level before() — mirrors the pattern from
bug-376-claude-js-hook-gsd-rewriter.test.cjs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1055): add installSurface/writesSharedSettings/permissionWriter/extendedHookEvents to runtime descriptors

Purely additive: four new fields on all 16 runtime capability.json descriptors,
validator extended with three new closed-vocab sets, registry regenerated.
Test fixtures (VALID_RUNTIME_CAP and makeRuntimeCap) updated to include the new
required fields so all 255 capability-registry tests continue to pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1076): drive per-event hook guards from extendedHookEvents descriptor

Replace hardcoded runtime-name checks (isQwen||runtime==='claude',
runtime==='claude', isGemini) in applySettingsJsonHooks with a single
descriptor-driven extendedEvents array derived from the new opts field.
Remove isQwen and isGemini derivations (no remaining uses after the three
guard blocks are migrated). Wire extendedHookEvents from the capability
registry in bin/install.js call site. Add behavioral regression test
(enh-1076-extended-hook-events-drive.test.cjs) confirming the drive is
purely descriptor-based and runtime-name-agnostic.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1055): drive resolveRuntimeConfigIntent from the runtime descriptor; retire hand-kept REGISTRY

- Rewrites src/runtime-config-adapter-registry.cts to require capability-registry.cjs
  and read installSurface / writesSharedSettings / permissionWriter from
  runtimes[id].runtime; deletes the hand-kept REGISTRY const (ADR-857 phase 5g drive 2).
- ALLOWED_CONFIG_RUNTIMES is now derived from descriptor entries that have installSurface.
- Fixes the configFormat parity gate in scripts/gen-capability-registry.cjs to read
  installSurface directly from capMap descriptor bodies, breaking the require cycle
  (adapter now requires the generated registry; gen-script must not require the adapter).
- Adds golden-master test tests/enh-1055-config-intent-descriptor-drive.test.cjs (41 tests)
  pinning all 16 runtimes' return shapes and the TypeError-on-unknown contract.
- Updates scripts/lint-test-file-count.allowlist.json (config module, +1 file).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1076): make hooksSurface descriptor load-bearing for the settings-json hook-skip

- Adds hooksSurface?: string to ApplySettingsJsonHooksOpts and destructuring in
  applySettingsJsonHooks (src/runtime-hooks-surface.cts).
- Replaces the hardcoded !isOpencode && !isKilo hook-skip guard with
  hooksSurface !== 'none'; removes the now-unused isOpencode/isKilo derivations
  (ADR-857 phase 5g drive 3).
- Passes hooksSurface from the runtime descriptor at the applySettingsJsonHooks
  call site in bin/install.js using the established
  _capabilityRegistry?.runtimes?.[runtime]?.runtime?.hooksSurface idiom.
- Extends tests/enh-1076-extended-hook-events-drive.test.cjs with two new suites
  proving: (a) hooksSurface:'none' writes no hooks regardless of runtime name;
  (b) hooksSurface:'settings-json' writes hooks even for 'opencode' (previously
  hardcoded to skip).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: record installSurface/writesSharedSettings/permissionWriter/extendedHookEvents descriptor axes in ADR-1016

Add Decision 7a documenting the four axes added in the 5f-completion pass,
update axis counts from "six" to "twelve", note 5f-completion drives as done
in Decision 8's ladder, update Out of scope to reflect #1055/#1076 are done
and 5g (InstallPlan capstone) remains the only open phase.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1055): parity gate must fire on configFormat↔installSurface mismatch (read installSurface at the descriptor level)

The test fixture makeRuntimeCapMap did not include installSurface in the runtime object,
so the gate's typeof r.installSurface !== 'string' guard always skipped the entry and never threw.
Added installSurface as an optional third parameter to makeRuntimeCapMap and passed the correct
installSurface values ('settings-json' for claude, 'codex-toml' for codex) to the two THROWS tests.
The gate implementation already reads r.installSurface correctly from the descriptor level.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1076): add installSurface↔hooksSurface + extendedHookEvents↔hookEvents consistency gates with rejection tests

GATE A: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES map in validateRuntimeBody enforces that a
runtime's hooksSurface is valid for its installSurface (e.g. profile-marker-only only allows none,
codex-toml only allows codex-hooks-json). Derived from the 16 real runtime descriptors.

GATE B: validateRuntimeBody checks that if extendedHookEvents contains Gemini agent-events
(BeforeAgent/AfterAgent/BeforeModel), hookEvents must be 'gemini'; if it contains Claude-family
events (SubagentStop/Stop/PreCompact/FileChanged), hookEvents must be 'claude'.

Added 10 rejection tests in suite 27 covering each gate + each new field validator.
All 16 real runtimes satisfy both gates (verified before coding).
Exports: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES, VALID_INSTALL_SURFACES,
VALID_EXTENDED_HOOK_EVENTS, VALID_PERMISSION_WRITERS, validateRuntimeBody.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1076): strengthen hooksSurface-drive assertions; defensive hooksSurface fallback; drop vacuous dup

1. bin/install.js: add explicit literal fallback for hooksSurface when the committed
   capability registry fails to load (opencode/kilo → 'none', all others → 'settings-json').
   The descriptor is always the source of truth in normal operation.

2. enh-1076 Suite 7: change SessionStart assertion from key-presence (hasOwnProperty)
   to at least-one-command (hasHooksFor), so the test fails if hooks are initialized-but-empty.
   ensureHooksDist() in before() guarantees hook files exist.

3. enh-1055 Test 2: remove vacuous duplicate suite that re-asserted intent.runtime === row.runtime
   already fully covered by Test 1's deepStrictEqual over all four fields.

4. capability-registry.test.cjs: fix stale comments in the grok-skip test that claimed the
   parity gate uses the adapter registry; gate reads purely from the descriptor (installSurface
   absent → typeof r.installSurface !== 'string' → soft-skip).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1082): materialize the InstallPlan — collect install-level descriptor axes into resolveInstallPlan; route install()/finishInstall() through it (ADR-58/5g)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: record 5g InstallPlan materialization (ADR-58 Accepted, ADR-1016 phase-5 complete)

ADR-1016 Decision 8 step 7 updated to DONE: resolveInstallPlan(runtime) in
runtime-config-adapter-registry collects install-level descriptor axes into
the typed InstallPlan consumed by install()/finishInstall(). Out-of-scope
section updated: 5g capstone is complete, phase 5 fully materialized.
ADR-1016 line ~20 updated: InstallPlan IS now materialized (both halves).
ADR-58 Implementation note added (2026-06-11): realized in
runtime-config-adapter-registry (co-located with adapter-selection).
CONTEXT.md Runtime Config Adapter Registry entry extended to document
resolveInstallPlan and both-halves realization.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1082): update install drift guard to the resolveInstallPlan seam (5g)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 22:36:47 -04:00
Tom Boucher
93c5ecd645 feat(#792): add devin-desktop runtime alias for windsurf (#1086)
windsurf now also answers to devin-desktop (CLI --devin-desktop) for the Windsurf→Devin Desktop rebrand; all paths unchanged. The .devin/skills/ workspace migration is split to #1085. Closes #792.
2026-06-11 22:32:49 -04:00
Tom Boucher
7c07fce70f fix(#381): make gsd_run launcher reachable in fresh-shell-per-block runtimes (#1084)
* fix(#381): make gsd_run launcher reachable in fresh-shell-per-block runtimes

On runtimes that execute each fenced bash block in a separate shell process
(e.g. Claude Code — documented behavior: each Bash command is a separate
process; inline shell functions and exported vars do not persist between
calls), the once-per-file gsd_run() function was undefined in every block
after the preamble block, and the call was swallowed by
`2>/dev/null || echo "{}"` into silent empty state.

Fix (budget-neutral session-level resolution):
- Ship gsd-core/bin/gsd_run, a POSIX sh wrapper that symlink-resolves its own
  location and execs the co-located gsd-tools.cjs. Exposed on PATH via the npm
  `bin` field (global installs) and shipped to local installs via the recursive
  gsd-core/ copy.
- The per-file launcher preamble now appends `export PATH='<bindir>':"$PATH"`
  to the file named by $CLAUDE_ENV_FILE (Claude Code's documented
  env-persistence mechanism) so later fresh-shell blocks resolve gsd_run from
  PATH. Guarded as a strict no-op when CLAUDE_ENV_FILE is unset; the inline
  gsd_run() definition remains the fallback for all other runtimes. The
  single-quoted dir neutralizes shell metacharacters at source time.
- Propagated via scripts/sync-runtime-launcher.cjs to all launcher-using files.
- XL workflow byte budget 93000 -> 93200 (the ~130B clause pushes plan-phase.md
  to 93135; legitimate content growth, ratchet-up per #717).

Regression tests (I)/(J) in runtime-launcher-parity.test.cjs cover wrapper
delegation and end-to-end PATH persistence (sourcing the env file with a
space-bearing install path).

Known limitation: an install path containing a literal single-quote yields a
malformed env-file line and falls back to the status quo (no regression);
rare on sanitized home directories.

Closes #381

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#381): add changeset for gsd_run fresh-shell reachability fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#381): scope test (J) bare-PATH execution to POSIX (Windows Git Bash exec bit)

Windows Git Bash (msys2) does not honor Node's chmod exec bit for
PATH-executing extension-less scripts, so the bare `gsd_run` command lookup
failed there even though the env-file PATH persistence was correct. The
env-file content assertions (the fix's actual cross-platform logic) still run
on every platform; only the final source-and-execute sub-step is gated to
non-win32. Global installs on Windows are covered by npm's generated bin shim.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 21:42:52 -04:00
Tom Boucher
8813ee5f95 feat(#429): HARD GATE on negative-grep literals echoed in plan <action> bodies (#1062)
Convert the planner's soft comment-text guideline into a plan-write-time
HARD GATE. When an acceptance criterion negative-greps for a literal
(`grep -c 'LIT' file == 0`) and that same literal appears verbatim in an
`<action>` body (JSDoc samples, head-comment references, "what NOT to do"
snippets), the executor's commit-time verify gate later fails on the
comment echo rather than a real regression — wasting cycles and training
the executor to distrust the gate.

`verify.plan-structure` (the `validate_plan` step) now scans for this:
- confidently-extracted (quoted) negative-grep literal echoed in an
  <action> → error (valid:false), failing plan creation
- unquoted/ambiguous grep target → warning (fallback policy)
- `<!-- planner-discipline-allow: LIT -->` escape hatch skips a literal
- positive-count gates (`== N`) and `!= 0`/`>= 0` are out of scope

Adds the `<comment_text_discipline>` block to gsd-planner.md, the full
rules + allowlist example to planner-antipatterns.md, and regression
fixtures for downstream incidents 12-04, 11-04, 12-02 (plus a boundary
case proving positive-count gate 11-02 is not flagged).

Closes #429

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 15:46:49 -04:00
Tom Boucher
58ed55683e feat(#1049): phase 5d — drive artifactLayout from the runtime descriptor (retire the 128-LOC switch) (#1053)
resolveRuntimeArtifactLayout now builds Layout from
registry.runtimes[id].runtime.artifactLayout[scope] — a loop dispatching each
ArtifactKind through the SAME 5 builders (commandsKind/agentsKind/skillsKind/
convertedCommandsKind/kimiAgentsKind, unchanged) by (kind, converter, nesting) —
replacing the hardcoded switch(runtime). Equivalence-preserving for all 16 runtimes
× {global, local} (Codex-verified, no divergence). -43 LOC; bin/install.js + the
converters + the install loop untouched. getInstallExports()[converterName]
resolution, configDir threading, scope default, unknown-runtime guard all preserved.

Driving the local scope surfaced a 5a gap: the old switch had no scope branch for 13
runtimes (cursor/gemini/codex/copilot/antigravity/windsurf/augment/trae/qwen/hermes/
codebuddy/opencode/kilo) → local == global for them, but 5a authored local:[].
Backfilled local=global for those 13 (descriptor-faithful; a fall-through shim would
wrongly give cline/kimi local=global). claude/cline/kimi scope-gating untouched.

validateArtifactKindEntry tightened: destSubpath/prefix/nesting/converter required
(ConverterName enum still open — 5e). New 39-case deep-equal golden equivalence test.

Closes #1049

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 14:08:46 -04:00
Tom Boucher
9e3b056b15 fix(#779): correct stale model-catalog model IDs verified against live providers (#1047)
Verify-first audit: gemini opus gemini-3-pro→gemini-3.1-pro-preview (undefined in gemini-cli source), codex sonnet gpt-5.3-codex→gpt-5.4 (deprecated per OpenAI); qwen3-coder-next verified valid, unchanged. Adds a regression guard + sourcing note. Closes #779.
2026-06-11 13:36:23 -04:00
Tom Boucher
1fab2e10ba fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver (#1045)
* fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver

Source agents/*.md (gsd-planner, gsd-executor, gsd-verifier, gsd-plan-checker,
gsd-intel-updater, gsd-debugger, …) called bare "gsd-tools …" in shell blocks.
On a shim-only install — where gsd-tools.cjs exists under the runtime home but
gsd-tools is NOT on PATH — those calls fail with "command not found" and the
agent silently skips init/state/validate/commit ceremony, deferring to the
orchestrator or bypassing GSD bookkeeping entirely.

were never migrated, so it persisted on Claude Code and every other runtime that
consumes the source agents directly. Only gsd-phase-researcher.md carried a
resolver — and a stale, claude-only truncated one.

Fix (all runtimes):
- Inject the canonical multi-runtime gsd_run preamble (byte-equal to
  _runtime-launcher.snippet.sh — claude/codex/cursor/gemini/copilot/windsurf/
  augment/trae/qwen/cline/opencode/kilo/hermes/antigravity homes) at the top of
  the first gsd_run block of all 12 gsd-tools-calling agents, and rewrite every
  command-position bare gsd-tools to gsd_run.
- Upgrade gsd-phase-researcher.md's stale resolver to the canonical one.
- Extend scripts/sync-runtime-launcher.cjs to maintain agents/ in parity (the
  sync caught and corrected a mis-placed preamble during development).
- Extend the bare-gsd-tools (#2851) and launcher-parity (#373) regression guards
  to agents/ so no runtime can silently regress.

Closes #1041

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1041): backfill changeset PR number to 1045

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:42:59 -04:00
Tom Boucher
e4dfa6b9ea fix(#1012): invoke fallow with its real CLI and wire the report normalizer (#1044)
* fix(#1012): invoke fallow with its real CLI and wire the report normalizer

The /gsd-code-review structural pre-pass invoked fallow with flags no published
fallow version accepts (--json, --profile, --stdin-files), so it failed on every
run and degraded silently per REQ-FALLOW-02 — the feature never delivered on any
fallow version. Three compounding defects:

1. Invalid flags. Real fallow audit uses --format json (not --json), -q/--quiet,
   --changed-since/--base for changed-files scoping (no file-list input), and
   --max-crap for thresholds. There is no --profile or --stdin-files.
2. Exit-code handling. fallow audit exits 1 when it FINDS issues (verdict=fail),
   0 when clean. The pre-pass treated any non-zero exit as a crash and discarded
   the output — i.e. it threw away exactly the findings it exists to surface.
   Success is now decided by whether a valid fallow JSON report was produced,
   not by the exit code.
3. Schema mismatch. normalizeFallowReport parsed a fictional top-level schema
   (unusedExports/duplicates/circularDependencies) fallow never shipped, and was
   dead code (the workflow embedded raw JSON; its tests asserted the fictional
   schema, one even calling a non-existent runFallowAudit and passing vacuously).

Fixes: align the invocation to fallow's documented agent-facing pattern; map the
profile preset (minimal/standard/strict) to --max-crap (50/30/15); scope phase
runs via --changed-since with a repo-scope fallback; rewrite the normalizer to
fallow's real schema (dead_code.unused_exports/unused_files/circular_dependencies
+ duplication.clone_groups) and wire it into the workflow so the reviewer
receives normalized findings; replace the fictional-schema fixtures and tests
with real-schema ones and delete the vacuous runFallowAudit test.

Closes #1012

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1012): backfill changeset PR number to 1044

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:36:50 -04:00
Tom Boucher
1e3ce6df05 fix(#222): orchestrator self-heal when synthesizer returns SUMMARY.md inline (#1042)
* fix(#222): orchestrator self-heal when synthesizer returns SUMMARY.md inline

The gsd-research-synthesizer agent intermittently hits an LLM false-refusal:
instead of writing .planning/research/SUMMARY.md with the Write tool, it returns
the SUMMARY.md content inline and fabricates a non-existent write restriction
(e.g. "the runtime is blocking file writes"). The shipped prompt hardening
(#240) is necessary but insufficient — the false-refusal recurs under some
context loads, and a drifting subagent then leaves gsd-roadmapper to fail with
"SUMMARY.md not found".

Adds an orchestrator-level self-heal to new-project.md and new-milestone.md:
after the synthesizer returns, verify .planning/research/SUMMARY.md exists; if it
is missing but the agent returned content inline, the orchestrator persists that
content with the Write tool (logging a warning) before spawning gsd-roadmapper;
if missing with no content, surface the error and stop rather than proceed
against a missing SUMMARY.md. This absorbs the failure mode deterministically
instead of depending on the subagent never drifting.

Closes #222

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#222): backfill changeset PR number to 1042

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:36:39 -04:00
Tom Boucher
fbd62cd84f feat(#1035): phase 5a — author 16 role:runtime capability descriptors (registry-only) (#1039)
Author capabilities/<runtime>/capability.json for all 16 runtimes (role:runtime),
populating the registry's runtimes index ({} -> 16). Each carries the 6 ADR-1016
axes (configHome structured, configFormat, artifactLayout structured, commandStyle,
hooksSurface + hookEvents, sandboxTier, supportTier), extracted from the live
modules (runtime-homes/runtime-slash/runtime-artifact-layout/runtime-config-adapter
+ CODEX_AGENT_SANDBOX). validateRuntimeBody tightened to enforce the closed
vocabularies + structured configHome/artifactLayout (rejects old string configHome;
env required; skillsHome recursively validated).

Registry-only: bin/install.js + src/*.cts untouched, zero behavior change. The 5b-5f
drive steps consume these one axis at a time (ADR-1016 §8).

Codex accuracy review caught + fixed: opencode/kilo register ZERO lifecycle hooks
(install.js skips the whole block) -> hooksSurface:none (configFormat stays
settings-json); kimi probe selects on <root>/skills existence -> probeExists:"skills".
ArtifactKind required-field strictness + ConverterName closing deferred to 5d/5e.
ADR-1016 (Proposed) amended to match (Decisions 1 + 5).

Closes #1035

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 09:34:07 -04:00
Tom Boucher
8bd5e07c58 feat(#1031): autonomous §3a.5 plan:pre ui-phase cutover — last inlined ui-phase site (step-only) (#1033)
Cut over autonomous.md §3a.5 (autonomous plan:pre ui-phase step) to the
loop.render-hooks plan:pre dispatch — completing the ui-phase migration begun
in #1026 (plan-phase.md §5.6). Step-only, non-blocking: autonomous is always
pipeline, so it fires active kind==step hooks and never runs the manual-only
plan:pre blocking gate.

Skip condition keys on "no active step hooks" (not empty activeHooks), so the
gate-only {ui_phase:false, ui_safety_gate:true} case skips silently with no
spurious warning — matching OLD §3a.5. Fires gsd-ui-phase under the identical
precondition (frontend + no UI-SPEC + workflow.ui_phase active), bare
${PHASE_NUM} args. Replaces the inline ui-safety-gate.cjs probe + config-get
with render-hooks + the ui.plan-gate check verb.

Codex caught the gate-only spurious-warning divergence on the first pass; fixed
+ re-confirmed equivalence-preserving. gsd-ui-phase skill, §5.6, §3d.5 untouched.

Closes #1031

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 08:31:35 -04:00
Tom Boucher
9ed8c7d574 feat(#1026): §5.6/ui-phase cutover — first gate dispatch (plan:pre step + blocking gate) (#1028)
Replace plan-phase.md §5.6 (a 6-branch inline UI gate) with a capability-driven
loop.render-hooks plan:pre dispatch — the FIRST gate dispatch in any workflow.
A step (ui-phase, when:workflow.ui_phase) + a new blocking gate
(when:workflow.ui_safety_gate). New ui.plan-gate check verb returns
{frontend, hasUiSpec, block}; the dispatch runs it unconditionally then fires
the active step (pipeline) or halts on the active blocking gate (manual). The
gate-handling (run check.query; halt if blocking+block) is the reusable
phase-6 template for blocking-gate cutovers.

Config semantics fixed per #1022 + maintainer call: ui_phase gates plan-time
UI-SPEC generation, ui_safety_gate gates the planning block. Common case + all
ui_phase=false cases are equivalence-preserving; the one intended change is
{ui_phase:true, ui_safety_gate:false} now auto-generating in pipelines.

Review found it broken twice (non-generic dispatch, phase-lookup divergence,
then the step-only check nested in a gate loop) — fixed; final Codex pass
verified all 8 (ui_phase,ui_safety_gate)x{pipeline,manual} cases correct.
gsd-ui-phase skill + autonomous §3a.5 untouched (§3a.5 deferred).
getRoadmapPhaseWithFallback mirrors cmdRoadmapGetPhase for lookup parity.

Closes #1026

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 00:58:21 -04:00
Tom Boucher
2ac6592096 feat(#1023): first phase-6 cutover — ui-review (verify:post) inline → loop.render-hooks dispatch (#1024)
Replace the inlined ui-review invocation in autonomous.md §3d.5 with a
loop.render-hooks verify:post dispatch — the first workflow to consume
render-hooks and fire a skill from it (closes the #1018 live-execution residual
as real wiring). Capability-driven, equivalence-preserving for the current
registry (only ui-review at verify:post, default on): fires gsd-ui-review under
the same precondition (UI-SPEC exists via consumes-gate + workflow.ui_review).

Gate findings (real pattern issues, fixed so every future cutover inherits them):
- bug-2643 static "Skill() references a real skill" check vs templated
  Skill(skill="gsd-${ref.skill}") dispatch → skip ${...}-templated names.
- Coverage moved, not lost: gen-capability-registry now validates
  steps[].ref.skill in skills + ref.agent in agents + rejects gsd- double-prefix.
- Tightened §3d.5 tests; markdown clarity (consumes rule, LLM-native JSON read,
  UI-REVIEW.md score hint).

gsd-ui-review skill + §3a.5/ui-phase untouched. §5.6/ui-phase cutover deferred
(#1022 step-can-halt-vs-gate model question).

Closes #1023

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 22:40:58 -04:00
Colin Johnson
76f42ddb4b feat(#1014): add Claude Fable 5 model config (#1015) 2026-06-10 20:32:26 -04:00
Jeremy McSpadden
092340d18a fix(#711): wire autonomous convergence flag (#729)
* fix(#711): wire autonomous convergence flag

* Update wise-ibex-tumble.md

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 16:10:07 -04:00
Jeremy McSpadden
f61b97276e fix(#724): block convergence on actionable review findings (#728)
* fix(#724): block convergence on actionable review findings

* merge: integrate clean next (#936 inline) onto author tip + re-apply cursor fixes (Mode field, REVIEWS.md extraction) and review hardening (#724)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 14:54:15 -04:00
Jeremy McSpadden
fb37fa7dd5 fix(#725): route Codex gsd-tools calls through shim (#731)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 14:44:43 -04:00
Joe
5e8a723089 feat(templates): add optional Business Context section to PROJECT.md template (#756)
* feat(templates): add optional Business Context section to PROJECT.md template

Adds an optional `## Business Context` section (Customer, Revenue model,
Success metric, Strategy notes) between Core Value and Requirements, for
monetized or customer-facing projects. Optional by default — an HTML comment
tells non-business projects to delete it; capped at four one-line fields to
stay a constraint reference, not a business plan. The milestone evolution
review in complete-milestone.md checks it only when the section is present.

Refs #72

* chore(changeset): set pr number for #72 fragment

* test(#72): add source-text-is-the-product exemption marker

Addresses review Minor #1 on PR #756. The contract test reads the
PROJECT.md template and complete-milestone workflow .md files and
asserts on their content (the local/no-source-grep pattern). Those
.md files ARE the product surface, so this is a valid
source-text-is-the-product case. Add the explicit // allow-test-rule
marker per RULESET.TESTS.no-source-grep.exemption so intent is
audit-traceable before the rule promotes to error (#453).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-10 14:13:01 -04:00
Tom Boucher
4c10eb2253 fix(#991): inject configured agent_skills into code-review family subagents (#1005)
* fix(#991): inject configured agent_skills into code-review family subagents

code-review.md, code-review-fix.md, and eval-review.md spawned their
subagents (gsd-code-reviewer / gsd-code-fixer / gsd-eval-auditor) without
querying or injecting the project-configured agent_skills, while ~20 sibling
workflows do. Subagents don't inherit the orchestrator's auto-loaded context,
so this injection is the only channel — reviewers/fixers/auditors silently ran
without the configured rule/skill context.

Mirror the established sibling idiom: add
`VAR=$(gsd_run query agent-skills <agent-type>)` in each workflow's initialize
step and interpolate `${VAR}` into every Agent() spawn of that type. This
covers all spawn sites, including code-review-fix.md's --auto loop which
re-spawns gsd-code-reviewer in addition to the two gsd-code-fixer spawns.

Regression test reads the workflow text (source-text-is-the-product) and
asserts each file queries agent-skills for every agent type it spawns and
interpolates the result at least once per spawn.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#991): add changeset for code-review agent_skills injection fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 14:07:23 -04:00
Tom Boucher
972a41a528 fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:) (#990)
* fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:)

Closes #967

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#967): backfill changeset pr number (990)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 11:10:40 -04:00
Tom Boucher
626575cbc5 fix(#978): parse --force in milestone complete dispatcher so the guard's documented override works (#982)
* fix(#978): parse --force in milestone complete dispatcher so the guard's documented override works

The dispatcher built `{ name, archivePhases }` but never parsed `--force`, so
`options.force` was always `undefined` and the guard inside `cmdMilestoneComplete`
(which tells users to "Re-run with --force to override") could never be bypassed.

Add `const force = args.includes('--force')` and pass it into the options object.
The guard already honors `options.force` — no changes to milestone.cts needed.

Closes #978

* chore(#978): backfill changeset pr number (982)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 10:40:06 -04:00
Tom Boucher
caca4d255c feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4) (#988)
* feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4)

Migrate the intel CLI command family from a hardcoded gsd-tools.cjs case arm to
a registry-dispatched Capability (commandFamilies mechanism, #961), mirroring
the graphify (#972) and audit (#984) cutovers. New src/intel-command-router.cts
exports routeIntelCommand reproducing all 9 subcommands verbatim (incl. the
status timeAgo non-raw post-processing), lazily requiring intel.cjs inside the
route fn. capabilities/intel/capability.json declares the intel command family;
commands-only (skills:[]), declares the existing intel.enabled gate (default
false — behavior unchanged).

Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing intel.test.cjs passes unchanged. Completes the first-party command-
family cutover sequence (graphify/audit/intel).

Closes #985

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#985): compute expected planningDir via path.join in intel cutover unit tests (Windows CI)

The intel-command-cutover unit-mock assertions hardcoded a POSIX `/.planning`
expectation while the router builds it with path.join(cwd, '.planning') →
backslashes on Windows, so the planningDir-arg assertions (query/status/diff/
snapshot/validate/update/api-surface) failed only on windows-latest CI. Compute
the expectation with path.join (cross-platform); production router unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 10:00:30 -04:00
Tom Boucher
9a03539c2d feat(#981): audit-uat + audit-open command cutover — commands-only capability (ADR-857 phase 4d-impl-3) (#984)
Migrate the audit-uat + audit-open CLI commands from hardcoded gsd-tools.cjs
case arms to a registry-dispatched Capability (commandFamilies mechanism, #961),
mirroring the graphify cutover (#972). New src/audit-command-router.cts exports
routeAuditUat/routeAuditOpen, each lazily requiring only its backing module
(uat.cjs/audit.cjs) inside the route fn — matching the old per-case lazy loads.
capabilities/audit/capability.json declares the two command families;
commands-only (skills:[]), no config gate, audit_review cluster untouched.

Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing audit regression tests pass unchanged.

Closes #981

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 08:34:49 -04:00