Commit Graph

90 Commits

Author SHA1 Message Date
Tom Boucher
b2c0086c1b fix(#1574): resolve review — copilot instruction file is .github/copilot-instructions.md
GitHub Copilot reads repository-wide instructions only from
.github/copilot-instructions.md (confirmed via GitHub Docs), not a root
copilot-instructions.md. Aligns getProjectInstructionFile with the installer
(runtime-config-adapter-registry installSurface 'copilot-instructions') and
cites the docs source in the doc-comment.
2026-06-22 09:56:15 -04:00
Tom Boucher
bf9bd1f4e0 fix(#1529): emit runtime-native instruction file from new-project 2026-06-22 09:32:29 -04:00
Joe
e12a2abfd8 feat(#441): add /gsd-capture --list-seeds for seed listing and audit (#722)
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit

Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed),
enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way
to browse or audit parked seeds on demand. This adds a read-only listing,
following the established --list → workflow pattern (per the approved scope on

- gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the
  seeds dir, returns { count, seeds[], summary } JSON with each seed's id,
  slug, status, scope, trigger_when, planted, title. Optional case-insensitive
  status filter. User-controlled content is sanitized (sanitizeForDisplay) and
  every path validated (requireSafePath); read-only. Independent of
  audit.scanSeeds, which only returns unimplemented seeds for the milestone surface.
- /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that
  renders the seed table.

Closes #441

* chore(#441): point changeset fragment at PR #722

* test(#441): allowlist list-seeds test in prompt-injection scan

The test asserts that list-seeds neutralizes injection payloads
(<system>, [INST]) embedded in seed content, so the fixtures legitimately
contain those patterns — same as the sibling security tests already on the
allowlist.

* fix(#441): use canonical /gsd:capture colon form in list-seeds workflow

Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use
the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is
retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new
list-seeds workflow used the hyphen form.

* docs(#441): sync help full.md + INVENTORY for --list-seeds

Adds the --list-seeds entry to the help reference (help/modes/full.md, per
bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow
in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json.

* docs(#441): add --list-seeds how-to + drop phantom statuses

Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers):

- USER-GUIDE.md Seeds section (how-to): extend the task to cover
  auditing parked seeds on demand via --list-seeds, including the
  status filter — kept task-oriented per Diataxis how-to mode.
- CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected
  from the list-seeds filter vocabulary; the system only produces
  dormant|active|triggered (src/audit.cts scanSeeds). Reference must
  be factually accurate and complete.

* fix(#441): guard non-scalar status frontmatter in cmdListSeeds

A seed with a bare `status:` line (extractFrontmatter yields {}) or a
`status: [a, b]` value (yields an array) crashed the whole audit list:
`(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string.
Coerce every frontmatter read through a `fmStr` helper (mirrors the existing
`typeof fm.id === 'string'` guard), so a non-scalar status falls back to
dormant and non-scalar scope/trigger_when/title can no longer leak a raw
array/object into the JSON contract. Title is now capped symmetrically.

Adds regression coverage for empty and array `status:` and non-scalar fields.

Refs #441

* docs(#441): align list-seeds workflow status vocabulary

The load_seeds step listed `implemented` as an example status filter, but the
real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds);
`implemented` has no producer. Matches the earlier CLI-TOOLS.md correction.

Refs #441

* refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds

Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported
deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested
in-process (review minor #1). No behavior change.

Filter comparison now matches the raw lowercased status (both sides already
normalized) instead of sanitizeForDisplay(status); sanitization is for output,
not matching (review nit #3).

* test(#441): add fast-check property coverage and count=1 boundary for list-seeds

Adds tests/list-seeds.property.test.cjs with four fast-check properties over
deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug
invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing
(review minor #1).

Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2).

* chore(#441): sync runtime launcher snippet into list-seeds workflow

Propagate the current _runtime-launcher.snippet.sh (with non-Claude
runtime home probes) into the new list-seeds.md workflow via
scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation.

* test(#441): record list-seeds.md in workflow size baseline (#1074)

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-22 00:59:41 -04:00
Behruz Nassre Esfahani
ac40f070ef feat(#1318): require external reviewers to verify plan claims against source (#1421)
* feat(#1318): require external reviewers to verify plan claims against source

/gsd-review built its external-reviewer prompt from plan text only and never
asked reviewers to open the repo and verify claims, so a grounded HIGH could be
outvoted by ungrounded LOWs. Add a concise, generic source-grounding block to
build_prompt's Review Instructions: treat yourself as running in the working
tree, open referenced files, cite path:line + mechanism, trace asserted
mechanisms, downgrade to an open question if you have no file access, and know
that grounded findings are weighted more heavily.

Also clarify that CodeRabbit (a diff-only reviewer that never receives the
prompt) must not be weighted as a grounded plan-level verdict in consensus
synthesis. Workflow stays under its size cap (baseline bumped deliberately).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1318): add changeset for reviewer source-grounding

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1318): mark changeset docs-exempt (internal reviewer-prompt wording)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#1318): document reviewer source-grounding in COMMANDS.md; drop docs-exempt

Review: a user-visible behavioral Changed warrants a docs touch, not a
docs-exempt. Add a sentence to the /gsd-review entry in docs/COMMANDS.md
(reviewers verify against source, cite file:line, grounded findings weighted
higher) and remove the changeset docs-exempt marker so lint:docs passes via
docs-updated. Also note the literal build_prompt test anchor.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1318): harden build_prompt fence extraction to be fence-run-aware

Addresses maintainer review on PR #1421 (required-before-merge).

The buildPromptReviewInstructions() test helper located the closing fence with
`src.indexOf('\n```')`, which terminates at the FIRST triple-backtick line — so
a build_prompt ```markdown block whose body embeds a fenced code example would
truncate mid-content (dropping the `## Review Instructions` section) and give a
spurious failure or false pass. Since this feature feeds source/plan content
(which routinely contains code fences) to reviewers, that is a live fragility.

Rewrite the extraction to be fence-run-aware, mirroring the CommonMark close
rule in src/markdown-sectionizer.cts stripFencedCode: parse the opener's
backtick run length, then close on the first line with >= that many backticks
and only trailing whitespace — so a shorter nested fence is treated as content.
Add a fail-first regression test (a 4-backtick outer fence wrapping a nested
```bash block) asserting the trailing `## Review Instructions` still extracts.

Test-only change; no production .cts touched. Verified: test file 7/7,
empirical fail-first proof the old indexOf logic truncated, full suite
4236/4236, eslint clean. Codex review: approve.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-22 00:47:49 -04:00
Enes Yağız
8748e95ed1 fix(#666): pr-branch silently ignored planning.sub_repos (#667)
* fix(pr-branch): handle sub_repos from config with git -C (#666)

Adds a `handle_sub_repos` step between `detect_state` and
`analyze_commits`. When `planning.sub_repos` is set in config, the
workflow now:

- Reads sub-repo paths via `gsd_run query config-get sub_repos`
- Skips the step entirely when the list is empty/null/[]
- Scans each repo with `git -C "$REPO" status --porcelain`
- Offers the user all/select/skip choices
- For selected repos: creates a PR branch, commits all staged/unstaged
  changes, pushes, and opens a companion PR via `gh pr create`

All git commands use `git -C "$REPO"` — never `cd "$REPO"` — because
shell state does not persist between agent-executed commands.

Closes #666

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: update changeset pr number to 667

* fix(pr-branch): address maintainer review — correct seam, behavioral tests, robustness

Resolves all three blockers and seven robustness issues raised in PR #667 review:

Blockers:
- Use `planning.sub_repos` (not top-level `sub_repos`) so config-get actually resolves
- Replace prose grep test with behavioral fixture tests using runGsdTools + local bare repo
- Extract sub-repo git work into new `cmdPrSubrepo` seam in src/commands.cts;
  never uses git add -A — stages explicit files only (universal-anti-patterns.md:44)

Robustness:
- Dirty-repo list persisted via mktemp/cat, not bash arrays (cross-block safe)
- Branch name embeds repo slug (${CURRENT_BRANCH}-${REPO_SAFE}-pr) to avoid collision
- push --set-upstream so gh pr create finds the branch
- Sub-repo base branch resolved via ls-remote with fallback to repo's default branch
- Remote slug parsed with /github\.com[:/]/ (handles SSH + HTTPS + .git-less URLs)
- rollback() cleans up branch on any mid-sequence failure
- node -e replaces jq (always available, no undeclared hard dep)

Refs: #666

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(pr-branch): security guard, push timeout, rollback fix, porcelain fix

Security (Blocker 1):
- Use security.cjs validatePath() in cmdPrSubrepo for symlink-safe workspace
  containment check — rejects ../escape, absolute paths, and symlink traversal
- Add negative regression test: '../escape' repo path must be rejected

Robustness:
- Push uses timeout: 60_000 ms (network op needs more than the 10 s default)
- Capture prevBranchName before checkout -b so rollback uses explicit name
  instead of git checkout - (fails on fresh single-branch repos)
- Porcelain path parse: line.trimStart().slice(2).trim() handles all XY
  combinations and the execGit global-trim edge case uniformly

Tests: 17/17 pass, lint: 0 errors

Refs: #666

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(pr-branch): move regression tests to commands.test.cjs, add core.quotePath=false

- Move cmdPrSubrepo behavioral + workflow source-invariant tests from
  standalone bug-666-*.test.cjs into tests/commands.test.cjs under
  describe('pr-subrepo') per TESTING-SUITES.md policy (no new bug-* files).
  Adds allow-test-rule: source-text-is-the-product see #666 for the
  workflow-source-invariant suite.
- Add -c core.quotePath=false to git status --porcelain call so non-ASCII
  filenames (e.g. café) are not C-escaped, keeping slice(2) parse correct.

* fix(pr-branch): remove obsolete regression tests for sub-repos handling

* fix(pr-branch): update workflow-size-baseline, add dirty-scan timeout

- Regenerate tests/workflow-size-baseline.json for pr-branch.md growth
  (+handle_sub_repos step, +timeout addition).
- Add { timeout: 10_000 } to the execFileSync git status --porcelain
  call in the handle_sub_repos dirty-scan (repo convention: every git
  subprocess is bounded, never hangs).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: regenerate INVENTORY-MANIFEST after rebase onto next

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): handle rename staging and split changedFiles from filesToStage

For git mv renames, the old path no longer exists in the worktree after
the move — staging it with git add fails. Split parsing into changedFiles
(both paths, for result.files) and filesToStage (new path only for
renames; old is already staged by git mv). Also adds porcelain tests
for staged renames, non-ASCII filenames, and a fast-check property test.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): rollback on push failure in cmdPrSubrepo

If push fails the branch only exists locally; rollback cleans it up so
the sub-repo is not left in a half-committed state.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): do not rollback after commit on push failure; add push-fail regression test

Post-commit push failures are network/auth/policy issues — the user's work
is already committed on the local branch. Calling rollback() at that point
force-deletes the only ref holding the commit (data loss). Leave the branch
in place and emit a retry instruction instead.

Adds a regression test (pre-receive hook that rejects all pushes) asserting
the branch and commit survive a push rejection so the failure path stays
covered going forward.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: regenerate INVENTORY-MANIFEST after rebase onto next

Rebased onto current next (#1267 retired core.cjs). Stale tsbuildinfo and
a leftover bin/lib/core.cjs build artifact were masking the drift — wiped
both, rebuilt clean, and regenerated the manifest. gen-inventory-manifest
--check now exits 0.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): validate sub-repo paths before git invocation in pr-branch.md

The handle_sub_repos workflow ran git -C on raw planning.sub_repos config
values at two points before the pr-subrepo seam's validatePath guard ever
ran: the dirty-scan detection (git status) and the base-branch resolution
(git ls-remote / remote show). A traversal entry could point git outside
the workspace; an embedded newline could inject a spurious record into
the newline-joined dirty-file output and into the shell-interpolated
commit message.

Adds a containment check + character allowlist to the dirty-scan node
script (reject before any execFileSync), and a defense-in-depth shell
case guard on the same value before the second, independent git -C
invocation in the base-branch resolution block.

Adds a behavioral test that extracts and executes the actual shipped
node script from pr-branch.md (not a mirror) against a real traversal
target and an embedded-newline entry, asserting neither reaches git or
the dirty-file output.

Also updates the stale cmdPrSubrepo doc comment: push failures no longer
delete the branch (see prior commit), only stage/commit failures do.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#666): make sub-repo traversal scan test genuinely fail-first

The outside repo's only change was an untracked file, which the ?? filter
excludes — so the repo looked clean even with the guard removed, making the
traversal assertion vacuous (it passed against a neutered guard). Commit the
file first, then modify it, so the outside repo has a tracked dirty change:
without the path guard it WOULD be reported dirty, so the test now fails-first.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#666): symlink-safe (realpath) sub-repo containment in pr-branch.md

Finding A from re-review: the workflow guard used path.resolve, which only
normalizes '..' textually and does not follow symlinks — so an in-tree symlink
whose name has no '..' or '/' (e.g. "evil" -> /outside) passed both the charset
filter and the resolve+startsWith check, letting git status / ls-remote /
remote show run against a directory outside the workspace. The pr-subrepo seam
already used fs.realpathSync (validatePath); this brings the workflow layer to
parity.

- dirty-scan: realpathSync the root once, and realpathSync each candidate before
  the containment check; skip on throw.
- base-branch resolution: replace the weak `case *..*|/*` guard with a realpath
  containment check that yields a validated absolute SUB_REPO_DIR, and run git -C
  against that instead of re-concatenating $ROOT/$REPO_REL.
- security test: add a symlink-escape entry and a positive control (legit in-root
  backend must still be reported). Confirmed fails-first — regressing the scan to
  path.resolve makes the symlink case leak.

Also fixes a misleading-fallback minor: the workflow now checks the seam's exit
status and skips the companion-PR step on failure, instead of printing
"branch pushed, open PR manually" after a real stage/commit/push failure.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#666): harden pr-branch sub-repo flow against round-12 edge cases

Pre-emptive hardening of the workflow changes from the symlink fix:

- continue-outside-loop: the "skip companion PR on seam failure" block used a
  bash `continue`, but the per-sub-repo iteration is prose-driven (the agent
  loops, not a literal `for`), so `continue` would warn and no-op. Reframed as
  prose-gated control flow keyed on $SUBREPO_EXIT — no bash loop assumption.
- Windows portability: the new symlink security case now degrades gracefully
  (try/catch around fs.symlinkSync; skip just the symlink assertion when symlink
  creation lacks privileges) so it doesn't hard-fail on Windows CI.

Verified: seam exits 1 on error / 0 on success (error() → process.exit(1),
propagated through the shim), so the $SUBREPO_EXIT check is meaningful; bash -n
clean on the touched blocks; commands 156/156; lint:ci green; manifest in sync.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-21 22:34:00 -04:00
Rezolv
195d356d7c Merge branch 'next' into feat/1346-enhance-verify-phase-project-a-check-vio 2026-06-21 17:41:37 -04:00
Behruz Nassre Esfahani
faac9331f2 feat(#1298): add validated worktree record-agent writer verb for wave manifests (#1448)
Closes #1298
2026-06-21 15:38:44 -04:00
Tom Boucher
2c718bf972 fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs (#1537)
* fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs

Generalizes the Codex-only #1515/#1519 fix to every non-Claude runtime, and
wires it into the real install path (where it was previously dead-on-arrival).

Root causes:
1. The runtime-default stamping lived only in `_applyRuntimeRewrites`, but the
   installer emits `gsd-core/workflows/*.md` via `copyWithPathReplacement`, which
   never calls it — so a real `--codex`/`--cursor`/etc. install emitted
   `--default claude` and worktrees-on. RUNTIME mis-resolved to claude and the
   workflow ran executors unisolated against the main checkout. (#1515/#1519 were
   also dead-on-arrival in real installs; this repairs them.)
2. Only `case 'codex'` was stamped; every other non-Claude runtime kept the
   Claude default.

Fix:
- New `_stampNonClaudeRuntimeDefaults(content, runtime)` (single shared helper)
  stamps `--default <runtime>` + `use_worktrees=false` for every `runtime !=
  claude`; called from both `_applyRuntimeRewrites` and, crucially,
  `copyWithPathReplacement` in bin/install.js (the real workflow emit path).
- Generalize the fail-closed worktree guard `= codex` -> `!= claude` in
  execute-phase/quick/diagnose-issues (worktree isolation is Claude-Code-only).
- Flip manager/autonomous inline-vs-background gating to `codex -> background,
  everything-else -> inline` (research: only Codex can background-nest the
  pipeline's subagents; all others run inline, which they support).

Worktree-capability determination is research-backed (official docs for all 14
non-Claude runtimes: none honor GSD's isolation="worktree" mechanism, only Codex
background-nests). New end-to-end real-install test asserts the EMITTED workflow
is stamped — the regression guard that would have caught the dead-on-arrival bug.

Closes #1521

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

* chore(#1521): backfill changeset PR number (#1537)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 13:48:47 -04:00
Tom Boucher
2436b76980 fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees (#1519)
* fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees

A Codex install with a runtime-neutral .planning/config.json resolved
RUNTIME=claude and enabled git worktree isolation, which Codex's
spawn_agent cannot honor. Two root causes:

1. Workflows read `config-get runtime` / `config-get workflow.use_worktrees`
   without `--raw`, so config-get's JSON-quoted output ("codex") was captured
   verbatim into the bash var and broke every `[ "$RUNTIME" = ... ]` check —
   the Codex fail-closed guard was dead even when runtime:codex was explicit,
   and Claude's own worktree degrade-check was dead too. Add `--raw` to those
   reads across execute-phase, autonomous, manager, diagnose-issues, quick.

2. The conversion engine emitted `--default claude` for every runtime. Stamp
   the codex-emitted workflows to `--default codex` (runtime) and
   `--default false` (use_worktrees) in _applyRuntimeRewrites case 'codex', so
   a neutral config on a Codex install resolves runtime=codex / worktrees off.

Also extend the Codex fail-closed worktree guard to quick.md and
diagnose-issues.md (they spawned isolation="worktree" with no runtime guard).

Regression test asserts source<->engine parity across all five workflows
(DEFECT.GENERATIVE-FIX) plus fast-check property coverage of the stamping.

Closes #1515

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

* chore(#1515): backfill changeset PR number (#1519)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 11:56:59 -04:00
Dave
c09c13f295 enhance(verify-phase): node-test causation control — prove the RED is content-caused (#1346)
The #1279 node-test machine-proof confirmed a known-bad subject drives the
negative test RED, but could not distinguish a genuine content-violation from a
deceptive test that reds merely because GSD_PROHIB_SUBJECT is set. Add an
optional fifth flat scalar `check_clean_fixture` (-> CheckDescriptor.cleanFixture)
threading a KNOWN-CLEAN control subject through projectProhibitions +
descriptorFromProjection. When present, the prover also runs the check against
the clean subject and requires GREEN, so fail-first is proven only when the check
is RED on the violation AND GREEN on the clean subject (content-dependent).

Opt-in and additive: absent a clean fixture the prover behaves exactly as
post-#1314 (no control, documented residual), preserving the zero-authoring
compose path; the lint-rule kind needs no analog (its subject IS the linted
file, no env indirection). Coverage: RED-first deceptive case, positive,
missing-clean fail-closed, round-trip read-back/emit, fast-check property
extended to the 5th scalar, and an end-to-end COMPOSE capstone (honest vs
deceptive). Docs: ADR-550 dated addendum, prohibition-probe reference,
spec-phase + verify-phase workflows.

Closes #1346

Claude-Session: https://claude.ai/code/session_01GsPRb8zvpcT7Eat6vZw8PX
2026-06-21 10:44:48 -04:00
Tom Boucher
b63a500fad fix(#1505): extract context_guard step to reference file; fix allow-test-rule see ref
- Extract execute-phase.md context_guard step prose to
  gsd-core/references/execute-phase-context-guard.md (@-ref lazy load),
  bringing execute-phase.md back under the ADR-857 phase-6 size ceiling
  (92914 < 93166 bytes)
- Fix allow-test-rule comment: add `see #1452` per ADR-456 lint rule
- Update feat-1452 tests to check the reference file for extracted content
- Register execute-phase-context-guard.md in INVENTORY-MANIFEST.json and
  INVENTORY.md Workflow References section
- Regenerate workflow-size-baseline.json after file shrinkage

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 18:38:20 -04:00
Tom Boucher
4eca5ac96c feat(#1452): add workflow.context_guard_mode to guard execute-phase against context exhaustion
Proactive checkpoint guard fires at each wave boundary before spawning agents.
Self-assesses context pressure against context-budget.md degradation tiers and
warns (warn, default) or auto-invokes /gsd:pause-work (auto) when POOR tier
(70%+) is detected. Config key validated; defaults to \"warn\".

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 18:10:49 -04:00
Tom Boucher
f5276b36b3 fix(#1369): refresh wave manifest and re-check base before each wave in execute-phase (#1492)
* fix(#1369): refresh wave manifest and re-check base before each wave in execute-phase

Two compounding issues caused wave N+1 worktrees to fork from the stale
pre-wave-N commit, immediately tripping the worktree_branch_check FATAL
guard in every executor:

1. worktree.base-check auto-degrade only ran once at initialize time.
   After wave N merges advanced orchestrator HEAD past origin/HEAD, new
   worktrees were still forked from origin/HEAD (Claude Code "fresh" base).

2. WAVE_WORKTREE_MANIFEST was never unset between waves, so wave N+1
   reused the consumed wave-N manifest file, which would have blocked the
   step 5.5 manifest guard (#3384) on subsequent waves.

Fix: add two safeguards in execute-phase.md —
- Step 0.5 (start of each wave): re-runs worktree.base-check; auto-degrades
  USE_WORKTREES=false for that wave when HEAD has diverged from origin/HEAD.
- Step 7c (end of each wave): unsets WAVE_WORKTREE_MANIFEST so wave N+1
  creates a fresh per-wave manifest; re-asserts worktree.set-baseref
  (idempotent) and re-evaluates base degradation after wave merges land.

17 regression tests added in tests/bug-1369-wave-stale-base.test.cjs.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1369): rename test to fix-NNN convention; update workflow size baseline

Rename tests/bug-1369-wave-stale-base.test.cjs → tests/fix-1369-wave-stale-base.test.cjs
to satisfy the lint-regression-test-names gate (new files cannot use bug-NNN prefix).

Update tests/workflow-size-baseline.json for execute-phase.md: 93157 → 97393
(LF-normalized byte count after adding step 0.5 inter-wave base re-check and
step 7c between-wave manifest reset).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1369): add issue reference to allow-test-rule comment

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1369): extract new execute-phase steps to references; satisfy ADR-857 cap

Step 0.5 (inter-wave worktree base re-check) and steps 7b–7c (pre-wave
dependency check + between-wave manifest reset/base refresh) added by this
PR grew execute-phase.md to 97393 bytes, violating the ADR-857 phase-6
architectural mandate that host-loop bodies remain strictly below the
pre-phase-6 baseline of 93166 bytes.

Extract both new blocks into dedicated reference files:
- gsd-core/references/execute-phase-wave-guard.md (step 0.5)
- gsd-core/references/execute-phase-between-wave-reset.md (steps 7b + 7c)

Replace inline prose with @-reference pointers. File now measures 92851
bytes (LF-normalized), satisfying the ADR-857 capstone conformance gate.

Also update tests/workflow-size-baseline.json to 92851 and add both new
reference files to docs/INVENTORY-MANIFEST.json.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1369): update regression tests to read from extracted reference files

Steps 0.5 and 7b+7c were moved to reference files to satisfy the ADR-857
size cap on execute-phase.md. Tests now check @-reference pointers in the
workflow for ordering and read content assertions from the reference files.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 14:22:59 -04:00
Tom Boucher
3a3b2135c2 chore(#1073): purge phantom pre-migration issue refs from source, tests, docs (#1471)
#2551/#3182/#2361 are pre-migration get-shit-done-redux issue numbers with no
equivalent in open-gsd/gsd-core; they mislead triage and manufacture phantom
blockers. Repoint to real successors (#717 byte-budget rework, #720) or rewrite
as prose referencing the discuss-phase/modes progressive-disclosure split.
Correct co-located 'line budget'/'<500 lines' framing to the byte-based reality
(#717). Add a CI guard (tests/no-phantom-issue-refs.test.cjs) that fails if a
phantom ref is reintroduced. SSH-key patterns (id_ed25519) left untouched.
No user-facing runtime behavior change.

Closes #1073
2026-06-20 13:36:57 -04:00
Tom Boucher
120f85164b feat(#1355): detect-and-warn guard for claude-code agent-teams (#1371)
* feat(#1355): detect-and-warn guard for claude-code agent-teams

GSD's multi-agent orchestration can stall under claude-code's experimental
agent-teams (a subagent's completion fails to route to the orchestrator). Per
the maintainer decision, the accepted scope is a read-only detector + one
non-fatal warning — NOT the declined run_in_background/TaskOutput conversion.

- New Teams Status Module (src/teams-status.cts → gsd-core/bin/lib/teams-status.cjs):
  pure resolveTeamsStatus({runtime, env}) + thin CLI cmdTeamsStatus reusing
  resolveRuntime. active = strictly-truthy env flag AND runtime === 'claude'.
- Wire `gsd-tools query teams-status [--active]` (read-only; no capability
  registration needed — conformance gates govern features, not query commands).
- One non-fatal warning in plan-phase.md before the first Agent spawn, gated on
  `query teams-status --active`; zero behavior change on non-claude/teams-off.
- Hermeticity: clear CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS in run-tests.cjs +
  SESSION_ENV_KEYS. Docs reference + CONTEXT.md glossary. Built lib gitignored.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1355): add changeset for teams-detect guard

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1355): bump plan-phase.md workflow size baseline (+407B for teams warning)

The non-fatal agent-teams warning block added to plan-phase.md grew it
92759 → 93166 bytes, past its committed per-file baseline ratchet. The growth
is small, deliberate, and still well under the workflow tier hard cap. Regenerate
the baseline via `npm run size:baseline` (only plan-phase.md changed).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1355): register teams-status.cjs in the inventory manifest

The new teams-status CLI module is a tracked surface; regenerate
docs/INVENTORY-MANIFEST.json (cli_modules family) via
gen-inventory-manifest.cjs --write so the inventory-manifest-sync gate passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 08:52:23 -04:00
Tom Boucher
1a678eb0e9 fix(#1359): migrate workflows off deprecated TaskOutput to Read(outputFile) (#1362)
* fix(#1359): migrate workflows off deprecated TaskOutput to Read(outputFile)

The map-codebase and docs-update workflows collected background sub-agent
results with the deprecated Claude Code `TaskOutput` tool using `block: true`,
which has a confirmed main-session hang after the agent completes
(anthropics/claude-code#20236).

Migrate the collection steps to the upstream-recommended pattern: keep
`run_in_background=true` on the Agent spawn, then `Read` each agent's
`outputFile` (from the `async_launched` result) once it reports completion.
Completion-marker contracts and on-disk verification are unchanged, and the
non-Claude runtime fallbacks (sequential_mapping / sequential_generation) are
preserved byte-for-byte. docs-update's timeout note no longer references the
unwired `workflow.subagent_timeout` key (it kept a literal before).

Regression coverage folded into tests/subagent-timeout.test.cjs (the owning
module for background-subagent collection). Workflow size baseline regenerated
for the justified prose growth.

Refs #1355 (same latent hang surface).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1359): add changeset for TaskOutput migration

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-16 21:55:45 -04:00
Rezolv
00c05eb717 Merge branch 'next' into feat/1279-fail-first-prover 2026-06-16 17:35:55 -04:00
Tom Boucher
9540fe43b9 fix: have executor self-report worktree metadata (#1349) 2026-06-16 14:19:08 -04:00
Dave
56d4a1bf39 enhance(#1279): project check_violation_fixture scalar — #1278 locate + #1279 proof compose end-to-end (#1346)
Delivers option (a) from the #1314 maintainer review: thread a fourth flat
scalar check_violation_fixture through the projection so a prohibition authored
at spec-phase machine-proves fail-first and greens through the deterministic
path alone — zero hand-authoring at verify time.

- src/probe-core.cts: Prohibition gains check_violation_fixture?; projectProhibitions
  emits it (both kinds) ONLY for a well-formed descriptor and ONLY when non-empty
  (blank/absent -> projects absent -> producer hard-gates, never a partial green).
- src/prohibition-enforcement.cts: descriptorFromProjection reads it back into
  violationFixture via the same numeric-coercion-safe scalar() normalizer.
- Tests (RED-first, proven non-vacuous by reverting both src edits): CHK-02(#1346)
  projection emit, CHK-08(#1346) read-back, CHK-03(D) example round-trip, the
  fast-check round-trip property extended to the 4th scalar (the contract trek-e
  blocked #1301 on), and a real-subprocess COMPOSE capstone greening end-to-end
  through project -> descriptorFromProjection -> default prover+runner.
- Docs flipped from 'hard-gates until #1346' to 'composes end-to-end': verify-phase.md,
  prohibition-probe.md, spec-phase.md authoring, ADR-550 addendum, changeset.
  #1346 now tracks only the node-test causation residual.

190 affected-suite tests green; eslint + tsc clean; size baseline regenerated.
2026-06-16 14:00:49 -04:00
Tom Boucher
6e242bd76a fix: allow quick worktree parent plan base (#1347) 2026-06-16 14:00:06 -04:00
Dave
33b6ee6f0c fix(#1279): node-test prover fail-closes on missing violationFixture + honest projection docs
Addresses the #1314 maintainer review (trek-e):

- Major 1 (fail-OPEN): defaultProveFailFirst's node-test branch only guarded
  `if (!fixture)`. A missing/typo'd/stale violationFixture made GSD_PROHIB_SUBJECT
  point at a missing file; an honest negative test threw ENOENT *inside its
  callback* (a failing test named distinctly from the file), which
  isNonVacuousNodeTestRed accepted as proof -> a green forged from a setup crash.
  Now requires fs.existsSync(path.resolve(cwd, fixture)) before spawning, symmetric
  with the lint-rule path's file-result guard. Regression test pins it (RED without
  the guard); a second test pins cwd-relative fixture resolution.

- Major 2 (misleading prose): the #1278 projection carries no violationFixture, so
  the deterministic-locate path always hard-gates (fail-closed) until a
  check_violation_fixture scalar is threaded through. verify-phase.md and
  prohibition-probe.md no longer read as if the projected path produces greens;
  the ADR-550 addendum records both items. Tracked as follow-up #1346.

- Documented residual: existence is necessary but not sufficient (a red caused by
  the env being set vs the subject's content); recorded as a constraint, in #1346.

- Nit: stale 'NOT attested fail-first' comment -> 'NOT machine-proven fail-first'.

64 tests pass; eslint + tsc clean; changeset valid.
2026-06-16 13:05:20 -04:00
Dave
49bef1927e Merge remote-tracking branch 'origin/next' into feat/1279-fail-first-prover
# Conflicts:
#	docs/adr/550-spec-phase-probe-contract.md
#	gsd-core/workflows/verify-phase.md
#	tests/workflow-size-baseline.json
2026-06-16 00:32:22 -04:00
Rezolv
e50ead7ad2 enhance(verify-phase): deterministic auto-locate of the prohibition check descriptor (#1278) (#1301)
* test(1278): RED-first descriptor parity + fail-closed guards + CHK-07 byte-stability (wave 1)

- CHK-03 (RED): extend PROB-14 parity in prohibition-probe.schema.test.cjs to carry the flat
  check_kind/check_target/check_rule scalars through project->write->parseMustHavesBlock; the
  non-droppable check_kind-presence assertion is the load-bearing RED trigger (fails because
  projectProhibitions strips check_* on the current build).
- CHK-07 (GREEN forward-guard): probe-core.test.cjs pins descriptor-less byte-stability +
  dispositionForProhibition fail-closed policy, with a t.todo marker forward-locking plan 01-02.
- CHK-06 (RED): prohibition-enforcement.test.cjs asserts descriptorFromProjection export +
  fail-closed on absent/partial/unknown descriptors via the projection adapter (RED until 01-03).
- No src/*.cts or .cjs edits; no new test files; lint-test-file-count clean.

* feat(1278): add optional flat-scalar check descriptor fields to Prohibition interface (wave 2)

- check_kind?/check_target?/check_rule? mirror CheckDescriptor.kind/target/rule (minus caller-attested failFirst, #1279)
- optional so existing Prohibition consumers compile unchanged

* feat(1278): project check descriptor as flat scalars in projectProhibitions (wave 2)

- emit check_kind/check_target (+ check_rule only for lint-rule with a rule) when descriptor well-formed
- under-specified/descriptor-less items project byte-identically (CHK-07); flat scalars ride existing parseMustHavesBlock continuation-KV path (no parser rewrite)
- add CHK-02 probe-core unit cases pinning the projection
- turns CHK-03 parity test GREEN; dispositionForProhibition untouched

* feat(1278): descriptorFromProjection read-back adapter feeds fail-closed locate (wave 3)

- Add descriptorFromProjection(projected) -> CheckDescriptor | null to
  src/prohibition-enforcement.cts: renames the projected flat scalars
  check_kind/check_target/check_rule -> {kind,target,rule?}, or null when
  the descriptor is absent/non-object (no check_kind key).
- failFirst is NEVER sourced from the projection (stays caller-attested; #1279).
- rule is set only when check_rule is a non-empty string; the adapter does NOT
  re-validate kind/target/rule — an under-specified descriptor reconstructs to
  one the EXISTING runProhibitionEnforcement LOCATE guard rejects (located:false,
  never green). The merged #1259 guard stays the single source of fail-closed truth.
- Turns the RED CHK-06 fail-closed tests (plan 01-01) GREEN end-to-end; CHK-03 /
  CHK-07 stay green. CheckDescriptor type, locate guard, dispositionForProhibition,
  and parseMustHavesBlock are unchanged (additive +36/-0).

* feat(1278): verify-phase locates prohibition check from projected descriptor (wave 3)

- request.check kind/target/rule sourced from projected check_kind/check_target/check_rule via descriptorFromProjection, not verifier invention (CHK-05)
- replaces the #1278 author-supplied / tracked-follow-up note with the delivered deterministic-locate behavior
- preserves fail-closed routing: absent/partial descriptor -> never green, hard-gate in both modes
- failFirst stays a verify-time caller attestation; #1279 bounds the remaining fail-first proof

* feat(1278): spec-phase captures wired-check descriptor on test-tier resolution (wave 3)

- Step 5.6 'Keep it' / verification: test path captures check_kind/check_target/check_rule, projected onto must_haves.prohibitions for verify-phase deterministic locate (CHK-04)
- SOFT capture: a test-tier prohibition without a descriptor is still allowed (no hard authoring block); stays fail-closed/flagged downstream
- --auto captures only an unambiguous descriptor, never fabricates a check path
- failFirst NOT captured at spec-phase (verify-time attestation; #1279)
- PROB-06 soft-gate + text-mode (PROB-09) behavior unchanged

* chore(1278): re-baseline workflow size for grown verify-phase + spec-phase prose (wave 3)

- spec-phase.md 28438 -> 30343 (+1905), verify-phase.md 35362 -> 36498 (+1136)
- regenerated via npm run size:baseline (no hand-picked numbers); growth is the #1278 deterministic-locate + descriptor-capture prose
- workflow-size-budget guard green (122/122)

* docs(1278): ratify optional check descriptor in dated ADR-550 addendum + type:Changed changeset

- Append dated 2026-06-15 ADR-550 addendum ratifying the D3 prohibition-item
  shape extension (optional flat-scalar check_kind/check_target/check_rule)
- Document flat-scalar rationale, deterministic projection/read-back,
  fail-closed on partial/invalid/absent, #1279/policy out-of-scope
- Add .changeset/1278-prohibition-check-descriptor.md (type: Changed)

* docs(1278): document optional check descriptor in prohibition-probe reference + FEATURES

- Add 'Optional wired-check descriptor (deterministic locate, #1278)' section
  to the prohibition-probe reference (flat-scalar keys, projection/read-back,
  fail-closed + backward-compat, failFirst stays attested)
- Add deterministic prohibition-check descriptor source entry to FEATURES.md
- No CONTEXT.md glossary change: descriptor reuses existing wired-check /
  verification:test vocabulary, no new glossary term introduced

* fix(1278): pass packaging gates — changeset pr field + retired slash-form fix

- Add required pr: 1278 to changeset (lint:changeset MISSING_PR hard requirement;
  plan's 'omit if unknown' was inaccurate — issue number per #1259 convention,
  updated to real PR number when opened) [Rule 3 - blocking]
- Fix retired /gsd-spec-phase -> /gsd:spec-phase at verify-phase.md:83 (wave-3
  prose; caught by slash-namespace invariant #3443/bug-2543, blocked CHK-09
  full-suite-green) [Rule 1 - bug]
- size:baseline + INVENTORY manifest verified in-sync post-build (no diff)

* docs(1278): add check descriptor + descriptorFromProjection to CONTEXT.md prohibition glossary

* fix(1278): harden descriptorFromProjection round-trip (numeric-coercion + stray-rule) per review

- MD-01/LW-01: narrow projected scalars to primitives + String()-coerce, so a
  numeric-looking check_target (parseMustHavesBlock coerces ^\d+$ to number)
  reconstructs as a string and locates instead of silently un-locating; no
  as-string type-lie, satisfies no-base-to-string.
- LW-02: attach rule only for the lint-rule kind (drop a stray node-test rule).
- LW-03: document the optional check_* keys in the reference Output schema.
RED->GREEN tests added in prohibition-enforcement.test.cjs.

* chore(1278): set changeset pr to 1301

* test(1278): add fast-check property for the check-descriptor round-trip + fail-closed (trek-e review)

RULESET.TESTS.property-based-testing: the projectProhibitions -> render ->
parseMustHavesBlock -> descriptorFromProjection chain is a bijective/transformation
contract. Adds 2 fc properties to tests/probe-core.property.test.cjs (no new file;
ratchet stays at 2 for probe-core):
- well-formed descriptors survive the round-trip across the full string domain
  incl. the numeric-coercion case (target/rule reconstruct as strings);
- under-specified/invalid descriptors (absent / target-less / rule-less /
  unknown-kind) are always fail-closed (never green, flagged, unlocated).
Stability is asserted at the descriptorFromProjection layer (the raw parse step is
intentionally lossy for numeric scalars; the shared parser is unchanged).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-16 00:01:03 -04:00
Dave
8cf4716c99 docs(#1279): ratify ADR-550 addendum + FEATURES/reference/verify-phase + changeset
- ADR-550 dated 2026-06-15 addendum: machine-proven fail-first REALIZED (D5d CLOSED); ratify violationFixture, GSD_PROHIB_SUBJECT, failFirst demotion
- FEATURES REQ-PROHIB-07 reads shipped (machine-proven, not caller-attested)
- prohibition-probe reference + verify-phase descriptor shape updated with violationFixture + convention
- tag GSD_PROHIB_SUBJECT + violationFixture PROPOSED/renamable (zero live consumers) as a PR-review flag
- add .changeset/1279-machine-proven-fail-first.md (type: Changed)
2026-06-15 22:57:47 -04:00
Tom Boucher
91bd82f9a6 enhance(execute): isolated-executor rejected/over-reaching run fails safe — never default recovery to main (#1292) (#1303)
* enhance(execute): isolated-executor rejected/over-reaching run fails safe (#1292)

When an isolated (worktree) executor run is rejected — the user declines to
merge it, the orchestrator surfaces recovery for a blocked/halted plan, or the
run over-reached the requested scope — the orchestrator must no longer
default/propose recovery by editing the primary checkout (`main`). Absent an
explicit guardrail, the LLM orchestrator could improvise "continue on main",
inverting the isolation contract at the moment it matters most.

Added an ISOLATED-RUN RECOVERY — FAIL SAFE policy: default to a safe halt that
offers a fresh, narrowly-scoped worktree or inspect/discard; editing the primary
checkout requires explicit, clearly-labeled confirmation and is never the
default/proposed option.

To respect the ADR-857 phase-6 host-loop size cap on execute-phase.md (it sits
just under the pre-phase-6 baseline), the policy is delivered as an extracted
reference fragment rather than inline:
- New `execute-phase/steps/worktree-recovery-policy.md` holds the recovery policy
  (the existing FAIL-CLOSED rule #48 for base/HEAD mismatches + the #1292
  fail-safe guardrail). No #48 behavior change — moved verbatim.
- execute-phase.md references the fragment at the worktree-spawn recovery point,
  the step-5.5 merge decision, and the stalled-agent "switch to inline execution"
  menu (which for an isolated run now follows the fail-safe policy). Net effect:
  execute-phase.md shrinks below its cap.
- quick.md carries the fail-safe guardrail inline at its post-return merge/discard
  decision (quick.md is not size-capped).

Scoped to the recovery offer only — no automatic scope-overreach detection
(explicitly out of scope per the issue) and no new config key. Adds content
regression tests, a USER-GUIDE note, and a workflow size-baseline update.

Closes #1292

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(changeset): Changed fragment for #1292 isolated-executor fail-safe recovery

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 20:53:38 -04:00
Tom Boucher
137760a655 fix(#1296): align config docs/prompts/schema with consumers (#1299)
* fix(#1296): align config docs/prompts/schema with consumers

The user-facing config surface disagreed with what the consumers actually do
(subset of the #1216 audit). No runtime consumption behavior changes.

- workflow.subagent_timeout: settings-advanced.md prompt + docs/CONFIGURATION.md
  said "seconds (default 600)" but the consumer (map-codebase.md) uses
  milliseconds (default 300000). Relabeled all four spots in settings-advanced.md
  (prompt, parse-default list, example, confirmation table) + the CONFIGURATION.md
  row.
- review.models.<cli>: settings-integrations.md, docs/CONFIGURATION.md (Integration
  Settings), and docs/CLI-TOOLS.md documented a shell command, but review.md injects
  the value into a --model/-m flag. Relabeled to a bare model id and reconciled the
  contradictory CONFIGURATION.md sections.
- workflow.test_command + workflow.build_command: consumed via config-get
  (test_command in verify-phase/execute-phase/audit-fix/post-merge-gate;
  build_command in post-merge-gate) and documented, but absent from validKeys so
  `config set` rejected them. Registered both in config-schema.manifest.json and
  documented them in references/planning-config.md (overview + complete reference).

Regression tests: behavioral config-set tests (tests/config.test.cjs) + doc-parity
content guards (tests/config-field-docs.test.cjs).

Deferred to other #1216 clusters: security-gate wiring, search_gitignored wiring,
mvp_mode, source_grounding_authority labeling, and config-set enum enforcement.

Closes #1296
Refs #1216

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(changeset): Fixed fragment for #1296 config-surface alignment

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 19:57:11 -04:00
Dave
0d783fa57d Merge remote-tracking branch 'origin/next' into feat/1259-test-tier-enforcement
# Conflicts:
#	tests/workflow-size-baseline.json
2026-06-15 15:42:23 -04:00
Tom Boucher
0a856f06cd enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier (#1271)
* enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier

Introduce a per-truth PRESENT_BEHAVIOR_UNVERIFIED state for must-haves that assert a state transition or a cancellation/cleanup/ordering invariant whose only evidence is symbol presence + wiring. Such truths are excluded from the verified_truths score, reported as a behavior_unverified count, recorded in an always-on behavior_unverified_items frontmatter list, and routed to the existing human_needed sink — so a clean N/N can no longer be reached on symbol presence alone.

The overall-status vocabulary and the src/verification.cts seam are unchanged (the new state is per-truth only); gaps_found keeps decision-tree precedence; override-passed truths still count toward verified_truths. Mirrors the calibration into the shipped verify-phase.md workflow (with an infra/foundation carve-out), the VERIFICATION.md templates, and docs (planning-artifacts.md, AGENTS.md). gsd-verifier.md kept under its 48KB LARGE cap; size baselines regenerated.

Closes #966

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#966): add changeset fragment for gsd-verifier behavior-unverified calibration

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 15:24:56 -04:00
Dave
6a9e6cb0ae fix(1259-01): address trek-e review — B1 (fatal/suppressed), B2 (bounded subprocess), M1/M2, minors
Maintainer CHANGES_REQUESTED (reviewed e08667e5, pre-portability-fix):
- B1: lint-rule no longer greens an unparseable target (eslintHasFatalError -> fail
  closed on any fatal/parse error) or an inline-suppressed violation (eslintJsonHasRule
  now scans suppressedMessages too). RED-first + real-runner repros.
- B2: both child spawns get a bounded timeout (30s node / 60s eslint) + 16MiB maxBuffer;
  timeout fails closed. Injectable timeoutMs (positive-only — 0/negative can't disable
  the bound) enables a fast 1.5s hang test.
- M1: scoped verify-phase.md — the check descriptor is author-supplied for now; filed
  #1278 for deterministic auto-locate of the descriptor (the locate half).
- M2: added tests/prohibition-enforcement.property.test.cjs (fast-check fail-closed
  invariants; within the <=2-file budget).
- m1: tapTestNames excludes # SKIP/# TODO; parseNodeTestSummary tracks # cancelled;
  isNonVacuousNodeTestPass requires cancelled===0.
- m2: scoped the determinism claim to the decision/parse layer (real runner is env-dependent).
- m3: filed #1279 for machine-proven fail-first (violation-fixture probe).
- n1: -- before target in both arg builders (option-injection). n2: dropped dead token.
- B3 (Windows npx) was already fixed in 2af76306 (pushed ~65s after the review).

Verified node 22 + 24; size baseline regenerated for the verify-phase note.
2026-06-15 15:19:04 -04:00
Dave
31b822b312 fix(1259-01): close adversarial review findings — genuine enforcement, honest fail-first scope
Adversarial pre-submission review found the injected-runCheck tests masked a
non-functional real runner. Fixes:

- BL-01 (false green on vacuous test): the node-test runner now parses the TAP
  summary and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail) AND a
  reported test named distinctly from the file — node --test counts an empty file
  as one passing test, so counts alone could not catch it.
- SF-01 (lint anchor never greened): the lint-rule runner now runs the project
  eslint as --format json and filters by ruleId, so plugin rules (local/*) load
  via the flat config — bare --rule cannot load a plugin. local/no-source-grep
  now genuinely greens (covered by a real, non-injected test).
- BL-02 (tautological fail-first): the runner no longer echoes the caller's
  failFirst as if confirmed. failFirst is documented as caller-ATTESTED; the
  producer requires attestation + a genuine non-vacuous pass. Machine-proven
  fail-first (needs a violation fixture) is flagged as a tracked follow-up in
  ADR-550, the changeset, FEATURES, the reference doc, and verify-phase.
- SF-02: added real-runner end-to-end tests (no injected runCheck) + pure,
  exported parse/filter helpers (parseNodeTestSummary, tapTestNames,
  eslintJsonHasRule, eslintFileResultCount) so the shipping branches are
  mutation-pinned.
- NIT-01/02: LOCATE guard rejects empty-string rule and unknown kinds.
- Hardening: spawn checks with NODE_TEST_CONTEXT/NODE_OPTIONS scrubbed so an
  ambient test-runner context cannot corrupt a verify-time result.
- Docs reconciled to the shipped behavior (no 'confirms fail-first' overclaim).
2026-06-15 13:38:15 -04:00
Dave
676436258a fix(1259-01): lint-rule real runner — keep rule id distinct from lint target
The default lint-rule runner passed check.target as BOTH the --rule id and the
eslint path, so it could never pass (eslint tried to lint a file named after the
rule). Add a distinct check.rule field (rule id) vs check.target (path to lint),
extract a pure exported buildLintArgs() so the mapping is mutation-testable
without spawning eslint, fail-closed on a lint-rule missing its rule id, and carry
the rule into enforcement evidence. Updates verify-phase descriptor docs.
2026-06-15 13:10:49 -04:00
Dave
ce01e1376b docs(1259-01): wire verify-phase consumer + ADR-550/FEATURES/reference + changeset
- verify-phase.md: replace test-tier 'fail-closed/deferred' bullet with the check prohibition-enforcement
  enforcement step (locate -> fail-first -> run -> evidence -> green-or-hard-gate); update determine_status tree
- ADR-550 addendum: mark D5d enforcement half LANDED (#1259); cite ADR-857 open-question §147 + D6 (core verify rail)
- FEATURES §146: enforcement wording + add REQ-PROHIB-07; keep REQ-PROHIB-06 intact
- references/prohibition-probe.md: test-tier enforced + hard-gates via check prohibition-enforcement
- changeset (type: Changed) with the D5 '2 no-source-grep invalid cases, not 96' correction
2026-06-15 12:47:39 -04:00
Rezolv
3556450b0d feat(spec-phase): surface zero-classification edge-probe requirements as unclassified candidates (#1110) (#1117)
Edge-probe now surfaces a zero-classification requirement (non-empty prose, no
shape cue matched, no `shapes` override) as a single soft `unclassified — review
manually` candidate instead of silently dropping it — the exact blind spot the
probe exists to catch. Dismissible like any edge; the `shapes: []` opt-out stays
silent; `TAXONOMY` (the closed 8 categories) is unchanged. Under `--auto` the
candidate is left `unresolved`, never auto-`backstop` (a missing shape is not
evidence an edge exists).

Closes #1110
2026-06-14 21:44:21 -04:00
Rezolv
395fb519e7 feat(spec-phase): prohibition probe — surface "must-NOT" constraints (#644) (#1149)
Adds the spec-time prohibition probe (spec-phase Step 5.6) — the second adapter of the probe-core resolution model. Surfaces unwritten must-NOT constraints as negative SPEC acceptance criteria with test/judgment verification tiers; fail-closed at verify time. Per ADR-550. Closes #644.
2026-06-14 21:29:11 -04:00
Tom Boucher
55eb1fd72e refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam (#1242)
* refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam

ADR-22's severity mapping, authority auto-upgrade, and rung>=3 hard-block lived only as prose in plan-review-convergence.md — untestable. Extracted into src/plan-drift-guard.cts (pure: AUTHORITY_RUNGS, getEffectiveAuthority, classifyDriftSeverity) + a gsd-tools drift-guard CLI seam (authority/severity), and rewired the workflow to call the seam deterministically instead of reasoning the decision in prose. 47 unit/e2e/structural tests cover the full severity table, the grep->intel auto-upgrade, and the rung>=3 HIGH hard-block.

Closes #1190

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1190): add changeset for ADR-22 drift-guard seam (#1242)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1190): register ADR-22 module in eslint-ignore + inventory manifest + docs-exempt changeset

Full-matrix CI surfaced new-module/command governance ripples beyond the lint-tests chain: (1) tsc-generated plan-drift-guard.cjs must be in the eslint ignore list (551-eslint-bin-lib-coverage); (2) docs/INVENTORY-MANIFEST.json must include the new module/command (regen via gen-inventory-manifest.cjs --write); (3) a type:Added changeset triggers docs-required — added a docs-exempt marker (internal seam, no user-facing surface).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 20:52:36 -04:00
Tom Boucher
0b3a2e5f9c feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15) (#1237)
* feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15)

ADR-15 designates /gsd-progress --next --auto --converge as the PRIMARY plan-convergence surface, but only the secondary surface (autonomous.md) was wired. next.md now parses --converge/--cross-ai into a plan strategy, gates on workflow.plan_review_convergence, forwards reviewer flags + --max-cycles, and routes Route-3 planning through /gsd:plan-review-convergence (mirroring autonomous.md); --auto chaining preserves converge mode. Adds argument-hint + help/full.md + COMMANDS.md + how-to parity and a structural regression test.

Closes #1190

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1190): add changeset for progress --converge surface (#1237)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1190): keep progress --converge docs skill-dep-clean + regen workflow size baseline

CI surfaced two ripples from the ADR-15 workflow edits: (1) lint-skill-deps + profile-closure flagged /gsd:plan-phase and /gsd:plan-review-convergence SlashCommand tokens in progress.md's flag docs as undeclared deps — reworded to plain prose since progress.md only advertises the flag (the real invocation lives in next.md); (2) the per-file workflow size baseline needed regenerating after the next.md/help edits.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 15:52:22 -04:00
Tom Boucher
bf634b95c3 feat(#1213): Capability State Writer — write-side inverse of the resolver (#1225)
* feat(#1213): Capability State Writer — write-side inverse of the resolver

Adds src/capability-writer.cts (setCapabilityState + cmdCapabilitySet) and the
`gsd-tools capability set` subcommand: the write-side inverse of the capability
resolver (ADR-1213). One desired capability state projects onto the substrates —
`enabled` drives the runtime surface (canonical on/off), `gates` drive federated
config keys (hook granularity), install profile is a read-only floor — then
re-resolves and reports divergence (assert-and-report), so "off means off" holds
as a write-time invariant. Adds batched setConfigValues; routes gsd:settings
capability hook-gates through the writer. Docs: CLI-TOOLS reference, how-to,
ADR-1213, CONTEXT.md term.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1213): add changeset for Capability State Writer (#1225)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 12:51:17 -04:00
Tom Boucher
7edd18fd2b feat(#1165): async external_job_waiting half-state + resume/pause contract (#1221)
Core half of #1105: a legal external_job_waiting deferred state so an async-dispatched Execute step (committing a .planning/async-jobs/<job>.json manifest, deferring SUMMARY.md) is not an illegal partial. execute-phase safe-resume, resume-project, and pause-work reconcile against the versioned scheduler-agnostic manifest stability contract without re-dispatching; the producer is the capability half (#1164). Closes #1165.
2026-06-14 12:23:07 -04:00
Tom Boucher
e9f9ae49c8 fix(#1146): single base-branch resolver across forking workflows (#1198)
* fix(#1146): single base-branch resolver across forking workflows

Replaces duplicated per-workflow bash detection that silently fell through
to :-main on repos where origin/HEAD is unset (git init+remote add+fetch
without set-head, most CI checkouts, many worktrees).

New CJS module git-base-branch.cjs exposes `gsd_run query git.base-branch`
with full precedence ladder: git.base_branch config override → origin/HEAD
symref → git remote show origin (authoritative) → local branch presence →
"main". All git subprocesses bounded with timeouts; degrades gracefully.

Wires execute-phase, quick, ship, complete-milestone, and pr-branch to the
single resolver. Removes 14 lines of duplicated detection bash across the
five workflows.

Includes 7 behavioral tests covering the full precedence ladder including the
key regression case (master repo, origin/HEAD unset → must return "master",
NOT "main") and an anti-regression guard that fails if any workflow
re-introduces the :-main/:-master fallback pattern.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(changeset): backfill PR number #1198

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1146): drop stray PR-body file from branch

pr-1146-body.md was committed during changeset backfill but must not
be tracked in the repo. Content preserved externally for PR body use.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#1146): add tests for flat base_branch config key and both-branch tie-break

Closes two mutation gaps identified in adversarial review:
- A2: flat {base_branch: ...} at config root (legacy key form) was covered
  by code but unguarded against mutation of lines 74-75 in resolver
- H: tier-4 tie-break when both main+master exist locally (main wins,
  per tryLocalBranch JSDoc) was documented but untested

9/9 tests pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#1146): allowlist workflow-literal guard as runtime-contract exemption

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1146): degrade gracefully when gsd_run unavailable in handle_branching bash blocks

handle_branching (execute-phase.md) and step 2.5 (quick.md) are extracted
and run verbatim by behavioral tests that lack the gsd_run preamble.
Adding a || fallback ladder (git symbolic-ref then echo main) keeps the
unified resolver as primary in real workflows while letting the test harness
succeed without gsd_run defined.

Also propagates updated runtime-launcher preamble to pr-branch.md (added in
origin/next MemPalace PR) and regenerates workflow-size-baseline.json.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 08:22:40 -04:00
Tom Boucher
a375c4b354 feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) (#1201)
* feat(#956): add MemPalace memory capability (ADR-857 feature plug-in)

Adds an opt-in, default-resilient ADR-857 feature capability that wires
MemPalace (local-first memory: MCP server + CLI) into the GSD loop:
deliberate recall before discuss/plan and verbatim + temporal-KG capture
at phase boundaries. Three memory modes (augment default; kg_backend and
replace forward-declared). Master gate mempalace.enabled (default off);
every hook onError:skip, zero gates; absent/disabled MemPalace => loop
unchanged. Transport is rendered-markdown only — MemPalace runs
out-of-process, no third-party code in gsd-core (ADR-857 §7).

Capability: capabilities/mempalace/ (manifest + 2 fragments), skills
commands/gsd/mempalace-{recall,capture}.md, agent
agents/gsd-mempalace-curator.md. Registration: ns-context router,
utility cluster, KNOWN_SKILLS, help full.md, model-catalog, copilot
install list, size baselines; regenerated capability-registry +
inventory manifest. ship:post wired into ship.md (wire-on-demand).

HELD on #1196: this capability also declares hooks at discuss:pre and
discuss:post, which are structurally un-wireable until the host-loop
conformance model covers the discuss phase (discuss-phase.md is not in
HOST_LOOP_FILES). The phase6-capstone-conformance gate therefore fails
on exactly those two orphaned points by design — see #1196. Once #1196
lands, rebase onto next and the gate goes green with no further change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#956): backfill changeset PR number (#1201)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 02:09:10 -04:00
Tom Boucher
b1e8a74708 fix(#1196): wire discuss loop step for capability hooks (#1199)
* fix(#1196): wire discuss loop step for capability hooks

discuss was contract-declared (gsd:loop-host marker, in POINT_ORDER and
LOOP_HOST_CONTRACT) but structurally unwireable: discuss-phase.md had no
`loop render-hooks` dispatch and was absent from the conformance gate's
HOST_LOOP_FILES, so capabilities could never wire discuss:pre/discuss:post.

- discuss-phase.md: add minimal discuss:pre (before analyze_phase) and
  discuss:post (after write_context) render-hooks dispatch steps that
  delegate consumption to a new shared reference (kept under the 32KB
  #2551 budget; no inline subagent dispatch token).
- references/loop-hook-dispatch.md: new canonical, point-agnostic contract
  for consuming `loop render-hooks --raw` activeHooks (contribution/step/
  gate) — single source for hook consumption across host loops.
- gen-loop-host-contract.cjs: derive HOST_LOOP_FILES from STEP_WORKFLOWS and
  export scanWiredPoints()/getWiredLoopPoints() (throws on a missing host
  file) — one source of truth for the host-loop file + wired-point set.
- phase6-capstone-conformance.test.cjs: consume the derived HOST_LOOP_FILES
  and shared scanWiredPoints (was a hand-maintained duplicate omitting
  discuss-phase.md + a duplicated regex).
- gen-capability-registry.cjs: add validateHooksWired() gen-time guard that
  rejects a capability hook declared at a valid-but-unwired loop point, with
  a clear remediation message — failure now surfaces at gen --check/--write
  time instead of deep in the full conformance suite.
- tests (capability-registry.test.cjs): regression + anti-pattern parity
  guards (every loop-host marker is in STEP_WORKFLOWS/HOST_LOOP_FILES;
  POINT_ORDER === flattened LOOP_HOST_CONTRACT) so no step can drift into
  the discuss-class gap again.
- docs/INVENTORY*: register the new reference.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1196): backfill changeset PR number (#1199)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 01:39:54 -04:00
Tom Boucher
b10e56818b feat(#1169): complete ADR-857 phase 6 — migrate features to Capabilities, revive dead gates, harden conformance gate (#1183)
* test(#1168): make phase-6 gate un-gameable — reject empty stubs + require loop shrink

The migration assertion previously checked only role==feature, so a registration-only stub (empty hooks, logic left inline) would turn the gate green while phase 6 stayed incomplete — the exact false-completion pattern this gate exists to prevent. Strengthen it: each ADR-named feature must OWN its behavior (>=1 hook, or a command family); and plan-phase.md/execute-phase.md must shrink strictly below their frozen pre-phase-6 sizes (94519/93166 LF bytes), which also defeats double-run gaming (declare a hook but keep the inline block -> file does not shrink -> red).

Gate now 5 pass / 4 fail (orphaned execute:wave:post, empty/unregistered features, config-key leaks, no shrink). Green is now reachable only by REAL migration. Refs #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate gap-analysis to a Capability (plan:post gate)

First real ADR-857 phase-6 migration (pattern-defining tracer). gap-analysis moves from an inline post_planning_gaps branch in plan-phase.md to a real plan:post gate Capability:

- capabilities/gap-analysis/capability.json: role:feature, plan:post gate (when=workflow.post_planning_gaps, blocking:false advisory), OWNS workflow.post_planning_gaps (federated out of central schema). - plan-phase.md: inline config-get + gsd_run gap-analysis block replaced with a plan:post render-hooks call site dispatching the gate; file shrinks 94519->93279. - src/check-command-router.cts: cmdGapAnalysisPlanPost runs the real gap analysis via gap-checker. - post_planning_gaps removed from central manifest; resolves via federated config (default true preserved). - tests/post-planning-gaps-2493: re-pointed to assert capability ownership.

Verified: gate 5 pass / 4 fail (gap-analysis cleared from migration, plan:post-orphan, config-leak, and plan-phase shrink checks); loadConfig still returns post_planning_gaps=true; check command runs real analysis; 392/392 in the config/registry/federation/router net. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate profile-pipeline to a command-family Capability

ADR-857 Decision 7: profile-pipeline becomes a command-family Capability (like audit/intel/graphify). capabilities/profile-pipeline/capability.json declares an 8-command family (scan-sessions, extract-messages, profile-sample, write-profile, profile-questionnaire, generate-dev-preferences, generate-claude-profile, generate-claude-md) backed by a new gsd-core/bin/lib/profile-pipeline-command-router.cjs; the inline case arms are removed from gsd-tools.cjs. Owns profile-pipeline.enabled (federated).

Verified: registry shows role:feature with commands.length=8; scan-sessions/profile-sample run live via the family; gate cleared profile-pipeline from the empty-stub failure (only tdd/schema-gate/drift remain); 296/296 registry+inventory+gsd-tools tests; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1167): wire execute:wave:post + implement ui.safety-gate check

Revives the second dead gate from #1167: ui.gates@execute:wave:post was declared but never dispatched AND its check.query (ui.safety-gate) was unimplemented. Adds the per-wave execute:wave:post render-hooks call site in execute-phase.md (fires after each wave's merge/cleanup, before the next forks) and implements cmdUiSafetyGate (frontend + UI-SPEC aware, mirrors cmdUiPlanGate) in check-command-router. +17 regression tests.

Verified: phase-6 orphaned-points conformance test now PASSES (gate 6 pass / 3 fail); ui-safety-gate routable in dot+hyphen forms; check-ui-safety-gate 17/17, check-ui-plan-gate 18/18; lint 0 errors. Refs #1167, #1168.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate drift (schema + codebase) to execute:wave:post gates

Removes the inline schema_drift_gate + codebase_drift_gate steps (77 lines) from execute-phase.md; drift becomes a Capability with two execute:wave:post gates (verify.schema-drift blocking, verify.codebase-drift advisory) dispatched via the per-wave render-hooks call site. check-command-router routes verify.schema-drift / verify.codebase-drift to the real detectors. Federates workflow.drift_threshold / drift_action / schema_drift_gate out of central.

Also fixes the execute:wave:post dispatch prose to run NON-blocking (advisory) gates too — the prior version only ran blocking gates, which would have silently dropped the codebase-drift advisory after its inline step was removed. Behavior preserved.

Verified: gate 7 pass / 2 fail (drift cleared from stub + config-leak; execute-phase.md 92297 < 93166 frozen -> shrink passes); both drift checks run real detection; loadConfig defaults preserved (threshold=3, action=warn, gate=true); drift-detection 56/56 + schema-drift 34/34; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate tdd to a Capability (plan:pre contribution + execute:post gate)

tdd becomes a real Capability: a plan:pre contribution injects the <tdd_mode_active> planner guidance (rendered from PLAN_PRE_HOOKS_JSON like security's contribution), and an execute:post gate (tdd.review-checkpoint, advisory) runs the real end-of-phase RED/GREEN review via a new check-command handler. Inline tdd_mode reads + the inline planner block + the tdd_review_checkpoint step are removed; workflow.tdd_mode is federated out of central. The MVP+TDD per-task RED-commit gate is preserved — TDD_MODE is now derived from the execute:post hooks (capId==tdd active), not an inline config-get.

BEHAVIOR CHANGE (documented, not silent): the --tdd CLI flag now persists workflow.tdd_mode=true via config-set instead of being per-invocation. Rationale: tdd is now a config-toggled Capability, and env vars do not persist across the workflow's separate bash blocks (config does), so an ephemeral override isn't cleanly achievable; --tdd therefore enables the tdd capability, consistent with how all capabilities are toggled.

Verified: gate 7 pass / 2 fail (tdd cleared from stub + config-leak; plan-phase + execute-phase both < frozen sizes); contribution injection + execute:post gate dispatch wired; MVP+TDD gate preserved; tdd.review-checkpoint runs real review; full unit suite 556/0; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate schema-gate to a plan:pre contribution Capability

The plan-time schema-push detection (former plan-phase.md §5.7) becomes a schema-gate Capability: a plan:pre contribution (into:planner, when:workflow.schema_push_detection) whose fragment carries the full ORM-detection + [BLOCKING] schema-push-task injection logic, rendered into the planner via the existing plan:pre render-hooks dispatch. The inline §5.7 block is removed (plan-phase.md 94519->90445). workflow.schema_push_detection is a new capability-owned (federated) key, default true. (The execute-side schema-drift gate was migrated separately into the drift capability.)

Verified: registry inlines the fragment (len 2704) so it is actually delivered at plan:pre; gate 8 pass / 1 fail — all 5 ADR-named features now real Capabilities, only the config-leak test remains (intel/security, next unit). Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): close the 3 capability config-key leaks — phase-6 gate now GREEN

Removes the last inline config-get reads of capability-owned keys from plan-phase.md. security_asvs_level/security_block_on now flow through the security plan:pre contribution via a new loop-resolver configValues mechanism (resolves declared config keys with the same 4-level precedence as activation and attaches them to the rendered hook); the §5.55 banner reads them from PLAN_PRE_HOOKS_JSON. intel.enabled becomes a real intel plan:pre step (ref.command: intel api-surface) dispatched via render-hooks; the inline intel branch is gone. gen-capability-registry now validates ref.command as a third dispatch shape.

Verified: phase-6 capstone conformance gate is FULLY GREEN (9/0); 3 leaks gone (grep=0); security configValues resolve to {2,medium}/default {1,high}; intel step present only when enabled; loop-render-hooks 62/0, capability-registry 287/0, capability-state/federated-config 113/0; lint 0 errors. Closes the migration half of #1169. Refs #1139, #1167, #1168.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): address adversarial review — restore schema-drift block, generic planner injection, uniform gate contract

Adversarial review caught 2 real regressions the green gate missed: (1) schema-drift no longer blocked — the execute:wave:post dispatch read GATE_RESULT.block but verify.schema-drift emitted drift_detected/blocking, and onError:skip wrongly bypassed positive blocks; (2) only tdd's plan:pre contribution was injected into the planner, dropping schema-gate's schema-push detection and security's threat-model guidance.

Fixes: (A) every gate check returns a uniform boolean 'block' under --raw (the dispatch form), with advisory gates (tdd/gap) carrying their report in 'message'; (B) gate-dispatch contract corrected at all sites — onError governs command errors only, a blocking gate's positive block always halts; (C) generic planner injection of all plan:pre contributions where into=='planner' (tdd + schema-gate + security incl configValues); (D) two new conformance assertions: planner contributions injected generically + every gate check.query returns boolean block under --raw.

Verified: gate 11/11; all 6 gate checks return boolean block under --raw; full suite 595/0; lint 0 errors. Refs #1167, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): restore MVP+TDD end-of-phase blocking escalation (2nd adversarial pass)

The migrated tdd execute:post gate is statically blocking:false, but the contract (references/execute-mvp-tdd.md + CONTEXT.md) requires the end-of-phase TDD review to ESCALATE from advisory to blocking when MVP_MODE && TDD_MODE && a TDD plan misses a RED/GREEN commit. The migration prose had downgraded this to a 'strong advisory recommendation' — silent loss of the blocking escalation. Restore it: the tdd-gate dispatch now refuses to mark the phase complete (Phase blocked message) under MVP+TDD when GATE_RESULT.block is true; advisory otherwise.

Also strengthen tests/execute-mvp-tdd-gate.test.cjs: hasBlockingEscalation previously matched any line with 'blocking'+'mvp+tdd' (so 'advisory (blocking: false) ... under MVP+TDD' was a false green); now it requires the real refusal semantics ('refuse to mark the phase complete' / 'phase blocked'). Caught by 2nd adversarial review pass.

Verified: execute-phase.md 92702 < 93166 frozen; mvp-tdd-gate + phase-6 gate 19/0; full suite green; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): restore MVP+TDD proceed-block, codebase auto-remap, schema skip-flag (3rd adversarial pass)

3rd adversarial pass found 4 more silent regressions: (1) the tdd MVP+TDD 'refuse to mark complete' was nullified by a downstream 'ALWAYS proceed regardless of gate results' line — proceed is now conditional (stops on an active MVP+TDD block); (2) the test now asserts the proceed is NOT an unconditional override; (3) codebase-drift auto-remap (spawn gsd-codebase-mapper when drift_action=auto-remap) was dropped — the execute:wave:post advisory dispatch now consumes spawn_mapper/directive; (4) GSD_SKIP_SCHEMA_CHECK bypass was lost from the gate path — cmdVerifySchemaDrift now honors the env var (block:false when set).

Verified: no unconditional proceed; GSD_SKIP_SCHEMA_CHECK=true -> block:false; gate 11/11 + mvp-tdd 9/9; full suite 569/0; lint 0; execute-phase.md 93109 < 93166. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): init.cts reads federated config keys from nested path (4th adversarial pass)

Config federation moved tdd_mode/research/nyquist_validation from flat config.<key> to nested config.workflow.<key>, but src/init.cts still read them flat — so init.plan-phase/init.execute-phase emitted tdd_mode:false / research_enabled:undefined / nyquist:undefined regardless of config (a public command-contract regression; the migrated loops use render-hooks so enforcement was unaffected). Read via config.workflow (type-safe Record cast). Now init reflects the same resolved values + federated defaults (research/nyquist default true) as the render-hooks path.

Verified: build clean; init.plan-phase emits tdd_mode:true/research:false/nyquist:false for set config, defaults true for empty; full suite 591/0; lint 0. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1169): add changeset for ADR-857 phase-6 completion (PR #1183)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): complete phase-6 migration fallout — restore TEXT_MODE, fix registry .claude leak, re-point stale workflow-contract tests

The capability migration left real regressions and stale consumer tests that
the per-module unit suite missed but the full cross-platform suite caught (27
failing tests):

Real source regressions (fixed):
- execute-phase.md lost its AskUserQuestion TEXT_MODE plain-text fallback when
  the inline schema_drift_gate step was removed — non-Claude runtimes would
  stall. Restored, and the execute:post gate-dispatch prose de-duplicated to
  cite the execute:wave:post contract (loop body shrinks below the frozen
  pre-phase-6 ceiling while keeping every onError/blocking nuance).
- capabilities/tdd inline fragment hardcoded `@~/.claude/gsd-core/references/tdd.md`,
  baked verbatim into the committed capability-registry.cjs and leaked the
  install path on 11 non-Claude runtimes (registry .cjs is copied, not
  path-converted). Made the fragment path-free; regenerated the registry. The
  phase-6 conformance gate now guards this (no ~/.claude install path in any
  capability source or the generated registry).
- plan-phase.md: removed a §5.7 stub re-added in error and routed Branch 2 to
  step 6 (schema-gate is a plan:pre capability, §5.7 is gone).

Stale workflow-contract tests re-pointed to the capability dispatch they now
must assert (behavior verified preserved in source first, assertions kept
equal-or-stronger): bug-621 + bug-2851 (gap-analysis via gsd_run render-hooks
plan:post + registry binding), feat-2527 (tdd_mode federated out of central),
phase6-planning + plan-phase-ui-redirect (§5.6 bounded by ## 6.),
plan-phase-drift-guard (intel when:intel.enabled skip branch).

profile-pipeline-command-router.cjs un-ignored from eslint (hand-written, no
TS source) + stale disable comments removed. Size baseline regenerated.

Verified: full suite 15140 tests / 0 fail; lint 0 errors; conformance gate green
legitimately. Refs #1139, #1167, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1169): add ADR-857 E2E content-test coverage for the 12 loop points + capability deliverables

Grounds the capability engine in behavioral E2E tests (drive the real
render-hooks/check CLI + the real registry, assert typed result content — no
source-grep), structured around what ADR-857 says to deliver. 207 tests; each
genuineness-checked (flip the expectation, confirm it fails).

Per-loop-point dispatch (7 files): empty-point negative-space across the 6
no-hook points; verify:post 3-step resolution+ordering+onError; plan:pre
contribution/configValues + ui.plan-gate + intel; plan:post gap-analysis;
execute:wave:post drift+ui gates via the check route (schema-drift block/skip,
codebase-drift threshold BVA, auto-remap); execute:post tdd.review-checkpoint
RED/GREEN; ship:pre security gate resolution + frontmatter-get predicate pieces.

ADR-deliverable coverage (4 files): predicate boundary held (edge/prohibition
probes stay core, not off-by-default Feature Capabilities — phase-6 exception);
core loop runs with zero capabilities (all 12 points empty, init bundles
resolve); contribution merge (multiple ordered <contribution from=> blocks);
federated-config key removal on uninstall.

federated-config allowlisted for its 3-file split (unit + integration +
lifecycle). Refs #1139, #1167, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): remove dead drifted converter dups + address adversarial review

Lint cleanup (root-caused, not waved off): src/runtime-artifact-conversion.cts
carried 11 agent-converter functions (+5 orphaned consts/helpers) that were
never exported, never called, and had silently DRIFTED from the live
hand-authored copies in bin/install.js (one even referenced an undefined
`claudeToCopilotTools`). Deleted the dead duplicates; install.js's live copies
are untouched (it never imported these). Lint now 0 errors / 0 warnings.

Adversarial-review (Codex) findings fixed:
- HIGH: execute-phase.md TDD_MODE used `jq ... || echo false`, silently
  disabling the MVP+TDD blocking gate on jq-less runtimes. Reverted to the
  `node -e` form (node is guaranteed; matches the file's other node-e usages) so
  a missing optional tool can no longer fail-open a blocking safety path.
- MEDIUM: federated-config-key-removal orphan-key test was vacuous (it skipped
  the orphan assertion). Now asserts the removed capability's key is genuinely
  not surfaced/validated after uninstall.
- LOW: phase-6 conformance leak regex broadened to catch absolute-home and
  Windows-backslash `.claude/(gsd-core|commands|agents|hooks)` paths, not only
  `~`/`$HOME` forward-slash forms.
- LOW: bug-2851 plan:post dispatch assertion now requires `--raw` (matched its
  stated contract).
- nit: plan-pre intel-step test duplicate assertion replaced with a distinct
  structured-output check.

Size baseline regenerated (execute-phase.md 93089 < 93166 frozen). Refs #1167, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1169): make runtime-homes-descriptor-drive titles environment-independent

The descriptor-equivalence test embedded the absolute golden config path
(`os.homedir()`-derived) directly in each `test(...)` title, so titles differed
between macOS (`/Users/x/.claude`) and Docker (`/home/gsdtest/.claude`). Every
test PASSES on both platforms (15885/0 leaf tests each), but gsd-test-summary
compares results by title and reported 29+29 false "only in Mac / only in
Docker" discrepancies for tests that actually pass everywhere.

Move the golden path out of the title and into the assertion message (still
shown on failure); titles are now byte-identical across platforms so the
cross-platform comparator matches them. No assertion logic or golden values
changed. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): derive TDD_MODE via gsd_run --active-cap, not node -e (fix prompt-injection CI gate)

The prior fix reverted execute-phase.md:181 from jq to `node -e` to close a
Codex HIGH (jq||echo-false silently disabling the MVP+TDD blocking gate on
jq-less runtimes) — but the CI prompt-injection scanner BLOCKS new `node -e` in
workflow markdown (inline code-exec = injection vector), turning the security
gate red. Both forms were wrong: node -e fails the scanner; jq fail-opens a
blocking safety gate; `config-get workflow.tdd_mode` is forbidden by the
conformance leak gate (tdd_mode is capability-owned).

Correct fix (what Codex recommended): a gsd_run-native boolean. Add an
`--active-cap <capId>` flag to `loop render-hooks <point>` that resolves hooks
the normal way and prints exactly `true`/`false` for whether a capId is active
— scanner-safe (canonical launcher, no inline code), node-reliable (no optional
jq to fail-open), and leak-free (render-hooks resolution, not config-get).
execute-phase.md:181 now `TDD_MODE=$(gsd_run loop render-hooks execute:post
--active-cap tdd)`. +5 behavioral tests for the flag.

Verified: prompt-injection-scan --diff origin/next → 0 findings; conformance
gate 13/13 (execute-phase.md 92934 < 93166); execute-mvp-tdd + tdd-mode +
loop-render-hooks 87/0; lint 0/0. Refs #1167, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 21:07:55 -04:00
Tom Boucher
70f35ebc7a feat(#1169): wire ship:pre security gate via render-hooks (ADR-857 phase 6)
Part of the ADR-857 phase-6 migration (#1169, epic #857) — not a shipped bug fix. The capability system (registry, render-hooks, gates, capability-state) lives only on next / 1.5.0-rc; npm latest is 1.4.5 and contains none of it, so no released user can hit this.

The security capability declares a blocking gates@ship:pre hook (when=workflow.security_enforcement, predicate SECURITY.md.threats_open==0), but ship.md never called `loop render-hooks ship:pre` and had no inline fallback, so the declared gate never fired. Wire it into ship.md preflight_checks via the render-hooks idiom — fail-closed: the ship blocks unless threats_open is exactly 0.

Add a phase-6 conformance assertion that every declared hook point has a render-hooks call site; execute:wave:post (ui_safety_gate) stays in KNOWN_UNWIRED pending its unimplemented ui.safety-gate check.

Part of #1169. Refs #857.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 11:09:21 -04:00
Tom Boucher
f116b76128 test(#1139): add phase 6 capstone conformance gate (#1158) 2026-06-13 02:12:16 -04:00
Tom Boucher
1b880fefd7 feat(#1137): migrate review verification hooks to capabilities (#1147)
* feat(#1137): migrate review verification hooks to capabilities

* chore(#1147): add changeset
2026-06-12 21:24:09 -04:00
Tom Boucher
44024aa535 fix(#1133): honor model_policy on the claude runtime (forward-port to next) (#1144)
Forward-port of the #1133 hotfix (commit 22f237a5, 1.4.5 hotfix line) onto
next. The hotfix was authored against src/core.cts (v1.4.4); on next the
resolver logic lives in src/model-resolver.cts (ADR-457 extraction, #888),
so the patch is re-applied there rather than cherry-picked.

resolveModelInternal step 2.5 now honors model_policy on the claude runtime:
the policy-resolved full model ID is mapped back to a Claude Code agent alias
via CLAUDE_POLICY_ID_TO_ALIAS (reverse of MODEL_ALIAS_MAP + claude-fable-5 ->
fable). Bare aliases (opus/sonnet/haiku/fable) pass through; an ID with no
Claude alias warns once to stderr (deduped by agentType::policyModel::tier)
and falls back to the configured tier alias. Non-claude runtimes return full
IDs verbatim (unchanged). resolveModelForTier is intentionally unchanged.

The warn-dedupe cache lives in model-resolver.cts; core.cts composes the
exported _resetRuntimeWarningCacheForTests to clear both that cache and the
config-loader warning cache (config-loader cannot import model-resolver --
circular dependency).

Ports the 6 #1133 tests (rewriting the old claude-no-op test that asserted
the bug) plus one added test covering the MODEL_ALIAS_MAP reverse-map path.

Forward-port of #1133

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 19:56:54 -04:00
Tom Boucher
4ab5c7b3f2 feat(#1135): migrate planning hooks to capabilities (#1141)
* feat(#1135): migrate planning hooks to capabilities

* chore(#1135): add phase 6 planning capabilities changeset

* fix(#1135): satisfy lint for agent hook rendering
2026-06-12 19:51:54 -04:00
Tom Boucher
827011b865 fix(#1098): guard generate-claude-md against clobbering hand-crafted files; redirect default to .claude/CLAUDE.md (#1118)
/gsd-new-project wrote a repo-root CLAUDE.md full of broad project docs,
overwriting/diluting a hand-crafted instruction file. --force was parsed but
silently dropped, and nothing guarded an existing non-GSD file.

- Guard: an existing instruction file with no `<!-- GSD:<section>-start` markers
  (hand-crafted) is left untouched; report action:"skipped". --force (now wired
  through CmdGenerateClaudeMdOptions) overwrites intentionally. The marker check
  uses /<!-- GSD:[a-z]+-start/ so a file merely documenting GSD syntax is safe.
- Redirect: the Claude-family default output is now ./.claude/CLAUDE.md (a valid
  auto-loaded project-memory location) instead of repo-root ./CLAUDE.md. Aligned
  across the handler default, config-defaults.manifest.json, buildNewProjectConfig,
  the config template, new-project.md, and cmdGenerateClaudeProfile; advisory
  read-CLAUDE.md hints in plan-phase/quick/profile-user updated. Codex still
  writes AGENTS.md.

Closes #1098

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 16:03:51 -04:00
Tom Boucher
fc37ae0c4f fix(#1115): capability-probe codex hook-trust bypass flag in /gsd:review; fail loud on empty output (#1122)
On codex-cli < 0.137.0 the review.md `codex exec` invocation passed
--dangerously-bypass-hook-trust (added in 0.137) unconditionally and discarded
stderr, so codex exited "unexpected argument" before reading the prompt and the
empty output file was treated as a completed review — a silent degraded review.

- Capability-probe the flag (`codex exec --help | grep`) and apply it via
  $CODEX_BYPASS_FLAG only when supported (works fine without it on older CLIs).
- Capture codex stderr to a .err file instead of /dev/null, and replace an empty
  output with a diagnostic so a broken reviewer is surfaced (mirrors Cursor/OpenCode).
- Update enh-773 enforcement test to require the capability gate + fail-loud guard
  instead of the unconditional flag.

Closes #1115

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 15:55:17 -04:00
Tom Boucher
1c073c1c81 fix(#1107): progress consults verification.status before reporting a phase complete (#1116)
/gsd-progress derived phase completeness from plan/summary counts only and never
consulted the verification.status query (the #651 seam), so a phase whose
VERIFICATION.md ended human_needed or gaps_found was reported complete and
routing skipped to the next phase. Add Step 1.7 (consult verification.status for
the current phase) and routing rows that send gaps_found to plan-phase --gaps
(Route V.gaps) and human_needed to verify-work (Route V.human) before the generic
complete row. passed/missing/unknown still route as complete so unverified phases
are not falsely blocked.

Closes #1107

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 15:55:01 -04:00