Commit Graph

174 Commits

Author SHA1 Message Date
Tom Boucher
cf2e66b39e feat(#1708): typed documentation-sourced #853 dispatch-flatten (ADR-1239 Phase B) (#1719)
* feat(#1708): typed documentation-sourced #853 dispatch-flatten

Graduate the #853 orchestrator-backgrounding decision from a scattered RUNTIME==='codex' prose check to a typed, documentation-sourced engine decision. Adds a backgroundDispatch dispatch sub-axis (sourced per host: codex+cursor documented true, 9 documented false, 5 undocumented), shouldFlattenDispatch(dispatch) (inline UNLESS background && backgroundDispatch, fail-closed), and a gsd_run query dispatch-should-flatten the plan/execute workflows call. Cursor is newly background-eligible per its docs (inline->background) — a documentation-justified behavior change. No RUNTIME-name residue for this decision.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1708): backgroundDispatch citations in matrix + CONTEXT note

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1708): address review findings on typed dispatch-flatten

Code/adversarial review: convert the manager.md/autonomous.md Compound Action preamble from hardcoded 'On Codex' to FLATTEN-based branching (the handlers already use the query; the preamble contradicted them and was wrong for cursor); make shouldFlattenDispatch null-safe + type-honest (accepts raw 'undocumented' registry values); make backgroundDispatch a required descriptor field (matching its siblings, all 16 carry it); strengthen the config.runtime behavioral test; update the bug-853 prose-pin test + comment. Security review clean; Codex confirmed no fail-open.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1708): backfill backgroundDispatch in role:runtime test fixtures

Making backgroundDispatch a required descriptor field broke role:runtime fixtures in capability-manifest-version/capability-registry/host-integration-descriptors tests that build a dispatch object without it (caught by full gsd-test, not scoped npm test). Backfill backgroundDispatch:false into the well-formed fixtures; the deliberately-malformed 'required-field' test fixture is left malformed by design.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1708): update fix-1521 dispatch-gating assertion to the FLATTEN gate

fix-1521 pinned the codex-specific run_in_background prose that #1708 graduated to the typed dispatch-should-flatten/FLATTEN gate. Update its assertions to verify FLATTEN=false gating (not a runtime name) + that the old RUNTIME===codex gate is gone. Caught by full gsd-test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1708): add changeset for typed dispatch-flatten

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1708): remove stray temp PR-body file

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1708): add issue ref to bug-853 allow-test-rule annotations

ADR-456 requires every allow-test-rule exemption to carry a see #NNN reference; the source-text-is-the-product annotations added when migrating the prose assertions lacked it (lint-tests CI gate). Add (see #1708).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 15:37:33 -04:00
Tom Boucher
fb5f89db10 feat(#1704): destSubpath write-confinement (ADR-1239 Phase B) (#1706)
* feat(#1679): confine install writes within configHome

ADR-1239 Phase B write-confinement: a pure assertDestWithinConfigHome(configDir, destSubpath) rejects a destSubpath that escapes configHome (path traversal / NUL byte) at plan-build time on BOTH the install and uninstall plan paths; surface.applySurface and installOpencodeFamilySkills route through it, and _copyStaged carries a defense-in-depth containment check. Security-load-bearing for the Phase C third-party-descriptor loader.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1704): add changeset for destSubpath write-confinement

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1704): fix windows path-portability in confinement test

The N1 'accepts a true child subpath' assertion compared against path.join (no drive resolution) while the helper uses path.resolve — on Windows that mismatches the C: drive prefix. Compute the expected via path.resolve to mirror the helper. Windows-CI-only failure (local gsd-test is Mac+Linux).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 13:20:25 -04:00
Tom Boucher
30d4b85de5 feat(#1684): negotiated host-integration interface (ADR-1239 Phase A) (#1690)
* feat(#1684): add negotiated host-integration interface module

ADR-1239 Phase A: a pure, additive, no-I/O module exposing PROTOCOL_VERSION, the 8-axis HOST_INTEGRATION_AXES closed vocabulary, the UNDOCUMENTED fail-closed sentinel, negotiateHostCapabilities (effective subset of host-declared and engine-known), a typed degradation ladder, and host-capability profiles.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1684): validate and document host-integration axes (16 runtimes)

Extend validateRuntimeBody to validate the 8 hostIntegration axes (closed enums + undocumented sentinel + dispatch struct + reserved-key guards) and the widened runtime vocabulary; author a documentation-sourced hostIntegration block in all 16 runtime descriptors; regenerate the registry. Every per-CLI value is documented (cited) or the explicit undocumented sentinel.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1684): add host-integration capability matrix and adr amendment

New per-CLI, per-axis citation reference (value/source/evidence for all 16 CLIs); ADR-1239 Phase-A-implemented amendment; CONTEXT.md glossary seam entry.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1684): harden dispatch negotiation edge cases

Code-review hardening: treat NaN/Infinity maxDepth as missing (fail-closed, +warning); reset nested/background when namedDispatch collapses to false (struct consistency); SAFE_DEFAULTS dispatch floor to read-only; warn on non-finite protocolVersion; symmetric undocumented warnings for dispatch fields. Pure module — no consumers; behaviour fail-closed throughout.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1684): register host-integration.cjs in lint-ignore and inventory

New tsc-generated bin/lib artifact: add to the eslint ignore list (ADR-457 — lint the .cts source), regenerate docs/INVENTORY-MANIFEST.json, and add the docs/INVENTORY.md CLI-modules row. Fixes the 3 gsd-test failures (551-eslint-bin-lib-coverage x2 + inventory-manifest-sync).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1684): add changeset fragment for host-integration interface

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1684): add how-to for sourcing a host's integration axes

Diataxis how-to guide for adding/updating a host's runtime.hostIntegration axes from authoritative docs, the undocumented-sentinel rule, validation, and extending the closed vocabulary. Completes the Step-5 doc quadrants (reference + explanation + how-to). Indexed in docs/README.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 11:33:06 -04:00
Tom Boucher
2b215b4163 feat(#1688): warn on stale model bake for static-frontmatter runtimes (#1692)
* docs(#1650): fix stale opencode install-path claim in core settings

* feat(#1688): warn on stale model bake for static-frontmatter runtimes

* chore(#1688): backfill changeset pr field with real PR number

* test(#1688): make resolveAgentDir assertions use path.join for windows

* docs(#1688): codify windows path-literal-in-assert anti-pattern + align test
2026-06-25 09:12:40 -04:00
github-actions[bot]
548a986ffc chore: sync next package version to 1.6.0 2026-06-24 23:09:44 +00:00
Behruz Nassre Esfahani
47906b052d fix(#1520): randomize mktemp temp paths on BSD/macOS (XXXXXX must be path-final) (#1550)
* fix(#1520): randomize mktemp temp paths on BSD/macOS (XXXXXX must be path-final)

BSD/macOS mktemp only substitutes the XXXXXX template when it is the final
path component. Templates like `...-XXXXXX.json` / `gsd-pr-body.XXXXXX.md`
return a LITERAL `XXXXXX` path (no randomization) on macOS, so concurrent
workflow runs collide on the same temp manifest/body file — one run can
overwrite or consume another's. Reproduced on macOS: the second call to the
suffixed template fails `mkstemp: File exists`.

Fix: use a suffixless `XXXXXX` template (so it IS the final component), then
rename to add the intended extension — portable across BSD + GNU userlands,
no GNU-only `--suffix` flag. Empty-file-then-write semantics are preserved at
every site.

Affected workflow temp files:
- execute-phase.md: gsd-worktree-wave-*.json (wave worktree manifest)
- quick.md:         gsd-quick-worktree-*.json
- spec-phase.md:    edge-probe-reqs-*.json
- ship.md:          gsd-pr-body-*.md
- profile-user.md:  gsd-profile-answers-*.json, gsd-profile-analysis-*.json

The execute-phase.md edit uses a compact intermediate var + trailing comment
to stay under the ADR-857 phase-6 size ceiling (93166); regenerated the
workflow size baseline accordingly. Validated on macOS: 20 concurrent calls
yield 20 unique randomized paths.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1520): add changeset fragment (Fixed)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1520): add fail-first workflow-prose guard for mktemp XXXXXX suffix

Repo-wide scan of gsd-core/workflows/**/*.md that fails on any mktemp
template whose XXXXXX run is followed by a filename suffix (the BSD/macOS
non-randomizing form). Fails on the six pre-fix instances and passes on
the fix, and locks the copy-paste-prone idiom out of future workflows.
Mirrors the bug-637 hardcoded-$HOME workflow guard.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1520): rename regression test to fix- prefix (regression-test-names lint)

New tests/bug-NNNN-*.test.cjs files are banned by the lint-regression-test-names
ratchet; use the fix- prefix (matches the fix-1445 precedent).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1520): add issue ref to allow-test-rule exemption (ADR-456 lint)

lint-allow-test-rule-refs requires every new `allow-test-rule:` comment to
carry a #NNN reference (don't allowlist). Add (#1520) to the source-text
exemption.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1520): abort touched mktemp chains on failure (|| exit 1)

Per review: the VAR=$(mktemp …) && mv … && VAR=… chains dropped the issue's
suggested failure guard. If mktemp fails, $VAR is empty and the subsequent
mv/write lands on an unintended relative path. Add `|| exit 1` to all six
touched chains so a mktemp failure aborts the snippet. Regenerated the
workflow size baseline for the slightly longer lines.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1520): rebase onto next — regen size baseline + describe rename

Resolve the workflow-size-baseline.json conflict from next advancing by
regenerating from the current workflow sizes. Also rename the test describe
from `bug #1520` to `#1520` (the file uses the fix- prefix) per review nit.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1520): harmonize profile-user mktemp to ${TMPDIR:-/tmp} (review nit)

The two profile-user.md temp sites this PR already rewrites kept a hardcoded
/tmp while the four sibling workflows use ${TMPDIR:-/tmp}. Harmonize for
consistency and macOS-correctness (some sandboxes have no writable /tmp).
Regenerated the workflow size baseline.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1520): regen size baseline after rebase onto next

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 17:27:45 -04:00
Alex V.
1a46109b97 enhance(#1579): deterministic gsd-tools query eval.score verb (#1583)
* feat(#1579): deterministic gsd-tools query eval.score verb

Split C of #1573 (pure code, lowest risk). Adds an eval.score query verb
(coverage*0.6 + infra*0.4; bands 80/60/40) mirroring the verify.* chain;
gsd-eval-auditor consumes it instead of doing weighted arithmetic in-prompt.
Non-breaking — additive only.

arXiv: 2601.15130 (Plausibility Trap/DPDM), 2507.10281 (Table Agent), 2508.15754 (TIR).

* fix(#1579): address review — domain guard, property test, glossary, SKIP_ROOT, inventory/baseline

- C3 input-domain: reject out-of-domain eval.score (require 0<=covered<=total; was emitting overall_score>100 / negatives)
- C1 property test: add tests/eval.property.test.cjs (fast-check) — determinism, band monotonicity, [0,100] bounds, never-throws
- C2 glossary: CONTEXT.md "Eval Scoring Module" entry (source-of-truth path + interface)
- C4: add `eval` to SKIP_ROOT_RESOLUTION (pure arithmetic; no .planning/ access)
- inventory: register generated eval.cjs/eval-command-router.cjs (INVENTORY-MANIFEST.json + INVENTORY.md rows)
- size: regen agent-size baseline for gsd-eval-auditor (reused gsd_run shim + eval.score step)
- eslint: ignore generated eval*.cjs (ADR-457 bin/lib migration coverage)

* fix(#1579): register eval family in alias-drift gates

Add EVAL_COMMAND_ALIASES/EVAL_SUBCOMMANDS to scripts/check-alias-drift.cjs
families and to familyArrayKeys in the manifest-coverage test, so the eval
family lands under the same drift guard as every sibling family
(state/verify/init/phase/phases/validate/roadmap). Addresses trek-e review.

check:alias-drift ok; feat-3251 coverage 9/9; eval suites 10/10.

* docs(#1579): use half-open verdict band ranges in CLI-TOOLS

overall_score is fractional and thresholds are >=80/>=60/>=40, so a score
in [79,80) is correctly NEEDS WORK despite the old '60-79' label. Relabel
bands as 60-<80 / 40-<60 / 0-<40 to match the code. Addresses trek-e nit.

* fix(#1579): validate eval.score CLI inputs

Reject unknown infra tokens and fractional counts, and pin the 80-point verdict boundary including rounding-before-banding behavior.
2026-06-24 17:16:13 -04:00
Alex V.
a63684c222 enhance(#1577): WebFetch/WebSearch injection isolation + opt-in blocking (#1585)
* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking

Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the
largest untrusted channel) in gsd-read-injection-scanner; shared
untrusted-input-boundary reference @-included by the 8 ingest agents
(randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring);
opt-in security.injection_blocking (default advisory — non-breaking).

arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth).

* fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized

- A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a
  circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already
  in the transcript. The prompt-level data/instruction boundary is the primary control.
- A2: registered security.injection_blocking in the config schema + defaults manifests (default
  false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads.
- A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention).
- A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale).
- A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input.
- Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) +
  drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as
  the noted pre-existing follow-up.

* fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate

The new reference quotes injection phrases ('ignore previous instructions',
'you are now…') as examples agents must NOT comply with, tripping the repo's
own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red
on HEAD). Allowlist it alongside the other security docs (security-model.md,
TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS
scanner test doesn't scan references/, so only the shell gate needed it.

Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15.

* fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer

trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so
no named web-ingress agent is uncovered, keeping the two justified additions
(gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10.
 - gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset.
 - gsd-assumptions-analyzer reads 5-15 codebase source files (external/source-
   document ingress per the boundary), though it has no web tools.
INGEST_AGENTS in the isolation test now asserts all 10; size baselines
regenerated (+60 bytes each, both well under the DEFAULT cap); changeset
reworded 8 -> 10.

Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39.

* docs(#1577): document security.injection_blocking + boundary seam

trek-e Major 2 + Minor:
 - docs/CONFIGURATION.md: add the top-level security.injection_blocking key to
   the Full Schema and a Security Settings subsection, distinguishing it from
   the workflow.security_* namespace; honest circuit-breaker-not-redactor
   framing matching ADR-1577 / security-model.
 - CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry.

Verified: lint:docs ok; config-field-docs + contributor-standards green.

* test(#1577): make read-injection property test git-text, not binary

trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as
degenerate-edge inputs. The NUL is what actually made git classify it binary
(git binary = NUL in first 8K). Replace both with text-safe escapes that keep
the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File
now diffs/blames line-by-line.

Verified: property test 2/2; no NUL/raw-noncharacter bytes remain.

* docs(#1577): align untrusted boundary docs

Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner.

* docs(#1577): align ADR ingest agent count

Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 17:07:23 -04:00
Tom Boucher
d101daff30 fix(#1516): expose adaptive model_profile in /gsd-new-project AI Models prompt (#1654)
* fix(#1516): expose adaptive model_profile in /gsd-new-project AI Models prompt

Both onboarding paths (Step 2a auto-mode + Step 5 interactive) enumerated only
4 profiles (Balanced/Quality/Budget/Inherit), omitting 'adaptive' even though the
model catalog (model-catalog.json profiles) and docs/CONFIGURATION.md register 5.
Mirrors the proven /gsd:settings two-question split (#3784): Q1 routes between
Adaptive/Standard-tier/Inherit; Q2 (conditional on Q1=Standard) picks
Quality/Balanced/Budget — keeping every AskUserQuestion within the 4-option cap.
Both config-new-project example payloads now list adaptive. Regression cases
folded into the owning tests/new-project-mvp-prompt.test.cjs (per the
lint-regression-test-names ban on new top-level bug-NNNN files): each AI Models
prompt makes adaptive reachable, all 5 profiles reachable, 4-option cap honored,
both example enums include adaptive, brace balance. Workflow size baseline bumped
(new-project.md 62324 -> 66138 bytes; still well under the XL hard cap).

* chore(#1516): backfill changeset pr ref to 1654
2026-06-24 14:46:44 -04:00
Jeremy McSpadden
77c7b4fc9d fix(#1522): enforce canonical verification before phase transition (#1548)
* fix: require fresh phase verification before transition

* no-mistakes(review): Fix canonical verification closeout gates

* no-mistakes(review): Fix verify-work frontmatter promotion command

* no-mistakes(review): Fix stale verification gates

* no-mistakes(review): Fix canonical verification routing gates

* no-mistakes(review): Fix verification dependency and runtime routing gates

* no-mistakes(review): Block stale verification bypasses

* fix: handle large init manager outputs in verification workflows

* chore: update changeset pr number

* fix(verify-work): use fresh verification.status for stale gate

The stale check after UAT used phase_completion.verification_status from
session-start INIT while human_needed promotion already queried fresh
verification.status. Align the stale gate with the canonical query so
mid-session verification refresh is not ignored.

* fix(init): skip roadmap-checked phases when selecting next_phase

Roadmap-only phases without a disk directory were still promoted to
next_phase when their checkbox was already checked. Exclude
checkboxComplete phases so progress routing does not point at work the
roadmap already marks done.

* fix: gaps_found not overridden by stale, transition uses canonical verification

- verification.cts: check gaps_found before stale so gap-closure routing
  is not masked by a newer summary mtime
- phase.cts: remove redundant findStaleVerificationSummary — readVerificationStatus
  already handles stale detection
- transition.md: replace raw grep on file content with verification.status query
  to avoid false-positive blocks from body text matching

* ci: retrigger tests after rebase

* fix(transition): replace gsd_run advisory check with awk frontmatter extraction

The runtime launcher is not defined until the update_roadmap_and_state step
bash block (~line 165). The early verify_completion block used gsd_run to
query verification.status, which violated the runtime-launcher-parity test:
'preamble appears AFTER the first gsd_run reference'.

Replace the gsd_run call with an awk-based frontmatter extractor that reads
only the status: field between the two --- fences. This avoids both the
preamble-ordering constraint and the original false-positive grep bug where
body text like 'previous_status: gaps_found' would match a full-text regex.

The phase.complete gate at update_roadmap_and_state is the canonical
enforcement point; this early check is advisory only.

Also update workflow-size-baseline.json for the updated transition.md size.

Fixes: runtime-launcher-parity test (B)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: re-check verification under planning lock in phase complete

Move readVerificationStatus into withPlanningLock so stale verification
cannot slip through when a SUMMARY.md is written between the gate and
the roadmap/state mutation. Return the blocked status from the lock
callback and emit the error after release to avoid leaving .lock behind.

* fix(transition): gate on canonical verification.status including stale

Replace awk frontmatter read with verification.status query so transition
blocks when summaries are newer than VERIFICATION.md, matching phase.complete
and other workflows (autonomous, progress, verify-work).

* Fix workflow verification gates for yolo transition and stale routing

Require VERIFY_STATUS passed before yolo/interactive transition advance.
Route stale verification recovery to verify-work, matching canonical projection.

* fix(transition): use verification.status query for stale-aware advisory check

The awk-based check read raw frontmatter status: passed, which misses the
stale case where summaries are newer than the VERIFICATION.md file even
though the frontmatter still says passed. The stale status is computed from
file modification times, not stored in frontmatter.

Move the preamble to the verify_completion bash block (the first block with
a gsd_run call) so gsd_run query verification.status can be used for the
advisory check. This gives the full readVerificationStatus logic including
mtime-based staleness detection, matching the enforcement gate at phase.complete.

Capture full JSON (VERIFY_JSON) so next_action can be included in the
advisory output alongside the status.

Also update workflow-size-baseline.json for the updated transition.md size.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* ci: trigger test matrix for 525b946

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix(transition): restore awk frontmatter extraction for pre-shim verification check

The gsd_run launcher shim is not defined until line ~163 of transition.md,
so the verification debt check at line ~80 cannot use gsd_run. Restore the
awk-based frontmatter extraction that correctly reads status without needing
the runtime, and restore the shim at its proper location before
phase.complete.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(#1522): clarify transition verification gate wording

* fix(#1522): update transition workflow size baseline

* fix(#1522): update workflow-size-baseline after rebase onto next

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix(#1522): guard findStaleVerificationSummary FS calls + thread opts.fs seam (review)

Address review blocker B1 on #1548: findStaleVerificationSummary ran fs.readdirSync
and two fs.statSync calls unguarded between readVerificationStatus's try/catch sections,
so a TOCTOU race (a SUMMARY listed by scanPhasePlans then removed before statSync) or any
FS error threw uncaught into callers NOT under the planning lock (init.manager /
init.progress / uat-predicate). Wrap the body in try/catch degrading to 'not stale', and
thread the injectable opts.fs seam (add statSync to FsLike, pass fsImpl from the caller)
for parity with readVerificationStatus's no-throw contract and testability. Also adds the
Verification Module glossary entry to CONTEXT.md (review B3).

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 13:19:10 -04:00
github-actions[bot]
2e406e8e05 chore: sync next package version to 1.6.0-rc.3 2026-06-24 03:52:20 +00:00
Tom Boucher
bcc5a6d1ba fix(#1634): honor capability hook matcher and node-prefix command (#1638)
* fix(#1634): honor capability hook matcher and node-prefix command

Capability hook install (applyCapabilitySharedEdits) wrote each settings.json
hook entry with no `matcher`, so a tool-scoped hook fired on every tool (a
fail-closed guard could then block the whole session), and emitted a bare
single-quoted script path so a .js-family hook from a git/tarball source
without +x failed with Permission denied on every matching call.

- Pass through an optional declared `matcher` (entry-level sibling of `hooks`);
  absent => omitted (match-all), so existing shipped capabilities are unchanged.
- Validate `matcher` in the declaration (non-empty string, no control chars).
- Emit `node <quoted-path>` for .js/.cjs/.mjs hooks (mirrors first-party);
  .sh and others keep the bare quoted path (unchanged).

Root cause: the manifest hook schema (validator rule C4) was {event, script}
only with no matcher, and applyCapabilitySharedEdits never read or wrote one;
the command used shellSingleQuote(absScript) with no node prefix.

Regression tests fail-first on both defects (matcher dropped; bare path) and
pass after the fix; #1460 command assertions updated for the node prefix.

* chore(#1634): backfill changeset pr:1638

* fix(#1634): resolve lint and windows CI failures

- validator: replace the control-character range regex with a char-code loop.
  The literal /[\x00-\x1f\x7f]/ tripped ESLint's no-control-regex rule; char
  codes are equally precise and lint-clean. Behavior unchanged (still rejects
  matchers containing ASCII control characters incl. DEL).
- test: gate the executable-bit precondition on POSIX. Windows fs does not
  honor POSIX write modes (a 0o644 write reads back as 0o666), so the
  precondition is meaningless there and failed the windows-latest lane. The
  node-prefix assertion — the actual fix — is platform-independent and still
  runs everywhere.

* docs(#1634): amend ADR-894 for optional lifecycle hook matcher

The `role: "feature"` `hooks[]` entry now carries an optional `matcher`
(settings.json tool-scoping pattern: exact/pipe/wildcard/regex). Document
the field in the §2 schema table and record a Grilling-amendments entry:
the install path projects a declared matcher onto the emitted settings.json
hook entry (absent = match-all, so shipped capabilities are unchanged), and
per-runtime matcher projection (ADR-857 D8) stays a separate concern. This
amendment ships with the fix that introduced the field rather than as a
follow-up.

* docs(#1634): record WINDOWS-POSIX-MODE-BIT-ASSERT defect in CONTEXT.md

Capture the CI failure pattern from #1634/PR #1638 so it is not repeated: a
test that writes a file with a POSIX mode and then asserts statSync().mode
& 0o777 === <octal> passes on macOS/Linux but fails on windows-latest
(Windows fs does not honor POSIX write modes — reads back 0o666). Added as a
machine-greppable DEFECT predicate (symptom/examples/detect/fix-forward/
prevention) next to DEFECT.WINDOWS-TEST-PORTABILITY, with the fix-forward:
gate the mode-bit precondition on process.platform !== 'win32' and keep the
platform-independent behavioral assertion running everywhere.
2026-06-23 21:49:43 -04:00
Tom Boucher
207d8f1697 fix(#1626): make the security gate severity-aware via per-threat severity (#1635)
workflow.security_block_on was documented as the minimum threat severity
that blocks advancement, but threats carried no severity and the auditor's
threats_open count (the SECURITY.md gate field) counted every open threat
regardless of severity — so the threshold had no effect, and the auditor's
block_on vocabulary (open/unregistered/none) did not even match the config
enum (critical/high/medium/low/none).

- planner: add a Severity column to the STRIDE threat register; assign
  severity per threat.
- auditor: read severity; reconcile the <config> block_on domain to the
  severity enum; redefine threats_open as the count of OPEN threats whose
  severity is at or above block_on (none => 0). Below-threshold opens are
  reported as non-blocking and excluded from threats_open.
- SECURITY.md template + planning-config.md reconciled.

No gate-check site changed: threats_open == 0 stays the gate everywhere;
only its computation is now severity-filtered.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 19:04:13 -04:00
Tom Boucher
f9d9dfb4bc fix(#1627): scale security rigor by ASVS level (planner disposition + auditor depth) (#1636)
workflow.security_asvs_level was display-only — the planner hardcoded
'mitigate if ASVS L1 requires it' and the auditor only echoed the level,
so L2/L3 behaved identically to L1.

- New reference gsd-core/references/security-asvs-levels.md defines L1
  (opportunistic), L2 (standard), L3 (comprehensive) for both planner
  threat disposition and auditor verification depth (higher = superset).
- planner: disposition now scales with the configured ASVS level (no
  hardcoded L1) + @-pointer to the reference.
- auditor: verification depth scales with asvs_level (L1 grep-presence,
  L2 boundary/vector check, L3 end-to-end trace + bypass check).
- planning-config.md + INVENTORY updated; planner kept under its 48K cap
  by extracting the goal-backward worked example to planner-guidance.md.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 18:48:58 -04:00
Tom Boucher
94be6d5b60 fix(#1625): resolve security config in secure-phase.md before auditor handoff (#1633) 2026-06-23 17:24:59 -04:00
Tom Boucher
fc2a7c0555 fix(#1615): install Windsurf slash workflows 2026-06-23 12:10:21 -04:00
Tom Boucher
c28cccbf85 Merge pull request #1595 from davesienkowski/feat/1592-plan-drift-precheck
feat(#1592): add plan:pre codebase-drift pre-check before planner runs
2026-06-23 10:56:29 -04:00
Tom Boucher
cbd21092a9 fix(#1614): install Antigravity skills flat 2026-06-23 10:21:06 -04:00
Tom Boucher
ba96c70b14 feat(#1602): deterministic coverage-metadata UAT routing for verify-work
Add an optional structured `coverage:` block to SUMMARY.md frontmatter and a
deterministic classifier that `verify-work` consumes to route deliverables to
auto-pass vs human-UAT — replacing the rejected #1598/#1599 post-hoc heuristic.

- New `src/coverage.cts` (→ bin/lib/coverage.cjs) parses the nested coverage
  block (extractFrontmatter can't — its `-` items are scalars-only; this is a
  focused parser, sibling of parseMustHavesBlock), validates each entry, and
  classifies into auto_passed vs present. Frozen MODE/PRESENT_REASON/ERROR_CODE
  typed-IR surface. Exposed via `uat classify-coverage --summary <f>`.
- Auto-pass is the narrow proven case only: strict-boolean human_judgment:false
  AND non-empty all-`pass` verification AND zero validation errors. Everything
  else — judgment, empty/failing verification, malformed entry — routes to the
  human (fail-safe). A malformed block falls back to legacy prose extraction and
  surfaces an error; an absent block is byte-identical to pre-#1602.
- execute-plan create_summary populates the block (fail-safe default
  human_judgment:true); verify-work extract_tests consumes it; create_uat_file
  marks auto-passed entries `source: automated`.
- Templates (summary + 3 variants), CONTEXT.md predicate + glossary, INVENTORY,
  eslint/gitignore registration, and Diataxis docs (COMMANDS reference +
  USER-GUIDE explanation) updated.
- Behavioral tests via the CLI (no source-grep); parser-robustness regressions
  for the null-entry/comment-header/mis-indent cases found in adversarial review.

Closes #1602

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 23:43:22 -04:00
Dave
d1f7ba82f2 feat(#1592): add plan:pre codebase-drift pre-check before planner runs
Add a non-blocking, warn-only codebase-drift gate at plan:pre so a stale
STRUCTURE.md is surfaced before /gsd:plan-phase spawns the planner, instead
of being discovered mid-execution by the existing execute:wave:post gate.
Gated on a dedicated workflow.plan_drift_precheck toggle (default true),
independent of schema_drift_gate. Never blocks planning, never spawns the
mapper agent at plan time.

Review feedback (#1595):
- Use a documented conventional-commit type (feat, not enhance) per
  CONTRIBUTING.md / gsd-validate-commit.sh.
- Normalize the plan_drift_precheck command references to the colon prose
  form (/gsd:plan-phase, /gsd:map-codebase) to match plan-phase.md §5.65;
  registry regenerated from capability.json.
- Make the test temp dirs hermetic: drain mkdtemp dirs in an after() hook
  via the helpers.cleanup() budget (local/no-raw-rmsync-in-tests-compliant).

Closes #1592

Claude-Session: https://claude.ai/code/session_016JBiXEAofvB3prJim29XMS
2026-06-22 20:30:59 -04:00
Tom Boucher
b2c0086c1b fix(#1574): resolve review — copilot instruction file is .github/copilot-instructions.md
GitHub Copilot reads repository-wide instructions only from
.github/copilot-instructions.md (confirmed via GitHub Docs), not a root
copilot-instructions.md. Aligns getProjectInstructionFile with the installer
(runtime-config-adapter-registry installSurface 'copilot-instructions') and
cites the docs source in the doc-comment.
2026-06-22 09:56:15 -04:00
Tom Boucher
bf9bd1f4e0 fix(#1529): emit runtime-native instruction file from new-project 2026-06-22 09:32:29 -04:00
github-actions[bot]
f20dc691c1 chore: sync next package version to 1.6.0-rc.2 2026-06-22 05:22:23 +00:00
Joe
e12a2abfd8 feat(#441): add /gsd-capture --list-seeds for seed listing and audit (#722)
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit

Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed),
enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way
to browse or audit parked seeds on demand. This adds a read-only listing,
following the established --list → workflow pattern (per the approved scope on

- gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the
  seeds dir, returns { count, seeds[], summary } JSON with each seed's id,
  slug, status, scope, trigger_when, planted, title. Optional case-insensitive
  status filter. User-controlled content is sanitized (sanitizeForDisplay) and
  every path validated (requireSafePath); read-only. Independent of
  audit.scanSeeds, which only returns unimplemented seeds for the milestone surface.
- /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that
  renders the seed table.

Closes #441

* chore(#441): point changeset fragment at PR #722

* test(#441): allowlist list-seeds test in prompt-injection scan

The test asserts that list-seeds neutralizes injection payloads
(<system>, [INST]) embedded in seed content, so the fixtures legitimately
contain those patterns — same as the sibling security tests already on the
allowlist.

* fix(#441): use canonical /gsd:capture colon form in list-seeds workflow

Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use
the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is
retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new
list-seeds workflow used the hyphen form.

* docs(#441): sync help full.md + INVENTORY for --list-seeds

Adds the --list-seeds entry to the help reference (help/modes/full.md, per
bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow
in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json.

* docs(#441): add --list-seeds how-to + drop phantom statuses

Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers):

- USER-GUIDE.md Seeds section (how-to): extend the task to cover
  auditing parked seeds on demand via --list-seeds, including the
  status filter — kept task-oriented per Diataxis how-to mode.
- CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected
  from the list-seeds filter vocabulary; the system only produces
  dormant|active|triggered (src/audit.cts scanSeeds). Reference must
  be factually accurate and complete.

* fix(#441): guard non-scalar status frontmatter in cmdListSeeds

A seed with a bare `status:` line (extractFrontmatter yields {}) or a
`status: [a, b]` value (yields an array) crashed the whole audit list:
`(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string.
Coerce every frontmatter read through a `fmStr` helper (mirrors the existing
`typeof fm.id === 'string'` guard), so a non-scalar status falls back to
dormant and non-scalar scope/trigger_when/title can no longer leak a raw
array/object into the JSON contract. Title is now capped symmetrically.

Adds regression coverage for empty and array `status:` and non-scalar fields.

Refs #441

* docs(#441): align list-seeds workflow status vocabulary

The load_seeds step listed `implemented` as an example status filter, but the
real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds);
`implemented` has no producer. Matches the earlier CLI-TOOLS.md correction.

Refs #441

* refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds

Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported
deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested
in-process (review minor #1). No behavior change.

Filter comparison now matches the raw lowercased status (both sides already
normalized) instead of sanitizeForDisplay(status); sanitization is for output,
not matching (review nit #3).

* test(#441): add fast-check property coverage and count=1 boundary for list-seeds

Adds tests/list-seeds.property.test.cjs with four fast-check properties over
deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug
invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing
(review minor #1).

Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2).

* chore(#441): sync runtime launcher snippet into list-seeds workflow

Propagate the current _runtime-launcher.snippet.sh (with non-Claude
runtime home probes) into the new list-seeds.md workflow via
scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation.

* test(#441): record list-seeds.md in workflow size baseline (#1074)

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-22 00:59:41 -04:00
Behruz Nassre Esfahani
ac40f070ef feat(#1318): require external reviewers to verify plan claims against source (#1421)
* feat(#1318): require external reviewers to verify plan claims against source

/gsd-review built its external-reviewer prompt from plan text only and never
asked reviewers to open the repo and verify claims, so a grounded HIGH could be
outvoted by ungrounded LOWs. Add a concise, generic source-grounding block to
build_prompt's Review Instructions: treat yourself as running in the working
tree, open referenced files, cite path:line + mechanism, trace asserted
mechanisms, downgrade to an open question if you have no file access, and know
that grounded findings are weighted more heavily.

Also clarify that CodeRabbit (a diff-only reviewer that never receives the
prompt) must not be weighted as a grounded plan-level verdict in consensus
synthesis. Workflow stays under its size cap (baseline bumped deliberately).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1318): add changeset for reviewer source-grounding

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1318): mark changeset docs-exempt (internal reviewer-prompt wording)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#1318): document reviewer source-grounding in COMMANDS.md; drop docs-exempt

Review: a user-visible behavioral Changed warrants a docs touch, not a
docs-exempt. Add a sentence to the /gsd-review entry in docs/COMMANDS.md
(reviewers verify against source, cite file:line, grounded findings weighted
higher) and remove the changeset docs-exempt marker so lint:docs passes via
docs-updated. Also note the literal build_prompt test anchor.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1318): harden build_prompt fence extraction to be fence-run-aware

Addresses maintainer review on PR #1421 (required-before-merge).

The buildPromptReviewInstructions() test helper located the closing fence with
`src.indexOf('\n```')`, which terminates at the FIRST triple-backtick line — so
a build_prompt ```markdown block whose body embeds a fenced code example would
truncate mid-content (dropping the `## Review Instructions` section) and give a
spurious failure or false pass. Since this feature feeds source/plan content
(which routinely contains code fences) to reviewers, that is a live fragility.

Rewrite the extraction to be fence-run-aware, mirroring the CommonMark close
rule in src/markdown-sectionizer.cts stripFencedCode: parse the opener's
backtick run length, then close on the first line with >= that many backticks
and only trailing whitespace — so a shorter nested fence is treated as content.
Add a fail-first regression test (a 4-backtick outer fence wrapping a nested
```bash block) asserting the trailing `## Review Instructions` still extracts.

Test-only change; no production .cts touched. Verified: test file 7/7,
empirical fail-first proof the old indexOf logic truncated, full suite
4236/4236, eslint clean. Codex review: approve.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-22 00:47:49 -04:00
Enes Yağız
8748e95ed1 fix(#666): pr-branch silently ignored planning.sub_repos (#667)
* fix(pr-branch): handle sub_repos from config with git -C (#666)

Adds a `handle_sub_repos` step between `detect_state` and
`analyze_commits`. When `planning.sub_repos` is set in config, the
workflow now:

- Reads sub-repo paths via `gsd_run query config-get sub_repos`
- Skips the step entirely when the list is empty/null/[]
- Scans each repo with `git -C "$REPO" status --porcelain`
- Offers the user all/select/skip choices
- For selected repos: creates a PR branch, commits all staged/unstaged
  changes, pushes, and opens a companion PR via `gh pr create`

All git commands use `git -C "$REPO"` — never `cd "$REPO"` — because
shell state does not persist between agent-executed commands.

Closes #666

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: update changeset pr number to 667

* fix(pr-branch): address maintainer review — correct seam, behavioral tests, robustness

Resolves all three blockers and seven robustness issues raised in PR #667 review:

Blockers:
- Use `planning.sub_repos` (not top-level `sub_repos`) so config-get actually resolves
- Replace prose grep test with behavioral fixture tests using runGsdTools + local bare repo
- Extract sub-repo git work into new `cmdPrSubrepo` seam in src/commands.cts;
  never uses git add -A — stages explicit files only (universal-anti-patterns.md:44)

Robustness:
- Dirty-repo list persisted via mktemp/cat, not bash arrays (cross-block safe)
- Branch name embeds repo slug (${CURRENT_BRANCH}-${REPO_SAFE}-pr) to avoid collision
- push --set-upstream so gh pr create finds the branch
- Sub-repo base branch resolved via ls-remote with fallback to repo's default branch
- Remote slug parsed with /github\.com[:/]/ (handles SSH + HTTPS + .git-less URLs)
- rollback() cleans up branch on any mid-sequence failure
- node -e replaces jq (always available, no undeclared hard dep)

Refs: #666

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(pr-branch): security guard, push timeout, rollback fix, porcelain fix

Security (Blocker 1):
- Use security.cjs validatePath() in cmdPrSubrepo for symlink-safe workspace
  containment check — rejects ../escape, absolute paths, and symlink traversal
- Add negative regression test: '../escape' repo path must be rejected

Robustness:
- Push uses timeout: 60_000 ms (network op needs more than the 10 s default)
- Capture prevBranchName before checkout -b so rollback uses explicit name
  instead of git checkout - (fails on fresh single-branch repos)
- Porcelain path parse: line.trimStart().slice(2).trim() handles all XY
  combinations and the execGit global-trim edge case uniformly

Tests: 17/17 pass, lint: 0 errors

Refs: #666

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(pr-branch): move regression tests to commands.test.cjs, add core.quotePath=false

- Move cmdPrSubrepo behavioral + workflow source-invariant tests from
  standalone bug-666-*.test.cjs into tests/commands.test.cjs under
  describe('pr-subrepo') per TESTING-SUITES.md policy (no new bug-* files).
  Adds allow-test-rule: source-text-is-the-product see #666 for the
  workflow-source-invariant suite.
- Add -c core.quotePath=false to git status --porcelain call so non-ASCII
  filenames (e.g. café) are not C-escaped, keeping slice(2) parse correct.

* fix(pr-branch): remove obsolete regression tests for sub-repos handling

* fix(pr-branch): update workflow-size-baseline, add dirty-scan timeout

- Regenerate tests/workflow-size-baseline.json for pr-branch.md growth
  (+handle_sub_repos step, +timeout addition).
- Add { timeout: 10_000 } to the execFileSync git status --porcelain
  call in the handle_sub_repos dirty-scan (repo convention: every git
  subprocess is bounded, never hangs).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: regenerate INVENTORY-MANIFEST after rebase onto next

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): handle rename staging and split changedFiles from filesToStage

For git mv renames, the old path no longer exists in the worktree after
the move — staging it with git add fails. Split parsing into changedFiles
(both paths, for result.files) and filesToStage (new path only for
renames; old is already staged by git mv). Also adds porcelain tests
for staged renames, non-ASCII filenames, and a fast-check property test.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): rollback on push failure in cmdPrSubrepo

If push fails the branch only exists locally; rollback cleans it up so
the sub-repo is not left in a half-committed state.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): do not rollback after commit on push failure; add push-fail regression test

Post-commit push failures are network/auth/policy issues — the user's work
is already committed on the local branch. Calling rollback() at that point
force-deletes the only ref holding the commit (data loss). Leave the branch
in place and emit a retry instruction instead.

Adds a regression test (pre-receive hook that rejects all pushes) asserting
the branch and commit survive a push rejection so the failure path stays
covered going forward.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: regenerate INVENTORY-MANIFEST after rebase onto next

Rebased onto current next (#1267 retired core.cjs). Stale tsbuildinfo and
a leftover bin/lib/core.cjs build artifact were masking the drift — wiped
both, rebuilt clean, and regenerated the manifest. gen-inventory-manifest
--check now exits 0.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#666): validate sub-repo paths before git invocation in pr-branch.md

The handle_sub_repos workflow ran git -C on raw planning.sub_repos config
values at two points before the pr-subrepo seam's validatePath guard ever
ran: the dirty-scan detection (git status) and the base-branch resolution
(git ls-remote / remote show). A traversal entry could point git outside
the workspace; an embedded newline could inject a spurious record into
the newline-joined dirty-file output and into the shell-interpolated
commit message.

Adds a containment check + character allowlist to the dirty-scan node
script (reject before any execFileSync), and a defense-in-depth shell
case guard on the same value before the second, independent git -C
invocation in the base-branch resolution block.

Adds a behavioral test that extracts and executes the actual shipped
node script from pr-branch.md (not a mirror) against a real traversal
target and an embedded-newline entry, asserting neither reaches git or
the dirty-file output.

Also updates the stale cmdPrSubrepo doc comment: push failures no longer
delete the branch (see prior commit), only stage/commit failures do.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#666): make sub-repo traversal scan test genuinely fail-first

The outside repo's only change was an untracked file, which the ?? filter
excludes — so the repo looked clean even with the guard removed, making the
traversal assertion vacuous (it passed against a neutered guard). Commit the
file first, then modify it, so the outside repo has a tracked dirty change:
without the path guard it WOULD be reported dirty, so the test now fails-first.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#666): symlink-safe (realpath) sub-repo containment in pr-branch.md

Finding A from re-review: the workflow guard used path.resolve, which only
normalizes '..' textually and does not follow symlinks — so an in-tree symlink
whose name has no '..' or '/' (e.g. "evil" -> /outside) passed both the charset
filter and the resolve+startsWith check, letting git status / ls-remote /
remote show run against a directory outside the workspace. The pr-subrepo seam
already used fs.realpathSync (validatePath); this brings the workflow layer to
parity.

- dirty-scan: realpathSync the root once, and realpathSync each candidate before
  the containment check; skip on throw.
- base-branch resolution: replace the weak `case *..*|/*` guard with a realpath
  containment check that yields a validated absolute SUB_REPO_DIR, and run git -C
  against that instead of re-concatenating $ROOT/$REPO_REL.
- security test: add a symlink-escape entry and a positive control (legit in-root
  backend must still be reported). Confirmed fails-first — regressing the scan to
  path.resolve makes the symlink case leak.

Also fixes a misleading-fallback minor: the workflow now checks the seam's exit
status and skips the companion-PR step on failure, instead of printing
"branch pushed, open PR manually" after a real stage/commit/push failure.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#666): harden pr-branch sub-repo flow against round-12 edge cases

Pre-emptive hardening of the workflow changes from the symlink fix:

- continue-outside-loop: the "skip companion PR on seam failure" block used a
  bash `continue`, but the per-sub-repo iteration is prose-driven (the agent
  loops, not a literal `for`), so `continue` would warn and no-op. Reframed as
  prose-gated control flow keyed on $SUBREPO_EXIT — no bash loop assumption.
- Windows portability: the new symlink security case now degrades gracefully
  (try/catch around fs.symlinkSync; skip just the symlink assertion when symlink
  creation lacks privileges) so it doesn't hard-fail on Windows CI.

Verified: seam exits 1 on error / 0 on success (error() → process.exit(1),
propagated through the shim), so the $SUBREPO_EXIT check is meaningful; bash -n
clean on the touched blocks; commands 156/156; lint:ci green; manifest in sync.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-21 22:34:00 -04:00
Tom Boucher
94e7e3f88f refactor(#1558): plan runtime artifact uninstall removal (#1564) 2026-06-21 21:55:29 -04:00
Tom Boucher
405ae9b3b7 refactor(#1557): add runtime artifact install plan module (#1560) 2026-06-21 20:40:13 -04:00
Rezolv
195d356d7c Merge branch 'next' into feat/1346-enhance-verify-phase-project-a-check-vio 2026-06-21 17:41:37 -04:00
Behruz Nassre Esfahani
faac9331f2 feat(#1298): add validated worktree record-agent writer verb for wave manifests (#1448)
Closes #1298
2026-06-21 15:38:44 -04:00
Tom Boucher
2c718bf972 fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs (#1537)
* fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs

Generalizes the Codex-only #1515/#1519 fix to every non-Claude runtime, and
wires it into the real install path (where it was previously dead-on-arrival).

Root causes:
1. The runtime-default stamping lived only in `_applyRuntimeRewrites`, but the
   installer emits `gsd-core/workflows/*.md` via `copyWithPathReplacement`, which
   never calls it — so a real `--codex`/`--cursor`/etc. install emitted
   `--default claude` and worktrees-on. RUNTIME mis-resolved to claude and the
   workflow ran executors unisolated against the main checkout. (#1515/#1519 were
   also dead-on-arrival in real installs; this repairs them.)
2. Only `case 'codex'` was stamped; every other non-Claude runtime kept the
   Claude default.

Fix:
- New `_stampNonClaudeRuntimeDefaults(content, runtime)` (single shared helper)
  stamps `--default <runtime>` + `use_worktrees=false` for every `runtime !=
  claude`; called from both `_applyRuntimeRewrites` and, crucially,
  `copyWithPathReplacement` in bin/install.js (the real workflow emit path).
- Generalize the fail-closed worktree guard `= codex` -> `!= claude` in
  execute-phase/quick/diagnose-issues (worktree isolation is Claude-Code-only).
- Flip manager/autonomous inline-vs-background gating to `codex -> background,
  everything-else -> inline` (research: only Codex can background-nest the
  pipeline's subagents; all others run inline, which they support).

Worktree-capability determination is research-backed (official docs for all 14
non-Claude runtimes: none honor GSD's isolation="worktree" mechanism, only Codex
background-nests). New end-to-end real-install test asserts the EMITTED workflow
is stamped — the regression guard that would have caught the dead-on-arrival bug.

Closes #1521

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

* chore(#1521): backfill changeset PR number (#1537)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 13:48:47 -04:00
Tom Boucher
2436b76980 fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees (#1519)
* fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees

A Codex install with a runtime-neutral .planning/config.json resolved
RUNTIME=claude and enabled git worktree isolation, which Codex's
spawn_agent cannot honor. Two root causes:

1. Workflows read `config-get runtime` / `config-get workflow.use_worktrees`
   without `--raw`, so config-get's JSON-quoted output ("codex") was captured
   verbatim into the bash var and broke every `[ "$RUNTIME" = ... ]` check —
   the Codex fail-closed guard was dead even when runtime:codex was explicit,
   and Claude's own worktree degrade-check was dead too. Add `--raw` to those
   reads across execute-phase, autonomous, manager, diagnose-issues, quick.

2. The conversion engine emitted `--default claude` for every runtime. Stamp
   the codex-emitted workflows to `--default codex` (runtime) and
   `--default false` (use_worktrees) in _applyRuntimeRewrites case 'codex', so
   a neutral config on a Codex install resolves runtime=codex / worktrees off.

Also extend the Codex fail-closed worktree guard to quick.md and
diagnose-issues.md (they spawned isolation="worktree" with no runtime guard).

Regression test asserts source<->engine parity across all five workflows
(DEFECT.GENERATIVE-FIX) plus fast-check property coverage of the stamping.

Closes #1515

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

* chore(#1515): backfill changeset PR number (#1519)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 11:56:59 -04:00
Dave
c09c13f295 enhance(verify-phase): node-test causation control — prove the RED is content-caused (#1346)
The #1279 node-test machine-proof confirmed a known-bad subject drives the
negative test RED, but could not distinguish a genuine content-violation from a
deceptive test that reds merely because GSD_PROHIB_SUBJECT is set. Add an
optional fifth flat scalar `check_clean_fixture` (-> CheckDescriptor.cleanFixture)
threading a KNOWN-CLEAN control subject through projectProhibitions +
descriptorFromProjection. When present, the prover also runs the check against
the clean subject and requires GREEN, so fail-first is proven only when the check
is RED on the violation AND GREEN on the clean subject (content-dependent).

Opt-in and additive: absent a clean fixture the prover behaves exactly as
post-#1314 (no control, documented residual), preserving the zero-authoring
compose path; the lint-rule kind needs no analog (its subject IS the linted
file, no env indirection). Coverage: RED-first deceptive case, positive,
missing-clean fail-closed, round-trip read-back/emit, fast-check property
extended to the 5th scalar, and an end-to-end COMPOSE capstone (honest vs
deceptive). Docs: ADR-550 dated addendum, prohibition-probe reference,
spec-phase + verify-phase workflows.

Closes #1346

Claude-Session: https://claude.ai/code/session_01GsPRb8zvpcT7Eat6vZw8PX
2026-06-21 10:44:48 -04:00
Tom Boucher
b63a500fad fix(#1505): extract context_guard step to reference file; fix allow-test-rule see ref
- Extract execute-phase.md context_guard step prose to
  gsd-core/references/execute-phase-context-guard.md (@-ref lazy load),
  bringing execute-phase.md back under the ADR-857 phase-6 size ceiling
  (92914 < 93166 bytes)
- Fix allow-test-rule comment: add `see #1452` per ADR-456 lint rule
- Update feat-1452 tests to check the reference file for extracted content
- Register execute-phase-context-guard.md in INVENTORY-MANIFEST.json and
  INVENTORY.md Workflow References section
- Regenerate workflow-size-baseline.json after file shrinkage

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 18:38:20 -04:00
Tom Boucher
4eca5ac96c feat(#1452): add workflow.context_guard_mode to guard execute-phase against context exhaustion
Proactive checkpoint guard fires at each wave boundary before spawning agents.
Self-assesses context pressure against context-budget.md degradation tiers and
warns (warn, default) or auto-invokes /gsd:pause-work (auto) when POOR tier
(70%+) is detected. Config key validated; defaults to \"warn\".

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 18:10:49 -04:00
github-actions[bot]
4441b20b4a chore: sync next package version to 1.6.0-rc.1 2026-06-20 19:58:40 +00:00
Tom Boucher
c330f70f65 feat(#1494): add workflow.mvp_mode to VALID_CONFIG_KEYS; document code_review_command and plan_chunked in planning-config.md (#1500)
* feat(#1494): add workflow.mvp_mode to VALID_CONFIG_KEYS; document code_review_command, plan_chunked, mvp_mode in planning-config.md

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: backfill PR number 1500 in changeset

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 15:32:07 -04:00
Tom Boucher
f5276b36b3 fix(#1369): refresh wave manifest and re-check base before each wave in execute-phase (#1492)
* fix(#1369): refresh wave manifest and re-check base before each wave in execute-phase

Two compounding issues caused wave N+1 worktrees to fork from the stale
pre-wave-N commit, immediately tripping the worktree_branch_check FATAL
guard in every executor:

1. worktree.base-check auto-degrade only ran once at initialize time.
   After wave N merges advanced orchestrator HEAD past origin/HEAD, new
   worktrees were still forked from origin/HEAD (Claude Code "fresh" base).

2. WAVE_WORKTREE_MANIFEST was never unset between waves, so wave N+1
   reused the consumed wave-N manifest file, which would have blocked the
   step 5.5 manifest guard (#3384) on subsequent waves.

Fix: add two safeguards in execute-phase.md —
- Step 0.5 (start of each wave): re-runs worktree.base-check; auto-degrades
  USE_WORKTREES=false for that wave when HEAD has diverged from origin/HEAD.
- Step 7c (end of each wave): unsets WAVE_WORKTREE_MANIFEST so wave N+1
  creates a fresh per-wave manifest; re-asserts worktree.set-baseref
  (idempotent) and re-evaluates base degradation after wave merges land.

17 regression tests added in tests/bug-1369-wave-stale-base.test.cjs.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1369): rename test to fix-NNN convention; update workflow size baseline

Rename tests/bug-1369-wave-stale-base.test.cjs → tests/fix-1369-wave-stale-base.test.cjs
to satisfy the lint-regression-test-names gate (new files cannot use bug-NNN prefix).

Update tests/workflow-size-baseline.json for execute-phase.md: 93157 → 97393
(LF-normalized byte count after adding step 0.5 inter-wave base re-check and
step 7c between-wave manifest reset).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1369): add issue reference to allow-test-rule comment

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1369): extract new execute-phase steps to references; satisfy ADR-857 cap

Step 0.5 (inter-wave worktree base re-check) and steps 7b–7c (pre-wave
dependency check + between-wave manifest reset/base refresh) added by this
PR grew execute-phase.md to 97393 bytes, violating the ADR-857 phase-6
architectural mandate that host-loop bodies remain strictly below the
pre-phase-6 baseline of 93166 bytes.

Extract both new blocks into dedicated reference files:
- gsd-core/references/execute-phase-wave-guard.md (step 0.5)
- gsd-core/references/execute-phase-between-wave-reset.md (steps 7b + 7c)

Replace inline prose with @-reference pointers. File now measures 92851
bytes (LF-normalized), satisfying the ADR-857 capstone conformance gate.

Also update tests/workflow-size-baseline.json to 92851 and add both new
reference files to docs/INVENTORY-MANIFEST.json.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1369): update regression tests to read from extracted reference files

Steps 0.5 and 7b+7c were moved to reference files to satisfy the ADR-857
size cap on execute-phase.md. Tests now check @-reference pointers in the
workflow for ordering and read content assertions from the reference files.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 14:22:59 -04:00
Tom Boucher
e627584f87 fix(#1453): rewrite stale get-shit-done paths in Codex skill mirror on upgrade (#1491)
* fix(#1453): clean up stale get-shit-done paths in Codex skill mirror on upgrade

Extends planLegacyCleanup in gsd-core/bin/lib/legacy-cleanup.cjs to scan
skills/gsd-* subdirectories for .md files that still embed the pre-rename
get-shit-done/ path (e.g. ~/.agents/skills/gsd-docs-update/SKILL.md).
These stale copies are removed by cleanupLegacyGsdCc during the next
install/upgrade so Codex can no longer discover and select them.

Adds 5 regression tests to tests/issue-607-legacy-cleanup.test.cjs covering
the stale path detection, non-flagging of fresh skills and user-owned dirs,
and the end-to-end ~/.agents/skills scenario from the issue.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1453): add gsd-allow-legacy-name markers to intentional legacy name references

Comments and test descriptions in legacy-cleanup.cjs and its test file
legitimately cite the old 'get-shit-done' directory name to explain what
the cleanup logic removes. Add the lint-exemption marker to each line so
lint-legacy-dir-name passes.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 13:37:38 -04:00
Tom Boucher
33ccf5f89d fix(#1367): project-local install uses flat gsd-<cmd>.md layout (fixes /gsd: colon namespace) (#1489)
* fix(#1367): project-local install uses flat gsd-<cmd>.md layout

Claude Code project-local installs now write command files as flat
gsd-<cmd>.md at .claude/commands/ level instead of commands/gsd/<cmd>.md
(subdirectory), so Claude Code registers /gsd-<cmd> (hyphen form)
matching hooks, statusline, and all cross-command references.

- capabilities/claude/capability.json: local destSubpath commands/gsd → commands
- bin/install.js else branch: flat gsd-<stem>.md loop with runtime rewrites
- bin/install.js uninstall (1c): remove flat files + legacy subdir cleanup
- bin/install.js writeManifest: record flat commands/gsd-<cmd>.md keys
- legacy migration: preserves dev-preferences.md across reinstall and uninstall
- gsd-core/bin/lib/capability-registry.cjs: regenerated
- 6 new regression tests (L0–L5) in bug-1367-*.test.cjs
- Updated E suite in bug-3683 + bug-1736, layout + surface + descriptor tests
- scripts/lint-regression-test-names.allowlist.json: grandfathered bug-1367 test

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1367): add issue reference to allow-test-rule comment

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 13:37:27 -04:00
Tom Boucher
fe0f9f904d fix(#1478,#1479,#1480): prohibit ungrounded baselines, error-suppressing fallbacks, and stale-artifact authority in planner verify blocks (#1482)
* fix(#1478,#1479,#1480): prohibit ungrounded baselines, error-suppressing fallbacks, and stale-artifact authority in verify blocks

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: add changeset for #1478/#1479/#1480 planner verify gate fix

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: correct changeset format

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1478,#1479,#1480): fix test contract violations from new planner dimensions

gsd-planner.md exceeded both the planner-decomposition 48K char limit and
the reachability-check 50K char limit after the new HARD RULE blocks were
added inline. The full rule details already exist in planner-antipatterns.md
(added in the same PR); replace the verbose inline blocks with a single
@-reference pointer to the antipatterns file, reducing the file from 50981
to 49130 chars (under both limits).

Also regenerate tests/agent-size-baseline.json to reflect the new sizes of
gsd-planner.md and gsd-plan-checker.md.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 13:37:03 -04:00
Tom Boucher
3a3b2135c2 chore(#1073): purge phantom pre-migration issue refs from source, tests, docs (#1471)
#2551/#3182/#2361 are pre-migration get-shit-done-redux issue numbers with no
equivalent in open-gsd/gsd-core; they mislead triage and manufacture phantom
blockers. Repoint to real successors (#717 byte-budget rework, #720) or rewrite
as prose referencing the discuss-phase/modes progressive-disclosure split.
Correct co-located 'line budget'/'<500 lines' framing to the byte-based reality
(#717). Add a CI guard (tests/no-phantom-issue-refs.test.cjs) that fails if a
phantom ref is reintroduced. SSH-key patterns (id_ed25519) left untouched.
No user-facing runtime behavior change.

Closes #1073
2026-06-20 13:36:57 -04:00
Tom Boucher
7c93d9e222 feat(#1463): add capability outdated (per-source update check); drop phantom slash-command docs
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 10:50:12 -04:00
Tom Boucher
08d1c57d6e fix(#1460): verify-or-reject capability --integrity per source; confine hook commands to the bundle
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-20 09:10:04 -04:00
Tom Boucher
e7855bc217 fix(#1459): user-owned consent store gates third-party capability activation; env/cwd in disclosure; loader validator parity (#1473) 2026-06-20 01:59:03 -04:00
Tom Boucher
c866ac1b24 fix(#1462): fail closed without data loss on a corrupt capability ledger; atomic ledger write (#1469) 2026-06-19 19:55:10 -04:00
Tom Boucher
34bc096ec2 feat(#1451): wire gsd capability install/update/remove/list/disable/enable management CLI (#1457)
* feat(#1451): wire gsd capability install/update/remove/list/disable/enable CLI

ADR-1244 D5/D6: the management command was built as a library (capability-lifecycle.cjs
install/upgrade/remove + capability-ledger.cjs) across Phases 3-5 but never wired to a
user-facing command — gsd-tools.cjs 'capability' only handled state/set. This adds the
six subcommands, dispatching to the existing lifecycle/ledger:

- install <spec> [--integrity] [--scope global|project] [--yes] [--shared-file <rel>]…
- update [<id>|--all] [--scope] [--yes] [--shared-file]  (re-resolves recorded source)
- remove <id> [--purge-data] [--scope]  (first-party rejected)
- list [--json]  (first-party + overlay, both scopes, JSON array)
- disable|enable <id>  (activation-state alias of capability set --off/--on)

Scope→runtimeDir mapping matches capability-loader exactly (global=$GSD_HOME||home,
project=project root; caps at <root>/.gsd/capabilities/<id>, ledger at <root>/.gsd-capabilities.json).
Consent is non-interactive: --yes grants; without it an executable install aborts after
printing the disclosure and writes nothing. Best-effort reconcile before each mutation.

Tests: tests/capability-cli.test.cjs (20 behavioral, real resolver via local specs,
GSD_HOME-sandboxed) — install consent/block/usage matrix, list, update round-trip,
remove round-trip + first-party guard, disable/enable, unknown subcommand.
Docs: docs/reference/gsd-capability-command.md reconciled to the real surface
(ledger paths, --shared-file, consent model, disable mechanism, outdated marked planned);
docs/COMMANDS.md gains the gsd capability entry.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1451): resolve adversarial-review findings + root-cause the --raw silent-output bug

Adversarial-review (Codex) fixes:
- capReadStrict passes a malformed strict_known_registries value THROUGH so the trust gate
  fail-closes on it (was silently downgrading to permissive)
- installCapability/upgradeCapability gain an expectedId guard + first-party-id rejection
  (capability-lifecycle.cts): an overlay can't shadow a first-party id, and 'update <id>' can't
  act on a different id if the recorded source was retargeted
- capability update: prints the consent disclosure, exits non-zero on --all partial failure,
  no longer masks the resolved id
- capability remove: ledger-first ordering so an overlay is removable even if it shadows a
  first-party name; first-party guard only fires for ids not in the ledger
- gsd-capability-command.md: disable/enable doc corrected (registry-known ids; overlay toggle
  not yet wired through this path)

Silent-output bug (root cause, not waved off as pre-existing):
- captureStdoutSyncWrites buffered fd-1 output and DISCARDED it on the throw path — any --raw
  command that emitted a result/error envelope then threw (to set a non-zero exit) lost ALL of
  stdout. Now it flushes the captured buffer before re-throwing (exit code preserved).
- cmdCapabilitySet threw via process.exit() (bypassing the capture wrapper entirely); now throws
  ExitError so the wrapper flushes — matches the repo's no-process-exit architecture.
- Regression test: capability disable <unknown> --raw must emit the JSON error envelope on stdout.

Verified: capability suite 165/165, @file/json-errors/phase 183/183, lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1451): address adversarial-review R2 — shared-file confinement, MCP no-clobber, config fail-closed

- confinedSharedFile(): realpath-confine every shared-config write/strip to the scope root (mirrors
  safeRmUnder), so a --shared-file whose parent is a symlink escaping the scope can't write outside it.
- mcpServers shared edits: never overwrite an UNOWNED entry — a name collision with the user's (or
  another capability's) server is skipped, so install/remove can't silently clobber user MCP config
  (hooks already append; the map-keyed mcpServers path was the gap).
- capReadStrict: a PRESENT-but-unparseable .planning/config.json now fails CLOSED (lockdown) instead
  of silently downgrading the strict_known_registries policy to permissive.
- Tests: symlink-escape shared-file writes nothing outside scope; colliding user mcpServers entry
  preserved; unparseable config blocks an external install. capability suite 83/83, lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1451): address code-review — aborted-status robustness + coverage + project-scoped strict doc

- install/update: handle an 'aborted' result independently of the requiresConsent flag so it can never
  fall through to the generic 'blocked: unknown reason' arm (aborted always means consent-needed per
  the lifecycle contract; latent today, hardened for future status additions).
- Clarify capResolveScope comment (project scope === already-resolved cwd) and document that
  strict_known_registries is a PROJECT-scoped policy (read regardless of --scope; no machine-wide
  allowlist) in gsd-capability-command.md.
- Tests: update --all over an empty ledger returns an empty result set (exit 0); a flag value that
  looks like another flag (--integrity --scope) is rejected, not swallowed. CLI suite 33/33, lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1451): FEATURES.md entry #147 + Added/Fixed changesets for the capability CLI

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1451): backfill changeset PR number → #1457

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 12:02:36 -04:00
Tom Boucher
1abebbf4fd feat(#1434): registry-driven dispatch for third-party capabilities (ADR-1244 Phase 5) (#1450)
ADR-1244 Phase 5 (D7). dispatchOverlayCapabilityCommand in gsd-tools.cjs dispatches an installed third-party capability command family via loadRegistry({includeInstalled}), gated on a committed ledger entry (consent) and confined to the capability's install root (defaultRequireFromInstallRoot: bare-.cjs basename + realpath containment, rejects ../ traversal + symlink escape); same own-property/function/sync/ExitError guards as the first-party path. capability-loader records _overlay.commandRoots only for accepted overlay caps with a committed, structurally-valid ledger entry (fail closed). First-party graphify/intel/audit unchanged (already on the registry seam). 3 Codex rounds converged + /security-review (no HIGH) + /code-review (Approve); gsd-test green both platforms; CI green.

Closes #1434.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 23:37:43 -04:00
Rezolv
dcceb1a004 fix(#1441): antigravity resolver prefers GSD-owned dir over first-existing (#1442)
* fix(#1441): antigravity resolver prefers GSD-owned dir over first-existing

resolveConfigHomeFromDescriptor's dot-home-nested probe returned the first
bare-existing candidate, so a CLI user (~/.gemini/antigravity-cli) who also
had the IDE's ~/.gemini/antigravity dir was silently shadowed to the legacy
dir (probed first). Regression from #217 — pre-#217 returned
~/.gemini/antigravity unconditionally.

Add an optional probeMarker (gsd-core/VERSION) to the dot-home-nested
descriptor: a two-pass probe prefers the candidate GSD installed into, then
bare existence, then probe[0]. Behavior is byte-identical when probeMarker is
absent (windsurf etc. unaffected). Adds detectAntigravityDirAmbiguity() for
installer/operator guidance on already-misinstalled users (auto-relocation
ruled out per ADR-0008's single-configDir migration bound).

Regression tests fail before / pass after: coexistence + marker-priority
cases, end-to-end through the registry descriptor.

Claude-Session: https://claude.ai/code/session_01JU2WB23JLE3QTysPXuFJjB

* chore(#1441): add changeset for antigravity resolver fix

Claude-Session: https://claude.ai/code/session_01JU2WB23JLE3QTysPXuFJjB
2026-06-18 21:47:18 -04:00
Tom Boucher
9219af3360 feat(#1433): capability trust gate + upgrade/compat (ADR-1244 Phase 4) (#1449)
ADR-1244 Phase 4 (D5 trust + D6 upgrade/compat). capability-trust.cjs (disclosure/consent, strict_known_registries, engines+compatVersions, reserved namespace) + capability-lifecycle.cjs (install/upgrade/remove/reconcile; ledger-as-commit-point _pending intent; atomic stage-then-swap; surgical marker-isolated shared-edit strip; owner-token lock) + capability-source promote/skipEnginesGate seams + loader pending-skip + config keys. No sandbox re-derived (consent+integrity+reversibility). 6 Codex adversarial rounds + /security-review (no HIGH) + /code-review; gsd-test green both platforms; CI green. Phase 5 (#1434) wires the CLI dispatch.

Closes #1433.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 21:40:37 -04:00