* feat(#1708): typed documentation-sourced #853 dispatch-flatten
Graduate the #853 orchestrator-backgrounding decision from a scattered RUNTIME==='codex' prose check to a typed, documentation-sourced engine decision. Adds a backgroundDispatch dispatch sub-axis (sourced per host: codex+cursor documented true, 9 documented false, 5 undocumented), shouldFlattenDispatch(dispatch) (inline UNLESS background && backgroundDispatch, fail-closed), and a gsd_run query dispatch-should-flatten the plan/execute workflows call. Cursor is newly background-eligible per its docs (inline->background) — a documentation-justified behavior change. No RUNTIME-name residue for this decision.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1708): backgroundDispatch citations in matrix + CONTEXT note
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1708): address review findings on typed dispatch-flatten
Code/adversarial review: convert the manager.md/autonomous.md Compound Action preamble from hardcoded 'On Codex' to FLATTEN-based branching (the handlers already use the query; the preamble contradicted them and was wrong for cursor); make shouldFlattenDispatch null-safe + type-honest (accepts raw 'undocumented' registry values); make backgroundDispatch a required descriptor field (matching its siblings, all 16 carry it); strengthen the config.runtime behavioral test; update the bug-853 prose-pin test + comment. Security review clean; Codex confirmed no fail-open.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1708): backfill backgroundDispatch in role:runtime test fixtures
Making backgroundDispatch a required descriptor field broke role:runtime fixtures in capability-manifest-version/capability-registry/host-integration-descriptors tests that build a dispatch object without it (caught by full gsd-test, not scoped npm test). Backfill backgroundDispatch:false into the well-formed fixtures; the deliberately-malformed 'required-field' test fixture is left malformed by design.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1708): update fix-1521 dispatch-gating assertion to the FLATTEN gate
fix-1521 pinned the codex-specific run_in_background prose that #1708 graduated to the typed dispatch-should-flatten/FLATTEN gate. Update its assertions to verify FLATTEN=false gating (not a runtime name) + that the old RUNTIME===codex gate is gone. Caught by full gsd-test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1708): add changeset for typed dispatch-flatten
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1708): remove stray temp PR-body file
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1708): add issue ref to bug-853 allow-test-rule annotations
ADR-456 requires every allow-test-rule exemption to carry a see #NNN reference; the source-text-is-the-product annotations added when migrating the prose assertions lacked it (lint-tests CI gate). Add (see #1708).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1520): randomize mktemp temp paths on BSD/macOS (XXXXXX must be path-final)
BSD/macOS mktemp only substitutes the XXXXXX template when it is the final
path component. Templates like `...-XXXXXX.json` / `gsd-pr-body.XXXXXX.md`
return a LITERAL `XXXXXX` path (no randomization) on macOS, so concurrent
workflow runs collide on the same temp manifest/body file — one run can
overwrite or consume another's. Reproduced on macOS: the second call to the
suffixed template fails `mkstemp: File exists`.
Fix: use a suffixless `XXXXXX` template (so it IS the final component), then
rename to add the intended extension — portable across BSD + GNU userlands,
no GNU-only `--suffix` flag. Empty-file-then-write semantics are preserved at
every site.
Affected workflow temp files:
- execute-phase.md: gsd-worktree-wave-*.json (wave worktree manifest)
- quick.md: gsd-quick-worktree-*.json
- spec-phase.md: edge-probe-reqs-*.json
- ship.md: gsd-pr-body-*.md
- profile-user.md: gsd-profile-answers-*.json, gsd-profile-analysis-*.json
The execute-phase.md edit uses a compact intermediate var + trailing comment
to stay under the ADR-857 phase-6 size ceiling (93166); regenerated the
workflow size baseline accordingly. Validated on macOS: 20 concurrent calls
yield 20 unique randomized paths.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1520): add changeset fragment (Fixed)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1520): add fail-first workflow-prose guard for mktemp XXXXXX suffix
Repo-wide scan of gsd-core/workflows/**/*.md that fails on any mktemp
template whose XXXXXX run is followed by a filename suffix (the BSD/macOS
non-randomizing form). Fails on the six pre-fix instances and passes on
the fix, and locks the copy-paste-prone idiom out of future workflows.
Mirrors the bug-637 hardcoded-$HOME workflow guard.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1520): rename regression test to fix- prefix (regression-test-names lint)
New tests/bug-NNNN-*.test.cjs files are banned by the lint-regression-test-names
ratchet; use the fix- prefix (matches the fix-1445 precedent).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1520): add issue ref to allow-test-rule exemption (ADR-456 lint)
lint-allow-test-rule-refs requires every new `allow-test-rule:` comment to
carry a #NNN reference (don't allowlist). Add (#1520) to the source-text
exemption.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1520): abort touched mktemp chains on failure (|| exit 1)
Per review: the VAR=$(mktemp …) && mv … && VAR=… chains dropped the issue's
suggested failure guard. If mktemp fails, $VAR is empty and the subsequent
mv/write lands on an unintended relative path. Add `|| exit 1` to all six
touched chains so a mktemp failure aborts the snippet. Regenerated the
workflow size baseline for the slightly longer lines.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1520): rebase onto next — regen size baseline + describe rename
Resolve the workflow-size-baseline.json conflict from next advancing by
regenerating from the current workflow sizes. Also rename the test describe
from `bug #1520` to `#1520` (the file uses the fix- prefix) per review nit.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1520): harmonize profile-user mktemp to ${TMPDIR:-/tmp} (review nit)
The two profile-user.md temp sites this PR already rewrites kept a hardcoded
/tmp while the four sibling workflows use ${TMPDIR:-/tmp}. Harmonize for
consistency and macOS-correctness (some sandboxes have no writable /tmp).
Regenerated the workflow size baseline.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1520): regen size baseline after rebase onto next
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(#1516): expose adaptive model_profile in /gsd-new-project AI Models prompt
Both onboarding paths (Step 2a auto-mode + Step 5 interactive) enumerated only
4 profiles (Balanced/Quality/Budget/Inherit), omitting 'adaptive' even though the
model catalog (model-catalog.json profiles) and docs/CONFIGURATION.md register 5.
Mirrors the proven /gsd:settings two-question split (#3784): Q1 routes between
Adaptive/Standard-tier/Inherit; Q2 (conditional on Q1=Standard) picks
Quality/Balanced/Budget — keeping every AskUserQuestion within the 4-option cap.
Both config-new-project example payloads now list adaptive. Regression cases
folded into the owning tests/new-project-mvp-prompt.test.cjs (per the
lint-regression-test-names ban on new top-level bug-NNNN files): each AI Models
prompt makes adaptive reachable, all 5 profiles reachable, 4-option cap honored,
both example enums include adaptive, brace balance. Workflow size baseline bumped
(new-project.md 62324 -> 66138 bytes; still well under the XL hard cap).
* chore(#1516): backfill changeset pr ref to 1654
* fix: require fresh phase verification before transition
* no-mistakes(review): Fix canonical verification closeout gates
* no-mistakes(review): Fix verify-work frontmatter promotion command
* no-mistakes(review): Fix stale verification gates
* no-mistakes(review): Fix canonical verification routing gates
* no-mistakes(review): Fix verification dependency and runtime routing gates
* no-mistakes(review): Block stale verification bypasses
* fix: handle large init manager outputs in verification workflows
* chore: update changeset pr number
* fix(verify-work): use fresh verification.status for stale gate
The stale check after UAT used phase_completion.verification_status from
session-start INIT while human_needed promotion already queried fresh
verification.status. Align the stale gate with the canonical query so
mid-session verification refresh is not ignored.
* fix(init): skip roadmap-checked phases when selecting next_phase
Roadmap-only phases without a disk directory were still promoted to
next_phase when their checkbox was already checked. Exclude
checkboxComplete phases so progress routing does not point at work the
roadmap already marks done.
* fix: gaps_found not overridden by stale, transition uses canonical verification
- verification.cts: check gaps_found before stale so gap-closure routing
is not masked by a newer summary mtime
- phase.cts: remove redundant findStaleVerificationSummary — readVerificationStatus
already handles stale detection
- transition.md: replace raw grep on file content with verification.status query
to avoid false-positive blocks from body text matching
* ci: retrigger tests after rebase
* fix(transition): replace gsd_run advisory check with awk frontmatter extraction
The runtime launcher is not defined until the update_roadmap_and_state step
bash block (~line 165). The early verify_completion block used gsd_run to
query verification.status, which violated the runtime-launcher-parity test:
'preamble appears AFTER the first gsd_run reference'.
Replace the gsd_run call with an awk-based frontmatter extractor that reads
only the status: field between the two --- fences. This avoids both the
preamble-ordering constraint and the original false-positive grep bug where
body text like 'previous_status: gaps_found' would match a full-text regex.
The phase.complete gate at update_roadmap_and_state is the canonical
enforcement point; this early check is advisory only.
Also update workflow-size-baseline.json for the updated transition.md size.
Fixes: runtime-launcher-parity test (B)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: re-check verification under planning lock in phase complete
Move readVerificationStatus into withPlanningLock so stale verification
cannot slip through when a SUMMARY.md is written between the gate and
the roadmap/state mutation. Return the blocked status from the lock
callback and emit the error after release to avoid leaving .lock behind.
* fix(transition): gate on canonical verification.status including stale
Replace awk frontmatter read with verification.status query so transition
blocks when summaries are newer than VERIFICATION.md, matching phase.complete
and other workflows (autonomous, progress, verify-work).
* Fix workflow verification gates for yolo transition and stale routing
Require VERIFY_STATUS passed before yolo/interactive transition advance.
Route stale verification recovery to verify-work, matching canonical projection.
* fix(transition): use verification.status query for stale-aware advisory check
The awk-based check read raw frontmatter status: passed, which misses the
stale case where summaries are newer than the VERIFICATION.md file even
though the frontmatter still says passed. The stale status is computed from
file modification times, not stored in frontmatter.
Move the preamble to the verify_completion bash block (the first block with
a gsd_run call) so gsd_run query verification.status can be used for the
advisory check. This gives the full readVerificationStatus logic including
mtime-based staleness detection, matching the enforcement gate at phase.complete.
Capture full JSON (VERIFY_JSON) so next_action can be included in the
advisory output alongside the status.
Also update workflow-size-baseline.json for the updated transition.md size.
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* ci: trigger test matrix for 525b946
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix(transition): restore awk frontmatter extraction for pre-shim verification check
The gsd_run launcher shim is not defined until line ~163 of transition.md,
so the verification debt check at line ~80 cannot use gsd_run. Restore the
awk-based frontmatter extraction that correctly reads status without needing
the runtime, and restore the shim at its proper location before
phase.complete.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(#1522): clarify transition verification gate wording
* fix(#1522): update transition workflow size baseline
* fix(#1522): update workflow-size-baseline after rebase onto next
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix(#1522): guard findStaleVerificationSummary FS calls + thread opts.fs seam (review)
Address review blocker B1 on #1548: findStaleVerificationSummary ran fs.readdirSync
and two fs.statSync calls unguarded between readVerificationStatus's try/catch sections,
so a TOCTOU race (a SUMMARY listed by scanPhasePlans then removed before statSync) or any
FS error threw uncaught into callers NOT under the planning lock (init.manager /
init.progress / uat-predicate). Wrap the body in try/catch degrading to 'not stale', and
thread the injectable opts.fs seam (add statSync to FsLike, pass fsImpl from the caller)
for parity with readVerificationStatus's no-throw contract and testability. Also adds the
Verification Module glossary entry to CONTEXT.md (review B3).
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
workflow.security_block_on was documented as the minimum threat severity
that blocks advancement, but threats carried no severity and the auditor's
threats_open count (the SECURITY.md gate field) counted every open threat
regardless of severity — so the threshold had no effect, and the auditor's
block_on vocabulary (open/unregistered/none) did not even match the config
enum (critical/high/medium/low/none).
- planner: add a Severity column to the STRIDE threat register; assign
severity per threat.
- auditor: read severity; reconcile the <config> block_on domain to the
severity enum; redefine threats_open as the count of OPEN threats whose
severity is at or above block_on (none => 0). Below-threshold opens are
reported as non-blocking and excluded from threats_open.
- SECURITY.md template + planning-config.md reconciled.
No gate-check site changed: threats_open == 0 stays the gate everywhere;
only its computation is now severity-filtered.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
workflow.security_asvs_level was display-only — the planner hardcoded
'mitigate if ASVS L1 requires it' and the auditor only echoed the level,
so L2/L3 behaved identically to L1.
- New reference gsd-core/references/security-asvs-levels.md defines L1
(opportunistic), L2 (standard), L3 (comprehensive) for both planner
threat disposition and auditor verification depth (higher = superset).
- planner: disposition now scales with the configured ASVS level (no
hardcoded L1) + @-pointer to the reference.
- auditor: verification depth scales with asvs_level (L1 grep-presence,
L2 boundary/vector check, L3 end-to-end trace + bypass check).
- planning-config.md + INVENTORY updated; planner kept under its 48K cap
by extracting the goal-backward worked example to planner-guidance.md.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Add an optional structured `coverage:` block to SUMMARY.md frontmatter and a
deterministic classifier that `verify-work` consumes to route deliverables to
auto-pass vs human-UAT — replacing the rejected #1598/#1599 post-hoc heuristic.
- New `src/coverage.cts` (→ bin/lib/coverage.cjs) parses the nested coverage
block (extractFrontmatter can't — its `-` items are scalars-only; this is a
focused parser, sibling of parseMustHavesBlock), validates each entry, and
classifies into auto_passed vs present. Frozen MODE/PRESENT_REASON/ERROR_CODE
typed-IR surface. Exposed via `uat classify-coverage --summary <f>`.
- Auto-pass is the narrow proven case only: strict-boolean human_judgment:false
AND non-empty all-`pass` verification AND zero validation errors. Everything
else — judgment, empty/failing verification, malformed entry — routes to the
human (fail-safe). A malformed block falls back to legacy prose extraction and
surfaces an error; an absent block is byte-identical to pre-#1602.
- execute-plan create_summary populates the block (fail-safe default
human_judgment:true); verify-work extract_tests consumes it; create_uat_file
marks auto-passed entries `source: automated`.
- Templates (summary + 3 variants), CONTEXT.md predicate + glossary, INVENTORY,
eslint/gitignore registration, and Diataxis docs (COMMANDS reference +
USER-GUIDE explanation) updated.
- Behavioral tests via the CLI (no source-grep); parser-robustness regressions
for the null-entry/comment-header/mis-indent cases found in adversarial review.
Closes#1602
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add a non-blocking, warn-only codebase-drift gate at plan:pre so a stale
STRUCTURE.md is surfaced before /gsd:plan-phase spawns the planner, instead
of being discovered mid-execution by the existing execute:wave:post gate.
Gated on a dedicated workflow.plan_drift_precheck toggle (default true),
independent of schema_drift_gate. Never blocks planning, never spawns the
mapper agent at plan time.
Review feedback (#1595):
- Use a documented conventional-commit type (feat, not enhance) per
CONTRIBUTING.md / gsd-validate-commit.sh.
- Normalize the plan_drift_precheck command references to the colon prose
form (/gsd:plan-phase, /gsd:map-codebase) to match plan-phase.md §5.65;
registry regenerated from capability.json.
- Make the test temp dirs hermetic: drain mkdtemp dirs in an after() hook
via the helpers.cleanup() budget (local/no-raw-rmsync-in-tests-compliant).
Closes#1592
Claude-Session: https://claude.ai/code/session_016JBiXEAofvB3prJim29XMS
GitHub Copilot reads repository-wide instructions only from
.github/copilot-instructions.md (confirmed via GitHub Docs), not a root
copilot-instructions.md. Aligns getProjectInstructionFile with the installer
(runtime-config-adapter-registry installSurface 'copilot-instructions') and
cites the docs source in the doc-comment.
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit
Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed),
enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way
to browse or audit parked seeds on demand. This adds a read-only listing,
following the established --list → workflow pattern (per the approved scope on
- gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the
seeds dir, returns { count, seeds[], summary } JSON with each seed's id,
slug, status, scope, trigger_when, planted, title. Optional case-insensitive
status filter. User-controlled content is sanitized (sanitizeForDisplay) and
every path validated (requireSafePath); read-only. Independent of
audit.scanSeeds, which only returns unimplemented seeds for the milestone surface.
- /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that
renders the seed table.
Closes#441
* chore(#441): point changeset fragment at PR #722
* test(#441): allowlist list-seeds test in prompt-injection scan
The test asserts that list-seeds neutralizes injection payloads
(<system>, [INST]) embedded in seed content, so the fixtures legitimately
contain those patterns — same as the sibling security tests already on the
allowlist.
* fix(#441): use canonical /gsd:capture colon form in list-seeds workflow
Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use
the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is
retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new
list-seeds workflow used the hyphen form.
* docs(#441): sync help full.md + INVENTORY for --list-seeds
Adds the --list-seeds entry to the help reference (help/modes/full.md, per
bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow
in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json.
* docs(#441): add --list-seeds how-to + drop phantom statuses
Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers):
- USER-GUIDE.md Seeds section (how-to): extend the task to cover
auditing parked seeds on demand via --list-seeds, including the
status filter — kept task-oriented per Diataxis how-to mode.
- CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected
from the list-seeds filter vocabulary; the system only produces
dormant|active|triggered (src/audit.cts scanSeeds). Reference must
be factually accurate and complete.
* fix(#441): guard non-scalar status frontmatter in cmdListSeeds
A seed with a bare `status:` line (extractFrontmatter yields {}) or a
`status: [a, b]` value (yields an array) crashed the whole audit list:
`(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string.
Coerce every frontmatter read through a `fmStr` helper (mirrors the existing
`typeof fm.id === 'string'` guard), so a non-scalar status falls back to
dormant and non-scalar scope/trigger_when/title can no longer leak a raw
array/object into the JSON contract. Title is now capped symmetrically.
Adds regression coverage for empty and array `status:` and non-scalar fields.
Refs #441
* docs(#441): align list-seeds workflow status vocabulary
The load_seeds step listed `implemented` as an example status filter, but the
real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds);
`implemented` has no producer. Matches the earlier CLI-TOOLS.md correction.
Refs #441
* refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds
Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported
deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested
in-process (review minor #1). No behavior change.
Filter comparison now matches the raw lowercased status (both sides already
normalized) instead of sanitizeForDisplay(status); sanitization is for output,
not matching (review nit #3).
* test(#441): add fast-check property coverage and count=1 boundary for list-seeds
Adds tests/list-seeds.property.test.cjs with four fast-check properties over
deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug
invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing
(review minor #1).
Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2).
* chore(#441): sync runtime launcher snippet into list-seeds workflow
Propagate the current _runtime-launcher.snippet.sh (with non-Claude
runtime home probes) into the new list-seeds.md workflow via
scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation.
* test(#441): record list-seeds.md in workflow size baseline (#1074)
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* feat(#1318): require external reviewers to verify plan claims against source
/gsd-review built its external-reviewer prompt from plan text only and never
asked reviewers to open the repo and verify claims, so a grounded HIGH could be
outvoted by ungrounded LOWs. Add a concise, generic source-grounding block to
build_prompt's Review Instructions: treat yourself as running in the working
tree, open referenced files, cite path:line + mechanism, trace asserted
mechanisms, downgrade to an open question if you have no file access, and know
that grounded findings are weighted more heavily.
Also clarify that CodeRabbit (a diff-only reviewer that never receives the
prompt) must not be weighted as a grounded plan-level verdict in consensus
synthesis. Workflow stays under its size cap (baseline bumped deliberately).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1318): add changeset for reviewer source-grounding
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1318): mark changeset docs-exempt (internal reviewer-prompt wording)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(#1318): document reviewer source-grounding in COMMANDS.md; drop docs-exempt
Review: a user-visible behavioral Changed warrants a docs touch, not a
docs-exempt. Add a sentence to the /gsd-review entry in docs/COMMANDS.md
(reviewers verify against source, cite file:line, grounded findings weighted
higher) and remove the changeset docs-exempt marker so lint:docs passes via
docs-updated. Also note the literal build_prompt test anchor.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1318): harden build_prompt fence extraction to be fence-run-aware
Addresses maintainer review on PR #1421 (required-before-merge).
The buildPromptReviewInstructions() test helper located the closing fence with
`src.indexOf('\n```')`, which terminates at the FIRST triple-backtick line — so
a build_prompt ```markdown block whose body embeds a fenced code example would
truncate mid-content (dropping the `## Review Instructions` section) and give a
spurious failure or false pass. Since this feature feeds source/plan content
(which routinely contains code fences) to reviewers, that is a live fragility.
Rewrite the extraction to be fence-run-aware, mirroring the CommonMark close
rule in src/markdown-sectionizer.cts stripFencedCode: parse the opener's
backtick run length, then close on the first line with >= that many backticks
and only trailing whitespace — so a shorter nested fence is treated as content.
Add a fail-first regression test (a 4-backtick outer fence wrapping a nested
```bash block) asserting the trailing `## Review Instructions` still extracts.
Test-only change; no production .cts touched. Verified: test file 7/7,
empirical fail-first proof the old indexOf logic truncated, full suite
4236/4236, eslint clean. Codex review: approve.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(pr-branch): handle sub_repos from config with git -C (#666)
Adds a `handle_sub_repos` step between `detect_state` and
`analyze_commits`. When `planning.sub_repos` is set in config, the
workflow now:
- Reads sub-repo paths via `gsd_run query config-get sub_repos`
- Skips the step entirely when the list is empty/null/[]
- Scans each repo with `git -C "$REPO" status --porcelain`
- Offers the user all/select/skip choices
- For selected repos: creates a PR branch, commits all staged/unstaged
changes, pushes, and opens a companion PR via `gh pr create`
All git commands use `git -C "$REPO"` — never `cd "$REPO"` — because
shell state does not persist between agent-executed commands.
Closes#666
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: update changeset pr number to 667
* fix(pr-branch): address maintainer review — correct seam, behavioral tests, robustness
Resolves all three blockers and seven robustness issues raised in PR #667 review:
Blockers:
- Use `planning.sub_repos` (not top-level `sub_repos`) so config-get actually resolves
- Replace prose grep test with behavioral fixture tests using runGsdTools + local bare repo
- Extract sub-repo git work into new `cmdPrSubrepo` seam in src/commands.cts;
never uses git add -A — stages explicit files only (universal-anti-patterns.md:44)
Robustness:
- Dirty-repo list persisted via mktemp/cat, not bash arrays (cross-block safe)
- Branch name embeds repo slug (${CURRENT_BRANCH}-${REPO_SAFE}-pr) to avoid collision
- push --set-upstream so gh pr create finds the branch
- Sub-repo base branch resolved via ls-remote with fallback to repo's default branch
- Remote slug parsed with /github\.com[:/]/ (handles SSH + HTTPS + .git-less URLs)
- rollback() cleans up branch on any mid-sequence failure
- node -e replaces jq (always available, no undeclared hard dep)
Refs: #666
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(pr-branch): security guard, push timeout, rollback fix, porcelain fix
Security (Blocker 1):
- Use security.cjs validatePath() in cmdPrSubrepo for symlink-safe workspace
containment check — rejects ../escape, absolute paths, and symlink traversal
- Add negative regression test: '../escape' repo path must be rejected
Robustness:
- Push uses timeout: 60_000 ms (network op needs more than the 10 s default)
- Capture prevBranchName before checkout -b so rollback uses explicit name
instead of git checkout - (fails on fresh single-branch repos)
- Porcelain path parse: line.trimStart().slice(2).trim() handles all XY
combinations and the execGit global-trim edge case uniformly
Tests: 17/17 pass, lint: 0 errors
Refs: #666
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(pr-branch): move regression tests to commands.test.cjs, add core.quotePath=false
- Move cmdPrSubrepo behavioral + workflow source-invariant tests from
standalone bug-666-*.test.cjs into tests/commands.test.cjs under
describe('pr-subrepo') per TESTING-SUITES.md policy (no new bug-* files).
Adds allow-test-rule: source-text-is-the-product see #666 for the
workflow-source-invariant suite.
- Add -c core.quotePath=false to git status --porcelain call so non-ASCII
filenames (e.g. café) are not C-escaped, keeping slice(2) parse correct.
* fix(pr-branch): remove obsolete regression tests for sub-repos handling
* fix(pr-branch): update workflow-size-baseline, add dirty-scan timeout
- Regenerate tests/workflow-size-baseline.json for pr-branch.md growth
(+handle_sub_repos step, +timeout addition).
- Add { timeout: 10_000 } to the execFileSync git status --porcelain
call in the handle_sub_repos dirty-scan (repo convention: every git
subprocess is bounded, never hangs).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: regenerate INVENTORY-MANIFEST after rebase onto next
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#666): handle rename staging and split changedFiles from filesToStage
For git mv renames, the old path no longer exists in the worktree after
the move — staging it with git add fails. Split parsing into changedFiles
(both paths, for result.files) and filesToStage (new path only for
renames; old is already staged by git mv). Also adds porcelain tests
for staged renames, non-ASCII filenames, and a fast-check property test.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#666): rollback on push failure in cmdPrSubrepo
If push fails the branch only exists locally; rollback cleans it up so
the sub-repo is not left in a half-committed state.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#666): do not rollback after commit on push failure; add push-fail regression test
Post-commit push failures are network/auth/policy issues — the user's work
is already committed on the local branch. Calling rollback() at that point
force-deletes the only ref holding the commit (data loss). Leave the branch
in place and emit a retry instruction instead.
Adds a regression test (pre-receive hook that rejects all pushes) asserting
the branch and commit survive a push rejection so the failure path stays
covered going forward.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: regenerate INVENTORY-MANIFEST after rebase onto next
Rebased onto current next (#1267 retired core.cjs). Stale tsbuildinfo and
a leftover bin/lib/core.cjs build artifact were masking the drift — wiped
both, rebuilt clean, and regenerated the manifest. gen-inventory-manifest
--check now exits 0.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#666): validate sub-repo paths before git invocation in pr-branch.md
The handle_sub_repos workflow ran git -C on raw planning.sub_repos config
values at two points before the pr-subrepo seam's validatePath guard ever
ran: the dirty-scan detection (git status) and the base-branch resolution
(git ls-remote / remote show). A traversal entry could point git outside
the workspace; an embedded newline could inject a spurious record into
the newline-joined dirty-file output and into the shell-interpolated
commit message.
Adds a containment check + character allowlist to the dirty-scan node
script (reject before any execFileSync), and a defense-in-depth shell
case guard on the same value before the second, independent git -C
invocation in the base-branch resolution block.
Adds a behavioral test that extracts and executes the actual shipped
node script from pr-branch.md (not a mirror) against a real traversal
target and an embedded-newline entry, asserting neither reaches git or
the dirty-file output.
Also updates the stale cmdPrSubrepo doc comment: push failures no longer
delete the branch (see prior commit), only stage/commit failures do.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#666): make sub-repo traversal scan test genuinely fail-first
The outside repo's only change was an untracked file, which the ?? filter
excludes — so the repo looked clean even with the guard removed, making the
traversal assertion vacuous (it passed against a neutered guard). Commit the
file first, then modify it, so the outside repo has a tracked dirty change:
without the path guard it WOULD be reported dirty, so the test now fails-first.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#666): symlink-safe (realpath) sub-repo containment in pr-branch.md
Finding A from re-review: the workflow guard used path.resolve, which only
normalizes '..' textually and does not follow symlinks — so an in-tree symlink
whose name has no '..' or '/' (e.g. "evil" -> /outside) passed both the charset
filter and the resolve+startsWith check, letting git status / ls-remote /
remote show run against a directory outside the workspace. The pr-subrepo seam
already used fs.realpathSync (validatePath); this brings the workflow layer to
parity.
- dirty-scan: realpathSync the root once, and realpathSync each candidate before
the containment check; skip on throw.
- base-branch resolution: replace the weak `case *..*|/*` guard with a realpath
containment check that yields a validated absolute SUB_REPO_DIR, and run git -C
against that instead of re-concatenating $ROOT/$REPO_REL.
- security test: add a symlink-escape entry and a positive control (legit in-root
backend must still be reported). Confirmed fails-first — regressing the scan to
path.resolve makes the symlink case leak.
Also fixes a misleading-fallback minor: the workflow now checks the seam's exit
status and skips the companion-PR step on failure, instead of printing
"branch pushed, open PR manually" after a real stage/commit/push failure.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#666): harden pr-branch sub-repo flow against round-12 edge cases
Pre-emptive hardening of the workflow changes from the symlink fix:
- continue-outside-loop: the "skip companion PR on seam failure" block used a
bash `continue`, but the per-sub-repo iteration is prose-driven (the agent
loops, not a literal `for`), so `continue` would warn and no-op. Reframed as
prose-gated control flow keyed on $SUBREPO_EXIT — no bash loop assumption.
- Windows portability: the new symlink security case now degrades gracefully
(try/catch around fs.symlinkSync; skip just the symlink assertion when symlink
creation lacks privileges) so it doesn't hard-fail on Windows CI.
Verified: seam exits 1 on error / 0 on success (error() → process.exit(1),
propagated through the shim), so the $SUBREPO_EXIT check is meaningful; bash -n
clean on the touched blocks; commands 156/156; lint:ci green; manifest in sync.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs
Generalizes the Codex-only #1515/#1519 fix to every non-Claude runtime, and
wires it into the real install path (where it was previously dead-on-arrival).
Root causes:
1. The runtime-default stamping lived only in `_applyRuntimeRewrites`, but the
installer emits `gsd-core/workflows/*.md` via `copyWithPathReplacement`, which
never calls it — so a real `--codex`/`--cursor`/etc. install emitted
`--default claude` and worktrees-on. RUNTIME mis-resolved to claude and the
workflow ran executors unisolated against the main checkout. (#1515/#1519 were
also dead-on-arrival in real installs; this repairs them.)
2. Only `case 'codex'` was stamped; every other non-Claude runtime kept the
Claude default.
Fix:
- New `_stampNonClaudeRuntimeDefaults(content, runtime)` (single shared helper)
stamps `--default <runtime>` + `use_worktrees=false` for every `runtime !=
claude`; called from both `_applyRuntimeRewrites` and, crucially,
`copyWithPathReplacement` in bin/install.js (the real workflow emit path).
- Generalize the fail-closed worktree guard `= codex` -> `!= claude` in
execute-phase/quick/diagnose-issues (worktree isolation is Claude-Code-only).
- Flip manager/autonomous inline-vs-background gating to `codex -> background,
everything-else -> inline` (research: only Codex can background-nest the
pipeline's subagents; all others run inline, which they support).
Worktree-capability determination is research-backed (official docs for all 14
non-Claude runtimes: none honor GSD's isolation="worktree" mechanism, only Codex
background-nests). New end-to-end real-install test asserts the EMITTED workflow
is stamped — the regression guard that would have caught the dead-on-arrival bug.
Closes#1521
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r
* chore(#1521): backfill changeset PR number (#1537)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees
A Codex install with a runtime-neutral .planning/config.json resolved
RUNTIME=claude and enabled git worktree isolation, which Codex's
spawn_agent cannot honor. Two root causes:
1. Workflows read `config-get runtime` / `config-get workflow.use_worktrees`
without `--raw`, so config-get's JSON-quoted output ("codex") was captured
verbatim into the bash var and broke every `[ "$RUNTIME" = ... ]` check —
the Codex fail-closed guard was dead even when runtime:codex was explicit,
and Claude's own worktree degrade-check was dead too. Add `--raw` to those
reads across execute-phase, autonomous, manager, diagnose-issues, quick.
2. The conversion engine emitted `--default claude` for every runtime. Stamp
the codex-emitted workflows to `--default codex` (runtime) and
`--default false` (use_worktrees) in _applyRuntimeRewrites case 'codex', so
a neutral config on a Codex install resolves runtime=codex / worktrees off.
Also extend the Codex fail-closed worktree guard to quick.md and
diagnose-issues.md (they spawned isolation="worktree" with no runtime guard).
Regression test asserts source<->engine parity across all five workflows
(DEFECT.GENERATIVE-FIX) plus fast-check property coverage of the stamping.
Closes#1515
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r
* chore(#1515): backfill changeset PR number (#1519)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The #1279 node-test machine-proof confirmed a known-bad subject drives the
negative test RED, but could not distinguish a genuine content-violation from a
deceptive test that reds merely because GSD_PROHIB_SUBJECT is set. Add an
optional fifth flat scalar `check_clean_fixture` (-> CheckDescriptor.cleanFixture)
threading a KNOWN-CLEAN control subject through projectProhibitions +
descriptorFromProjection. When present, the prover also runs the check against
the clean subject and requires GREEN, so fail-first is proven only when the check
is RED on the violation AND GREEN on the clean subject (content-dependent).
Opt-in and additive: absent a clean fixture the prover behaves exactly as
post-#1314 (no control, documented residual), preserving the zero-authoring
compose path; the lint-rule kind needs no analog (its subject IS the linted
file, no env indirection). Coverage: RED-first deceptive case, positive,
missing-clean fail-closed, round-trip read-back/emit, fast-check property
extended to the 5th scalar, and an end-to-end COMPOSE capstone (honest vs
deceptive). Docs: ADR-550 dated addendum, prohibition-probe reference,
spec-phase + verify-phase workflows.
Closes#1346
Claude-Session: https://claude.ai/code/session_01GsPRb8zvpcT7Eat6vZw8PX
Proactive checkpoint guard fires at each wave boundary before spawning agents.
Self-assesses context pressure against context-budget.md degradation tiers and
warns (warn, default) or auto-invokes /gsd:pause-work (auto) when POOR tier
(70%+) is detected. Config key validated; defaults to \"warn\".
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1369): refresh wave manifest and re-check base before each wave in execute-phase
Two compounding issues caused wave N+1 worktrees to fork from the stale
pre-wave-N commit, immediately tripping the worktree_branch_check FATAL
guard in every executor:
1. worktree.base-check auto-degrade only ran once at initialize time.
After wave N merges advanced orchestrator HEAD past origin/HEAD, new
worktrees were still forked from origin/HEAD (Claude Code "fresh" base).
2. WAVE_WORKTREE_MANIFEST was never unset between waves, so wave N+1
reused the consumed wave-N manifest file, which would have blocked the
step 5.5 manifest guard (#3384) on subsequent waves.
Fix: add two safeguards in execute-phase.md —
- Step 0.5 (start of each wave): re-runs worktree.base-check; auto-degrades
USE_WORKTREES=false for that wave when HEAD has diverged from origin/HEAD.
- Step 7c (end of each wave): unsets WAVE_WORKTREE_MANIFEST so wave N+1
creates a fresh per-wave manifest; re-asserts worktree.set-baseref
(idempotent) and re-evaluates base degradation after wave merges land.
17 regression tests added in tests/bug-1369-wave-stale-base.test.cjs.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1369): rename test to fix-NNN convention; update workflow size baseline
Rename tests/bug-1369-wave-stale-base.test.cjs → tests/fix-1369-wave-stale-base.test.cjs
to satisfy the lint-regression-test-names gate (new files cannot use bug-NNN prefix).
Update tests/workflow-size-baseline.json for execute-phase.md: 93157 → 97393
(LF-normalized byte count after adding step 0.5 inter-wave base re-check and
step 7c between-wave manifest reset).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1369): add issue reference to allow-test-rule comment
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1369): extract new execute-phase steps to references; satisfy ADR-857 cap
Step 0.5 (inter-wave worktree base re-check) and steps 7b–7c (pre-wave
dependency check + between-wave manifest reset/base refresh) added by this
PR grew execute-phase.md to 97393 bytes, violating the ADR-857 phase-6
architectural mandate that host-loop bodies remain strictly below the
pre-phase-6 baseline of 93166 bytes.
Extract both new blocks into dedicated reference files:
- gsd-core/references/execute-phase-wave-guard.md (step 0.5)
- gsd-core/references/execute-phase-between-wave-reset.md (steps 7b + 7c)
Replace inline prose with @-reference pointers. File now measures 92851
bytes (LF-normalized), satisfying the ADR-857 capstone conformance gate.
Also update tests/workflow-size-baseline.json to 92851 and add both new
reference files to docs/INVENTORY-MANIFEST.json.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1369): update regression tests to read from extracted reference files
Steps 0.5 and 7b+7c were moved to reference files to satisfy the ADR-857
size cap on execute-phase.md. Tests now check @-reference pointers in the
workflow for ordering and read content assertions from the reference files.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
#2551/#3182/#2361 are pre-migration get-shit-done-redux issue numbers with no
equivalent in open-gsd/gsd-core; they mislead triage and manufacture phantom
blockers. Repoint to real successors (#717 byte-budget rework, #720) or rewrite
as prose referencing the discuss-phase/modes progressive-disclosure split.
Correct co-located 'line budget'/'<500 lines' framing to the byte-based reality
(#717). Add a CI guard (tests/no-phantom-issue-refs.test.cjs) that fails if a
phantom ref is reintroduced. SSH-key patterns (id_ed25519) left untouched.
No user-facing runtime behavior change.
Closes#1073
* feat(#1355): detect-and-warn guard for claude-code agent-teams
GSD's multi-agent orchestration can stall under claude-code's experimental
agent-teams (a subagent's completion fails to route to the orchestrator). Per
the maintainer decision, the accepted scope is a read-only detector + one
non-fatal warning — NOT the declined run_in_background/TaskOutput conversion.
- New Teams Status Module (src/teams-status.cts → gsd-core/bin/lib/teams-status.cjs):
pure resolveTeamsStatus({runtime, env}) + thin CLI cmdTeamsStatus reusing
resolveRuntime. active = strictly-truthy env flag AND runtime === 'claude'.
- Wire `gsd-tools query teams-status [--active]` (read-only; no capability
registration needed — conformance gates govern features, not query commands).
- One non-fatal warning in plan-phase.md before the first Agent spawn, gated on
`query teams-status --active`; zero behavior change on non-claude/teams-off.
- Hermeticity: clear CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS in run-tests.cjs +
SESSION_ENV_KEYS. Docs reference + CONTEXT.md glossary. Built lib gitignored.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1355): add changeset for teams-detect guard
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1355): bump plan-phase.md workflow size baseline (+407B for teams warning)
The non-fatal agent-teams warning block added to plan-phase.md grew it
92759 → 93166 bytes, past its committed per-file baseline ratchet. The growth
is small, deliberate, and still well under the workflow tier hard cap. Regenerate
the baseline via `npm run size:baseline` (only plan-phase.md changed).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1355): register teams-status.cjs in the inventory manifest
The new teams-status CLI module is a tracked surface; regenerate
docs/INVENTORY-MANIFEST.json (cli_modules family) via
gen-inventory-manifest.cjs --write so the inventory-manifest-sync gate passes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1359): migrate workflows off deprecated TaskOutput to Read(outputFile)
The map-codebase and docs-update workflows collected background sub-agent
results with the deprecated Claude Code `TaskOutput` tool using `block: true`,
which has a confirmed main-session hang after the agent completes
(anthropics/claude-code#20236).
Migrate the collection steps to the upstream-recommended pattern: keep
`run_in_background=true` on the Agent spawn, then `Read` each agent's
`outputFile` (from the `async_launched` result) once it reports completion.
Completion-marker contracts and on-disk verification are unchanged, and the
non-Claude runtime fallbacks (sequential_mapping / sequential_generation) are
preserved byte-for-byte. docs-update's timeout note no longer references the
unwired `workflow.subagent_timeout` key (it kept a literal before).
Regression coverage folded into tests/subagent-timeout.test.cjs (the owning
module for background-subagent collection). Workflow size baseline regenerated
for the justified prose growth.
Refs #1355 (same latent hang surface).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1359): add changeset for TaskOutput migration
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Delivers option (a) from the #1314 maintainer review: thread a fourth flat
scalar check_violation_fixture through the projection so a prohibition authored
at spec-phase machine-proves fail-first and greens through the deterministic
path alone — zero hand-authoring at verify time.
- src/probe-core.cts: Prohibition gains check_violation_fixture?; projectProhibitions
emits it (both kinds) ONLY for a well-formed descriptor and ONLY when non-empty
(blank/absent -> projects absent -> producer hard-gates, never a partial green).
- src/prohibition-enforcement.cts: descriptorFromProjection reads it back into
violationFixture via the same numeric-coercion-safe scalar() normalizer.
- Tests (RED-first, proven non-vacuous by reverting both src edits): CHK-02(#1346)
projection emit, CHK-08(#1346) read-back, CHK-03(D) example round-trip, the
fast-check round-trip property extended to the 4th scalar (the contract trek-e
blocked #1301 on), and a real-subprocess COMPOSE capstone greening end-to-end
through project -> descriptorFromProjection -> default prover+runner.
- Docs flipped from 'hard-gates until #1346' to 'composes end-to-end': verify-phase.md,
prohibition-probe.md, spec-phase.md authoring, ADR-550 addendum, changeset.
#1346 now tracks only the node-test causation residual.
190 affected-suite tests green; eslint + tsc clean; size baseline regenerated.
Addresses the #1314 maintainer review (trek-e):
- Major 1 (fail-OPEN): defaultProveFailFirst's node-test branch only guarded
`if (!fixture)`. A missing/typo'd/stale violationFixture made GSD_PROHIB_SUBJECT
point at a missing file; an honest negative test threw ENOENT *inside its
callback* (a failing test named distinctly from the file), which
isNonVacuousNodeTestRed accepted as proof -> a green forged from a setup crash.
Now requires fs.existsSync(path.resolve(cwd, fixture)) before spawning, symmetric
with the lint-rule path's file-result guard. Regression test pins it (RED without
the guard); a second test pins cwd-relative fixture resolution.
- Major 2 (misleading prose): the #1278 projection carries no violationFixture, so
the deterministic-locate path always hard-gates (fail-closed) until a
check_violation_fixture scalar is threaded through. verify-phase.md and
prohibition-probe.md no longer read as if the projected path produces greens;
the ADR-550 addendum records both items. Tracked as follow-up #1346.
- Documented residual: existence is necessary but not sufficient (a red caused by
the env being set vs the subject's content); recorded as a constraint, in #1346.
- Nit: stale 'NOT attested fail-first' comment -> 'NOT machine-proven fail-first'.
64 tests pass; eslint + tsc clean; changeset valid.
* test(1278): RED-first descriptor parity + fail-closed guards + CHK-07 byte-stability (wave 1)
- CHK-03 (RED): extend PROB-14 parity in prohibition-probe.schema.test.cjs to carry the flat
check_kind/check_target/check_rule scalars through project->write->parseMustHavesBlock; the
non-droppable check_kind-presence assertion is the load-bearing RED trigger (fails because
projectProhibitions strips check_* on the current build).
- CHK-07 (GREEN forward-guard): probe-core.test.cjs pins descriptor-less byte-stability +
dispositionForProhibition fail-closed policy, with a t.todo marker forward-locking plan 01-02.
- CHK-06 (RED): prohibition-enforcement.test.cjs asserts descriptorFromProjection export +
fail-closed on absent/partial/unknown descriptors via the projection adapter (RED until 01-03).
- No src/*.cts or .cjs edits; no new test files; lint-test-file-count clean.
* feat(1278): add optional flat-scalar check descriptor fields to Prohibition interface (wave 2)
- check_kind?/check_target?/check_rule? mirror CheckDescriptor.kind/target/rule (minus caller-attested failFirst, #1279)
- optional so existing Prohibition consumers compile unchanged
* feat(1278): project check descriptor as flat scalars in projectProhibitions (wave 2)
- emit check_kind/check_target (+ check_rule only for lint-rule with a rule) when descriptor well-formed
- under-specified/descriptor-less items project byte-identically (CHK-07); flat scalars ride existing parseMustHavesBlock continuation-KV path (no parser rewrite)
- add CHK-02 probe-core unit cases pinning the projection
- turns CHK-03 parity test GREEN; dispositionForProhibition untouched
* feat(1278): descriptorFromProjection read-back adapter feeds fail-closed locate (wave 3)
- Add descriptorFromProjection(projected) -> CheckDescriptor | null to
src/prohibition-enforcement.cts: renames the projected flat scalars
check_kind/check_target/check_rule -> {kind,target,rule?}, or null when
the descriptor is absent/non-object (no check_kind key).
- failFirst is NEVER sourced from the projection (stays caller-attested; #1279).
- rule is set only when check_rule is a non-empty string; the adapter does NOT
re-validate kind/target/rule — an under-specified descriptor reconstructs to
one the EXISTING runProhibitionEnforcement LOCATE guard rejects (located:false,
never green). The merged #1259 guard stays the single source of fail-closed truth.
- Turns the RED CHK-06 fail-closed tests (plan 01-01) GREEN end-to-end; CHK-03 /
CHK-07 stay green. CheckDescriptor type, locate guard, dispositionForProhibition,
and parseMustHavesBlock are unchanged (additive +36/-0).
* feat(1278): verify-phase locates prohibition check from projected descriptor (wave 3)
- request.check kind/target/rule sourced from projected check_kind/check_target/check_rule via descriptorFromProjection, not verifier invention (CHK-05)
- replaces the #1278 author-supplied / tracked-follow-up note with the delivered deterministic-locate behavior
- preserves fail-closed routing: absent/partial descriptor -> never green, hard-gate in both modes
- failFirst stays a verify-time caller attestation; #1279 bounds the remaining fail-first proof
* feat(1278): spec-phase captures wired-check descriptor on test-tier resolution (wave 3)
- Step 5.6 'Keep it' / verification: test path captures check_kind/check_target/check_rule, projected onto must_haves.prohibitions for verify-phase deterministic locate (CHK-04)
- SOFT capture: a test-tier prohibition without a descriptor is still allowed (no hard authoring block); stays fail-closed/flagged downstream
- --auto captures only an unambiguous descriptor, never fabricates a check path
- failFirst NOT captured at spec-phase (verify-time attestation; #1279)
- PROB-06 soft-gate + text-mode (PROB-09) behavior unchanged
* chore(1278): re-baseline workflow size for grown verify-phase + spec-phase prose (wave 3)
- spec-phase.md 28438 -> 30343 (+1905), verify-phase.md 35362 -> 36498 (+1136)
- regenerated via npm run size:baseline (no hand-picked numbers); growth is the #1278 deterministic-locate + descriptor-capture prose
- workflow-size-budget guard green (122/122)
* docs(1278): ratify optional check descriptor in dated ADR-550 addendum + type:Changed changeset
- Append dated 2026-06-15 ADR-550 addendum ratifying the D3 prohibition-item
shape extension (optional flat-scalar check_kind/check_target/check_rule)
- Document flat-scalar rationale, deterministic projection/read-back,
fail-closed on partial/invalid/absent, #1279/policy out-of-scope
- Add .changeset/1278-prohibition-check-descriptor.md (type: Changed)
* docs(1278): document optional check descriptor in prohibition-probe reference + FEATURES
- Add 'Optional wired-check descriptor (deterministic locate, #1278)' section
to the prohibition-probe reference (flat-scalar keys, projection/read-back,
fail-closed + backward-compat, failFirst stays attested)
- Add deterministic prohibition-check descriptor source entry to FEATURES.md
- No CONTEXT.md glossary change: descriptor reuses existing wired-check /
verification:test vocabulary, no new glossary term introduced
* fix(1278): pass packaging gates — changeset pr field + retired slash-form fix
- Add required pr: 1278 to changeset (lint:changeset MISSING_PR hard requirement;
plan's 'omit if unknown' was inaccurate — issue number per #1259 convention,
updated to real PR number when opened) [Rule 3 - blocking]
- Fix retired /gsd-spec-phase -> /gsd:spec-phase at verify-phase.md:83 (wave-3
prose; caught by slash-namespace invariant #3443/bug-2543, blocked CHK-09
full-suite-green) [Rule 1 - bug]
- size:baseline + INVENTORY manifest verified in-sync post-build (no diff)
* docs(1278): add check descriptor + descriptorFromProjection to CONTEXT.md prohibition glossary
* fix(1278): harden descriptorFromProjection round-trip (numeric-coercion + stray-rule) per review
- MD-01/LW-01: narrow projected scalars to primitives + String()-coerce, so a
numeric-looking check_target (parseMustHavesBlock coerces ^\d+$ to number)
reconstructs as a string and locates instead of silently un-locating; no
as-string type-lie, satisfies no-base-to-string.
- LW-02: attach rule only for the lint-rule kind (drop a stray node-test rule).
- LW-03: document the optional check_* keys in the reference Output schema.
RED->GREEN tests added in prohibition-enforcement.test.cjs.
* chore(1278): set changeset pr to 1301
* test(1278): add fast-check property for the check-descriptor round-trip + fail-closed (trek-e review)
RULESET.TESTS.property-based-testing: the projectProhibitions -> render ->
parseMustHavesBlock -> descriptorFromProjection chain is a bijective/transformation
contract. Adds 2 fc properties to tests/probe-core.property.test.cjs (no new file;
ratchet stays at 2 for probe-core):
- well-formed descriptors survive the round-trip across the full string domain
incl. the numeric-coercion case (target/rule reconstruct as strings);
- under-specified/invalid descriptors (absent / target-less / rule-less /
unknown-kind) are always fail-closed (never green, flagged, unlocated).
Stability is asserted at the descriptorFromProjection layer (the raw parse step is
intentionally lossy for numeric scalars; the shared parser is unchanged).
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* enhance(execute): isolated-executor rejected/over-reaching run fails safe (#1292)
When an isolated (worktree) executor run is rejected — the user declines to
merge it, the orchestrator surfaces recovery for a blocked/halted plan, or the
run over-reached the requested scope — the orchestrator must no longer
default/propose recovery by editing the primary checkout (`main`). Absent an
explicit guardrail, the LLM orchestrator could improvise "continue on main",
inverting the isolation contract at the moment it matters most.
Added an ISOLATED-RUN RECOVERY — FAIL SAFE policy: default to a safe halt that
offers a fresh, narrowly-scoped worktree or inspect/discard; editing the primary
checkout requires explicit, clearly-labeled confirmation and is never the
default/proposed option.
To respect the ADR-857 phase-6 host-loop size cap on execute-phase.md (it sits
just under the pre-phase-6 baseline), the policy is delivered as an extracted
reference fragment rather than inline:
- New `execute-phase/steps/worktree-recovery-policy.md` holds the recovery policy
(the existing FAIL-CLOSED rule #48 for base/HEAD mismatches + the #1292
fail-safe guardrail). No #48 behavior change — moved verbatim.
- execute-phase.md references the fragment at the worktree-spawn recovery point,
the step-5.5 merge decision, and the stalled-agent "switch to inline execution"
menu (which for an isolated run now follows the fail-safe policy). Net effect:
execute-phase.md shrinks below its cap.
- quick.md carries the fail-safe guardrail inline at its post-return merge/discard
decision (quick.md is not size-capped).
Scoped to the recovery offer only — no automatic scope-overreach detection
(explicitly out of scope per the issue) and no new config key. Adds content
regression tests, a USER-GUIDE note, and a workflow size-baseline update.
Closes#1292
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(changeset): Changed fragment for #1292 isolated-executor fail-safe recovery
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1296): align config docs/prompts/schema with consumers
The user-facing config surface disagreed with what the consumers actually do
(subset of the #1216 audit). No runtime consumption behavior changes.
- workflow.subagent_timeout: settings-advanced.md prompt + docs/CONFIGURATION.md
said "seconds (default 600)" but the consumer (map-codebase.md) uses
milliseconds (default 300000). Relabeled all four spots in settings-advanced.md
(prompt, parse-default list, example, confirmation table) + the CONFIGURATION.md
row.
- review.models.<cli>: settings-integrations.md, docs/CONFIGURATION.md (Integration
Settings), and docs/CLI-TOOLS.md documented a shell command, but review.md injects
the value into a --model/-m flag. Relabeled to a bare model id and reconciled the
contradictory CONFIGURATION.md sections.
- workflow.test_command + workflow.build_command: consumed via config-get
(test_command in verify-phase/execute-phase/audit-fix/post-merge-gate;
build_command in post-merge-gate) and documented, but absent from validKeys so
`config set` rejected them. Registered both in config-schema.manifest.json and
documented them in references/planning-config.md (overview + complete reference).
Regression tests: behavioral config-set tests (tests/config.test.cjs) + doc-parity
content guards (tests/config-field-docs.test.cjs).
Deferred to other #1216 clusters: security-gate wiring, search_gitignored wiring,
mvp_mode, source_grounding_authority labeling, and config-set enum enforcement.
Closes#1296
Refs #1216
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(changeset): Fixed fragment for #1296 config-surface alignment
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier
Introduce a per-truth PRESENT_BEHAVIOR_UNVERIFIED state for must-haves that assert a state transition or a cancellation/cleanup/ordering invariant whose only evidence is symbol presence + wiring. Such truths are excluded from the verified_truths score, reported as a behavior_unverified count, recorded in an always-on behavior_unverified_items frontmatter list, and routed to the existing human_needed sink — so a clean N/N can no longer be reached on symbol presence alone.
The overall-status vocabulary and the src/verification.cts seam are unchanged (the new state is per-truth only); gaps_found keeps decision-tree precedence; override-passed truths still count toward verified_truths. Mirrors the calibration into the shipped verify-phase.md workflow (with an infra/foundation carve-out), the VERIFICATION.md templates, and docs (planning-artifacts.md, AGENTS.md). gsd-verifier.md kept under its 48KB LARGE cap; size baselines regenerated.
Closes#966
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#966): add changeset fragment for gsd-verifier behavior-unverified calibration
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Maintainer CHANGES_REQUESTED (reviewed e08667e5, pre-portability-fix):
- B1: lint-rule no longer greens an unparseable target (eslintHasFatalError -> fail
closed on any fatal/parse error) or an inline-suppressed violation (eslintJsonHasRule
now scans suppressedMessages too). RED-first + real-runner repros.
- B2: both child spawns get a bounded timeout (30s node / 60s eslint) + 16MiB maxBuffer;
timeout fails closed. Injectable timeoutMs (positive-only — 0/negative can't disable
the bound) enables a fast 1.5s hang test.
- M1: scoped verify-phase.md — the check descriptor is author-supplied for now; filed
#1278 for deterministic auto-locate of the descriptor (the locate half).
- M2: added tests/prohibition-enforcement.property.test.cjs (fast-check fail-closed
invariants; within the <=2-file budget).
- m1: tapTestNames excludes # SKIP/# TODO; parseNodeTestSummary tracks # cancelled;
isNonVacuousNodeTestPass requires cancelled===0.
- m2: scoped the determinism claim to the decision/parse layer (real runner is env-dependent).
- m3: filed #1279 for machine-proven fail-first (violation-fixture probe).
- n1: -- before target in both arg builders (option-injection). n2: dropped dead token.
- B3 (Windows npx) was already fixed in 2af76306 (pushed ~65s after the review).
Verified node 22 + 24; size baseline regenerated for the verify-phase note.
Adversarial pre-submission review found the injected-runCheck tests masked a
non-functional real runner. Fixes:
- BL-01 (false green on vacuous test): the node-test runner now parses the TAP
summary and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail) AND a
reported test named distinctly from the file — node --test counts an empty file
as one passing test, so counts alone could not catch it.
- SF-01 (lint anchor never greened): the lint-rule runner now runs the project
eslint as --format json and filters by ruleId, so plugin rules (local/*) load
via the flat config — bare --rule cannot load a plugin. local/no-source-grep
now genuinely greens (covered by a real, non-injected test).
- BL-02 (tautological fail-first): the runner no longer echoes the caller's
failFirst as if confirmed. failFirst is documented as caller-ATTESTED; the
producer requires attestation + a genuine non-vacuous pass. Machine-proven
fail-first (needs a violation fixture) is flagged as a tracked follow-up in
ADR-550, the changeset, FEATURES, the reference doc, and verify-phase.
- SF-02: added real-runner end-to-end tests (no injected runCheck) + pure,
exported parse/filter helpers (parseNodeTestSummary, tapTestNames,
eslintJsonHasRule, eslintFileResultCount) so the shipping branches are
mutation-pinned.
- NIT-01/02: LOCATE guard rejects empty-string rule and unknown kinds.
- Hardening: spawn checks with NODE_TEST_CONTEXT/NODE_OPTIONS scrubbed so an
ambient test-runner context cannot corrupt a verify-time result.
- Docs reconciled to the shipped behavior (no 'confirms fail-first' overclaim).
The default lint-rule runner passed check.target as BOTH the --rule id and the
eslint path, so it could never pass (eslint tried to lint a file named after the
rule). Add a distinct check.rule field (rule id) vs check.target (path to lint),
extract a pure exported buildLintArgs() so the mapping is mutation-testable
without spawning eslint, fail-closed on a lint-rule missing its rule id, and carry
the rule into enforcement evidence. Updates verify-phase descriptor docs.
Edge-probe now surfaces a zero-classification requirement (non-empty prose, no
shape cue matched, no `shapes` override) as a single soft `unclassified — review
manually` candidate instead of silently dropping it — the exact blind spot the
probe exists to catch. Dismissible like any edge; the `shapes: []` opt-out stays
silent; `TAXONOMY` (the closed 8 categories) is unchanged. Under `--auto` the
candidate is left `unresolved`, never auto-`backstop` (a missing shape is not
evidence an edge exists).
Closes#1110
Adds the spec-time prohibition probe (spec-phase Step 5.6) — the second adapter of the probe-core resolution model. Surfaces unwritten must-NOT constraints as negative SPEC acceptance criteria with test/judgment verification tiers; fail-closed at verify time. Per ADR-550. Closes#644.
* refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam
ADR-22's severity mapping, authority auto-upgrade, and rung>=3 hard-block lived only as prose in plan-review-convergence.md — untestable. Extracted into src/plan-drift-guard.cts (pure: AUTHORITY_RUNGS, getEffectiveAuthority, classifyDriftSeverity) + a gsd-tools drift-guard CLI seam (authority/severity), and rewired the workflow to call the seam deterministically instead of reasoning the decision in prose. 47 unit/e2e/structural tests cover the full severity table, the grep->intel auto-upgrade, and the rung>=3 HIGH hard-block.
Closes#1190
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1190): add changeset for ADR-22 drift-guard seam (#1242)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1190): register ADR-22 module in eslint-ignore + inventory manifest + docs-exempt changeset
Full-matrix CI surfaced new-module/command governance ripples beyond the lint-tests chain: (1) tsc-generated plan-drift-guard.cjs must be in the eslint ignore list (551-eslint-bin-lib-coverage); (2) docs/INVENTORY-MANIFEST.json must include the new module/command (regen via gen-inventory-manifest.cjs --write); (3) a type:Added changeset triggers docs-required — added a docs-exempt marker (internal seam, no user-facing surface).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15)
ADR-15 designates /gsd-progress --next --auto --converge as the PRIMARY plan-convergence surface, but only the secondary surface (autonomous.md) was wired. next.md now parses --converge/--cross-ai into a plan strategy, gates on workflow.plan_review_convergence, forwards reviewer flags + --max-cycles, and routes Route-3 planning through /gsd:plan-review-convergence (mirroring autonomous.md); --auto chaining preserves converge mode. Adds argument-hint + help/full.md + COMMANDS.md + how-to parity and a structural regression test.
Closes#1190
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1190): add changeset for progress --converge surface (#1237)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1190): keep progress --converge docs skill-dep-clean + regen workflow size baseline
CI surfaced two ripples from the ADR-15 workflow edits: (1) lint-skill-deps + profile-closure flagged /gsd:plan-phase and /gsd:plan-review-convergence SlashCommand tokens in progress.md's flag docs as undeclared deps — reworded to plain prose since progress.md only advertises the flag (the real invocation lives in next.md); (2) the per-file workflow size baseline needed regenerating after the next.md/help edits.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1213): Capability State Writer — write-side inverse of the resolver
Adds src/capability-writer.cts (setCapabilityState + cmdCapabilitySet) and the
`gsd-tools capability set` subcommand: the write-side inverse of the capability
resolver (ADR-1213). One desired capability state projects onto the substrates —
`enabled` drives the runtime surface (canonical on/off), `gates` drive federated
config keys (hook granularity), install profile is a read-only floor — then
re-resolves and reports divergence (assert-and-report), so "off means off" holds
as a write-time invariant. Adds batched setConfigValues; routes gsd:settings
capability hook-gates through the writer. Docs: CLI-TOOLS reference, how-to,
ADR-1213, CONTEXT.md term.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1213): add changeset for Capability State Writer (#1225)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Core half of #1105: a legal external_job_waiting deferred state so an async-dispatched Execute step (committing a .planning/async-jobs/<job>.json manifest, deferring SUMMARY.md) is not an illegal partial. execute-phase safe-resume, resume-project, and pause-work reconcile against the versioned scheduler-agnostic manifest stability contract without re-dispatching; the producer is the capability half (#1164). Closes#1165.
* fix(#1146): single base-branch resolver across forking workflows
Replaces duplicated per-workflow bash detection that silently fell through
to :-main on repos where origin/HEAD is unset (git init+remote add+fetch
without set-head, most CI checkouts, many worktrees).
New CJS module git-base-branch.cjs exposes `gsd_run query git.base-branch`
with full precedence ladder: git.base_branch config override → origin/HEAD
symref → git remote show origin (authoritative) → local branch presence →
"main". All git subprocesses bounded with timeouts; degrades gracefully.
Wires execute-phase, quick, ship, complete-milestone, and pr-branch to the
single resolver. Removes 14 lines of duplicated detection bash across the
five workflows.
Includes 7 behavioral tests covering the full precedence ladder including the
key regression case (master repo, origin/HEAD unset → must return "master",
NOT "main") and an anti-regression guard that fails if any workflow
re-introduces the :-main/:-master fallback pattern.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(changeset): backfill PR number #1198
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1146): drop stray PR-body file from branch
pr-1146-body.md was committed during changeset backfill but must not
be tracked in the repo. Content preserved externally for PR body use.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#1146): add tests for flat base_branch config key and both-branch tie-break
Closes two mutation gaps identified in adversarial review:
- A2: flat {base_branch: ...} at config root (legacy key form) was covered
by code but unguarded against mutation of lines 74-75 in resolver
- H: tier-4 tie-break when both main+master exist locally (main wins,
per tryLocalBranch JSDoc) was documented but untested
9/9 tests pass.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#1146): allowlist workflow-literal guard as runtime-contract exemption
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1146): degrade gracefully when gsd_run unavailable in handle_branching bash blocks
handle_branching (execute-phase.md) and step 2.5 (quick.md) are extracted
and run verbatim by behavioral tests that lack the gsd_run preamble.
Adding a || fallback ladder (git symbolic-ref then echo main) keeps the
unified resolver as primary in real workflows while letting the test harness
succeed without gsd_run defined.
Also propagates updated runtime-launcher preamble to pr-branch.md (added in
origin/next MemPalace PR) and regenerates workflow-size-baseline.json.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#956): add MemPalace memory capability (ADR-857 feature plug-in)
Adds an opt-in, default-resilient ADR-857 feature capability that wires
MemPalace (local-first memory: MCP server + CLI) into the GSD loop:
deliberate recall before discuss/plan and verbatim + temporal-KG capture
at phase boundaries. Three memory modes (augment default; kg_backend and
replace forward-declared). Master gate mempalace.enabled (default off);
every hook onError:skip, zero gates; absent/disabled MemPalace => loop
unchanged. Transport is rendered-markdown only — MemPalace runs
out-of-process, no third-party code in gsd-core (ADR-857 §7).
Capability: capabilities/mempalace/ (manifest + 2 fragments), skills
commands/gsd/mempalace-{recall,capture}.md, agent
agents/gsd-mempalace-curator.md. Registration: ns-context router,
utility cluster, KNOWN_SKILLS, help full.md, model-catalog, copilot
install list, size baselines; regenerated capability-registry +
inventory manifest. ship:post wired into ship.md (wire-on-demand).
HELD on #1196: this capability also declares hooks at discuss:pre and
discuss:post, which are structurally un-wireable until the host-loop
conformance model covers the discuss phase (discuss-phase.md is not in
HOST_LOOP_FILES). The phase6-capstone-conformance gate therefore fails
on exactly those two orphaned points by design — see #1196. Once #1196
lands, rebase onto next and the gate goes green with no further change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(#956): backfill changeset PR number (#1201)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>