954e963d2b341b10eac00f5bdc33cabfaaab1aa2
82 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2c718bf972 |
fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs (#1537)
* fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs Generalizes the Codex-only #1515/#1519 fix to every non-Claude runtime, and wires it into the real install path (where it was previously dead-on-arrival). Root causes: 1. The runtime-default stamping lived only in `_applyRuntimeRewrites`, but the installer emits `gsd-core/workflows/*.md` via `copyWithPathReplacement`, which never calls it — so a real `--codex`/`--cursor`/etc. install emitted `--default claude` and worktrees-on. RUNTIME mis-resolved to claude and the workflow ran executors unisolated against the main checkout. (#1515/#1519 were also dead-on-arrival in real installs; this repairs them.) 2. Only `case 'codex'` was stamped; every other non-Claude runtime kept the Claude default. Fix: - New `_stampNonClaudeRuntimeDefaults(content, runtime)` (single shared helper) stamps `--default <runtime>` + `use_worktrees=false` for every `runtime != claude`; called from both `_applyRuntimeRewrites` and, crucially, `copyWithPathReplacement` in bin/install.js (the real workflow emit path). - Generalize the fail-closed worktree guard `= codex` -> `!= claude` in execute-phase/quick/diagnose-issues (worktree isolation is Claude-Code-only). - Flip manager/autonomous inline-vs-background gating to `codex -> background, everything-else -> inline` (research: only Codex can background-nest the pipeline's subagents; all others run inline, which they support). Worktree-capability determination is research-backed (official docs for all 14 non-Claude runtimes: none honor GSD's isolation="worktree" mechanism, only Codex background-nests). New end-to-end real-install test asserts the EMITTED workflow is stamped — the regression guard that would have caught the dead-on-arrival bug. Closes #1521 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r * chore(#1521): backfill changeset PR number (#1537) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
2436b76980 |
fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees (#1519)
* fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees A Codex install with a runtime-neutral .planning/config.json resolved RUNTIME=claude and enabled git worktree isolation, which Codex's spawn_agent cannot honor. Two root causes: 1. Workflows read `config-get runtime` / `config-get workflow.use_worktrees` without `--raw`, so config-get's JSON-quoted output ("codex") was captured verbatim into the bash var and broke every `[ "$RUNTIME" = ... ]` check — the Codex fail-closed guard was dead even when runtime:codex was explicit, and Claude's own worktree degrade-check was dead too. Add `--raw` to those reads across execute-phase, autonomous, manager, diagnose-issues, quick. 2. The conversion engine emitted `--default claude` for every runtime. Stamp the codex-emitted workflows to `--default codex` (runtime) and `--default false` (use_worktrees) in _applyRuntimeRewrites case 'codex', so a neutral config on a Codex install resolves runtime=codex / worktrees off. Also extend the Codex fail-closed worktree guard to quick.md and diagnose-issues.md (they spawned isolation="worktree" with no runtime guard). Regression test asserts source<->engine parity across all five workflows (DEFECT.GENERATIVE-FIX) plus fast-check property coverage of the stamping. Closes #1515 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r * chore(#1515): backfill changeset PR number (#1519) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b63a500fad |
fix(#1505): extract context_guard step to reference file; fix allow-test-rule see ref
- Extract execute-phase.md context_guard step prose to gsd-core/references/execute-phase-context-guard.md (@-ref lazy load), bringing execute-phase.md back under the ADR-857 phase-6 size ceiling (92914 < 93166 bytes) - Fix allow-test-rule comment: add `see #1452` per ADR-456 lint rule - Update feat-1452 tests to check the reference file for extracted content - Register execute-phase-context-guard.md in INVENTORY-MANIFEST.json and INVENTORY.md Workflow References section - Regenerate workflow-size-baseline.json after file shrinkage Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
4eca5ac96c |
feat(#1452): add workflow.context_guard_mode to guard execute-phase against context exhaustion
Proactive checkpoint guard fires at each wave boundary before spawning agents. Self-assesses context pressure against context-budget.md degradation tiers and warns (warn, default) or auto-invokes /gsd:pause-work (auto) when POOR tier (70%+) is detected. Config key validated; defaults to \"warn\". Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
f5276b36b3 |
fix(#1369): refresh wave manifest and re-check base before each wave in execute-phase (#1492)
* fix(#1369): refresh wave manifest and re-check base before each wave in execute-phase Two compounding issues caused wave N+1 worktrees to fork from the stale pre-wave-N commit, immediately tripping the worktree_branch_check FATAL guard in every executor: 1. worktree.base-check auto-degrade only ran once at initialize time. After wave N merges advanced orchestrator HEAD past origin/HEAD, new worktrees were still forked from origin/HEAD (Claude Code "fresh" base). 2. WAVE_WORKTREE_MANIFEST was never unset between waves, so wave N+1 reused the consumed wave-N manifest file, which would have blocked the step 5.5 manifest guard (#3384) on subsequent waves. Fix: add two safeguards in execute-phase.md — - Step 0.5 (start of each wave): re-runs worktree.base-check; auto-degrades USE_WORKTREES=false for that wave when HEAD has diverged from origin/HEAD. - Step 7c (end of each wave): unsets WAVE_WORKTREE_MANIFEST so wave N+1 creates a fresh per-wave manifest; re-asserts worktree.set-baseref (idempotent) and re-evaluates base degradation after wave merges land. 17 regression tests added in tests/bug-1369-wave-stale-base.test.cjs. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1369): rename test to fix-NNN convention; update workflow size baseline Rename tests/bug-1369-wave-stale-base.test.cjs → tests/fix-1369-wave-stale-base.test.cjs to satisfy the lint-regression-test-names gate (new files cannot use bug-NNN prefix). Update tests/workflow-size-baseline.json for execute-phase.md: 93157 → 97393 (LF-normalized byte count after adding step 0.5 inter-wave base re-check and step 7c between-wave manifest reset). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1369): add issue reference to allow-test-rule comment Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1369): extract new execute-phase steps to references; satisfy ADR-857 cap Step 0.5 (inter-wave worktree base re-check) and steps 7b–7c (pre-wave dependency check + between-wave manifest reset/base refresh) added by this PR grew execute-phase.md to 97393 bytes, violating the ADR-857 phase-6 architectural mandate that host-loop bodies remain strictly below the pre-phase-6 baseline of 93166 bytes. Extract both new blocks into dedicated reference files: - gsd-core/references/execute-phase-wave-guard.md (step 0.5) - gsd-core/references/execute-phase-between-wave-reset.md (steps 7b + 7c) Replace inline prose with @-reference pointers. File now measures 92851 bytes (LF-normalized), satisfying the ADR-857 capstone conformance gate. Also update tests/workflow-size-baseline.json to 92851 and add both new reference files to docs/INVENTORY-MANIFEST.json. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1369): update regression tests to read from extracted reference files Steps 0.5 and 7b+7c were moved to reference files to satisfy the ADR-857 size cap on execute-phase.md. Tests now check @-reference pointers in the workflow for ordering and read content assertions from the reference files. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
3a3b2135c2 |
chore(#1073): purge phantom pre-migration issue refs from source, tests, docs (#1471)
#2551/#3182/#2361 are pre-migration get-shit-done-redux issue numbers with no equivalent in open-gsd/gsd-core; they mislead triage and manufacture phantom blockers. Repoint to real successors (#717 byte-budget rework, #720) or rewrite as prose referencing the discuss-phase/modes progressive-disclosure split. Correct co-located 'line budget'/'<500 lines' framing to the byte-based reality (#717). Add a CI guard (tests/no-phantom-issue-refs.test.cjs) that fails if a phantom ref is reintroduced. SSH-key patterns (id_ed25519) left untouched. No user-facing runtime behavior change. Closes #1073 |
||
|
|
120f85164b |
feat(#1355): detect-and-warn guard for claude-code agent-teams (#1371)
* feat(#1355): detect-and-warn guard for claude-code agent-teams GSD's multi-agent orchestration can stall under claude-code's experimental agent-teams (a subagent's completion fails to route to the orchestrator). Per the maintainer decision, the accepted scope is a read-only detector + one non-fatal warning — NOT the declined run_in_background/TaskOutput conversion. - New Teams Status Module (src/teams-status.cts → gsd-core/bin/lib/teams-status.cjs): pure resolveTeamsStatus({runtime, env}) + thin CLI cmdTeamsStatus reusing resolveRuntime. active = strictly-truthy env flag AND runtime === 'claude'. - Wire `gsd-tools query teams-status [--active]` (read-only; no capability registration needed — conformance gates govern features, not query commands). - One non-fatal warning in plan-phase.md before the first Agent spawn, gated on `query teams-status --active`; zero behavior change on non-claude/teams-off. - Hermeticity: clear CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS in run-tests.cjs + SESSION_ENV_KEYS. Docs reference + CONTEXT.md glossary. Built lib gitignored. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1355): add changeset for teams-detect guard Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1355): bump plan-phase.md workflow size baseline (+407B for teams warning) The non-fatal agent-teams warning block added to plan-phase.md grew it 92759 → 93166 bytes, past its committed per-file baseline ratchet. The growth is small, deliberate, and still well under the workflow tier hard cap. Regenerate the baseline via `npm run size:baseline` (only plan-phase.md changed). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1355): register teams-status.cjs in the inventory manifest The new teams-status CLI module is a tracked surface; regenerate docs/INVENTORY-MANIFEST.json (cli_modules family) via gen-inventory-manifest.cjs --write so the inventory-manifest-sync gate passes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1a678eb0e9 |
fix(#1359): migrate workflows off deprecated TaskOutput to Read(outputFile) (#1362)
* fix(#1359): migrate workflows off deprecated TaskOutput to Read(outputFile) The map-codebase and docs-update workflows collected background sub-agent results with the deprecated Claude Code `TaskOutput` tool using `block: true`, which has a confirmed main-session hang after the agent completes (anthropics/claude-code#20236). Migrate the collection steps to the upstream-recommended pattern: keep `run_in_background=true` on the Agent spawn, then `Read` each agent's `outputFile` (from the `async_launched` result) once it reports completion. Completion-marker contracts and on-disk verification are unchanged, and the non-Claude runtime fallbacks (sequential_mapping / sequential_generation) are preserved byte-for-byte. docs-update's timeout note no longer references the unwired `workflow.subagent_timeout` key (it kept a literal before). Regression coverage folded into tests/subagent-timeout.test.cjs (the owning module for background-subagent collection). Workflow size baseline regenerated for the justified prose growth. Refs #1355 (same latent hang surface). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1359): add changeset for TaskOutput migration Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
00c05eb717 | Merge branch 'next' into feat/1279-fail-first-prover | ||
|
|
9540fe43b9 | fix: have executor self-report worktree metadata (#1349) | ||
|
|
56d4a1bf39 |
enhance(#1279): project check_violation_fixture scalar — #1278 locate + #1279 proof compose end-to-end (#1346)
Delivers option (a) from the #1314 maintainer review: thread a fourth flat scalar check_violation_fixture through the projection so a prohibition authored at spec-phase machine-proves fail-first and greens through the deterministic path alone — zero hand-authoring at verify time. - src/probe-core.cts: Prohibition gains check_violation_fixture?; projectProhibitions emits it (both kinds) ONLY for a well-formed descriptor and ONLY when non-empty (blank/absent -> projects absent -> producer hard-gates, never a partial green). - src/prohibition-enforcement.cts: descriptorFromProjection reads it back into violationFixture via the same numeric-coercion-safe scalar() normalizer. - Tests (RED-first, proven non-vacuous by reverting both src edits): CHK-02(#1346) projection emit, CHK-08(#1346) read-back, CHK-03(D) example round-trip, the fast-check round-trip property extended to the 4th scalar (the contract trek-e blocked #1301 on), and a real-subprocess COMPOSE capstone greening end-to-end through project -> descriptorFromProjection -> default prover+runner. - Docs flipped from 'hard-gates until #1346' to 'composes end-to-end': verify-phase.md, prohibition-probe.md, spec-phase.md authoring, ADR-550 addendum, changeset. #1346 now tracks only the node-test causation residual. 190 affected-suite tests green; eslint + tsc clean; size baseline regenerated. |
||
|
|
6e242bd76a | fix: allow quick worktree parent plan base (#1347) | ||
|
|
33b6ee6f0c |
fix(#1279): node-test prover fail-closes on missing violationFixture + honest projection docs
Addresses the #1314 maintainer review (trek-e): - Major 1 (fail-OPEN): defaultProveFailFirst's node-test branch only guarded `if (!fixture)`. A missing/typo'd/stale violationFixture made GSD_PROHIB_SUBJECT point at a missing file; an honest negative test threw ENOENT *inside its callback* (a failing test named distinctly from the file), which isNonVacuousNodeTestRed accepted as proof -> a green forged from a setup crash. Now requires fs.existsSync(path.resolve(cwd, fixture)) before spawning, symmetric with the lint-rule path's file-result guard. Regression test pins it (RED without the guard); a second test pins cwd-relative fixture resolution. - Major 2 (misleading prose): the #1278 projection carries no violationFixture, so the deterministic-locate path always hard-gates (fail-closed) until a check_violation_fixture scalar is threaded through. verify-phase.md and prohibition-probe.md no longer read as if the projected path produces greens; the ADR-550 addendum records both items. Tracked as follow-up #1346. - Documented residual: existence is necessary but not sufficient (a red caused by the env being set vs the subject's content); recorded as a constraint, in #1346. - Nit: stale 'NOT attested fail-first' comment -> 'NOT machine-proven fail-first'. 64 tests pass; eslint + tsc clean; changeset valid. |
||
|
|
49bef1927e |
Merge remote-tracking branch 'origin/next' into feat/1279-fail-first-prover
# Conflicts: # docs/adr/550-spec-phase-probe-contract.md # gsd-core/workflows/verify-phase.md # tests/workflow-size-baseline.json |
||
|
|
e50ead7ad2 |
enhance(verify-phase): deterministic auto-locate of the prohibition check descriptor (#1278) (#1301)
* test(1278): RED-first descriptor parity + fail-closed guards + CHK-07 byte-stability (wave 1) - CHK-03 (RED): extend PROB-14 parity in prohibition-probe.schema.test.cjs to carry the flat check_kind/check_target/check_rule scalars through project->write->parseMustHavesBlock; the non-droppable check_kind-presence assertion is the load-bearing RED trigger (fails because projectProhibitions strips check_* on the current build). - CHK-07 (GREEN forward-guard): probe-core.test.cjs pins descriptor-less byte-stability + dispositionForProhibition fail-closed policy, with a t.todo marker forward-locking plan 01-02. - CHK-06 (RED): prohibition-enforcement.test.cjs asserts descriptorFromProjection export + fail-closed on absent/partial/unknown descriptors via the projection adapter (RED until 01-03). - No src/*.cts or .cjs edits; no new test files; lint-test-file-count clean. * feat(1278): add optional flat-scalar check descriptor fields to Prohibition interface (wave 2) - check_kind?/check_target?/check_rule? mirror CheckDescriptor.kind/target/rule (minus caller-attested failFirst, #1279) - optional so existing Prohibition consumers compile unchanged * feat(1278): project check descriptor as flat scalars in projectProhibitions (wave 2) - emit check_kind/check_target (+ check_rule only for lint-rule with a rule) when descriptor well-formed - under-specified/descriptor-less items project byte-identically (CHK-07); flat scalars ride existing parseMustHavesBlock continuation-KV path (no parser rewrite) - add CHK-02 probe-core unit cases pinning the projection - turns CHK-03 parity test GREEN; dispositionForProhibition untouched * feat(1278): descriptorFromProjection read-back adapter feeds fail-closed locate (wave 3) - Add descriptorFromProjection(projected) -> CheckDescriptor | null to src/prohibition-enforcement.cts: renames the projected flat scalars check_kind/check_target/check_rule -> {kind,target,rule?}, or null when the descriptor is absent/non-object (no check_kind key). - failFirst is NEVER sourced from the projection (stays caller-attested; #1279). - rule is set only when check_rule is a non-empty string; the adapter does NOT re-validate kind/target/rule — an under-specified descriptor reconstructs to one the EXISTING runProhibitionEnforcement LOCATE guard rejects (located:false, never green). The merged #1259 guard stays the single source of fail-closed truth. - Turns the RED CHK-06 fail-closed tests (plan 01-01) GREEN end-to-end; CHK-03 / CHK-07 stay green. CheckDescriptor type, locate guard, dispositionForProhibition, and parseMustHavesBlock are unchanged (additive +36/-0). * feat(1278): verify-phase locates prohibition check from projected descriptor (wave 3) - request.check kind/target/rule sourced from projected check_kind/check_target/check_rule via descriptorFromProjection, not verifier invention (CHK-05) - replaces the #1278 author-supplied / tracked-follow-up note with the delivered deterministic-locate behavior - preserves fail-closed routing: absent/partial descriptor -> never green, hard-gate in both modes - failFirst stays a verify-time caller attestation; #1279 bounds the remaining fail-first proof * feat(1278): spec-phase captures wired-check descriptor on test-tier resolution (wave 3) - Step 5.6 'Keep it' / verification: test path captures check_kind/check_target/check_rule, projected onto must_haves.prohibitions for verify-phase deterministic locate (CHK-04) - SOFT capture: a test-tier prohibition without a descriptor is still allowed (no hard authoring block); stays fail-closed/flagged downstream - --auto captures only an unambiguous descriptor, never fabricates a check path - failFirst NOT captured at spec-phase (verify-time attestation; #1279) - PROB-06 soft-gate + text-mode (PROB-09) behavior unchanged * chore(1278): re-baseline workflow size for grown verify-phase + spec-phase prose (wave 3) - spec-phase.md 28438 -> 30343 (+1905), verify-phase.md 35362 -> 36498 (+1136) - regenerated via npm run size:baseline (no hand-picked numbers); growth is the #1278 deterministic-locate + descriptor-capture prose - workflow-size-budget guard green (122/122) * docs(1278): ratify optional check descriptor in dated ADR-550 addendum + type:Changed changeset - Append dated 2026-06-15 ADR-550 addendum ratifying the D3 prohibition-item shape extension (optional flat-scalar check_kind/check_target/check_rule) - Document flat-scalar rationale, deterministic projection/read-back, fail-closed on partial/invalid/absent, #1279/policy out-of-scope - Add .changeset/1278-prohibition-check-descriptor.md (type: Changed) * docs(1278): document optional check descriptor in prohibition-probe reference + FEATURES - Add 'Optional wired-check descriptor (deterministic locate, #1278)' section to the prohibition-probe reference (flat-scalar keys, projection/read-back, fail-closed + backward-compat, failFirst stays attested) - Add deterministic prohibition-check descriptor source entry to FEATURES.md - No CONTEXT.md glossary change: descriptor reuses existing wired-check / verification:test vocabulary, no new glossary term introduced * fix(1278): pass packaging gates — changeset pr field + retired slash-form fix - Add required pr: 1278 to changeset (lint:changeset MISSING_PR hard requirement; plan's 'omit if unknown' was inaccurate — issue number per #1259 convention, updated to real PR number when opened) [Rule 3 - blocking] - Fix retired /gsd-spec-phase -> /gsd:spec-phase at verify-phase.md:83 (wave-3 prose; caught by slash-namespace invariant #3443/bug-2543, blocked CHK-09 full-suite-green) [Rule 1 - bug] - size:baseline + INVENTORY manifest verified in-sync post-build (no diff) * docs(1278): add check descriptor + descriptorFromProjection to CONTEXT.md prohibition glossary * fix(1278): harden descriptorFromProjection round-trip (numeric-coercion + stray-rule) per review - MD-01/LW-01: narrow projected scalars to primitives + String()-coerce, so a numeric-looking check_target (parseMustHavesBlock coerces ^\d+$ to number) reconstructs as a string and locates instead of silently un-locating; no as-string type-lie, satisfies no-base-to-string. - LW-02: attach rule only for the lint-rule kind (drop a stray node-test rule). - LW-03: document the optional check_* keys in the reference Output schema. RED->GREEN tests added in prohibition-enforcement.test.cjs. * chore(1278): set changeset pr to 1301 * test(1278): add fast-check property for the check-descriptor round-trip + fail-closed (trek-e review) RULESET.TESTS.property-based-testing: the projectProhibitions -> render -> parseMustHavesBlock -> descriptorFromProjection chain is a bijective/transformation contract. Adds 2 fc properties to tests/probe-core.property.test.cjs (no new file; ratchet stays at 2 for probe-core): - well-formed descriptors survive the round-trip across the full string domain incl. the numeric-coercion case (target/rule reconstruct as strings); - under-specified/invalid descriptors (absent / target-less / rule-less / unknown-kind) are always fail-closed (never green, flagged, unlocated). Stability is asserted at the descriptorFromProjection layer (the raw parse step is intentionally lossy for numeric scalars; the shared parser is unchanged). --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
8cf4716c99 |
docs(#1279): ratify ADR-550 addendum + FEATURES/reference/verify-phase + changeset
- ADR-550 dated 2026-06-15 addendum: machine-proven fail-first REALIZED (D5d CLOSED); ratify violationFixture, GSD_PROHIB_SUBJECT, failFirst demotion - FEATURES REQ-PROHIB-07 reads shipped (machine-proven, not caller-attested) - prohibition-probe reference + verify-phase descriptor shape updated with violationFixture + convention - tag GSD_PROHIB_SUBJECT + violationFixture PROPOSED/renamable (zero live consumers) as a PR-review flag - add .changeset/1279-machine-proven-fail-first.md (type: Changed) |
||
|
|
91bd82f9a6 |
enhance(execute): isolated-executor rejected/over-reaching run fails safe — never default recovery to main (#1292) (#1303)
* enhance(execute): isolated-executor rejected/over-reaching run fails safe (#1292) When an isolated (worktree) executor run is rejected — the user declines to merge it, the orchestrator surfaces recovery for a blocked/halted plan, or the run over-reached the requested scope — the orchestrator must no longer default/propose recovery by editing the primary checkout (`main`). Absent an explicit guardrail, the LLM orchestrator could improvise "continue on main", inverting the isolation contract at the moment it matters most. Added an ISOLATED-RUN RECOVERY — FAIL SAFE policy: default to a safe halt that offers a fresh, narrowly-scoped worktree or inspect/discard; editing the primary checkout requires explicit, clearly-labeled confirmation and is never the default/proposed option. To respect the ADR-857 phase-6 host-loop size cap on execute-phase.md (it sits just under the pre-phase-6 baseline), the policy is delivered as an extracted reference fragment rather than inline: - New `execute-phase/steps/worktree-recovery-policy.md` holds the recovery policy (the existing FAIL-CLOSED rule #48 for base/HEAD mismatches + the #1292 fail-safe guardrail). No #48 behavior change — moved verbatim. - execute-phase.md references the fragment at the worktree-spawn recovery point, the step-5.5 merge decision, and the stalled-agent "switch to inline execution" menu (which for an isolated run now follows the fail-safe policy). Net effect: execute-phase.md shrinks below its cap. - quick.md carries the fail-safe guardrail inline at its post-return merge/discard decision (quick.md is not size-capped). Scoped to the recovery offer only — no automatic scope-overreach detection (explicitly out of scope per the issue) and no new config key. Adds content regression tests, a USER-GUIDE note, and a workflow size-baseline update. Closes #1292 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(changeset): Changed fragment for #1292 isolated-executor fail-safe recovery Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
137760a655 |
fix(#1296): align config docs/prompts/schema with consumers (#1299)
* fix(#1296): align config docs/prompts/schema with consumers The user-facing config surface disagreed with what the consumers actually do (subset of the #1216 audit). No runtime consumption behavior changes. - workflow.subagent_timeout: settings-advanced.md prompt + docs/CONFIGURATION.md said "seconds (default 600)" but the consumer (map-codebase.md) uses milliseconds (default 300000). Relabeled all four spots in settings-advanced.md (prompt, parse-default list, example, confirmation table) + the CONFIGURATION.md row. - review.models.<cli>: settings-integrations.md, docs/CONFIGURATION.md (Integration Settings), and docs/CLI-TOOLS.md documented a shell command, but review.md injects the value into a --model/-m flag. Relabeled to a bare model id and reconciled the contradictory CONFIGURATION.md sections. - workflow.test_command + workflow.build_command: consumed via config-get (test_command in verify-phase/execute-phase/audit-fix/post-merge-gate; build_command in post-merge-gate) and documented, but absent from validKeys so `config set` rejected them. Registered both in config-schema.manifest.json and documented them in references/planning-config.md (overview + complete reference). Regression tests: behavioral config-set tests (tests/config.test.cjs) + doc-parity content guards (tests/config-field-docs.test.cjs). Deferred to other #1216 clusters: security-gate wiring, search_gitignored wiring, mvp_mode, source_grounding_authority labeling, and config-set enum enforcement. Closes #1296 Refs #1216 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(changeset): Fixed fragment for #1296 config-surface alignment Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0d783fa57d |
Merge remote-tracking branch 'origin/next' into feat/1259-test-tier-enforcement
# Conflicts: # tests/workflow-size-baseline.json |
||
|
|
0a856f06cd |
enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier (#1271)
* enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier Introduce a per-truth PRESENT_BEHAVIOR_UNVERIFIED state for must-haves that assert a state transition or a cancellation/cleanup/ordering invariant whose only evidence is symbol presence + wiring. Such truths are excluded from the verified_truths score, reported as a behavior_unverified count, recorded in an always-on behavior_unverified_items frontmatter list, and routed to the existing human_needed sink — so a clean N/N can no longer be reached on symbol presence alone. The overall-status vocabulary and the src/verification.cts seam are unchanged (the new state is per-truth only); gaps_found keeps decision-tree precedence; override-passed truths still count toward verified_truths. Mirrors the calibration into the shipped verify-phase.md workflow (with an infra/foundation carve-out), the VERIFICATION.md templates, and docs (planning-artifacts.md, AGENTS.md). gsd-verifier.md kept under its 48KB LARGE cap; size baselines regenerated. Closes #966 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#966): add changeset fragment for gsd-verifier behavior-unverified calibration Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6a9e6cb0ae |
fix(1259-01): address trek-e review — B1 (fatal/suppressed), B2 (bounded subprocess), M1/M2, minors
Maintainer CHANGES_REQUESTED (reviewed |
||
|
|
31b822b312 |
fix(1259-01): close adversarial review findings — genuine enforcement, honest fail-first scope
Adversarial pre-submission review found the injected-runCheck tests masked a non-functional real runner. Fixes: - BL-01 (false green on vacuous test): the node-test runner now parses the TAP summary and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail) AND a reported test named distinctly from the file — node --test counts an empty file as one passing test, so counts alone could not catch it. - SF-01 (lint anchor never greened): the lint-rule runner now runs the project eslint as --format json and filters by ruleId, so plugin rules (local/*) load via the flat config — bare --rule cannot load a plugin. local/no-source-grep now genuinely greens (covered by a real, non-injected test). - BL-02 (tautological fail-first): the runner no longer echoes the caller's failFirst as if confirmed. failFirst is documented as caller-ATTESTED; the producer requires attestation + a genuine non-vacuous pass. Machine-proven fail-first (needs a violation fixture) is flagged as a tracked follow-up in ADR-550, the changeset, FEATURES, the reference doc, and verify-phase. - SF-02: added real-runner end-to-end tests (no injected runCheck) + pure, exported parse/filter helpers (parseNodeTestSummary, tapTestNames, eslintJsonHasRule, eslintFileResultCount) so the shipping branches are mutation-pinned. - NIT-01/02: LOCATE guard rejects empty-string rule and unknown kinds. - Hardening: spawn checks with NODE_TEST_CONTEXT/NODE_OPTIONS scrubbed so an ambient test-runner context cannot corrupt a verify-time result. - Docs reconciled to the shipped behavior (no 'confirms fail-first' overclaim). |
||
|
|
676436258a |
fix(1259-01): lint-rule real runner — keep rule id distinct from lint target
The default lint-rule runner passed check.target as BOTH the --rule id and the eslint path, so it could never pass (eslint tried to lint a file named after the rule). Add a distinct check.rule field (rule id) vs check.target (path to lint), extract a pure exported buildLintArgs() so the mapping is mutation-testable without spawning eslint, fail-closed on a lint-rule missing its rule id, and carry the rule into enforcement evidence. Updates verify-phase descriptor docs. |
||
|
|
ce01e1376b |
docs(1259-01): wire verify-phase consumer + ADR-550/FEATURES/reference + changeset
- verify-phase.md: replace test-tier 'fail-closed/deferred' bullet with the check prohibition-enforcement enforcement step (locate -> fail-first -> run -> evidence -> green-or-hard-gate); update determine_status tree - ADR-550 addendum: mark D5d enforcement half LANDED (#1259); cite ADR-857 open-question §147 + D6 (core verify rail) - FEATURES §146: enforcement wording + add REQ-PROHIB-07; keep REQ-PROHIB-06 intact - references/prohibition-probe.md: test-tier enforced + hard-gates via check prohibition-enforcement - changeset (type: Changed) with the D5 '2 no-source-grep invalid cases, not 96' correction |
||
|
|
3556450b0d |
feat(spec-phase): surface zero-classification edge-probe requirements as unclassified candidates (#1110) (#1117)
Edge-probe now surfaces a zero-classification requirement (non-empty prose, no shape cue matched, no `shapes` override) as a single soft `unclassified — review manually` candidate instead of silently dropping it — the exact blind spot the probe exists to catch. Dismissible like any edge; the `shapes: []` opt-out stays silent; `TAXONOMY` (the closed 8 categories) is unchanged. Under `--auto` the candidate is left `unresolved`, never auto-`backstop` (a missing shape is not evidence an edge exists). Closes #1110 |
||
|
|
395fb519e7 |
feat(spec-phase): prohibition probe — surface "must-NOT" constraints (#644) (#1149)
Adds the spec-time prohibition probe (spec-phase Step 5.6) — the second adapter of the probe-core resolution model. Surfaces unwritten must-NOT constraints as negative SPEC acceptance criteria with test/judgment verification tiers; fail-closed at verify time. Per ADR-550. Closes #644. |
||
|
|
55eb1fd72e |
refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam (#1242)
* refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam ADR-22's severity mapping, authority auto-upgrade, and rung>=3 hard-block lived only as prose in plan-review-convergence.md — untestable. Extracted into src/plan-drift-guard.cts (pure: AUTHORITY_RUNGS, getEffectiveAuthority, classifyDriftSeverity) + a gsd-tools drift-guard CLI seam (authority/severity), and rewired the workflow to call the seam deterministically instead of reasoning the decision in prose. 47 unit/e2e/structural tests cover the full severity table, the grep->intel auto-upgrade, and the rung>=3 HIGH hard-block. Closes #1190 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1190): add changeset for ADR-22 drift-guard seam (#1242) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1190): register ADR-22 module in eslint-ignore + inventory manifest + docs-exempt changeset Full-matrix CI surfaced new-module/command governance ripples beyond the lint-tests chain: (1) tsc-generated plan-drift-guard.cjs must be in the eslint ignore list (551-eslint-bin-lib-coverage); (2) docs/INVENTORY-MANIFEST.json must include the new module/command (regen via gen-inventory-manifest.cjs --write); (3) a type:Added changeset triggers docs-required — added a docs-exempt marker (internal seam, no user-facing surface). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0b3a2e5f9c |
feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15) (#1237)
* feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15) ADR-15 designates /gsd-progress --next --auto --converge as the PRIMARY plan-convergence surface, but only the secondary surface (autonomous.md) was wired. next.md now parses --converge/--cross-ai into a plan strategy, gates on workflow.plan_review_convergence, forwards reviewer flags + --max-cycles, and routes Route-3 planning through /gsd:plan-review-convergence (mirroring autonomous.md); --auto chaining preserves converge mode. Adds argument-hint + help/full.md + COMMANDS.md + how-to parity and a structural regression test. Closes #1190 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1190): add changeset for progress --converge surface (#1237) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1190): keep progress --converge docs skill-dep-clean + regen workflow size baseline CI surfaced two ripples from the ADR-15 workflow edits: (1) lint-skill-deps + profile-closure flagged /gsd:plan-phase and /gsd:plan-review-convergence SlashCommand tokens in progress.md's flag docs as undeclared deps — reworded to plain prose since progress.md only advertises the flag (the real invocation lives in next.md); (2) the per-file workflow size baseline needed regenerating after the next.md/help edits. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
bf634b95c3 |
feat(#1213): Capability State Writer — write-side inverse of the resolver (#1225)
* feat(#1213): Capability State Writer — write-side inverse of the resolver Adds src/capability-writer.cts (setCapabilityState + cmdCapabilitySet) and the `gsd-tools capability set` subcommand: the write-side inverse of the capability resolver (ADR-1213). One desired capability state projects onto the substrates — `enabled` drives the runtime surface (canonical on/off), `gates` drive federated config keys (hook granularity), install profile is a read-only floor — then re-resolves and reports divergence (assert-and-report), so "off means off" holds as a write-time invariant. Adds batched setConfigValues; routes gsd:settings capability hook-gates through the writer. Docs: CLI-TOOLS reference, how-to, ADR-1213, CONTEXT.md term. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1213): add changeset for Capability State Writer (#1225) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
7edd18fd2b |
feat(#1165): async external_job_waiting half-state + resume/pause contract (#1221)
Core half of #1105: a legal external_job_waiting deferred state so an async-dispatched Execute step (committing a .planning/async-jobs/<job>.json manifest, deferring SUMMARY.md) is not an illegal partial. execute-phase safe-resume, resume-project, and pause-work reconcile against the versioned scheduler-agnostic manifest stability contract without re-dispatching; the producer is the capability half (#1164). Closes #1165. |
||
|
|
e9f9ae49c8 |
fix(#1146): single base-branch resolver across forking workflows (#1198)
* fix(#1146): single base-branch resolver across forking workflows Replaces duplicated per-workflow bash detection that silently fell through to :-main on repos where origin/HEAD is unset (git init+remote add+fetch without set-head, most CI checkouts, many worktrees). New CJS module git-base-branch.cjs exposes `gsd_run query git.base-branch` with full precedence ladder: git.base_branch config override → origin/HEAD symref → git remote show origin (authoritative) → local branch presence → "main". All git subprocesses bounded with timeouts; degrades gracefully. Wires execute-phase, quick, ship, complete-milestone, and pr-branch to the single resolver. Removes 14 lines of duplicated detection bash across the five workflows. Includes 7 behavioral tests covering the full precedence ladder including the key regression case (master repo, origin/HEAD unset → must return "master", NOT "main") and an anti-regression guard that fails if any workflow re-introduces the :-main/:-master fallback pattern. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(changeset): backfill PR number #1198 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1146): drop stray PR-body file from branch pr-1146-body.md was committed during changeset backfill but must not be tracked in the repo. Content preserved externally for PR body use. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#1146): add tests for flat base_branch config key and both-branch tie-break Closes two mutation gaps identified in adversarial review: - A2: flat {base_branch: ...} at config root (legacy key form) was covered by code but unguarded against mutation of lines 74-75 in resolver - H: tier-4 tie-break when both main+master exist locally (main wins, per tryLocalBranch JSDoc) was documented but untested 9/9 tests pass. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#1146): allowlist workflow-literal guard as runtime-contract exemption Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#1146): degrade gracefully when gsd_run unavailable in handle_branching bash blocks handle_branching (execute-phase.md) and step 2.5 (quick.md) are extracted and run verbatim by behavioral tests that lack the gsd_run preamble. Adding a || fallback ladder (git symbolic-ref then echo main) keeps the unified resolver as primary in real workflows while letting the test harness succeed without gsd_run defined. Also propagates updated runtime-launcher preamble to pr-branch.md (added in origin/next MemPalace PR) and regenerates workflow-size-baseline.json. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a375c4b354 |
feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) (#1201)
* feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) Adds an opt-in, default-resilient ADR-857 feature capability that wires MemPalace (local-first memory: MCP server + CLI) into the GSD loop: deliberate recall before discuss/plan and verbatim + temporal-KG capture at phase boundaries. Three memory modes (augment default; kg_backend and replace forward-declared). Master gate mempalace.enabled (default off); every hook onError:skip, zero gates; absent/disabled MemPalace => loop unchanged. Transport is rendered-markdown only — MemPalace runs out-of-process, no third-party code in gsd-core (ADR-857 §7). Capability: capabilities/mempalace/ (manifest + 2 fragments), skills commands/gsd/mempalace-{recall,capture}.md, agent agents/gsd-mempalace-curator.md. Registration: ns-context router, utility cluster, KNOWN_SKILLS, help full.md, model-catalog, copilot install list, size baselines; regenerated capability-registry + inventory manifest. ship:post wired into ship.md (wire-on-demand). HELD on #1196: this capability also declares hooks at discuss:pre and discuss:post, which are structurally un-wireable until the host-loop conformance model covers the discuss phase (discuss-phase.md is not in HOST_LOOP_FILES). The phase6-capstone-conformance gate therefore fails on exactly those two orphaned points by design — see #1196. Once #1196 lands, rebase onto next and the gate goes green with no further change. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#956): backfill changeset PR number (#1201) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b1e8a74708 |
fix(#1196): wire discuss loop step for capability hooks (#1199)
* fix(#1196): wire discuss loop step for capability hooks discuss was contract-declared (gsd:loop-host marker, in POINT_ORDER and LOOP_HOST_CONTRACT) but structurally unwireable: discuss-phase.md had no `loop render-hooks` dispatch and was absent from the conformance gate's HOST_LOOP_FILES, so capabilities could never wire discuss:pre/discuss:post. - discuss-phase.md: add minimal discuss:pre (before analyze_phase) and discuss:post (after write_context) render-hooks dispatch steps that delegate consumption to a new shared reference (kept under the 32KB #2551 budget; no inline subagent dispatch token). - references/loop-hook-dispatch.md: new canonical, point-agnostic contract for consuming `loop render-hooks --raw` activeHooks (contribution/step/ gate) — single source for hook consumption across host loops. - gen-loop-host-contract.cjs: derive HOST_LOOP_FILES from STEP_WORKFLOWS and export scanWiredPoints()/getWiredLoopPoints() (throws on a missing host file) — one source of truth for the host-loop file + wired-point set. - phase6-capstone-conformance.test.cjs: consume the derived HOST_LOOP_FILES and shared scanWiredPoints (was a hand-maintained duplicate omitting discuss-phase.md + a duplicated regex). - gen-capability-registry.cjs: add validateHooksWired() gen-time guard that rejects a capability hook declared at a valid-but-unwired loop point, with a clear remediation message — failure now surfaces at gen --check/--write time instead of deep in the full conformance suite. - tests (capability-registry.test.cjs): regression + anti-pattern parity guards (every loop-host marker is in STEP_WORKFLOWS/HOST_LOOP_FILES; POINT_ORDER === flattened LOOP_HOST_CONTRACT) so no step can drift into the discuss-class gap again. - docs/INVENTORY*: register the new reference. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1196): backfill changeset PR number (#1199) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b10e56818b |
feat(#1169): complete ADR-857 phase 6 — migrate features to Capabilities, revive dead gates, harden conformance gate (#1183)
* test(#1168): make phase-6 gate un-gameable — reject empty stubs + require loop shrink The migration assertion previously checked only role==feature, so a registration-only stub (empty hooks, logic left inline) would turn the gate green while phase 6 stayed incomplete — the exact false-completion pattern this gate exists to prevent. Strengthen it: each ADR-named feature must OWN its behavior (>=1 hook, or a command family); and plan-phase.md/execute-phase.md must shrink strictly below their frozen pre-phase-6 sizes (94519/93166 LF bytes), which also defeats double-run gaming (declare a hook but keep the inline block -> file does not shrink -> red). Gate now 5 pass / 4 fail (orphaned execute:wave:post, empty/unregistered features, config-key leaks, no shrink). Green is now reachable only by REAL migration. Refs #1168, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): migrate gap-analysis to a Capability (plan:post gate) First real ADR-857 phase-6 migration (pattern-defining tracer). gap-analysis moves from an inline post_planning_gaps branch in plan-phase.md to a real plan:post gate Capability: - capabilities/gap-analysis/capability.json: role:feature, plan:post gate (when=workflow.post_planning_gaps, blocking:false advisory), OWNS workflow.post_planning_gaps (federated out of central schema). - plan-phase.md: inline config-get + gsd_run gap-analysis block replaced with a plan:post render-hooks call site dispatching the gate; file shrinks 94519->93279. - src/check-command-router.cts: cmdGapAnalysisPlanPost runs the real gap analysis via gap-checker. - post_planning_gaps removed from central manifest; resolves via federated config (default true preserved). - tests/post-planning-gaps-2493: re-pointed to assert capability ownership. Verified: gate 5 pass / 4 fail (gap-analysis cleared from migration, plan:post-orphan, config-leak, and plan-phase shrink checks); loadConfig still returns post_planning_gaps=true; check command runs real analysis; 392/392 in the config/registry/federation/router net. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): migrate profile-pipeline to a command-family Capability ADR-857 Decision 7: profile-pipeline becomes a command-family Capability (like audit/intel/graphify). capabilities/profile-pipeline/capability.json declares an 8-command family (scan-sessions, extract-messages, profile-sample, write-profile, profile-questionnaire, generate-dev-preferences, generate-claude-profile, generate-claude-md) backed by a new gsd-core/bin/lib/profile-pipeline-command-router.cjs; the inline case arms are removed from gsd-tools.cjs. Owns profile-pipeline.enabled (federated). Verified: registry shows role:feature with commands.length=8; scan-sessions/profile-sample run live via the family; gate cleared profile-pipeline from the empty-stub failure (only tdd/schema-gate/drift remain); 296/296 registry+inventory+gsd-tools tests; lint 0 errors. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1167): wire execute:wave:post + implement ui.safety-gate check Revives the second dead gate from #1167: ui.gates@execute:wave:post was declared but never dispatched AND its check.query (ui.safety-gate) was unimplemented. Adds the per-wave execute:wave:post render-hooks call site in execute-phase.md (fires after each wave's merge/cleanup, before the next forks) and implements cmdUiSafetyGate (frontend + UI-SPEC aware, mirrors cmdUiPlanGate) in check-command-router. +17 regression tests. Verified: phase-6 orphaned-points conformance test now PASSES (gate 6 pass / 3 fail); ui-safety-gate routable in dot+hyphen forms; check-ui-safety-gate 17/17, check-ui-plan-gate 18/18; lint 0 errors. Refs #1167, #1168. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): migrate drift (schema + codebase) to execute:wave:post gates Removes the inline schema_drift_gate + codebase_drift_gate steps (77 lines) from execute-phase.md; drift becomes a Capability with two execute:wave:post gates (verify.schema-drift blocking, verify.codebase-drift advisory) dispatched via the per-wave render-hooks call site. check-command-router routes verify.schema-drift / verify.codebase-drift to the real detectors. Federates workflow.drift_threshold / drift_action / schema_drift_gate out of central. Also fixes the execute:wave:post dispatch prose to run NON-blocking (advisory) gates too — the prior version only ran blocking gates, which would have silently dropped the codebase-drift advisory after its inline step was removed. Behavior preserved. Verified: gate 7 pass / 2 fail (drift cleared from stub + config-leak; execute-phase.md 92297 < 93166 frozen -> shrink passes); both drift checks run real detection; loadConfig defaults preserved (threshold=3, action=warn, gate=true); drift-detection 56/56 + schema-drift 34/34; lint 0 errors. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): migrate tdd to a Capability (plan:pre contribution + execute:post gate) tdd becomes a real Capability: a plan:pre contribution injects the <tdd_mode_active> planner guidance (rendered from PLAN_PRE_HOOKS_JSON like security's contribution), and an execute:post gate (tdd.review-checkpoint, advisory) runs the real end-of-phase RED/GREEN review via a new check-command handler. Inline tdd_mode reads + the inline planner block + the tdd_review_checkpoint step are removed; workflow.tdd_mode is federated out of central. The MVP+TDD per-task RED-commit gate is preserved — TDD_MODE is now derived from the execute:post hooks (capId==tdd active), not an inline config-get. BEHAVIOR CHANGE (documented, not silent): the --tdd CLI flag now persists workflow.tdd_mode=true via config-set instead of being per-invocation. Rationale: tdd is now a config-toggled Capability, and env vars do not persist across the workflow's separate bash blocks (config does), so an ephemeral override isn't cleanly achievable; --tdd therefore enables the tdd capability, consistent with how all capabilities are toggled. Verified: gate 7 pass / 2 fail (tdd cleared from stub + config-leak; plan-phase + execute-phase both < frozen sizes); contribution injection + execute:post gate dispatch wired; MVP+TDD gate preserved; tdd.review-checkpoint runs real review; full unit suite 556/0; lint 0 errors. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): migrate schema-gate to a plan:pre contribution Capability The plan-time schema-push detection (former plan-phase.md §5.7) becomes a schema-gate Capability: a plan:pre contribution (into:planner, when:workflow.schema_push_detection) whose fragment carries the full ORM-detection + [BLOCKING] schema-push-task injection logic, rendered into the planner via the existing plan:pre render-hooks dispatch. The inline §5.7 block is removed (plan-phase.md 94519->90445). workflow.schema_push_detection is a new capability-owned (federated) key, default true. (The execute-side schema-drift gate was migrated separately into the drift capability.) Verified: registry inlines the fragment (len 2704) so it is actually delivered at plan:pre; gate 8 pass / 1 fail — all 5 ADR-named features now real Capabilities, only the config-leak test remains (intel/security, next unit). Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): close the 3 capability config-key leaks — phase-6 gate now GREEN Removes the last inline config-get reads of capability-owned keys from plan-phase.md. security_asvs_level/security_block_on now flow through the security plan:pre contribution via a new loop-resolver configValues mechanism (resolves declared config keys with the same 4-level precedence as activation and attaches them to the rendered hook); the §5.55 banner reads them from PLAN_PRE_HOOKS_JSON. intel.enabled becomes a real intel plan:pre step (ref.command: intel api-surface) dispatched via render-hooks; the inline intel branch is gone. gen-capability-registry now validates ref.command as a third dispatch shape. Verified: phase-6 capstone conformance gate is FULLY GREEN (9/0); 3 leaks gone (grep=0); security configValues resolve to {2,medium}/default {1,high}; intel step present only when enabled; loop-render-hooks 62/0, capability-registry 287/0, capability-state/federated-config 113/0; lint 0 errors. Closes the migration half of #1169. Refs #1139, #1167, #1168. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): address adversarial review — restore schema-drift block, generic planner injection, uniform gate contract Adversarial review caught 2 real regressions the green gate missed: (1) schema-drift no longer blocked — the execute:wave:post dispatch read GATE_RESULT.block but verify.schema-drift emitted drift_detected/blocking, and onError:skip wrongly bypassed positive blocks; (2) only tdd's plan:pre contribution was injected into the planner, dropping schema-gate's schema-push detection and security's threat-model guidance. Fixes: (A) every gate check returns a uniform boolean 'block' under --raw (the dispatch form), with advisory gates (tdd/gap) carrying their report in 'message'; (B) gate-dispatch contract corrected at all sites — onError governs command errors only, a blocking gate's positive block always halts; (C) generic planner injection of all plan:pre contributions where into=='planner' (tdd + schema-gate + security incl configValues); (D) two new conformance assertions: planner contributions injected generically + every gate check.query returns boolean block under --raw. Verified: gate 11/11; all 6 gate checks return boolean block under --raw; full suite 595/0; lint 0 errors. Refs #1167, #1168, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): restore MVP+TDD end-of-phase blocking escalation (2nd adversarial pass) The migrated tdd execute:post gate is statically blocking:false, but the contract (references/execute-mvp-tdd.md + CONTEXT.md) requires the end-of-phase TDD review to ESCALATE from advisory to blocking when MVP_MODE && TDD_MODE && a TDD plan misses a RED/GREEN commit. The migration prose had downgraded this to a 'strong advisory recommendation' — silent loss of the blocking escalation. Restore it: the tdd-gate dispatch now refuses to mark the phase complete (Phase blocked message) under MVP+TDD when GATE_RESULT.block is true; advisory otherwise. Also strengthen tests/execute-mvp-tdd-gate.test.cjs: hasBlockingEscalation previously matched any line with 'blocking'+'mvp+tdd' (so 'advisory (blocking: false) ... under MVP+TDD' was a false green); now it requires the real refusal semantics ('refuse to mark the phase complete' / 'phase blocked'). Caught by 2nd adversarial review pass. Verified: execute-phase.md 92702 < 93166 frozen; mvp-tdd-gate + phase-6 gate 19/0; full suite green; lint 0 errors. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): restore MVP+TDD proceed-block, codebase auto-remap, schema skip-flag (3rd adversarial pass) 3rd adversarial pass found 4 more silent regressions: (1) the tdd MVP+TDD 'refuse to mark complete' was nullified by a downstream 'ALWAYS proceed regardless of gate results' line — proceed is now conditional (stops on an active MVP+TDD block); (2) the test now asserts the proceed is NOT an unconditional override; (3) codebase-drift auto-remap (spawn gsd-codebase-mapper when drift_action=auto-remap) was dropped — the execute:wave:post advisory dispatch now consumes spawn_mapper/directive; (4) GSD_SKIP_SCHEMA_CHECK bypass was lost from the gate path — cmdVerifySchemaDrift now honors the env var (block:false when set). Verified: no unconditional proceed; GSD_SKIP_SCHEMA_CHECK=true -> block:false; gate 11/11 + mvp-tdd 9/9; full suite 569/0; lint 0; execute-phase.md 93109 < 93166. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): init.cts reads federated config keys from nested path (4th adversarial pass) Config federation moved tdd_mode/research/nyquist_validation from flat config.<key> to nested config.workflow.<key>, but src/init.cts still read them flat — so init.plan-phase/init.execute-phase emitted tdd_mode:false / research_enabled:undefined / nyquist:undefined regardless of config (a public command-contract regression; the migrated loops use render-hooks so enforcement was unaffected). Read via config.workflow (type-safe Record cast). Now init reflects the same resolved values + federated defaults (research/nyquist default true) as the render-hooks path. Verified: build clean; init.plan-phase emits tdd_mode:true/research:false/nyquist:false for set config, defaults true for empty; full suite 591/0; lint 0. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1169): add changeset for ADR-857 phase-6 completion (PR #1183) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): complete phase-6 migration fallout — restore TEXT_MODE, fix registry .claude leak, re-point stale workflow-contract tests The capability migration left real regressions and stale consumer tests that the per-module unit suite missed but the full cross-platform suite caught (27 failing tests): Real source regressions (fixed): - execute-phase.md lost its AskUserQuestion TEXT_MODE plain-text fallback when the inline schema_drift_gate step was removed — non-Claude runtimes would stall. Restored, and the execute:post gate-dispatch prose de-duplicated to cite the execute:wave:post contract (loop body shrinks below the frozen pre-phase-6 ceiling while keeping every onError/blocking nuance). - capabilities/tdd inline fragment hardcoded `@~/.claude/gsd-core/references/tdd.md`, baked verbatim into the committed capability-registry.cjs and leaked the install path on 11 non-Claude runtimes (registry .cjs is copied, not path-converted). Made the fragment path-free; regenerated the registry. The phase-6 conformance gate now guards this (no ~/.claude install path in any capability source or the generated registry). - plan-phase.md: removed a §5.7 stub re-added in error and routed Branch 2 to step 6 (schema-gate is a plan:pre capability, §5.7 is gone). Stale workflow-contract tests re-pointed to the capability dispatch they now must assert (behavior verified preserved in source first, assertions kept equal-or-stronger): bug-621 + bug-2851 (gap-analysis via gsd_run render-hooks plan:post + registry binding), feat-2527 (tdd_mode federated out of central), phase6-planning + plan-phase-ui-redirect (§5.6 bounded by ## 6.), plan-phase-drift-guard (intel when:intel.enabled skip branch). profile-pipeline-command-router.cjs un-ignored from eslint (hand-written, no TS source) + stale disable comments removed. Size baseline regenerated. Verified: full suite 15140 tests / 0 fail; lint 0 errors; conformance gate green legitimately. Refs #1139, #1167, #1168, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1169): add ADR-857 E2E content-test coverage for the 12 loop points + capability deliverables Grounds the capability engine in behavioral E2E tests (drive the real render-hooks/check CLI + the real registry, assert typed result content — no source-grep), structured around what ADR-857 says to deliver. 207 tests; each genuineness-checked (flip the expectation, confirm it fails). Per-loop-point dispatch (7 files): empty-point negative-space across the 6 no-hook points; verify:post 3-step resolution+ordering+onError; plan:pre contribution/configValues + ui.plan-gate + intel; plan:post gap-analysis; execute:wave:post drift+ui gates via the check route (schema-drift block/skip, codebase-drift threshold BVA, auto-remap); execute:post tdd.review-checkpoint RED/GREEN; ship:pre security gate resolution + frontmatter-get predicate pieces. ADR-deliverable coverage (4 files): predicate boundary held (edge/prohibition probes stay core, not off-by-default Feature Capabilities — phase-6 exception); core loop runs with zero capabilities (all 12 points empty, init bundles resolve); contribution merge (multiple ordered <contribution from=> blocks); federated-config key removal on uninstall. federated-config allowlisted for its 3-file split (unit + integration + lifecycle). Refs #1139, #1167, #1168, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): remove dead drifted converter dups + address adversarial review Lint cleanup (root-caused, not waved off): src/runtime-artifact-conversion.cts carried 11 agent-converter functions (+5 orphaned consts/helpers) that were never exported, never called, and had silently DRIFTED from the live hand-authored copies in bin/install.js (one even referenced an undefined `claudeToCopilotTools`). Deleted the dead duplicates; install.js's live copies are untouched (it never imported these). Lint now 0 errors / 0 warnings. Adversarial-review (Codex) findings fixed: - HIGH: execute-phase.md TDD_MODE used `jq ... || echo false`, silently disabling the MVP+TDD blocking gate on jq-less runtimes. Reverted to the `node -e` form (node is guaranteed; matches the file's other node-e usages) so a missing optional tool can no longer fail-open a blocking safety path. - MEDIUM: federated-config-key-removal orphan-key test was vacuous (it skipped the orphan assertion). Now asserts the removed capability's key is genuinely not surfaced/validated after uninstall. - LOW: phase-6 conformance leak regex broadened to catch absolute-home and Windows-backslash `.claude/(gsd-core|commands|agents|hooks)` paths, not only `~`/`$HOME` forward-slash forms. - LOW: bug-2851 plan:post dispatch assertion now requires `--raw` (matched its stated contract). - nit: plan-pre intel-step test duplicate assertion replaced with a distinct structured-output check. Size baseline regenerated (execute-phase.md 93089 < 93166 frozen). Refs #1167, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1169): make runtime-homes-descriptor-drive titles environment-independent The descriptor-equivalence test embedded the absolute golden config path (`os.homedir()`-derived) directly in each `test(...)` title, so titles differed between macOS (`/Users/x/.claude`) and Docker (`/home/gsdtest/.claude`). Every test PASSES on both platforms (15885/0 leaf tests each), but gsd-test-summary compares results by title and reported 29+29 false "only in Mac / only in Docker" discrepancies for tests that actually pass everywhere. Move the golden path out of the title and into the assertion message (still shown on failure); titles are now byte-identical across platforms so the cross-platform comparator matches them. No assertion logic or golden values changed. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): derive TDD_MODE via gsd_run --active-cap, not node -e (fix prompt-injection CI gate) The prior fix reverted execute-phase.md:181 from jq to `node -e` to close a Codex HIGH (jq||echo-false silently disabling the MVP+TDD blocking gate on jq-less runtimes) — but the CI prompt-injection scanner BLOCKS new `node -e` in workflow markdown (inline code-exec = injection vector), turning the security gate red. Both forms were wrong: node -e fails the scanner; jq fail-opens a blocking safety gate; `config-get workflow.tdd_mode` is forbidden by the conformance leak gate (tdd_mode is capability-owned). Correct fix (what Codex recommended): a gsd_run-native boolean. Add an `--active-cap <capId>` flag to `loop render-hooks <point>` that resolves hooks the normal way and prints exactly `true`/`false` for whether a capId is active — scanner-safe (canonical launcher, no inline code), node-reliable (no optional jq to fail-open), and leak-free (render-hooks resolution, not config-get). execute-phase.md:181 now `TDD_MODE=$(gsd_run loop render-hooks execute:post --active-cap tdd)`. +5 behavioral tests for the flag. Verified: prompt-injection-scan --diff origin/next → 0 findings; conformance gate 13/13 (execute-phase.md 92934 < 93166); execute-mvp-tdd + tdd-mode + loop-render-hooks 87/0; lint 0/0. Refs #1167, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
70f35ebc7a |
feat(#1169): wire ship:pre security gate via render-hooks (ADR-857 phase 6)
Part of the ADR-857 phase-6 migration (#1169, epic #857) — not a shipped bug fix. The capability system (registry, render-hooks, gates, capability-state) lives only on next / 1.5.0-rc; npm latest is 1.4.5 and contains none of it, so no released user can hit this. The security capability declares a blocking gates@ship:pre hook (when=workflow.security_enforcement, predicate SECURITY.md.threats_open==0), but ship.md never called `loop render-hooks ship:pre` and had no inline fallback, so the declared gate never fired. Wire it into ship.md preflight_checks via the render-hooks idiom — fail-closed: the ship blocks unless threats_open is exactly 0. Add a phase-6 conformance assertion that every declared hook point has a render-hooks call site; execute:wave:post (ui_safety_gate) stays in KNOWN_UNWIRED pending its unimplemented ui.safety-gate check. Part of #1169. Refs #857. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f116b76128 | test(#1139): add phase 6 capstone conformance gate (#1158) | ||
|
|
1b880fefd7 |
feat(#1137): migrate review verification hooks to capabilities (#1147)
* feat(#1137): migrate review verification hooks to capabilities * chore(#1147): add changeset |
||
|
|
44024aa535 |
fix(#1133): honor model_policy on the claude runtime (forward-port to next) (#1144)
Forward-port of the #1133 hotfix (commit 22f237a5, 1.4.5 hotfix line) onto next. The hotfix was authored against src/core.cts (v1.4.4); on next the resolver logic lives in src/model-resolver.cts (ADR-457 extraction, #888), so the patch is re-applied there rather than cherry-picked. resolveModelInternal step 2.5 now honors model_policy on the claude runtime: the policy-resolved full model ID is mapped back to a Claude Code agent alias via CLAUDE_POLICY_ID_TO_ALIAS (reverse of MODEL_ALIAS_MAP + claude-fable-5 -> fable). Bare aliases (opus/sonnet/haiku/fable) pass through; an ID with no Claude alias warns once to stderr (deduped by agentType::policyModel::tier) and falls back to the configured tier alias. Non-claude runtimes return full IDs verbatim (unchanged). resolveModelForTier is intentionally unchanged. The warn-dedupe cache lives in model-resolver.cts; core.cts composes the exported _resetRuntimeWarningCacheForTests to clear both that cache and the config-loader warning cache (config-loader cannot import model-resolver -- circular dependency). Ports the 6 #1133 tests (rewriting the old claude-no-op test that asserted the bug) plus one added test covering the MODEL_ALIAS_MAP reverse-map path. Forward-port of #1133 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4ab5c7b3f2 |
feat(#1135): migrate planning hooks to capabilities (#1141)
* feat(#1135): migrate planning hooks to capabilities * chore(#1135): add phase 6 planning capabilities changeset * fix(#1135): satisfy lint for agent hook rendering |
||
|
|
827011b865 |
fix(#1098): guard generate-claude-md against clobbering hand-crafted files; redirect default to .claude/CLAUDE.md (#1118)
/gsd-new-project wrote a repo-root CLAUDE.md full of broad project docs, overwriting/diluting a hand-crafted instruction file. --force was parsed but silently dropped, and nothing guarded an existing non-GSD file. - Guard: an existing instruction file with no `<!-- GSD:<section>-start` markers (hand-crafted) is left untouched; report action:"skipped". --force (now wired through CmdGenerateClaudeMdOptions) overwrites intentionally. The marker check uses /<!-- GSD:[a-z]+-start/ so a file merely documenting GSD syntax is safe. - Redirect: the Claude-family default output is now ./.claude/CLAUDE.md (a valid auto-loaded project-memory location) instead of repo-root ./CLAUDE.md. Aligned across the handler default, config-defaults.manifest.json, buildNewProjectConfig, the config template, new-project.md, and cmdGenerateClaudeProfile; advisory read-CLAUDE.md hints in plan-phase/quick/profile-user updated. Codex still writes AGENTS.md. Closes #1098 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
fc37ae0c4f |
fix(#1115): capability-probe codex hook-trust bypass flag in /gsd:review; fail loud on empty output (#1122)
On codex-cli < 0.137.0 the review.md `codex exec` invocation passed --dangerously-bypass-hook-trust (added in 0.137) unconditionally and discarded stderr, so codex exited "unexpected argument" before reading the prompt and the empty output file was treated as a completed review — a silent degraded review. - Capability-probe the flag (`codex exec --help | grep`) and apply it via $CODEX_BYPASS_FLAG only when supported (works fine without it on older CLIs). - Capture codex stderr to a .err file instead of /dev/null, and replace an empty output with a diagnostic so a broken reviewer is surfaced (mirrors Cursor/OpenCode). - Update enh-773 enforcement test to require the capability gate + fail-loud guard instead of the unconditional flag. Closes #1115 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1c073c1c81 |
fix(#1107): progress consults verification.status before reporting a phase complete (#1116)
/gsd-progress derived phase completeness from plan/summary counts only and never consulted the verification.status query (the #651 seam), so a phase whose VERIFICATION.md ended human_needed or gaps_found was reported complete and routing skipped to the next phase. Add Step 1.7 (consult verification.status for the current phase) and routing rows that send gaps_found to plan-phase --gaps (Route V.gaps) and human_needed to verify-work (Route V.human) before the generic complete row. passed/missing/unknown still route as complete so unverified phases are not falsely blocked. Closes #1107 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3e836fef0d |
feat(spec-phase): spec-completeness edge-probe (#550) (#584)
* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/
Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.
Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.
Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
source-checkout-gated build fallback) instead of LLM re-derivation; the engine
capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.
Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.
* test(#550): RED — status×verification re-cut + probe-core engine specs
Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):
status: resolved | dismissed | unresolved (lifecycle, shared)
verification: explicit | backstop | null (only when resolved)
- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
to be extracted — validateResolution(r, validators), validateRequirement,
analyzeCoverage(items, resolutions?, validators), byVerification rollup,
runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
{resolved,backstop}; coverage gains byVerification.{explicit,backstop};
proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
COUNT preserved on every fixture (closed set = resolved+dismissed; doc
line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
rewritten to the two-axis model.
Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).
* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)
Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.
probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
items[] (core never assumes propose is deterministic — edge resolves via
LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
{categories, verification, requiredFieldsByVerification} (ADR-550 #5)
edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.
* chore(#550): register probe-core.cjs artifact in ledgers + inventory
New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:
- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).
probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.
* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]
trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.
Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.
* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard
Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:
- validateResolution now enforces the 'verification is null unless resolved'
invariant for EVERY status (not just resolved): a dismissed/unresolved
resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
(was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
contract; a new test locks that an all-dismissed run is NOT affirmatively covered
(byVerification is the honest gate).
Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.
* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics
Re-review #5 (trek-e) clarity edits:
- Decision 5: annotate that only contract item (a) ships on #584 (the edge
adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
parenthetical — the blessed/implemented semantics are count-preserved = the
CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
carrying the per-tier resolved-status breakdown. The old parenthetical
contradicted the shipped count.
* test(#550): cover runProbeCli structural-guard numeric-count branch
Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.
* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)
The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.
* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)
A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.
* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)
templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.
* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)
The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.
* docs(#550): add how-to for resolving edge-coverage findings (B1)
Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.
* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)
trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.
* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)
trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.
* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)
trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.
* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe
Rebased onto next (
|
||
|
|
77a671ec53 |
feat(#791): migrate antigravity workspace base dir .agent → .agents (#1090)
Fresh antigravity workspace installs write under the canonical .agents/ (plural) base; legacy .agent/ stays recognized (dual-read). Global ~/.gemini/antigravity/ path unchanged. Closes #791. |
||
|
|
7c07fce70f |
fix(#381): make gsd_run launcher reachable in fresh-shell-per-block runtimes (#1084)
* fix(#381): make gsd_run launcher reachable in fresh-shell-per-block runtimes On runtimes that execute each fenced bash block in a separate shell process (e.g. Claude Code — documented behavior: each Bash command is a separate process; inline shell functions and exported vars do not persist between calls), the once-per-file gsd_run() function was undefined in every block after the preamble block, and the call was swallowed by `2>/dev/null || echo "{}"` into silent empty state. Fix (budget-neutral session-level resolution): - Ship gsd-core/bin/gsd_run, a POSIX sh wrapper that symlink-resolves its own location and execs the co-located gsd-tools.cjs. Exposed on PATH via the npm `bin` field (global installs) and shipped to local installs via the recursive gsd-core/ copy. - The per-file launcher preamble now appends `export PATH='<bindir>':"$PATH"` to the file named by $CLAUDE_ENV_FILE (Claude Code's documented env-persistence mechanism) so later fresh-shell blocks resolve gsd_run from PATH. Guarded as a strict no-op when CLAUDE_ENV_FILE is unset; the inline gsd_run() definition remains the fallback for all other runtimes. The single-quoted dir neutralizes shell metacharacters at source time. - Propagated via scripts/sync-runtime-launcher.cjs to all launcher-using files. - XL workflow byte budget 93000 -> 93200 (the ~130B clause pushes plan-phase.md to 93135; legitimate content growth, ratchet-up per #717). Regression tests (I)/(J) in runtime-launcher-parity.test.cjs cover wrapper delegation and end-to-end PATH persistence (sourcing the env file with a space-bearing install path). Known limitation: an install path containing a literal single-quote yields a malformed env-file line and falls back to the status quo (no regression); rare on sanitized home directories. Closes #381 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#381): add changeset for gsd_run fresh-shell reachability fix Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#381): scope test (J) bare-PATH execution to POSIX (Windows Git Bash exec bit) Windows Git Bash (msys2) does not honor Node's chmod exec bit for PATH-executing extension-less scripts, so the bare `gsd_run` command lookup failed there even though the env-file PATH persistence was correct. The env-file content assertions (the fix's actual cross-platform logic) still run on every platform; only the final source-and-execute sub-step is gated to non-win32. Global installs on Windows are covered by npm's generated bin shim. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9e3b056b15 |
fix(#779): correct stale model-catalog model IDs verified against live providers (#1047)
Verify-first audit: gemini opus gemini-3-pro→gemini-3.1-pro-preview (undefined in gemini-cli source), codex sonnet gpt-5.3-codex→gpt-5.4 (deprecated per OpenAI); qwen3-coder-next verified valid, unchanged. Adds a regression guard + sourcing note. Closes #779. |
||
|
|
e4dfa6b9ea |
fix(#1012): invoke fallow with its real CLI and wire the report normalizer (#1044)
* fix(#1012): invoke fallow with its real CLI and wire the report normalizer The /gsd-code-review structural pre-pass invoked fallow with flags no published fallow version accepts (--json, --profile, --stdin-files), so it failed on every run and degraded silently per REQ-FALLOW-02 — the feature never delivered on any fallow version. Three compounding defects: 1. Invalid flags. Real fallow audit uses --format json (not --json), -q/--quiet, --changed-since/--base for changed-files scoping (no file-list input), and --max-crap for thresholds. There is no --profile or --stdin-files. 2. Exit-code handling. fallow audit exits 1 when it FINDS issues (verdict=fail), 0 when clean. The pre-pass treated any non-zero exit as a crash and discarded the output — i.e. it threw away exactly the findings it exists to surface. Success is now decided by whether a valid fallow JSON report was produced, not by the exit code. 3. Schema mismatch. normalizeFallowReport parsed a fictional top-level schema (unusedExports/duplicates/circularDependencies) fallow never shipped, and was dead code (the workflow embedded raw JSON; its tests asserted the fictional schema, one even calling a non-existent runFallowAudit and passing vacuously). Fixes: align the invocation to fallow's documented agent-facing pattern; map the profile preset (minimal/standard/strict) to --max-crap (50/30/15); scope phase runs via --changed-since with a repo-scope fallback; rewrite the normalizer to fallow's real schema (dead_code.unused_exports/unused_files/circular_dependencies + duplication.clone_groups) and wire it into the workflow so the reviewer receives normalized findings; replace the fictional-schema fixtures and tests with real-schema ones and delete the vacuous runFallowAudit test. Closes #1012 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1012): backfill changeset PR number to 1044 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1e3ce6df05 |
fix(#222): orchestrator self-heal when synthesizer returns SUMMARY.md inline (#1042)
* fix(#222): orchestrator self-heal when synthesizer returns SUMMARY.md inline The gsd-research-synthesizer agent intermittently hits an LLM false-refusal: instead of writing .planning/research/SUMMARY.md with the Write tool, it returns the SUMMARY.md content inline and fabricates a non-existent write restriction (e.g. "the runtime is blocking file writes"). The shipped prompt hardening (#240) is necessary but insufficient — the false-refusal recurs under some context loads, and a drifting subagent then leaves gsd-roadmapper to fail with "SUMMARY.md not found". Adds an orchestrator-level self-heal to new-project.md and new-milestone.md: after the synthesizer returns, verify .planning/research/SUMMARY.md exists; if it is missing but the agent returned content inline, the orchestrator persists that content with the Write tool (logging a warning) before spawning gsd-roadmapper; if missing with no content, surface the error and stop rather than proceed against a missing SUMMARY.md. This absorbs the failure mode deterministically instead of depending on the subagent never drifting. Closes #222 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#222): backfill changeset PR number to 1042 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8bd5e07c58 |
feat(#1031): autonomous §3a.5 plan:pre ui-phase cutover — last inlined ui-phase site (step-only) (#1033)
Cut over autonomous.md §3a.5 (autonomous plan:pre ui-phase step) to the loop.render-hooks plan:pre dispatch — completing the ui-phase migration begun in #1026 (plan-phase.md §5.6). Step-only, non-blocking: autonomous is always pipeline, so it fires active kind==step hooks and never runs the manual-only plan:pre blocking gate. Skip condition keys on "no active step hooks" (not empty activeHooks), so the gate-only {ui_phase:false, ui_safety_gate:true} case skips silently with no spurious warning — matching OLD §3a.5. Fires gsd-ui-phase under the identical precondition (frontend + no UI-SPEC + workflow.ui_phase active), bare ${PHASE_NUM} args. Replaces the inline ui-safety-gate.cjs probe + config-get with render-hooks + the ui.plan-gate check verb. Codex caught the gate-only spurious-warning divergence on the first pass; fixed + re-confirmed equivalence-preserving. gsd-ui-phase skill, §5.6, §3d.5 untouched. Closes #1031 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9ed8c7d574 |
feat(#1026): §5.6/ui-phase cutover — first gate dispatch (plan:pre step + blocking gate) (#1028)
Replace plan-phase.md §5.6 (a 6-branch inline UI gate) with a capability-driven
loop.render-hooks plan:pre dispatch — the FIRST gate dispatch in any workflow.
A step (ui-phase, when:workflow.ui_phase) + a new blocking gate
(when:workflow.ui_safety_gate). New ui.plan-gate check verb returns
{frontend, hasUiSpec, block}; the dispatch runs it unconditionally then fires
the active step (pipeline) or halts on the active blocking gate (manual). The
gate-handling (run check.query; halt if blocking+block) is the reusable
phase-6 template for blocking-gate cutovers.
Config semantics fixed per #1022 + maintainer call: ui_phase gates plan-time
UI-SPEC generation, ui_safety_gate gates the planning block. Common case + all
ui_phase=false cases are equivalence-preserving; the one intended change is
{ui_phase:true, ui_safety_gate:false} now auto-generating in pipelines.
Review found it broken twice (non-generic dispatch, phase-lookup divergence,
then the step-only check nested in a gate loop) — fixed; final Codex pass
verified all 8 (ui_phase,ui_safety_gate)x{pipeline,manual} cases correct.
gsd-ui-phase skill + autonomous §3a.5 untouched (§3a.5 deferred).
getRoadmapPhaseWithFallback mirrors cmdRoadmapGetPhase for lookup parity.
Closes #1026
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|