5b5e5bcf2ea14c2aa9deb0dc2b81cfcb7e5cdbc0
26 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
91bd82f9a6 |
enhance(execute): isolated-executor rejected/over-reaching run fails safe — never default recovery to main (#1292) (#1303)
* enhance(execute): isolated-executor rejected/over-reaching run fails safe (#1292) When an isolated (worktree) executor run is rejected — the user declines to merge it, the orchestrator surfaces recovery for a blocked/halted plan, or the run over-reached the requested scope — the orchestrator must no longer default/propose recovery by editing the primary checkout (`main`). Absent an explicit guardrail, the LLM orchestrator could improvise "continue on main", inverting the isolation contract at the moment it matters most. Added an ISOLATED-RUN RECOVERY — FAIL SAFE policy: default to a safe halt that offers a fresh, narrowly-scoped worktree or inspect/discard; editing the primary checkout requires explicit, clearly-labeled confirmation and is never the default/proposed option. To respect the ADR-857 phase-6 host-loop size cap on execute-phase.md (it sits just under the pre-phase-6 baseline), the policy is delivered as an extracted reference fragment rather than inline: - New `execute-phase/steps/worktree-recovery-policy.md` holds the recovery policy (the existing FAIL-CLOSED rule #48 for base/HEAD mismatches + the #1292 fail-safe guardrail). No #48 behavior change — moved verbatim. - execute-phase.md references the fragment at the worktree-spawn recovery point, the step-5.5 merge decision, and the stalled-agent "switch to inline execution" menu (which for an isolated run now follows the fail-safe policy). Net effect: execute-phase.md shrinks below its cap. - quick.md carries the fail-safe guardrail inline at its post-return merge/discard decision (quick.md is not size-capped). Scoped to the recovery offer only — no automatic scope-overreach detection (explicitly out of scope per the issue) and no new config key. Adds content regression tests, a USER-GUIDE note, and a workflow size-baseline update. Closes #1292 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(changeset): Changed fragment for #1292 isolated-executor fail-safe recovery Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
137760a655 |
fix(#1296): align config docs/prompts/schema with consumers (#1299)
* fix(#1296): align config docs/prompts/schema with consumers The user-facing config surface disagreed with what the consumers actually do (subset of the #1216 audit). No runtime consumption behavior changes. - workflow.subagent_timeout: settings-advanced.md prompt + docs/CONFIGURATION.md said "seconds (default 600)" but the consumer (map-codebase.md) uses milliseconds (default 300000). Relabeled all four spots in settings-advanced.md (prompt, parse-default list, example, confirmation table) + the CONFIGURATION.md row. - review.models.<cli>: settings-integrations.md, docs/CONFIGURATION.md (Integration Settings), and docs/CLI-TOOLS.md documented a shell command, but review.md injects the value into a --model/-m flag. Relabeled to a bare model id and reconciled the contradictory CONFIGURATION.md sections. - workflow.test_command + workflow.build_command: consumed via config-get (test_command in verify-phase/execute-phase/audit-fix/post-merge-gate; build_command in post-merge-gate) and documented, but absent from validKeys so `config set` rejected them. Registered both in config-schema.manifest.json and documented them in references/planning-config.md (overview + complete reference). Regression tests: behavioral config-set tests (tests/config.test.cjs) + doc-parity content guards (tests/config-field-docs.test.cjs). Deferred to other #1216 clusters: security-gate wiring, search_gitignored wiring, mvp_mode, source_grounding_authority labeling, and config-set enum enforcement. Closes #1296 Refs #1216 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(changeset): Fixed fragment for #1296 config-surface alignment Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0d783fa57d |
Merge remote-tracking branch 'origin/next' into feat/1259-test-tier-enforcement
# Conflicts: # tests/workflow-size-baseline.json |
||
|
|
0a856f06cd |
enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier (#1271)
* enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier Introduce a per-truth PRESENT_BEHAVIOR_UNVERIFIED state for must-haves that assert a state transition or a cancellation/cleanup/ordering invariant whose only evidence is symbol presence + wiring. Such truths are excluded from the verified_truths score, reported as a behavior_unverified count, recorded in an always-on behavior_unverified_items frontmatter list, and routed to the existing human_needed sink — so a clean N/N can no longer be reached on symbol presence alone. The overall-status vocabulary and the src/verification.cts seam are unchanged (the new state is per-truth only); gaps_found keeps decision-tree precedence; override-passed truths still count toward verified_truths. Mirrors the calibration into the shipped verify-phase.md workflow (with an infra/foundation carve-out), the VERIFICATION.md templates, and docs (planning-artifacts.md, AGENTS.md). gsd-verifier.md kept under its 48KB LARGE cap; size baselines regenerated. Closes #966 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#966): add changeset fragment for gsd-verifier behavior-unverified calibration Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6a9e6cb0ae |
fix(1259-01): address trek-e review — B1 (fatal/suppressed), B2 (bounded subprocess), M1/M2, minors
Maintainer CHANGES_REQUESTED (reviewed |
||
|
|
e08667e5aa |
test(1259-01): regenerate workflow-size baseline for verify-phase.md growth (+1459 B)
verify-phase.md grew 30875 -> 32334 from the test-tier enforcement consumer step + descriptor docs. Well under the 40960 DEFAULT hard cap; regenerate the per-file baseline ratchet (the growth is the load-bearing consumer wiring, not bloat). |
||
|
|
3556450b0d |
feat(spec-phase): surface zero-classification edge-probe requirements as unclassified candidates (#1110) (#1117)
Edge-probe now surfaces a zero-classification requirement (non-empty prose, no shape cue matched, no `shapes` override) as a single soft `unclassified — review manually` candidate instead of silently dropping it — the exact blind spot the probe exists to catch. Dismissible like any edge; the `shapes: []` opt-out stays silent; `TAXONOMY` (the closed 8 categories) is unchanged. Under `--auto` the candidate is left `unresolved`, never auto-`backstop` (a missing shape is not evidence an edge exists). Closes #1110 |
||
|
|
395fb519e7 |
feat(spec-phase): prohibition probe — surface "must-NOT" constraints (#644) (#1149)
Adds the spec-time prohibition probe (spec-phase Step 5.6) — the second adapter of the probe-core resolution model. Surfaces unwritten must-NOT constraints as negative SPEC acceptance criteria with test/judgment verification tiers; fail-closed at verify time. Per ADR-550. Closes #644. |
||
|
|
55eb1fd72e |
refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam (#1242)
* refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam ADR-22's severity mapping, authority auto-upgrade, and rung>=3 hard-block lived only as prose in plan-review-convergence.md — untestable. Extracted into src/plan-drift-guard.cts (pure: AUTHORITY_RUNGS, getEffectiveAuthority, classifyDriftSeverity) + a gsd-tools drift-guard CLI seam (authority/severity), and rewired the workflow to call the seam deterministically instead of reasoning the decision in prose. 47 unit/e2e/structural tests cover the full severity table, the grep->intel auto-upgrade, and the rung>=3 HIGH hard-block. Closes #1190 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1190): add changeset for ADR-22 drift-guard seam (#1242) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1190): register ADR-22 module in eslint-ignore + inventory manifest + docs-exempt changeset Full-matrix CI surfaced new-module/command governance ripples beyond the lint-tests chain: (1) tsc-generated plan-drift-guard.cjs must be in the eslint ignore list (551-eslint-bin-lib-coverage); (2) docs/INVENTORY-MANIFEST.json must include the new module/command (regen via gen-inventory-manifest.cjs --write); (3) a type:Added changeset triggers docs-required — added a docs-exempt marker (internal seam, no user-facing surface). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0b3a2e5f9c |
feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15) (#1237)
* feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15) ADR-15 designates /gsd-progress --next --auto --converge as the PRIMARY plan-convergence surface, but only the secondary surface (autonomous.md) was wired. next.md now parses --converge/--cross-ai into a plan strategy, gates on workflow.plan_review_convergence, forwards reviewer flags + --max-cycles, and routes Route-3 planning through /gsd:plan-review-convergence (mirroring autonomous.md); --auto chaining preserves converge mode. Adds argument-hint + help/full.md + COMMANDS.md + how-to parity and a structural regression test. Closes #1190 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1190): add changeset for progress --converge surface (#1237) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1190): keep progress --converge docs skill-dep-clean + regen workflow size baseline CI surfaced two ripples from the ADR-15 workflow edits: (1) lint-skill-deps + profile-closure flagged /gsd:plan-phase and /gsd:plan-review-convergence SlashCommand tokens in progress.md's flag docs as undeclared deps — reworded to plain prose since progress.md only advertises the flag (the real invocation lives in next.md); (2) the per-file workflow size baseline needed regenerating after the next.md/help edits. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
bf634b95c3 |
feat(#1213): Capability State Writer — write-side inverse of the resolver (#1225)
* feat(#1213): Capability State Writer — write-side inverse of the resolver Adds src/capability-writer.cts (setCapabilityState + cmdCapabilitySet) and the `gsd-tools capability set` subcommand: the write-side inverse of the capability resolver (ADR-1213). One desired capability state projects onto the substrates — `enabled` drives the runtime surface (canonical on/off), `gates` drive federated config keys (hook granularity), install profile is a read-only floor — then re-resolves and reports divergence (assert-and-report), so "off means off" holds as a write-time invariant. Adds batched setConfigValues; routes gsd:settings capability hook-gates through the writer. Docs: CLI-TOOLS reference, how-to, ADR-1213, CONTEXT.md term. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1213): add changeset for Capability State Writer (#1225) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
7edd18fd2b |
feat(#1165): async external_job_waiting half-state + resume/pause contract (#1221)
Core half of #1105: a legal external_job_waiting deferred state so an async-dispatched Execute step (committing a .planning/async-jobs/<job>.json manifest, deferring SUMMARY.md) is not an illegal partial. execute-phase safe-resume, resume-project, and pause-work reconcile against the versioned scheduler-agnostic manifest stability contract without re-dispatching; the producer is the capability half (#1164). Closes #1165. |
||
|
|
e9f9ae49c8 |
fix(#1146): single base-branch resolver across forking workflows (#1198)
* fix(#1146): single base-branch resolver across forking workflows Replaces duplicated per-workflow bash detection that silently fell through to :-main on repos where origin/HEAD is unset (git init+remote add+fetch without set-head, most CI checkouts, many worktrees). New CJS module git-base-branch.cjs exposes `gsd_run query git.base-branch` with full precedence ladder: git.base_branch config override → origin/HEAD symref → git remote show origin (authoritative) → local branch presence → "main". All git subprocesses bounded with timeouts; degrades gracefully. Wires execute-phase, quick, ship, complete-milestone, and pr-branch to the single resolver. Removes 14 lines of duplicated detection bash across the five workflows. Includes 7 behavioral tests covering the full precedence ladder including the key regression case (master repo, origin/HEAD unset → must return "master", NOT "main") and an anti-regression guard that fails if any workflow re-introduces the :-main/:-master fallback pattern. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(changeset): backfill PR number #1198 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1146): drop stray PR-body file from branch pr-1146-body.md was committed during changeset backfill but must not be tracked in the repo. Content preserved externally for PR body use. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#1146): add tests for flat base_branch config key and both-branch tie-break Closes two mutation gaps identified in adversarial review: - A2: flat {base_branch: ...} at config root (legacy key form) was covered by code but unguarded against mutation of lines 74-75 in resolver - H: tier-4 tie-break when both main+master exist locally (main wins, per tryLocalBranch JSDoc) was documented but untested 9/9 tests pass. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#1146): allowlist workflow-literal guard as runtime-contract exemption Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#1146): degrade gracefully when gsd_run unavailable in handle_branching bash blocks handle_branching (execute-phase.md) and step 2.5 (quick.md) are extracted and run verbatim by behavioral tests that lack the gsd_run preamble. Adding a || fallback ladder (git symbolic-ref then echo main) keeps the unified resolver as primary in real workflows while letting the test harness succeed without gsd_run defined. Also propagates updated runtime-launcher preamble to pr-branch.md (added in origin/next MemPalace PR) and regenerates workflow-size-baseline.json. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a375c4b354 |
feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) (#1201)
* feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) Adds an opt-in, default-resilient ADR-857 feature capability that wires MemPalace (local-first memory: MCP server + CLI) into the GSD loop: deliberate recall before discuss/plan and verbatim + temporal-KG capture at phase boundaries. Three memory modes (augment default; kg_backend and replace forward-declared). Master gate mempalace.enabled (default off); every hook onError:skip, zero gates; absent/disabled MemPalace => loop unchanged. Transport is rendered-markdown only — MemPalace runs out-of-process, no third-party code in gsd-core (ADR-857 §7). Capability: capabilities/mempalace/ (manifest + 2 fragments), skills commands/gsd/mempalace-{recall,capture}.md, agent agents/gsd-mempalace-curator.md. Registration: ns-context router, utility cluster, KNOWN_SKILLS, help full.md, model-catalog, copilot install list, size baselines; regenerated capability-registry + inventory manifest. ship:post wired into ship.md (wire-on-demand). HELD on #1196: this capability also declares hooks at discuss:pre and discuss:post, which are structurally un-wireable until the host-loop conformance model covers the discuss phase (discuss-phase.md is not in HOST_LOOP_FILES). The phase6-capstone-conformance gate therefore fails on exactly those two orphaned points by design — see #1196. Once #1196 lands, rebase onto next and the gate goes green with no further change. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#956): backfill changeset PR number (#1201) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b1e8a74708 |
fix(#1196): wire discuss loop step for capability hooks (#1199)
* fix(#1196): wire discuss loop step for capability hooks discuss was contract-declared (gsd:loop-host marker, in POINT_ORDER and LOOP_HOST_CONTRACT) but structurally unwireable: discuss-phase.md had no `loop render-hooks` dispatch and was absent from the conformance gate's HOST_LOOP_FILES, so capabilities could never wire discuss:pre/discuss:post. - discuss-phase.md: add minimal discuss:pre (before analyze_phase) and discuss:post (after write_context) render-hooks dispatch steps that delegate consumption to a new shared reference (kept under the 32KB #2551 budget; no inline subagent dispatch token). - references/loop-hook-dispatch.md: new canonical, point-agnostic contract for consuming `loop render-hooks --raw` activeHooks (contribution/step/ gate) — single source for hook consumption across host loops. - gen-loop-host-contract.cjs: derive HOST_LOOP_FILES from STEP_WORKFLOWS and export scanWiredPoints()/getWiredLoopPoints() (throws on a missing host file) — one source of truth for the host-loop file + wired-point set. - phase6-capstone-conformance.test.cjs: consume the derived HOST_LOOP_FILES and shared scanWiredPoints (was a hand-maintained duplicate omitting discuss-phase.md + a duplicated regex). - gen-capability-registry.cjs: add validateHooksWired() gen-time guard that rejects a capability hook declared at a valid-but-unwired loop point, with a clear remediation message — failure now surfaces at gen --check/--write time instead of deep in the full conformance suite. - tests (capability-registry.test.cjs): regression + anti-pattern parity guards (every loop-host marker is in STEP_WORKFLOWS/HOST_LOOP_FILES; POINT_ORDER === flattened LOOP_HOST_CONTRACT) so no step can drift into the discuss-class gap again. - docs/INVENTORY*: register the new reference. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1196): backfill changeset PR number (#1199) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b10e56818b |
feat(#1169): complete ADR-857 phase 6 — migrate features to Capabilities, revive dead gates, harden conformance gate (#1183)
* test(#1168): make phase-6 gate un-gameable — reject empty stubs + require loop shrink The migration assertion previously checked only role==feature, so a registration-only stub (empty hooks, logic left inline) would turn the gate green while phase 6 stayed incomplete — the exact false-completion pattern this gate exists to prevent. Strengthen it: each ADR-named feature must OWN its behavior (>=1 hook, or a command family); and plan-phase.md/execute-phase.md must shrink strictly below their frozen pre-phase-6 sizes (94519/93166 LF bytes), which also defeats double-run gaming (declare a hook but keep the inline block -> file does not shrink -> red). Gate now 5 pass / 4 fail (orphaned execute:wave:post, empty/unregistered features, config-key leaks, no shrink). Green is now reachable only by REAL migration. Refs #1168, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): migrate gap-analysis to a Capability (plan:post gate) First real ADR-857 phase-6 migration (pattern-defining tracer). gap-analysis moves from an inline post_planning_gaps branch in plan-phase.md to a real plan:post gate Capability: - capabilities/gap-analysis/capability.json: role:feature, plan:post gate (when=workflow.post_planning_gaps, blocking:false advisory), OWNS workflow.post_planning_gaps (federated out of central schema). - plan-phase.md: inline config-get + gsd_run gap-analysis block replaced with a plan:post render-hooks call site dispatching the gate; file shrinks 94519->93279. - src/check-command-router.cts: cmdGapAnalysisPlanPost runs the real gap analysis via gap-checker. - post_planning_gaps removed from central manifest; resolves via federated config (default true preserved). - tests/post-planning-gaps-2493: re-pointed to assert capability ownership. Verified: gate 5 pass / 4 fail (gap-analysis cleared from migration, plan:post-orphan, config-leak, and plan-phase shrink checks); loadConfig still returns post_planning_gaps=true; check command runs real analysis; 392/392 in the config/registry/federation/router net. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): migrate profile-pipeline to a command-family Capability ADR-857 Decision 7: profile-pipeline becomes a command-family Capability (like audit/intel/graphify). capabilities/profile-pipeline/capability.json declares an 8-command family (scan-sessions, extract-messages, profile-sample, write-profile, profile-questionnaire, generate-dev-preferences, generate-claude-profile, generate-claude-md) backed by a new gsd-core/bin/lib/profile-pipeline-command-router.cjs; the inline case arms are removed from gsd-tools.cjs. Owns profile-pipeline.enabled (federated). Verified: registry shows role:feature with commands.length=8; scan-sessions/profile-sample run live via the family; gate cleared profile-pipeline from the empty-stub failure (only tdd/schema-gate/drift remain); 296/296 registry+inventory+gsd-tools tests; lint 0 errors. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1167): wire execute:wave:post + implement ui.safety-gate check Revives the second dead gate from #1167: ui.gates@execute:wave:post was declared but never dispatched AND its check.query (ui.safety-gate) was unimplemented. Adds the per-wave execute:wave:post render-hooks call site in execute-phase.md (fires after each wave's merge/cleanup, before the next forks) and implements cmdUiSafetyGate (frontend + UI-SPEC aware, mirrors cmdUiPlanGate) in check-command-router. +17 regression tests. Verified: phase-6 orphaned-points conformance test now PASSES (gate 6 pass / 3 fail); ui-safety-gate routable in dot+hyphen forms; check-ui-safety-gate 17/17, check-ui-plan-gate 18/18; lint 0 errors. Refs #1167, #1168. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): migrate drift (schema + codebase) to execute:wave:post gates Removes the inline schema_drift_gate + codebase_drift_gate steps (77 lines) from execute-phase.md; drift becomes a Capability with two execute:wave:post gates (verify.schema-drift blocking, verify.codebase-drift advisory) dispatched via the per-wave render-hooks call site. check-command-router routes verify.schema-drift / verify.codebase-drift to the real detectors. Federates workflow.drift_threshold / drift_action / schema_drift_gate out of central. Also fixes the execute:wave:post dispatch prose to run NON-blocking (advisory) gates too — the prior version only ran blocking gates, which would have silently dropped the codebase-drift advisory after its inline step was removed. Behavior preserved. Verified: gate 7 pass / 2 fail (drift cleared from stub + config-leak; execute-phase.md 92297 < 93166 frozen -> shrink passes); both drift checks run real detection; loadConfig defaults preserved (threshold=3, action=warn, gate=true); drift-detection 56/56 + schema-drift 34/34; lint 0 errors. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): migrate tdd to a Capability (plan:pre contribution + execute:post gate) tdd becomes a real Capability: a plan:pre contribution injects the <tdd_mode_active> planner guidance (rendered from PLAN_PRE_HOOKS_JSON like security's contribution), and an execute:post gate (tdd.review-checkpoint, advisory) runs the real end-of-phase RED/GREEN review via a new check-command handler. Inline tdd_mode reads + the inline planner block + the tdd_review_checkpoint step are removed; workflow.tdd_mode is federated out of central. The MVP+TDD per-task RED-commit gate is preserved — TDD_MODE is now derived from the execute:post hooks (capId==tdd active), not an inline config-get. BEHAVIOR CHANGE (documented, not silent): the --tdd CLI flag now persists workflow.tdd_mode=true via config-set instead of being per-invocation. Rationale: tdd is now a config-toggled Capability, and env vars do not persist across the workflow's separate bash blocks (config does), so an ephemeral override isn't cleanly achievable; --tdd therefore enables the tdd capability, consistent with how all capabilities are toggled. Verified: gate 7 pass / 2 fail (tdd cleared from stub + config-leak; plan-phase + execute-phase both < frozen sizes); contribution injection + execute:post gate dispatch wired; MVP+TDD gate preserved; tdd.review-checkpoint runs real review; full unit suite 556/0; lint 0 errors. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): migrate schema-gate to a plan:pre contribution Capability The plan-time schema-push detection (former plan-phase.md §5.7) becomes a schema-gate Capability: a plan:pre contribution (into:planner, when:workflow.schema_push_detection) whose fragment carries the full ORM-detection + [BLOCKING] schema-push-task injection logic, rendered into the planner via the existing plan:pre render-hooks dispatch. The inline §5.7 block is removed (plan-phase.md 94519->90445). workflow.schema_push_detection is a new capability-owned (federated) key, default true. (The execute-side schema-drift gate was migrated separately into the drift capability.) Verified: registry inlines the fragment (len 2704) so it is actually delivered at plan:pre; gate 8 pass / 1 fail — all 5 ADR-named features now real Capabilities, only the config-leak test remains (intel/security, next unit). Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): close the 3 capability config-key leaks — phase-6 gate now GREEN Removes the last inline config-get reads of capability-owned keys from plan-phase.md. security_asvs_level/security_block_on now flow through the security plan:pre contribution via a new loop-resolver configValues mechanism (resolves declared config keys with the same 4-level precedence as activation and attaches them to the rendered hook); the §5.55 banner reads them from PLAN_PRE_HOOKS_JSON. intel.enabled becomes a real intel plan:pre step (ref.command: intel api-surface) dispatched via render-hooks; the inline intel branch is gone. gen-capability-registry now validates ref.command as a third dispatch shape. Verified: phase-6 capstone conformance gate is FULLY GREEN (9/0); 3 leaks gone (grep=0); security configValues resolve to {2,medium}/default {1,high}; intel step present only when enabled; loop-render-hooks 62/0, capability-registry 287/0, capability-state/federated-config 113/0; lint 0 errors. Closes the migration half of #1169. Refs #1139, #1167, #1168. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): address adversarial review — restore schema-drift block, generic planner injection, uniform gate contract Adversarial review caught 2 real regressions the green gate missed: (1) schema-drift no longer blocked — the execute:wave:post dispatch read GATE_RESULT.block but verify.schema-drift emitted drift_detected/blocking, and onError:skip wrongly bypassed positive blocks; (2) only tdd's plan:pre contribution was injected into the planner, dropping schema-gate's schema-push detection and security's threat-model guidance. Fixes: (A) every gate check returns a uniform boolean 'block' under --raw (the dispatch form), with advisory gates (tdd/gap) carrying their report in 'message'; (B) gate-dispatch contract corrected at all sites — onError governs command errors only, a blocking gate's positive block always halts; (C) generic planner injection of all plan:pre contributions where into=='planner' (tdd + schema-gate + security incl configValues); (D) two new conformance assertions: planner contributions injected generically + every gate check.query returns boolean block under --raw. Verified: gate 11/11; all 6 gate checks return boolean block under --raw; full suite 595/0; lint 0 errors. Refs #1167, #1168, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): restore MVP+TDD end-of-phase blocking escalation (2nd adversarial pass) The migrated tdd execute:post gate is statically blocking:false, but the contract (references/execute-mvp-tdd.md + CONTEXT.md) requires the end-of-phase TDD review to ESCALATE from advisory to blocking when MVP_MODE && TDD_MODE && a TDD plan misses a RED/GREEN commit. The migration prose had downgraded this to a 'strong advisory recommendation' — silent loss of the blocking escalation. Restore it: the tdd-gate dispatch now refuses to mark the phase complete (Phase blocked message) under MVP+TDD when GATE_RESULT.block is true; advisory otherwise. Also strengthen tests/execute-mvp-tdd-gate.test.cjs: hasBlockingEscalation previously matched any line with 'blocking'+'mvp+tdd' (so 'advisory (blocking: false) ... under MVP+TDD' was a false green); now it requires the real refusal semantics ('refuse to mark the phase complete' / 'phase blocked'). Caught by 2nd adversarial review pass. Verified: execute-phase.md 92702 < 93166 frozen; mvp-tdd-gate + phase-6 gate 19/0; full suite green; lint 0 errors. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): restore MVP+TDD proceed-block, codebase auto-remap, schema skip-flag (3rd adversarial pass) 3rd adversarial pass found 4 more silent regressions: (1) the tdd MVP+TDD 'refuse to mark complete' was nullified by a downstream 'ALWAYS proceed regardless of gate results' line — proceed is now conditional (stops on an active MVP+TDD block); (2) the test now asserts the proceed is NOT an unconditional override; (3) codebase-drift auto-remap (spawn gsd-codebase-mapper when drift_action=auto-remap) was dropped — the execute:wave:post advisory dispatch now consumes spawn_mapper/directive; (4) GSD_SKIP_SCHEMA_CHECK bypass was lost from the gate path — cmdVerifySchemaDrift now honors the env var (block:false when set). Verified: no unconditional proceed; GSD_SKIP_SCHEMA_CHECK=true -> block:false; gate 11/11 + mvp-tdd 9/9; full suite 569/0; lint 0; execute-phase.md 93109 < 93166. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): init.cts reads federated config keys from nested path (4th adversarial pass) Config federation moved tdd_mode/research/nyquist_validation from flat config.<key> to nested config.workflow.<key>, but src/init.cts still read them flat — so init.plan-phase/init.execute-phase emitted tdd_mode:false / research_enabled:undefined / nyquist:undefined regardless of config (a public command-contract regression; the migrated loops use render-hooks so enforcement was unaffected). Read via config.workflow (type-safe Record cast). Now init reflects the same resolved values + federated defaults (research/nyquist default true) as the render-hooks path. Verified: build clean; init.plan-phase emits tdd_mode:true/research:false/nyquist:false for set config, defaults true for empty; full suite 591/0; lint 0. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1169): add changeset for ADR-857 phase-6 completion (PR #1183) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): complete phase-6 migration fallout — restore TEXT_MODE, fix registry .claude leak, re-point stale workflow-contract tests The capability migration left real regressions and stale consumer tests that the per-module unit suite missed but the full cross-platform suite caught (27 failing tests): Real source regressions (fixed): - execute-phase.md lost its AskUserQuestion TEXT_MODE plain-text fallback when the inline schema_drift_gate step was removed — non-Claude runtimes would stall. Restored, and the execute:post gate-dispatch prose de-duplicated to cite the execute:wave:post contract (loop body shrinks below the frozen pre-phase-6 ceiling while keeping every onError/blocking nuance). - capabilities/tdd inline fragment hardcoded `@~/.claude/gsd-core/references/tdd.md`, baked verbatim into the committed capability-registry.cjs and leaked the install path on 11 non-Claude runtimes (registry .cjs is copied, not path-converted). Made the fragment path-free; regenerated the registry. The phase-6 conformance gate now guards this (no ~/.claude install path in any capability source or the generated registry). - plan-phase.md: removed a §5.7 stub re-added in error and routed Branch 2 to step 6 (schema-gate is a plan:pre capability, §5.7 is gone). Stale workflow-contract tests re-pointed to the capability dispatch they now must assert (behavior verified preserved in source first, assertions kept equal-or-stronger): bug-621 + bug-2851 (gap-analysis via gsd_run render-hooks plan:post + registry binding), feat-2527 (tdd_mode federated out of central), phase6-planning + plan-phase-ui-redirect (§5.6 bounded by ## 6.), plan-phase-drift-guard (intel when:intel.enabled skip branch). profile-pipeline-command-router.cjs un-ignored from eslint (hand-written, no TS source) + stale disable comments removed. Size baseline regenerated. Verified: full suite 15140 tests / 0 fail; lint 0 errors; conformance gate green legitimately. Refs #1139, #1167, #1168, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1169): add ADR-857 E2E content-test coverage for the 12 loop points + capability deliverables Grounds the capability engine in behavioral E2E tests (drive the real render-hooks/check CLI + the real registry, assert typed result content — no source-grep), structured around what ADR-857 says to deliver. 207 tests; each genuineness-checked (flip the expectation, confirm it fails). Per-loop-point dispatch (7 files): empty-point negative-space across the 6 no-hook points; verify:post 3-step resolution+ordering+onError; plan:pre contribution/configValues + ui.plan-gate + intel; plan:post gap-analysis; execute:wave:post drift+ui gates via the check route (schema-drift block/skip, codebase-drift threshold BVA, auto-remap); execute:post tdd.review-checkpoint RED/GREEN; ship:pre security gate resolution + frontmatter-get predicate pieces. ADR-deliverable coverage (4 files): predicate boundary held (edge/prohibition probes stay core, not off-by-default Feature Capabilities — phase-6 exception); core loop runs with zero capabilities (all 12 points empty, init bundles resolve); contribution merge (multiple ordered <contribution from=> blocks); federated-config key removal on uninstall. federated-config allowlisted for its 3-file split (unit + integration + lifecycle). Refs #1139, #1167, #1168, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): remove dead drifted converter dups + address adversarial review Lint cleanup (root-caused, not waved off): src/runtime-artifact-conversion.cts carried 11 agent-converter functions (+5 orphaned consts/helpers) that were never exported, never called, and had silently DRIFTED from the live hand-authored copies in bin/install.js (one even referenced an undefined `claudeToCopilotTools`). Deleted the dead duplicates; install.js's live copies are untouched (it never imported these). Lint now 0 errors / 0 warnings. Adversarial-review (Codex) findings fixed: - HIGH: execute-phase.md TDD_MODE used `jq ... || echo false`, silently disabling the MVP+TDD blocking gate on jq-less runtimes. Reverted to the `node -e` form (node is guaranteed; matches the file's other node-e usages) so a missing optional tool can no longer fail-open a blocking safety path. - MEDIUM: federated-config-key-removal orphan-key test was vacuous (it skipped the orphan assertion). Now asserts the removed capability's key is genuinely not surfaced/validated after uninstall. - LOW: phase-6 conformance leak regex broadened to catch absolute-home and Windows-backslash `.claude/(gsd-core|commands|agents|hooks)` paths, not only `~`/`$HOME` forward-slash forms. - LOW: bug-2851 plan:post dispatch assertion now requires `--raw` (matched its stated contract). - nit: plan-pre intel-step test duplicate assertion replaced with a distinct structured-output check. Size baseline regenerated (execute-phase.md 93089 < 93166 frozen). Refs #1167, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1169): make runtime-homes-descriptor-drive titles environment-independent The descriptor-equivalence test embedded the absolute golden config path (`os.homedir()`-derived) directly in each `test(...)` title, so titles differed between macOS (`/Users/x/.claude`) and Docker (`/home/gsdtest/.claude`). Every test PASSES on both platforms (15885/0 leaf tests each), but gsd-test-summary compares results by title and reported 29+29 false "only in Mac / only in Docker" discrepancies for tests that actually pass everywhere. Move the golden path out of the title and into the assertion message (still shown on failure); titles are now byte-identical across platforms so the cross-platform comparator matches them. No assertion logic or golden values changed. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): derive TDD_MODE via gsd_run --active-cap, not node -e (fix prompt-injection CI gate) The prior fix reverted execute-phase.md:181 from jq to `node -e` to close a Codex HIGH (jq||echo-false silently disabling the MVP+TDD blocking gate on jq-less runtimes) — but the CI prompt-injection scanner BLOCKS new `node -e` in workflow markdown (inline code-exec = injection vector), turning the security gate red. Both forms were wrong: node -e fails the scanner; jq fail-opens a blocking safety gate; `config-get workflow.tdd_mode` is forbidden by the conformance leak gate (tdd_mode is capability-owned). Correct fix (what Codex recommended): a gsd_run-native boolean. Add an `--active-cap <capId>` flag to `loop render-hooks <point>` that resolves hooks the normal way and prints exactly `true`/`false` for whether a capId is active — scanner-safe (canonical launcher, no inline code), node-reliable (no optional jq to fail-open), and leak-free (render-hooks resolution, not config-get). execute-phase.md:181 now `TDD_MODE=$(gsd_run loop render-hooks execute:post --active-cap tdd)`. +5 behavioral tests for the flag. Verified: prompt-injection-scan --diff origin/next → 0 findings; conformance gate 13/13 (execute-phase.md 92934 < 93166); execute-mvp-tdd + tdd-mode + loop-render-hooks 87/0; lint 0/0. Refs #1167, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
70f35ebc7a |
feat(#1169): wire ship:pre security gate via render-hooks (ADR-857 phase 6)
Part of the ADR-857 phase-6 migration (#1169, epic #857) — not a shipped bug fix. The capability system (registry, render-hooks, gates, capability-state) lives only on next / 1.5.0-rc; npm latest is 1.4.5 and contains none of it, so no released user can hit this. The security capability declares a blocking gates@ship:pre hook (when=workflow.security_enforcement, predicate SECURITY.md.threats_open==0), but ship.md never called `loop render-hooks ship:pre` and had no inline fallback, so the declared gate never fired. Wire it into ship.md preflight_checks via the render-hooks idiom — fail-closed: the ship blocks unless threats_open is exactly 0. Add a phase-6 conformance assertion that every declared hook point has a render-hooks call site; execute:wave:post (ui_safety_gate) stays in KNOWN_UNWIRED pending its unimplemented ui.safety-gate check. Part of #1169. Refs #857. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f116b76128 | test(#1139): add phase 6 capstone conformance gate (#1158) | ||
|
|
1b880fefd7 |
feat(#1137): migrate review verification hooks to capabilities (#1147)
* feat(#1137): migrate review verification hooks to capabilities * chore(#1147): add changeset |
||
|
|
44024aa535 |
fix(#1133): honor model_policy on the claude runtime (forward-port to next) (#1144)
Forward-port of the #1133 hotfix (commit 22f237a5, 1.4.5 hotfix line) onto next. The hotfix was authored against src/core.cts (v1.4.4); on next the resolver logic lives in src/model-resolver.cts (ADR-457 extraction, #888), so the patch is re-applied there rather than cherry-picked. resolveModelInternal step 2.5 now honors model_policy on the claude runtime: the policy-resolved full model ID is mapped back to a Claude Code agent alias via CLAUDE_POLICY_ID_TO_ALIAS (reverse of MODEL_ALIAS_MAP + claude-fable-5 -> fable). Bare aliases (opus/sonnet/haiku/fable) pass through; an ID with no Claude alias warns once to stderr (deduped by agentType::policyModel::tier) and falls back to the configured tier alias. Non-claude runtimes return full IDs verbatim (unchanged). resolveModelForTier is intentionally unchanged. The warn-dedupe cache lives in model-resolver.cts; core.cts composes the exported _resetRuntimeWarningCacheForTests to clear both that cache and the config-loader warning cache (config-loader cannot import model-resolver -- circular dependency). Ports the 6 #1133 tests (rewriting the old claude-no-op test that asserted the bug) plus one added test covering the MODEL_ALIAS_MAP reverse-map path. Forward-port of #1133 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4ab5c7b3f2 |
feat(#1135): migrate planning hooks to capabilities (#1141)
* feat(#1135): migrate planning hooks to capabilities * chore(#1135): add phase 6 planning capabilities changeset * fix(#1135): satisfy lint for agent hook rendering |
||
|
|
827011b865 |
fix(#1098): guard generate-claude-md against clobbering hand-crafted files; redirect default to .claude/CLAUDE.md (#1118)
/gsd-new-project wrote a repo-root CLAUDE.md full of broad project docs, overwriting/diluting a hand-crafted instruction file. --force was parsed but silently dropped, and nothing guarded an existing non-GSD file. - Guard: an existing instruction file with no `<!-- GSD:<section>-start` markers (hand-crafted) is left untouched; report action:"skipped". --force (now wired through CmdGenerateClaudeMdOptions) overwrites intentionally. The marker check uses /<!-- GSD:[a-z]+-start/ so a file merely documenting GSD syntax is safe. - Redirect: the Claude-family default output is now ./.claude/CLAUDE.md (a valid auto-loaded project-memory location) instead of repo-root ./CLAUDE.md. Aligned across the handler default, config-defaults.manifest.json, buildNewProjectConfig, the config template, new-project.md, and cmdGenerateClaudeProfile; advisory read-CLAUDE.md hints in plan-phase/quick/profile-user updated. Codex still writes AGENTS.md. Closes #1098 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
fc37ae0c4f |
fix(#1115): capability-probe codex hook-trust bypass flag in /gsd:review; fail loud on empty output (#1122)
On codex-cli < 0.137.0 the review.md `codex exec` invocation passed --dangerously-bypass-hook-trust (added in 0.137) unconditionally and discarded stderr, so codex exited "unexpected argument" before reading the prompt and the empty output file was treated as a completed review — a silent degraded review. - Capability-probe the flag (`codex exec --help | grep`) and apply it via $CODEX_BYPASS_FLAG only when supported (works fine without it on older CLIs). - Capture codex stderr to a .err file instead of /dev/null, and replace an empty output with a diagnostic so a broken reviewer is surfaced (mirrors Cursor/OpenCode). - Update enh-773 enforcement test to require the capability gate + fail-loud guard instead of the unconditional flag. Closes #1115 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1c073c1c81 |
fix(#1107): progress consults verification.status before reporting a phase complete (#1116)
/gsd-progress derived phase completeness from plan/summary counts only and never consulted the verification.status query (the #651 seam), so a phase whose VERIFICATION.md ended human_needed or gaps_found was reported complete and routing skipped to the next phase. Add Step 1.7 (consult verification.status for the current phase) and routing rows that send gaps_found to plan-phase --gaps (Route V.gaps) and human_needed to verify-work (Route V.human) before the generic complete row. passed/missing/unknown still route as complete so unverified phases are not falsely blocked. Closes #1107 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3e836fef0d |
feat(spec-phase): spec-completeness edge-probe (#550) (#584)
* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/
Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.
Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.
Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
source-checkout-gated build fallback) instead of LLM re-derivation; the engine
capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.
Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.
* test(#550): RED — status×verification re-cut + probe-core engine specs
Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):
status: resolved | dismissed | unresolved (lifecycle, shared)
verification: explicit | backstop | null (only when resolved)
- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
to be extracted — validateResolution(r, validators), validateRequirement,
analyzeCoverage(items, resolutions?, validators), byVerification rollup,
runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
{resolved,backstop}; coverage gains byVerification.{explicit,backstop};
proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
COUNT preserved on every fixture (closed set = resolved+dismissed; doc
line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
rewritten to the two-axis model.
Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).
* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)
Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.
probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
items[] (core never assumes propose is deterministic — edge resolves via
LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
{categories, verification, requiredFieldsByVerification} (ADR-550 #5)
edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.
* chore(#550): register probe-core.cjs artifact in ledgers + inventory
New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:
- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).
probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.
* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]
trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.
Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.
* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard
Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:
- validateResolution now enforces the 'verification is null unless resolved'
invariant for EVERY status (not just resolved): a dismissed/unresolved
resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
(was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
contract; a new test locks that an all-dismissed run is NOT affirmatively covered
(byVerification is the honest gate).
Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.
* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics
Re-review #5 (trek-e) clarity edits:
- Decision 5: annotate that only contract item (a) ships on #584 (the edge
adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
parenthetical — the blessed/implemented semantics are count-preserved = the
CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
carrying the per-tier resolved-status breakdown. The old parenthetical
contradicted the shipped count.
* test(#550): cover runProbeCli structural-guard numeric-count branch
Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.
* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)
The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.
* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)
A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.
* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)
templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.
* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)
The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.
* docs(#550): add how-to for resolving edge-coverage findings (B1)
Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.
* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)
trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.
* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)
trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.
* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)
trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.
* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe
Rebased onto next (
|
||
|
|
74d7bc8239 |
test(#1074): add additive per-file workflow size baseline guard (PR 1/3) (#1089)
* test(#1074): add additive per-file workflow size baseline guard (PR 1/3) Introduces a committed per-file size baseline scheme alongside (not replacing) the existing tier anti-creep tests. Green by construction — the baseline records current sizes, so both schemes pass side by side during migration. - scripts/lib/allowlist-ratchet.cjs: add assertFileBaseline (third pure helper, same injected-fail style) — per-file growth/shrink/add/remove diff vs baseline. - scripts/workflow-size.cjs: single source of truth for LF-normalized byte counting (#683) + workflow enumeration, shared by the guard and the generator so they can never measure differently. Lives in scripts/ root (NOT scripts/lib/) because it is dev/CI-only tooling — scripts/lib/ is bundled into the installed runtime, scripts/ root is not, so this keeps it out of the shipped payload. - scripts/update-size-baseline.cjs + npm run size:baseline: regenerate the snapshot (sorted keys, trailing newline, idempotent). - tests/workflow-size-baseline.json: generated snapshot (88 workflows). - tests/workflow-size-budget.test.cjs: import the shared counter (drops the duplicated local byteCount) and add the per-file baseline describe block. - Tests for the helper, the shared module, and the generator (incl. round-trip and fault-injection cases). Refs #1074. Part 1 of 3; PR 2 swaps enforcement, PR 3 covers the agent test. * test(#1074): regenerate workflow baseline after Update-branch merge with next The 'Update branch' merge (652a916b) pulled in next's update.md change (#1090) without regenerating the snapshot, leaving the per-file baseline stale by one file. Re-ran `npm run size:baseline` so the committed baseline matches the merged workflow files. Refs #1074. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |