3eb1cede26c03412e762c326f5ef2ee00056ebcf
307 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bd570618d4 |
feat(#2632): executor actuals and the closed estimate-calibration loop (#2672)
* feat(#2632): record executor actuals and close the estimate calibration loop * fix(#2632): calibrate against the raw projection so the loop converges * test(#2632): add closed-loop convergence guard and codify the feedback-loop rule * fix(#2632): pair calibration samples per plan; atomic write; amend adr * chore(#2632): backfill changeset pr to 2672 * fix(#2632): retry renameSync on transient windows errnos and clean up the temp |
||
|
|
89b673e40d |
feat(#2631): planner emits estimate and plan-checker surfaces the over-budget flag (#2670)
* test(#2631): failing-first planner estimate emission and over-budget surfacing * feat(#2631): emit plan estimate and surface the over-budget split recommendation * fix(#2631): extract sizing prose to references to fit planner and plan-phase caps * fix(#2631): move estimate check to plan-checker; fix template regex and caps * fix(#2631): restore ALWAYS split literal and keep gsd_run after the launcher preamble * fix(#2631): invoke estimate-check after the launcher preamble in plan-checker * fix(#2631): stop double-applying calibration; repair COMMANDS table and stale reference * chore(#2631): backfill changeset pr to 2670 * chore(#2631): backfill changeset pr to 2670 |
||
|
|
0ad3c5dd14 |
fix(#2279): refresh date stamps on map-codebase Update runs (#2550)
* fix(#2279): reword date stamping to overwrite existing dates on Update runs The map-codebase agent and workflow instructions only said to replace [YYYY-MM-DD] placeholders, but Update-path files already contain concrete dates from the prior run. Reword to SET the date stamps unconditionally, overwriting whatever date is already there. Closes #2279 * docs(#2279): backfill changeset PR number (2550) |
||
|
|
77bf21b3a6 |
fix(#1995): widen worktree branch regex to accept agent-<id> namespace (#2548)
* test(#1995): regression test for agent-<id> branch namespace Add failing-first tests proving that normalizeCleanupManifestEntry and planWorktreeRecordAgent reject Claude Code's current agent-<id> isolation branches (only worktree-agent-<id> is accepted). Boundary tests cover both namespaces plus rejection cases. * fix(#1995): widen worktree branch regex to accept agent-<id> namespace Claude Code's isolation="worktree" branch naming changed from worktree-agent-<id> to agent-<id>. Widen the regex in all 7 locations from ^worktree-agent-[A-Za-z0-9._/-]+$ to ^(worktree-)?agent-[A-Za-z0-9._/-]+$ so both namespaces are accepted. Introduce a shared WORKTREE_AGENT_BRANCH_RE constant in src/worktree-safety.cts to prevent future drift. Closes #1995 * fix(#1995): update workflow guards, test assertions, and baselines Widen the branch-check regex in execute-phase.md and execute-plan.md. Update all test assertions that checked for ^worktree-agent- to expect the widened ^(worktree-)?agent- pattern. Regenerate golden-install-parity fixtures, agent-size-baseline, and workflow-size-baseline. Closes #1995 * fix(#1995): update extractCwdGuardBash sanity check for widened regex The e2e test's sanity check verified the extracted bash block contained 'worktree-agent-'. After widening to '(worktree-)?agent-', update the check to match the new pattern. * fix(#1995): widen missed workflow-guard branch check + changeset + lint fixes - hooks/gsd-workflow-guard.js: widen startsWith('worktree-agent-') to /^(worktree-)?agent-/ regex — same defect class, was missed in prior commit - tests/worktree.test.cjs: fix indentation regression from prior edit - Add .changeset/1995-worktree-agent-branch-namespace.md (pr:0 placeholder) Found by orthogonal code review (Step 4). * fix(#1995): regenerate golden + size baselines for workflow-guard change * docs(#1995): backfill changeset PR number (2548) |
||
|
|
c5e0371775 |
feat(#1951): reversibility tagging — gate one-way-door decisions (#2471)
* test(#1951): add failing-first tests for reversibility tagging Red phase for issue #1951 (reversibility tagging: classify decisions by undo cost, gate one-way doors behind a checkpoint:decision). Tests assert, per the issue's acceptance criteria: - discuss-phase CONTEXT.md template records a **Reversibility:** field with a rationale on captured decisions, and states it is optional - gsd-planner @-references planner-reversibility.md and stays under the 49152-char agent cap (LARGE_CAP, tests/agent-size-budget.test.cjs) - a one-way rating inserts a checkpoint:decision before the dependent task; reversible inserts none; costly is flagged but never blocks - the taxonomy defaults to reversible when unsure (checkpoint-fatigue guard) and inserting a checkpoint implies autonomous: false - docs/reference/plan-md.md documents <reversibility> as optional with all three ratings - --no-reversibility-gates parses to REVERSIBILITY_GATES=false, is injected into the planner prompt, and is advertised in the command argument-hint and help full mode (argument-hint parity) - the override suppresses the gate but still persists the rating - cmdVerifyPlanStructure accepts every rating and the absent case (additive-validator guarantee, behavioral via runGsdTools) - parity: thinking-models-planning.md #4 adopts the canonical three-level taxonomy and the binary REVERSIBLE/IRREVERSIBLE vocabulary is gone - no content loss from the planner extraction made to fit under the cap Prose-contract assertions are Red until the implementation lands. The behavioral validator assertions pass immediately — regression guards proving the validator already accepts unknown optional tags. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1951): reversibility tagging — gate one-way-door decisions Classify planning decisions by what undoing them would cost, and give a one-way door a human beat before the agent walks through it (issue #1951, The Pragmatic Programmer Topic 15 'Reversibility'; Bezos's one-way/two-way door framing). Acceptance criteria met: - discuss-phase records an optional reversibility rating with a rationale on <decisions> entries in the phase CONTEXT.md template. Unrated decisions are treated as reversible, so existing phases are unaffected. - a one-way rating makes gsd-planner insert a checkpoint:decision before the task that implements the decision, reusing the existing checkpoint mechanism -- no new checkpoint machinery. - reversible ratings trigger no checkpoint; costly ratings are flagged in the plan but never block. - the rating persists on the task as the optional <reversibility rating=> element. cmdVerifyPlanStructure accepts every rating and the absent case; the structural validator does not reject unknown optional tags. - --no-reversibility-gates (REVERSIBILITY_GATES=false) suppresses checkpoint insertion for intentionally-unattended runs while still recording ratings -- the override changes what stops the run, not what the plan remembers. Single taxonomy, not two: references/thinking-models-planning.md #4 already shipped a binary REVERSIBLE/IRREVERSIBLE classification and is loaded by both gsd-planner and gsd-plan-checker. It is rewritten onto the canonical three-level vocabulary and now points at planner-reversibility.md as the taxonomy owner, with a parity test that fails if the surfaces diverge (DEFECT.GENERATIVE-FIX-DIVERGENCE). agents/gsd-planner.md sat 47 chars under the 49152 LARGE_CAP, so the checkpoint DO/DON'T guidance was relocated verbatim into planner-antipatterns.md -- already @-referenced from the same section for the same topic, so the planner still loads it and nothing was dropped. A test guards the relocation against content loss. Files: gsd-core/references/planner-reversibility.md (NEW, canonical taxonomy + emission rules + anti-patterns), gsd-planner.md, plan-phase workflow/command/help (flag wiring + parity), plan-md.md schema, discuss-phase context template, CONTEXT.md glossary, INVENTORY + manifest, size baselines, install goldens, plugin skills regen, changeset. Closes #1951 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1951): address orthogonal review findings Two isolated reviewers (correctness + security), neither of which authored the change. Every finding fixed: Security — the rationale is untrusted input (ADR-1577). It originates in conversation and flows CONTEXT.md -> planner -> PLAN.md -> executor, each hop an LLM reading the previous hop's output, with no validation on the path. planner-reversibility.md and the discuss-phase template now state it is data and never instructions, and name the </reversibility> early-termination hazard explicitly -- a rationale that closes its own element injects sibling structure the executor reads as real tasks. Four tests guard it. Correctness 1 — nothing machine-enforced the feature's own promise: a task rated one-way with no preceding checkpoint:decision validated as fully clean, so a planner error silently reopened the gap this feature exists to close. cmdVerifyPlanStructure now warns on an ungated one-way rating. A warning, not an error: <reversibility> stays additive and the plan stays valid. Four tests cover ungated (warns), gated (silent), still-valid, and reversible/costly never flagged. Correctness 2 — pass-always test. The --no-reversibility-gates parse test substring-matched the whole workflow file, and plan-phase.md prose mentions both tokens in one sentence, so it passed with the bash conditional deleted: it was testing the documentation, not the parser. Now scoped to the fenced bash blocks and matched as one physical line, with a negative control confirming prose alone cannot satisfy it. Correctness 3 — costly had no itemized emission rule, only one-way did, so two agents could diverge on whether to tag costly at all. Correctness 4 — template convention break: the example ratings were bare while every sibling field uses [...] to signal substitution, inviting an LLM to copy one-way/costly forward as boilerplate. Now bracketed. Correctness 5 — latent false-green: .includes('reversible') also matches inside irreversible/irreversibility, which appear in anti-pattern prose, so a surface that dropped the real taxonomy entry would still pass. Now word-boundary matched. ADR-857 phase-6 ceiling — the first gsd-test run caught plan-phase.md 1216 bytes over its frozen 94519 ceiling (it had 49 bytes of headroom on next). The ceiling may only rise for privileged host machinery, and reversibility gating is optional-feature logic, so the wiring was slimmed to its minimum and the explanatory prose moved to the reference files the planner already loads. plan-phase.md is now 94400 bytes -- 119 under the ceiling and 70 bytes SMALLER than on next, so the host loop shrank while gaining the feature, which is what phase 6 ratchets toward. The tracer contract (tests/tracer-bullet.test.cjs) is unchanged. Lint — fixed an unnecessary non-null assertion in verify.cts and a CRLF-fragile bare \n regex in the new test (DEFECT.WINDOWS-CRLF-TEST- PORTABILITY, the #1658/#1668/#2206/#2449/#2450 class). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture must carry the common task elements The gated-one-way fixture built a checkpoint:decision task from the abbreviated skeleton in gsd-planner.md, which shows only the checkpoint-specific elements (<decision>/<context>/<resume-signal>). cmdVerifyPlanStructure requires <name> and <action> on EVERY task regardless of type, so the fixture failed validation for reasons that had nothing to do with reversibility: errors: ["Task missing <name> element", "Task 'unnamed' missing <action>"] Caught by gsd-test on 14d14a39 (2 failures, both this fixture). The canonical shape is in tests/verify.test.cjs:266 — a checkpoint task carries <name>/<files>/<action>/<verify> like any other. Fixture corrected to match. Verified behaviorally against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Not a product defect: the validator's every-task contract is intentional and pre-existing, and docs/reference/plan-md.md scopes its required-element list to type=auto/tracer only because those are the elements a planner must author, not because checkpoints are exempt from <name>. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1951): backfill changeset pr number to 2471 * fix(#1951): CodeQL incomplete-sanitization + prompt-injection scan collision Both CI failures were real defects in code this PR added, not false positives. CodeQL js/incomplete-sanitization (high), reversibility-tagging.test.cjs:46 — the namesRating helper built its regex with `rating.replace(/[-]/g, '\\-')`, which escapes the hyphen but not backslash, so the escape was incomplete. It was also unnecessary: `-` carries no special meaning outside a character class. Replaced with a complete metacharacter escape (backslash included). Word-boundary behavior verified unchanged across all three ratings — notably that "irreversible" prose still does not satisfy a "reversible" match, which is the false-green this helper exists to prevent. Prompt injection scan — the checkpoint fixture used the human-verification child element inside <verify>. That tag name is a fake-instruction-boundary pattern in scripts/prompt-injection-scan.sh, and the scan runs over changed files, so copying the shape from tests/verify.test.cjs (unflagged only because it is not in this diff) tripped the gate. Switched to the documented plain-prose <verify> form. The first attempt at that fix failed the same gate a second time: the comment explaining the collision quoted the offending tag literally. The comment now names it in prose instead — the scanner does not care whether a match is code or commentary, which is the whole point of the DEFECT.PROMPT-INJECTION-SCAN-COLLISION note in CLAUDE.md. Verified locally before push: scan reports 0 findings across 57 changed files, eslint clean, and both fixtures still validate as designed (gated one-way silent, ungated one-way warns, neither errors). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): record measured cost and halve gsd-tools spawns The Windows shard 1/3 job timeout was traced to the sharding layer, not to this PR's assertions — see #2472. Two contributing factors were this file's own, and are fixed here. 1. tests/test-timings.json had no entry for reversibility-tagging.test.cjs, so scripts/run-tests.cjs weighted it at the table's median fallback (~315ms) for LPT chunk packing. It actually measures 5595ms — an 18x under-weight. Recorded the measured value from the green gsd-test run (max across the node22/node24 lanes, per gen-test-timings.cjs's convention). Only this one entry: a full regen churns 634 entries of run-to-run drift, and the table is explicitly advisory and un-gated, so a 637-line diff does not belong in a feature PR. 2. Each verifyPlan() spawns gsd-tools, which dominates this file's cost. Spawns cut from 9 to 6 with no coverage lost: - the ungated-one-way warning and its stays-valid assertion now share one plan instead of building the same plan twice; - the reversible/costly never-flagged-as-ungated test was strictly subsumed by the additive suite, which already runs those two ratings ungated and asserts no /reversibilit/ warning at all — and the gate warning's text contains both "reversibility" and "one-way", so the broader assertion catches it. It only re-spawned gsd-tools twice to prove the same thing. Both are symptom fixes. The shard imbalance itself (19/11/10 minutes against a 20-minute cap, from a cost-blind round-robin partition that also reshuffles downstream files whenever one is inserted) is tracked in #2472. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture adopts the #2444 type-branched contract Surfaced by rebasing onto next, which gained #2444 (branch plan-structure validation on task type=checkpoint:*) while this PR was in review. cmdVerifyPlanStructure no longer applies one required-element set to every task. A checkpoint:decision now requires <name> + <resume-signal> + <decision> + <options>, and is exempt from the <action>/<verify>/<done>/ <files> set that auto and tracer tasks carry. The gated-one-way fixture predated that split and failed on the new requirement: errors: ["Task 'Task 0: Confirm the on-disk format' missing <options>"] Fixture rewritten to mirror the checkpoint:decision contract exactly — real <options> with two <option> children — rather than padding it with fields checkpoints no longer need. That also drops the plain-prose <verify> the earlier revision carried purely to dodge the prompt-injection scan; a checkpoint task has no <verify> requirement at all, so the workaround is moot. Verified against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
517bae8d6d |
fix(#2372): widen decision-coverage-plan to all planner-canonical tags, drop misleading "(or body)" (#2443)
* fix(#2372): widen decision-coverage scan to planner-canonical tags, fix message Bug: check.decision-coverage-plan's remediation message told the user to cite decisions "(or body)" but extractPlanDesignatedSections only scanned <objective>/<tasks>/<task>/<action>. A decision cited in <read_first>, <behavior>, <verify>, <acceptance_criteria>, or <done> was invisible to the gate — false BLOCKING coverage gap, plus the message's own fix-hint sent the user to "the body" where re-citing still failed. Two-part fix (must change together — that drift was the bug): 1. Widen XML_DECISION_TAGS_RE in src/check-command-router.cts to also match <read_first>, <behavior>, <verify>, <acceptance_criteria>, <done>. These are all planner-canonical tags the planner is told to use (plan-phase.md:830-862, plan-phase.md:772). The body negative- lookahead mirrors the opening-tag set so each tag's body is captured independently. 2. Correct buildPlanMessage to name ONLY the surfaces the extractor actually scans (front-matter must_haves/truths/objective, designated markdown headings, and the nine planner-canonical tag bodies). The misleading "(or body)" clause is gone. Also updates the planner's documented contract (agents/gsd-planner.md:69) and user-facing docs (docs/CONFIGURATION.md, docs/USER-GUIDE.md) to reflect the wider scan. Regression tests in tests/decisions.test.cjs cover each newly-scanned tag body, a control (no citation still uncovered), and a message/extractor parity assertion that names every scanned surface — so the two cannot drift apart again. Out of scope (per triage): cmdDecisionCoverageVerify/buildVerifyMessage is a separate command (decision-coverage-verify) checking shipped artifacts, not plan citations — untouched. * chore(#2372): regenerate agent-size-baseline + golden-install-parity fixtures gsd-planner.md grew 49172 → 49294 (+122 chars) from the widened decision- coverage contract (5 new scanned tag names + heading clarification). Growth is justified: the contract surface is itself the fix — the prior text under-described what the gate scans, which was the bug. Updates: - tests/agent-size-baseline.json (gsd-planner.md: 49172 → 49294) - 17 tests/fixtures/golden-install-parity/*.json (one hash per runtime) - tests/fixtures/install-tree/*.json (regenerated by gen:golden) * fix(#2372): per-tag matching — outer-tag citations survive inner-tag nesting Code review (subagent) flagged a Medium edge-case regression from the single-alternation regex: when a newly-scanned tag nests inside another scanned tag, the alternation's negative lookahead halts the outer tag's body at the inner tag — losing any D-NN citation in the outer tag's prefix prose. Concretely: <action>per D-05 <verify>npm test</verify></action> → 3-tag alternation (old): captured 'per D-05 <verify>npm test</verify>' as <action> body → D-05 caught → 9-tag alternation (bug): captured 'npm test' only (from <verify>); D-05 in <action> prefix LOST Switches extractXmlTagBodies to per-tag matching: each tag gets its own regex whose negative-lookahead tempers only against the SAME tag's reopening. So <verify> inside <action> is absorbed into <action>'s body (D-05 caught) AND <verify> is matched separately on its own pass. Per-tag preserves both: - the reporter's case (sibling tags inside <read_first>) - nested-tag citations in outer-tag prefix prose - ReDoS safety (each per-tag regex keeps the #2128 body tempering) Also adds the reviewer's other requested edge-case tests: - non-scanned tag (<name>) bearing D-NN must NOT count - self-closing form <read_first /> safely ignored - attribute form <verify type="...">D-NN</verify> (canonical planner shape) - CRLF newlines in tag body do not break capture * chore(changeset): backfill pr:2443 in .changeset/noble-elks-chatter.md |
||
|
|
d16a66479a |
feat(#1950): broken-windows ledger — cross-phase defect register gating ship (#2441)
* feat(#1950): broken-windows ledger — cross-phase defect register gating ship Adds a new capability (#1950) that operationalizes GSD's no-defer discipline as a tracked, enforced artifact: accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths across phases, and /gsd-ship blocks while any entry is open. Implementation: - src/broken-windows.cts → gsd-core/bin/lib/broken-windows.cjs: typed IR + I/O entry points (parseLedger/renderLedger/appendWindow/markWaived/markFixed + cmdWindowsStatus/Append/Waive/MarkFixed). Frozen REASON enum for typed error assertions. Windows-safe atomic rename with retry on transient EPERM/EBUSY/EACCES. - gsd-tools.cjs: new subcommand (status | append | waive | fixed), wired via routeWindows + HOST_COMMAND_ROUTERS.windows. - capabilities/broken-windows/capability.json: one ship:pre gate with artifact-frontmatter-equals predicate on WINDOWS.md open_count == 0. activationKey windows.enabled (default true) + sibling windows.enforce (default true, separate so tracking can precede enforcement). - gsd-core/workflows/ship.md: capId==broken-windows branch in preflight, sibling to security — reads gsd_run windows status --raw, fails closed on open_count > 0 or unreadable ledger. - agents/gsd-executor.md: extends the existing ## Known Stubs instruction to also append to WINDOWS.md via gsd_run windows append (best-effort, never blocks execution). - agents/gsd-verifier.md: new Step 8b — record unmet truths + human-verify items in WINDOWS.md. - gsd-core/workflows/progress.md: surfaces open + waived counts. - docs/COMMANDS.md + CONTEXT.md glossary entry + docs/INVENTORY.md: document the gate, waiver mechanism, and new module. - tests/broken-windows.test.cjs: pure + CLI behavioral coverage + fast-check roundtrip property; fail-closed on malformed ledger; security boundary on path traversal in --file. Backward-compatible: a project with no .planning/WINDOWS.md reports open_count: 0 and ships cleanly. Disable enforcement per-project with gsd config-set windows.enforce false (tracking continues, gate stays open). * chore(#1950): ratchet size baselines, defer verifier integration - Workflow size baseline: ship.md 25575→27928, progress.md 31789→32632 (broken-windows preflight branch + open-windows surface). - Agent size baseline: gsd-executor.md 46644→47951 (Known Stubs → also appends to WINDOWS.md). gsd-verifier.md unchanged. - LARGE_CAP (49152) preempted the planned verifier integration (gsd-verifier.md was at 49140 pre-PR — 12 bytes of headroom, not the documented 'real headroom'). Verifier integration deferred to a follow-up PR that extracts the VERIFICATION.md template (lines 739-859) to gsd-core/references/ — a pre-existing cap-tightness defect this PR exposed but does not expand scope to fix. Verifier integration is not in the issue's acceptance criteria (executor writes is; unmet-truths recording was an enhancement, not a gate). * fix(#1950): gate default-off, rename to workflow.windows_enforce, regen goldens Test-failure-driven fixes after first gsd-test run on db8733c8f failed 44 cases (pre-existing structural tests encoded 'ship:pre has 1 gate' / 'all caps off → empty hooks'): - capability manifest: rename windows.enabled+windows.enforce (default true) → single federated key workflow.windows_enforce (default FALSE, opt-in). Matches security's workflow.security_enforce convention and makes the adr857 all-caps-off test pass without modification (the test's buildAllFalseConfig handles workflow.* out of the box). Default-OFF keeps the gate out of the registry's default ship:pre resolution so existing loop-hooks-ship-pre-e2e structural assertions (exactly 1 gate, capId 'security') stay valid; users opt in via gsd config-set workflow.windows_enforce true. - drop activationKey (security doesn't have one either; workflow.* key doubles as the activation toggle). - regenerate docs/reference/capability-matrix.md to include broken-windows (capability-matrix-sync test). - regenerate tests/fixtures/golden-install-parity/*.json (18 runtimes) — installer now emits the new capability + lib file. - update CONTEXT.md, docs/COMMANDS.md, docs/FEATURES.md, ship.md, agents/gsd-executor.md to use the new key name and /gsd:colon slash syntax (slash-command-namespace test). - restore accidentally-regressed /gsd:capture in progress.md. Tracking-only by default; enforcement is opt-in. Acceptance criterion '/gsd-ship fails while any ledger entry is open' is met when workflow.windows_enforce=true (test fixture enables it). * test(#1950): update ship:pre structural invariants for 2-gate registry - loop-hooks-ship-pre-e2e: the registry now declares 2 gates at ship:pre (security + broken-windows), regardless of activation. Activation tests above still pin security-only or empty behavior via fixtures; these structural tests pin the REGISTRY shape, which has 2 gates as of #1950. - workflow-size-baseline: ship.md 27928→27945 (workflow.windows_enforce rename added 17 bytes). * fix(#1950): review H1+H2+M1+M2+M3 — fence-injection, EACCES fail-closed, cleanup, strict line, stryker Adversarial isolated review (Step 6.3) found 2 HIGH findings that block the PR and 3 mediums. All addressed: H1 (HIGH): description containing the markdown 3-backtick fence would terminate the ledger's JSON code block early inside JSON.stringify output (JSON doesn't escape backticks), corrupting the file and bricking the next parse. Fix: use a 4-backtick fence (json ... ) which JSON.stringify cannot produce on its own, AND validate that no entry text field contains a 4-backtick run (reject at append time with new WINDOWS_INVALID_TEXT reason code). Locked by a regression test. H2 (HIGH): readLedgerOrNull swallowed ALL fs errors as 'no ledger', silently returning open_count:0 on EACCES/EPERM/EIO. The ship gate would then pass on an unreadable ledger — the precise vector the workflow doc claims is impossible. Fix: only ENOENT returns null; every other fs error propagates as WINDOWS_LEDGER_MALFORMED so the gate blocks and the operator sees a real diagnostic. Locked by a regression test that chmod 000s a ledger with open_count=1 and asserts the result is never a false-green 0. M1: writeLedgerAtomic left an orphaned .tmp file on rename failure. Wrapped renameWithRetry in try/catch with best-effort unlink. M2: validateLine silently coerced 'abc' → NaN → null, hiding type drift. Removed the line === 0 special case (was undocumented) and made the error message match the strict check. Now any non-positive- integer line value throws, including strings. M3: tests/broken-windows.test.cjs (with its fast-check property test) was not in stryker.config.mjs DEFAULT_TEST_CMD — Stryker would mutate src/broken-windows.cts but no test would catch the mutations, producing false surviving-mutant scores. Added to the list. L1 (dead throw e after error()), L7 (line boundary tests, H1/H2 regression tests, 4-backtick CLI test) also addressed. * docs(#1950): inline concurrency + busy-wait notes (review L2+L3) * fix(#1950): regen goldens against latest gsd-tools; correct --line 0 boundary test gsd-test v4 caught two issues: - goldens I regenerated earlier (commit 526682084) predated the L1 routeWindows catch-block cleanup (commit dd844d565). Regenerated via 'npm run gen:golden' against current HEAD so the install parity hash for gsd-tools.cjs matches. - 'append --line boundary' test expected --line 0 to succeed with null entry.line, but the M2 fix correctly rejects 0 (lines are 1-indexed; 0 is not a valid source line). Updated the boundary test to assert --line 0 fails alongside -1 and 'abc'. * chore(#1950): regen goldens after rebase onto next * chore(#1950): quick.md baseline 50699→50993 (correct resolution from next rebase) * chore(changeset): backfill pr:2441 in .changeset/broken-windows-ledger.md * fix(#1950): renderTable escapes backslash before pipe (CodeQL incomplete-sanitization) CodeQL flagged the markdown-table cell escaper: String(s ?? '').replace(/\|/g, '\\|') — it escapes pipe but not backslash first. A description containing '\|' would render as '\\|' which markdown parses as 'literal backslash' + 'cell separator', splitting the column. Fix: escape backslash FIRST (each \ → \\), then pipe (each | → \|). Now a description with '\|' renders as '\\\\|' (literal '\\' + escaped pipe), which markdown renders as a single '\|' inside the cell. The JSON code block (the parse source-of-truth) was already correctly escaped via JSON.stringify; only the display-only table was affected. Locked by a regression test that: 1. Verifies the JSON block reparses with the description intact. 2. Walks the rendered table row counting unescaped pipes — must be exactly 11 (the row separators for 10 cells), proving no in-cell pipe added a split. |
||
|
|
d0bacc2517 |
fix(#2351): replace hardcoded timeout with portable run-with-timeout (#2426)
* fix(#2351): replace hardcoded gnu timeout with portable run-with-timeout Stock macOS ships neither `timeout` nor `gtimeout` (GNU coreutils). The 10 hardcoded `timeout <n> <cmd>` calls across the workflow/agent/reference gates exited 127 ("command not found") on such hosts, and the gates — which only distinguish 0/124/other — misreported a passing build or test as a FAILURE. Fix: a single Node-based `gsd_run run-with-timeout <secs> [--] <cmd…>` verb in gsd-tools.cjs. Coreutils-independent (stock macOS AND Windows), keeps GNU `timeout`'s exit-code contract (124 timeout, passthrough, 127/126 ENOENT/EACCES, 128+signum on signal), inherits stdio so pipes/redirects work, and reaps the whole process group so a watch-mode runner cannot outlive its budget. Runs before gsd-tools' flag parsing so the wrapped argv stays opaque. Hardened per adversarial review: - On timeout, SIGKILL the group SYNCHRONOUSLY before resolving — a descendant that traps SIGTERM was otherwise orphaned holding stdout, hanging captured gates (the exact watch-mode hang the feature prevents). - Forward SIGINT/SIGTERM to the child tree instead of dying and orphaning it. - Reject blank/whitespace <seconds> (was a silent unbounded run); clamp the timer to the 32-bit setTimeout ceiling (was a spurious immediate timeout). - Lint detector: catch GNU long options / `-k5` / `$((...))`; anchor to command position so prose "timeout 30 seconds" no longer false-positives. Resolution lives once in the CLI; all 10 sites call the shared verb. A parity guard (scripts/lint-portable-timeout.cjs, wired into lint:ci) fails the build if a bare `timeout`/`gtimeout` execution reappears (the portable `command -v timeout` probe form is intentionally allowed). Also fixes the identical bug in the zh-CN checkpoints translation, updates the tests that asserted the old strings, trims a redundant phrase in gsd-verifier.md to keep it under its size hard cap, and refreshes the size baselines + golden install-parity fixtures. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2351): add changeset (#2426) * chore: regenerate golden/size baseline after rebase onto next --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1720aacf0c |
feat(#1949): <precondition> task element — Design by Contract (#2422)
* test(#1949): add failing-first tests for <precondition> element Red phase for issue #1949 (Design by Contract: <precondition> element asserted before task execution). Tests assert: - docs/reference/plan-md.md documents the new <precondition> element - agents/gsd-planner.md @-references planner-preconditions.md and stays under the 49152-char cap (progressive-disclosure requirement) - gsd-core/references/planner-preconditions.md exists and documents the three emission cases mandated by the issue (user_setup / prior-phase artifact / env-var) and the contract triad mapping - agents/gsd-executor.md asserts <precondition> before task execution and routes unmet preconditions through existing checkpoint machinery - cmdVerifyPlanStructure (behavioral via runGsdTools) accepts plans both with and without <precondition> — the additive-validation guarantee - Parity assertion: plan-md.md and planner-preconditions.md agree on the canonical tag spelling (DEFECT.GENERATIVE-FIX-DIVERGENCE guard) Most prose-contract assertions are Red until the implementation lands. The behavioral validator assertions pass immediately (regression guards proving the validator already accepts unknown optional tags). * feat(#1949): <precondition> task element — Design by Contract Add an optional <precondition> element to <task> in PLAN.md (issue #1949, The Pragmatic Programmer Topic 23). The front-of-task side of the plan contract — preconditions (before) ↔ postconditions (<verify>/<done>/ <acceptance_criteria>, after) ↔ invariants (must_haves.truths, across the whole plan). Together with the tracer-bullet proposal (#1945), this closes both ends of the 'outrunning your headlights' failure mode for an autonomous AI executor. Acceptance criteria met: - <precondition> is an optional element on <task>; plans that omit it validate unchanged (cmdVerifyPlanStructure checks for presence of required tags, does not reject unknown optional tags). - gsd-executor evaluates the precondition before any other task work. Unmet halts execution with a checkpoint:human-verify and no partial commit; met or absent produces no visible change to execution flow. Unmet is never auto-approved under AUTO_CFG=true — a missing prerequisite is a fact the executor cannot establish on its own. - gsd-planner emits <precondition> in exactly the three cases the issue mandates: user_setup consumption, prior-phase artifact dependency, and env-var/runtime-config dependency. - Tests cover met, unmet, and absent preconditions plus the additive- validator guarantee. Files: - gsd-core/references/planner-preconditions.md (NEW): full emission rules, the three cases with worked examples, format guidance, anti-patterns, the contract triad mapping, and the executor assertion contract. Progressive disclosure. - agents/gsd-planner.md: slim <precondition> note in Task Anatomy with @-reference to the new file. To stay under the 49152-char agent-file cap (27-char headroom before this change), the inline <comment_text_discipline> and <region_scoped_negative_gate> summaries are compressed to one-line pointers — their full rules already live in planner-antipatterns.md, so no content is lost. - agents/gsd-executor.md: new step 0 'Precondition check' in the execute_tasks loop, before the type dispatch, routing unmet through checkpoint_return_format. - docs/reference/plan-md.md: new Preconditions section in the schema reference, with the canonical example and the three emission cases. - CONTEXT.md: Precondition glossary entry as a sibling of Tracer Bullet. - docs/INVENTORY.md + INVENTORY-MANIFEST.json: row for the new references/planner-preconditions.md (regen via gen-inventory-manifest). - tests/precondition-element.test.cjs: failing-first tests covering schema docs, planner emission contract, executor assertion contract, reference-file presence + the three cases, behavioral additive- validator guarantee, and a parity assertion (DEFECT.GENERATIVE-FIX- DIVERGENCE guard). - .changeset/quick-hawks-bark.md: Added fragment. Companion to #1945 (tracer bullets). * chore(#1949): regen agent-size baseline + install-tree goldens Documented baseline regenerations required by the feat(#1949) prose changes (RULESET.AGENT_SIZE_BUDGET + golden-install-parity): - npm run size:baseline — locks in the new gsd-executor.md size (+1050 bytes: the precondition-check step 0 block). gsd-planner.md is net smaller (-142 bytes: compressed two inline summary blocks whose full rules already lived in planner-antipatterns.md to make room for the slim <precondition> pointer). No hard-cap breach. - npm run gen:golden — pick up the new references/planner-preconditions.md + the two changed agent files across all 18 runtime install trees. Both regens are CI-mandated after intentional agent/reference changes; see CLAUDE.md 'RULESET.AGENT_SIZE_BUDGET' and the comments in tests/golden-install-parity.test.cjs. * fix(#1949): bound <precondition> checks to read-only (security review) Apply the security-review finding (LOW, isolated /security-review subagent): the executor's 'run the cheapest check' phrasing for a plan-author-controlled prose line was broader than ideal — a hostile plan author could craft a <precondition> whose 'cheapest check' is side-effecting (curl to an attacker host under the guise of verification, rm -rf before checking, secret emission). The risk is inherited from GSD's existing plan-trust model (<verify>, <action>, <done> already direct the executor to run arbitrary shell), so <precondition> does not materially expand it. But the new prose actively directs execution ('run the check') rather than passively consuming the element, so the bound is worth making explicit. Tightened across all four surfaces that describe the check shape: - agents/gsd-executor.md step 0: 'Verify with read-only checks only — file existence, env var presence (no value output), idempotent GET /health-style pings. Do NOT run commands with side effects (writes, network POSTs, secret emission) as the check; if a side-effecting check seems required, halt and surface via checkpoint instead.' - gsd-core/references/planner-preconditions.md Format section: same bound, plus the halt-and-surface escape hatch. - docs/reference/plan-md.md Preconditions section: mirrored. - CONTEXT.md Precondition glossary entry: mirrored. Regenerated agent-size baseline (executor grew 46186 -> 46440; still under the 49152 cap) and install-tree goldens. * chore(#1949): backfill changeset pr number 2422 Per CONTRIBUTING.md changeset workflow + feature-builder directive Step 8.7: backfill the placeholder pr:0 with the real PR number immediately after gh pr create returns. Avoids the fail_invalid_fragment gate. * fix(#1949): cite [#1949] on allow-test-rule exemption (ADR-456) CI's lint:ci runs lint-allow-test-rule-refs which per ADR-456 requires every // allow-test-rule: exemption on a NEW test file to carry an issue reference (#NNN or URL). My earlier push omitted it. Local 'npm run lint' (eslint) does NOT run this check — only 'npm run lint:ci' does. CLAUDE.md explicitly warns: 'lint:ci ≠ lint — CI runs lint:ci; a local pass is not the gate.' I should have run lint:ci before pushing; correcting now. Pattern matches the companion feature's test file: tests/tracer-bullet.test.cjs:1 // allow-test-rule: source-text-is-the-product [#1945] |
||
|
|
dd5a2211c9 |
enhance(#1964): semantic knowledge-base recall via MemPalace (keyword fallback) (#2416)
* test(#1964): add failing-first semantic-recall contract tests Epic #1957 Phase 3C (final). Source-text-is-the-product contract tests: semantic recall via MemPalace (top-k meaning-similar prior resolutions, catches same-root-cause/different-wording cases), indexing resolved sessions at archive, graceful degradation to keyword matching when MemPalace is absent, knowledge-base.md stays the durable plain-text source of truth, agent Phase 0 / Matching Logic is semantic-first (the stale 'keyword overlap, not semantic similarity' claim must go), and no new embedding/vector infra (reuse MemPalace). Failing-first: reference, the Matching Logic reframe, the Phase 0 consolidation, and the archive indexing step do not yet exist. * feat(#1964): semantic knowledge-base recall via MemPalace (keyword fallback) Epic #1957 Phase 3C (FINAL). Replaces keyword-overlap matching with semantic recall: at Phase 0 the debugger queries MemPalace with the current symptoms and surfaces the top-k meaning-similar prior resolutions, catching the same-root-cause/different-wording cases keyword overlap missed (the self-noted 'keyword overlap, not semantic similarity' limitation). Resolved sessions are indexed into MemPalace at archive (symptoms + root_cause(s) + fix + recurrence guard). knowledge-base.md remains the durable plain-text source of truth; when MemPalace is absent the debugger falls back to keyword-overlap matching (logged, never a silent skip). No new embedding/vector infrastructure — MemPalace is reused. Size-neutral agent edits: the Matching Logic section reframed (keyword-only -> semantic-first + keyword-fallback + @-include); Phase 0's three keyword bullets consolidated into one semantic-first bullet; one MemPalace-indexing step added at archive. Agent at 57222 B (122 B headroom — final phase). Full rules in gsd-core/references/debugger-semantic-recall.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1964): address orthogonal review (invocation mechanism, index Resolution-not-symptoms + redaction, fallback detail) - HIGH: the 'query MemPalace' instruction was WHAT-level only; the agent has no MCP tools. Added an Invocation section naming the Bash CLI (mempalace search --wing <wing>) + MCP-when-registered + wing resolution (config.mempalace.wing -> project_code -> project dir), matching every other MemPalace integration. Without this the feature silently degraded to keyword matching even when MemPalace was present. - MEDIUM (security x2): index the agent-authored Resolution summary (root_cause + fix + recurrence_guard), NOT raw user-supplied Symptoms — excludes attacker-controlled prose from the cross-session index AND reduces secret/PII leakage. Redact secret-shaped values before indexing. Stated the write order (KB append + commit MUST succeed before indexing). - LOW: restored 'identifiers' + 'case-insensitive' to the keyword fallback; added a test asserting the fallback mechanics survived the Phase 0 consolidation (Error patterns field, 2+ token overlap, identifiers, case-insensitive). * chore(#1964): ratchet agent-size baseline downward (leaner archive bullet shrank gsd-debugger.md 57222->57197) * chore(#1964): backfill changeset pr number (PR #2416) |
||
|
|
c67f301867 |
feat(#1963): emit blameless-postmortem Prevention block at resolution (#2410)
* test(#1963): add failing-first prevention/postmortem contract tests Epic #1957 Phase 3B. Source-text-is-the-product contract tests: blameless 5-Whys that BRANCHES per Phase 2A RCA (not a single-cause chain; treats agent error as 'why was that possible?'), the 'why wasn't this caught?' question, the recurrence-guard taxonomy (regression test / assertion / lint rule / KB pattern), the KB-entry why_not_caught + recurrence_guard fields with backward compat, the session-manager prevention summary line, and the Zawinski scope-boundary (a block, not a subsystem). Failing-first: reference, archive_session edit, KB schema extension, and session-manager summary do not yet exist. * feat(#1963): emit blameless-postmortem Prevention block at resolution Epic #1957 Phase 3B. At archive_session the debugger now produces a Prevention block with three blame-free components: a branching 5-Whys causal chain (branches per Phase 2A RCA, not a single chain; 'agent error' prompts 'why was that possible?', never blame), a 'why wasn't this caught?' answer naming the missed gate (test/typecheck/lint/review/verify), and a concrete recurrence guard (regression test / assertion / lint rule / KB pattern). The knowledge-base entry gains two structured fields (why_not_caught + recurrence_guard) so future Phase-0 recall surfaces the prior prevention, not just the prior fix. Additive: old entries without the fields still load. The session-manager compact summary surfaces a one-line prevention summary. Full rules extracted to gsd-core/references/debugger-prevention.md (slim archive_session step + 2 KB fields kept in the agent). INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1963): address orthogonal review (CRITICAL append-template drift + Phase-0 consumption + parity test) - CRITICAL: the archive_session KB append template omitted Why not caught + Recurrence guard (only the Entry Format had them) — the feature's core deliverable silently did not happen. Added both fields to the append template the agent actually follows (nearest-instruction wins). - HIGH: Phase 0 (KB read) only surfaced root_cause + fix; the new fields were dead data. Extended the Phase 0 Evidence line to consume why_not_caught + recurrence_guard when present (absent on old entries — backward compat holds). - MEDIUM: added a cross-section parity test (every Entry-Format field must also appear in the append template — the guard that would have caught the Critical) + a Phase-0-consumption assertion. - MEDIUM: the 'branches per Phase 2A' claim is now wired — reuses reasoning_checkpoint.candidate_causes across the four categories. - MEDIUM: recurrence-guard taxonomy gains type refinement + config-default change; LOW: added 'build' gate to both surfaces for parity. - NIT: compact-summary fallback shape ('no gate existed'); verify the guard artifact exists before recording it. * test(#1963): anchor Phase-0 consumption test on the specific heading The regex /Phase 0[\s\S]{0,1200}/ matched the first 'Phase 0' in the file (in knowledge_base_protocol prose), not the Phase 0 block in investigation_loop. Anchor on '**Phase 0: Check knowledge base**' and widen to 1500 chars. * chore(#1963): backfill changeset pr number (PR #2410) |
||
|
|
36a311c5bb |
enhance(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries) (#2409)
* test(#1962): add failing-first repro-hardening contract tests Epic #1957 Phase 3A. Source-text-is-the-product contract tests: PBT shrinking (fast-check/Hypothesis, minimized seed, manual-minimization degradation), the four oracle types (specified/derived/metamorphic/implicit with implicit flagged weakest), boundary neighbors (off-by-one/min-max/empty-singleton tied to the equivalence class), oracle_type in DEBUG Resolution, and the Phase 1A tie-in (minimized seed + real oracle => the mutation guardrail bites). Failing-first: reference, agent cross-refs, and template field do not yet exist. * feat(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries) Epic #1957 Phase 3A. Extends Minimal Reproduction (shrinking) and Test-First Debugging (oracle classification + boundary neighbors): - Shrinking: wrap an input-space failing input in a property (fast-check JS/TS, Hypothesis Python) and store the MINIMIZED counterexample as the regression seed; degrade to manual minimization when no PBT framework is present. - Oracle classification: state specified / derived (contract/model) / metamorphic / implicit (crash, weakest) before writing the assertion; record under Resolution.oracle_type; never default to implicit silently. - Boundary neighbors: off-by-one, min/max, empty/singleton around the fixed defect's equivalence class. Together they turn the regression test into a root-cause check — what the Phase 1A mutation guardrail needs to bite. Full rules extracted to gsd-core/references/ debugger-repro-hardening.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md + DEBUG template updated. * fix(#1962): address orthogonal review (bounding, provenance, oracle scope, sufficient-triple) - HIGH: added a 'Bound the property/shrink run' section (60s timeout, degrade- to-manual on timeout, do-not-raise-default-run-limits, argv-not-shell) — the gauntlet violation the sibling references already honored. - Medium: test-provenance caveat (the failing input often comes from the bug report — author the generator from a sanitized description, cross-ref debugger-fix-acceptance.md). - Medium: oracle scope note — the 4 types cover deterministic bugs; non- deterministic failures re-route to stability-stress per bug-taxonomy. - Medium: Phase 1A tie-in corrected — seed+oracle is necessary not sufficient; boundary neighbors close the adjacent-input escape; the sufficient triple is seed+oracle+neighbors. - Low: preserve the original noisy repro as a secondary reference; operationalize 'equivalence class' (the predicate the fix draws). Nit: degradation reworded. * chore(#1962): backfill changeset pr number (PR #2409) --------- Co-authored-by: sim <sim@local> |
||
|
|
6baa2a8182 |
feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger (#2407)
* test(#1961): add failing-first bug-taxonomy routing contract tests Epic #1957 Phase 2B. Source-text-is-the-product contract tests (3 taxonomy classes, explicit class->technique routing table, Bohrbug->repro+SBFL+bisect, Heisenbug->record-replay/stability+SKIP-SBFL, Concurrency->atomicity/order/ deadlock checklist, bug_class in DEBUG Current Focus, supersede-not-append) plus a routing-table specification object pinning the documented decisions (SBFL forbidden on Heisenbug is the load-bearing 1B/2B seam). Failing-first: reference, Phase 1.75, and routing-table reframe do not yet exist. * feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger Epic #1957 Phase 2B (reliability-critical). Adds Phase 1.75: classify the failure as Bohrbug / Heisenbug-Mandelbug / Concurrency, then route the investigation technique via an explicit class->technique table (Kernighan: no opaque heuristic). Bohrbug -> reproduction + SBFL (Phase 1.25) + git bisect; Heisenbug/Mandelbug -> record-replay (rr) + stability-stress + statistical sampling, with SBFL explicitly SKIPPED (a flaky spectrum poisons the Ochiai ranking — the load-bearing 1B/2B seam); Concurrency -> the atomicity/order/deadlock checklist first. Reframes (supersedes, not appends — Zawinski) the flat 'Technique Selection by situation' table into a class-routed table; the 11 techniques remain as routed targets. bug_class recorded in Current Focus (DEBUG template); common-bug- patterns catalog cross-referenced to the taxonomy. Full rules extracted to gsd-core/references/debugger-bug-taxonomy.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1961): address orthogonal review (phase-name drift, General lane, revoke framing, row-scoped tests, bounding) - HIGH: reference said 'Phase 1B' (epic shorthand); corrected to the deployed 'Phase 1.25' (matches the agent + SBFL reference). - HIGH: 6 of 11 techniques (Rubber duck, Delta, Working backwards, Differential, Comment-out, Follow-the-indirection) were orphaned by the situation-table reframe. Added a 'General (any class, situation-cued)' lane to BOTH the reference routing table and the agent's Technique Selection table that re-homes them — supersede-not-append now holds. - MEDIUM: the SBFL-skip is structurally retroactive (Phase 1.25 runs before Phase 1.75 classification), so reframed the table column from 'Do NOT use' to 'Revoke if already run' + an explicit 'retroactive revocation, not proactive skip' note stating the ordering honestly. - MEDIUM: contract tests are now row-scoped (parse the table by class, assert per-row) instead of presence-only; added a guard that the previously- orphaned techniques now have a General-lane route. - LOW: pinned the canonical bug_class value form (lowercase-kebab: bohrbug|heisenbug-mandelbug|concurrency; prose may use title-case). - NIT: added a 'Bound the Heisenbug-chase runs' note (rr/stability/sampling timeouts) per the unbounded-subprocess gauntlet. * chore(#1961): backfill changeset pr number (PR #2407) |
||
|
|
f8b16d1874 |
enhance(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger (#2405)
* test(#1960): add failing-first RCA-branching contract + schema-invariant tests Epic #1957 Phase 2A. Source-text-is-the-product contract tests (fishbone >=2 categories, AND-gate, multi-cause root_cause, backward compat, reasoning checkpoint candidate_causes+and_gate fields, debugger-philosophy single-cause note, DEBUG template) plus behavioral schema-invariant checks on two fixtures: two contributing causes (AND-gate yes) -> both recorded; single-cause (AND-gate no) -> one root_cause, identical to today. Failing-first: reference, agent edits, and template note do not yet exist. * feat(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger Epic #1957 Phase 2A. Guards against 5-Whys single-cause bias: before committing root_cause, the debugger enumerates candidate causes across >=2 Ishikawa categories (code/config/environment/data) and explicitly answers an AND-gate question. When the AND-gate fires, every contributing cause is recorded, so a multi-cause fix no longer recurs via the unaddressed second cause. Resolution.root_cause may hold one OR a small set (additive; single-cause sessions are byte-identical to today). The Structured Reasoning Checkpoint gains candidate_causes + and_gate fields; debugger-philosophy.md adds the single-cause-bias trap. Full rules extracted to gsd-core/references/debugger-rca-branching.md (slim Phase 2 routing + 2 checkpoint fields kept in the agent). INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md + DEBUG template updated. * fix(#1960): address orthogonal review (AND-gate self-consistency, parity guard, narrowed claim, ripples) - Reference: the collapse rule now enforces AND-gate self-consistency — and_gate=yes with a single confirmed cause is flagged as incomplete (return to Phase 3); a race/timing note clarifies such bugs bridge categories; the 'byte-identical' backward-compat claim narrowed to 'root_cause shape unchanged; reasoning_checkpoint gains 2 fields in every session'. - DEBUG.md: stale 'five-field' mirror prose -> seven-field (parallel-surface drift the reviewer flagged); new debug-session-management parity test pins the field-count claim to the gsd-debugger.md YAML keys (CRLF-safe). - Scalar-assuming consumers of set-valued root_cause updated: session-manager compact summaries (319/332), diagnose-only return (1062), archive entry (1216), ROOT CAUSE FOUND return (1322). - Test: added the AND-gate-yes/single-cause invariant + fixture; rephrased the fixture describe block honestly as a schema-invariant specification. - Phase 2 bullet phrasing clarified ('at hypothesis formation, before the Phase 4 commit'). * test(#1960): parity regex accepts word-form count ('seven-field' or '7-field') * test(#1960): parity regex counts array-valued YAML keys (no inline value) * chore(#1960): backfill changeset pr number (PR #2405) |
||
|
|
56a5c6404c |
feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger (#2403)
* test(#1959): add failing-first SBFL contract + Ochiai correctness tests Epic #1957 Phase 1B. Source-text-is-the-product contract tests (Ochiai formula documented, Tarantula fallback, top-N seeding, no-coverage skip logged, ranking->Evidence, Bohrbug gating) plus a behavioral Ochiai formula-correctness section: bound [0,1], max-score invariant, a known-fault fixture proving the fault ranks #1 (criterion 2), clean degradation on zero failing tests, and two fast-check properties. Failing-first: reference file and agent routing do not yet exist. * feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger Epic #1957 Phase 1B. When a runnable test suite with per-test coverage exists (>=1 failing AND >=1 passing test), the debugger computes an Ochiai suspiciousness ranking over the coverage spectrum and seeds the top-N suspicious locations into Evidence as first-class hypothesis candidates, narrowing the search space deterministically before LLM reasoning. Tarantula documented as fallback. Degrades cleanly (logged, never silent) when there is no test suite, no failing tests, or no per-test coverage, and is explicitly not trusted on flaky/Heisenbug spectra (pairs with Phase 2B bug-taxonomy). Full rules extracted to gsd-core/references/debugger-sbfl.md (slim Phase 1.25 routing kept in the agent to respect the size cap). No new coverage framework — reuses the project's existing test/coverage runner. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * test(#1959): bound property generators to valid coverage counts The [0,1] property generated failedExec independently of totalFailed, but Ochiai's score is only bounded by 1 under the coverage invariant failedExec <= totalFailed (a failing test that executed s is one of the totalFailed failing tests). Out-of-domain inputs (failedExec=100, totalFailed=5) make the formula correctly return >1. Bound failedExec by totalFailed via fc.chain so the property tests the real domain. Also cleaned up the ranking property (removed dead code). * fix(#1959): address orthogonal review (monotonicity property, degradation row, coverage bounding) - Replace vacuous ranking property (true-by-sort-construction) with a non-trivial monotonicity property: holding totalFailed + passedExec fixed, ochiai is non-decreasing in failedExec. An inverted formula would fail it. - Add the missing 'no passing tests' degradation row (preconditions require >=1 passing test; Tarantula would divide by totalPassed=0). - Bound the coverage subprocess (CLAUDE.md gauntlet): cap the coverage run, degrade-to-skip on timeout, never hang the debug session. - Reword 'discard the ranking' -> 'mark the Evidence entry as revoked (do not delete)' per Kernighan auditability. * test(#1959): bound monotonicity-property generator to valid coverage (failedExecA <= totalFailed) * chore(#1959): backfill changeset pr number (PR #2403) |
||
|
|
5e52350736 |
feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger (#2396)
* test(#1958): add failing-first guardrail contract tests Epic #1957 Phase 1A. Adds source-text-is-the-product tests asserting the 5-signal fix-acceptance guardrail contract (target test, mutation check, no-op/deletion detector, adjacent tests, revert-and-reconfirm), graceful degradation, FIX REJECTED BY GUARDRAIL return path, per-signal debug-file recording, and subprocess bounding. Failing-first: reference file and agent sections do not yet exist. * feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger Epic #1957 Phase 1A. Prevents accepting a fix that merely greens the test (Goodhart defense / APR overfitting). Adds a 5-signal gate run before fix acceptance: target test, mutation check (Stryker), no-op/behavior-deleting detector, adjacent/held-out tests, revert-and-reconfirm. Degrades gracefully when Stryker or a test suite is absent (each skip logged, never a silent pass), records per-signal results under Resolution.verification, and returns a FIX REJECTED BY GUARDRAIL outcome the session-manager surfaces for revise / accept-as-debt / abandon. Full rules extracted to gsd-core/references/debugger-fix-acceptance.md (slim routing kept in the agent to respect the agent-size cap). Debug template + INVENTORY + manifest + agent-size baseline + AGENTS.md updated. * test(#1958): correct newline-tolerant assertion + regen install-parity goldens The revert-and-reconfirm assertion collapsed whitespace before matching so markdown line-wrapping does not break it. Regenerated the golden-install-parity and install-tree fixtures (npm run gen:golden) to absorb the intentional gsd-debugger.md / gsd-debug-session-manager.md / DEBUG.md / new reference-file changes to the installed artifact tree. * fix(#1958): tighten guardrail per orthogonal review Addresses the isolated reviewer's findings: - signal 5 now states its recorded-repro dependency and routes the no-repro case to the degradation row; revert mechanism specified (git stash / git revert -n); minimality flag tied to diff structure, not revert-ability. - bounded-subprocesses section now bounds the git subprocess (5-30s) too, requires argv-array argument passing, and scopes Stryker to the driving regression test (a mutant killed only by a non-driving test is a finding). - new test-provenance (security) clause: the driving test must be agent-authored; bug-report repro scripts are DATA, never executed verbatim. - tightened 3 contract assertions to bind to specific clauses (guardrail_verdict field, deletion-reject-unless-RCA, 60s+git bounding). - Goodhart framing softened to 'partially-independent'; DEBUG.md template verification field notes the nested map shape. * chore(#1958): backfill changeset pr number (PR #2396) * fix(#1958): add issue ref to allow-test-rule annotation (ADR-456) CI lint-allow-test-rule-refs requires every allow-test-rule exemption to carry a 'see #NNN' issue ref per ADR-456. The new test file's annotation lacked it; this adds (see #1958). |
||
|
|
f74442310d |
fix(#2257): auto-resume debug on non-terminal session-manager return (#2300)
The /gsd-debug orchestrator handled the gsd-debug-session-manager return with only two literal-string checks (DEBUG SESSION COMPLETE, ABANDONED) and no else branch, so a usable-but-non-terminal progress summary (the manager's own turn/context budget exhausted mid-loop, with a valid on-disk checkpoint) fell through to the user as if the debug were complete. Same gap at the continue subcommand. Callee side (agents/gsd-debug-session-manager.md): add an explicit non-terminal CONTINUE_REQUIRED return marker, distinct from the two terminal shapes and from a genuine user-input checkpoint. Orchestrator (gsd-core/workflows/debug.md Sections 4 and 1c): classify returns exhaustively — recognized terminal markers behave as before, anything else is non-terminal and auto-resumes by re-spawning the session manager from the same slug/checkpoint. Anti-loop guard: after two consecutive no-progress resumes (unchanged next_action/updated), emit a blocker report instead of looping. Regression test (source-text contract guard, fix-2196 idiom) asserts both sections' non-terminal/auto-resume branch, the CONTINUE_REQUIRED marker, and the anti-loop bound. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
315d94f6d4 |
feat(#1945): tracer-first planning default + executor feedback gate (#2294)
* feat(#1945): tracer-first planning default + executor feedback gate Make "thin end-to-end slice first, verify, then expand" the default planning + execution discipline instead of the opt-in --mvp mode. - gsd-planner: first-class `type="tracer"` task; every plan LEADS with one production-quality end-to-end tracer slice by default; --no-tracer restores horizontal layers; --mvp/--tdd compose on top. - gsd-executor + execute-plan: post-tracer feedback gate — autonomous runs halt-on-fail before expansion, interactive runs emit checkpoint:human-verify after the tracer. - --no-tracer flag wired through plan-phase workflow/command/help/skill. - CONTEXT.md glossary defines tracer bullet vs prototype; docs + references reconciled. - tests/tracer-bullet.test.cjs: prose-contract + behavioral (verify plan-structure accepts tracer) coverage. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1945): backfill changeset PR number to 2294 --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3592697bed |
fix(#2107): orchestrator honors gate="blocking-human" checkpoints in auto-mode (#2113)
* fix(execute-phase): honor gate="blocking-human" in auto-mode checkpoint handling The package-legitimacy gate (#2827) spans two layers. gsd-executor refuses to auto-approve a gate="blocking-human" checkpoint and escalates it so a human can vet the package. execute-phase's checkpoint_handling step then dispatched purely on checkpoint *type* and never read gate -- so under --auto/--chain it auto-approved the checkpoint the executor had just refused to auto-approve. Net effect: the slopsquatting defence was inert in exactly the unattended mode where it matters. An [ASSUMED]/[SUS] package reached install with no human ever seeing the prompt. - gsd-core/workflows/execute-phase.md: carve out gate="blocking-human" (and the package-legitimacy what-built markers) ahead of every auto-mode branch. - gsd-core/references/checkpoints.md: document the gate attribute and its two values. blocking-human previously appeared nowhere outside gsd-executor.md, so no planner had a documented way to author a non-auto-approvable checkpoint. - tests/package-legitimacy-gate.test.cjs: the existing regression test asserted the executor half only, which is why it stayed green while the gate was open. Now asserts the orchestrator half too. * chore(changeset): link to issue #2107 * chore(changeset): backfill PR number 2113 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa * test(#2107): refresh golden-install-parity hashes for edited gsd-core files The golden fixtures pin content hashes for gsd-core/references/checkpoints.md and gsd-core/workflows/execute-phase.md, both edited by this fix. Regenerated via UPDATE_GOLDEN=1; only those two keys change across all 17 runtime fixtures. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa * fix(#2107): keep the carve-out inside the ADR-857 host-loop budget The ADR-857 phase-6 ratchet pins execute-phase.md below 93600 LF bytes so optional-feature logic keeps migrating out of the host loop. The carve-out first landed 623 bytes over that ceiling. Move the two-layer rationale (why gsd-executor escalates these checkpoints) into references/checkpoints.md, where the gate is now documented, and reduce the workflow to the operative rule. execute-phase.md is 93589 bytes, under the ceiling; the gate token and both <what-built> marker strings are kept because the orchestrator matches on them. Refresh the two baselines the edit invalidates: golden-install-parity fixtures (only the checkpoints.md and execute-phase.md hashes move) and workflow-size-baseline.json (one line). The ADR-857 ceiling itself is untouched. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa * fix(#2107): executor honors blocking-human on the decision branch + gate transport Review found the fix incomplete one layer down. Two executor-layer gaps: 1. Blocker — agents/gsd-executor.md auto-mode dispatch gated checkpoint:human-verify on gate="blocking-human" but the checkpoint:decision branch below auto-selected the first option with no gate check. The executor resolves a decision itself (auto-selects and continues) without returning it, so the orchestrator carve-out never runs for it. A planner following the new checkpoints.md rule 6 ("gate a decision whose default would be wrong to assume") would have it silently auto-selected under --auto/--chain — the exact #2107 harm, one checkpoint type over. The decision branch now STOPs and returns for an explicit human decision when gate="blocking-human". 2. Major (transport) — checkpoint_return_format carried no field conveying the gate to the freshly-spawned orchestrator, so recognition of the proactive pre-install checkpoint rested on freeform prose. Added a **Gate:** field to the return format and re-pointed the execute-phase carve-out at it ("If the returned Gate: is blocking-human"). Net byte-negative: execute-phase.md drops 93589 -> 93583, widening ADR-857 headroom from 11 to 17 bytes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#2107): cover decision carve-out + gate transport, de-vacuum conditional tests - New: 'auto mode does not auto-select a blocking-human decision checkpoint' asserts the executor decision branch STOPs on blocking-human. Verified red on the pre-fix executor (2 fail), green with the fix (27 pass). - New: 'checkpoint_return_format transports the gate ...' asserts the **Gate:** field carries blocking-human across the executor->orchestrator boundary. - New: 'auto-select rule for decision is conditional' — orchestrator-side mirror of the human-verify conditional test, for the execute-phase decision branch. - Fix vacuous test: both conditional tests now assert the anchor matched (length > 0) before iterating, so anchor drift can no longer pass with zero assertions. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#2107): refresh golden + size baselines for executor + execute-phase edits Regenerated via UPDATE_GOLDEN=1 and update-size-baseline.cjs. Only the gsd-executor.md and gsd-core/workflows/execute-phase.md hashes move across the runtime fixtures (35 ins / 35 del, no keys added or removed); checkpoints.md is unchanged this round. Size baselines: gsd-executor.md 43607 -> 43973, execute-phase.md 93589 -> 93583 (still under the ADR-857 ceiling). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
b5ce72f729 |
fix(#2119): single SECURITY.md writer — auditor is return-only (#2154)
* fix(2119): single SECURITY.md writer — auditor is return-only The gsd-security-auditor held Write/Edit and was instructed to write SECURITY.md (no <N>- prefix, no template frontmatter), while the orchestrator's Step 6 also wrote the correct padded <N>-SECURITY.md from templates/SECURITY.md. Two writers, two naming conventions, two shapes — the auditor's unprefixed file was invisible to the workflow's *-SECURITY.md glob detector and unparseable for the threats_open gate. Fix (option 1 from the issue): make the auditor return-only. - Remove Write/Edit from auditor's tools - Rewrite all 'Write SECURITY.md' instructions to 'Return structured verdict' with threats_open count - Add explicit constraint in workflow Step 5 spawn prompt - Update existing test (was asserting Write in tools — now asserts absence) - Add new regression test for single-writer contract - Update docs/AGENTS.md stale Tools/Produces rows - Regenerate golden fixtures + agent size baseline * docs(changeset): backfill PR number (#2154) * chore(#2119): regenerate pi/qwen golden fixtures after next merge The single-writer change edits gsd-core/workflows/secure-phase.md and agents/gsd-security-auditor.md; pi.json (added on next) and qwen.json (merge straggler) were the only runtime fixtures still holding pre-change hashes for those files. All other runtimes already reflect the change. Regenerated via the sanctioned gen-golden-install-parity script. * merge origin/next — regenerate goldens + baseline for merged state * fix slash-command syntax: /gsd-secure-phase → /gsd:secure-phase (#2154 CI fix) |
||
|
|
c1756d0cd5 |
chore(#1867): register ui-consideration probe + regen install cascade (SHIP-01)
Ship-safe registration + regenerated snapshots for the #1867 UI-consideration probe (Phase 3, SHIP-01): - CONTEXT.md: PROBE.ui.{verification,axis,seam} predicates + ui-consideration -probe added to PROBE.family (machine-canon for the 3rd adapter, MIXED axis). - agents/gsd-ui-{researcher,checker}.md: one @-include of references/ui-consideration-probe.md each (both under the 24576 agent cap). - docs/INVENTORY.md + INVENTORY-MANIFEST.json: register the reference doc and the compiled ui-consideration-probe.cjs (inventory-manifest-sync green). - tests/fixtures/golden-install-parity/*.json (16 runtimes): recaptured against a clean full build — folds in the deferred Phase-1 (ref doc, plan-phase lift) and Phase-2 (ui-phase step, UI-SPEC section) install-surface changes. - tests/agent-size-baseline.json: ratcheted the two grown UI agents. - .changeset/vivid-orcas-chatter.md: type Added (pr updated at PR-open). Inventory/golden/size gates green; lint:ci + lint:docs + lint:changeset green. The plan-phase.md PRE_PHASE6 ceiling stays RED pending #1852 (unchanged). Claude-Session: https://claude.ai/code/session_01BKt4hgNZwXSeJYJtYAQUSS |
||
|
|
f15c25867d | docs(#1578): rebase B2+B3 agent prompt changes | ||
|
|
68a5258d45 |
fix(#2017): grant mcp__plugin_context7_context7__* for plugin-marketplace context7 (8 agents) (#2029)
* fix(#2017: grant mcp__plugin_context7_context7__* for plugin-marketplace context7 The 8 context7-using agents granted only mcp__context7__* (standalone server form). Claude Code's plugin-marketplace context7 install names tools mcp__plugin_context7_context7__*, so the grant never matched and every researcher/planner/executor silently lost doc lookup (fell back to WebSearch). - 8 agents: add mcp__plugin_context7_context7__* alongside mcp__context7__*. - scripts/research-profiles.cjs: update the researcher profile tools to match. - tests/context7-plugin-grant-parity.test.cjs: regression guard — no agent grants the standalone form without the plugin form. Closes #2017 * docs(#2017): backfill changeset pr 2029 |
||
|
|
9f0d785b61 |
fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows (#2027)
* fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows gsd-executor.md referenced sdk/src/query/QUERY-HANDLERS.md and reapply-patches.md referenced sdk/dist/cli.js — both retired with the SDK (ADR-0174). AI runtimes that locate doc refs via filesystem search ran find /, which on Git Bash for Windows traverses the whole drive (14h+, orphaned find.exe, 4M+ handles). - agents/gsd-executor.md: drop dead QUERY-HANDLERS.md ref. - workflows/reapply-patches.md: drop dead sdk/dist/cli.js clause. - tests/no-dead-sdk-refs.test.cjs: regression guard — no sdk/src|dist|handlers file refs in agents/workflows/references markdown. Closes #2020 * docs(#2020): backfill changeset pr 2027 |
||
|
|
ed79902509 |
feat(#2007): implement mempalace memory_mode kg_backend and replace routing (#2010)
Wire the two forward-declared mempalace.memory_mode modes so they actually route recall/capture instead of silently behaving as `augment`: - kg_backend: the palace temporal KG is the primary knowledge-graph source; native .planning/graphs/ is the fallback. Non-KG drawer recall stays additive. - replace: recall resolves through the palace as the source of truth; native artifacts are the fallback. Every mode stays onError:skip and default-resilient — an unreachable palace degrades to native memory and GSD keeps writing .planning/graphs/, so no memory is lost. Cross-mode .planning/graphs/ migration remains a documented open question (PRD/ADR §17), out of scope here. Surfaces updated (instruction-only contract): recall/capture commands (+ generated skills), discuss/wave fragments, curator agent, capability.json schema. Docs: how-to Step 3, CONFIGURATION, FEATURES, CONTEXT glossary. Regenerated capability-registry, golden install-parity fixtures (mempalace hashes only), agent-size-baseline. Added a routing-contract + cross-surface parity test. Incidental (folded per no-defer rule): removed pre-existing unused imports (spawnSync in capability-registry.test.cjs; fs in issue-498-package-identity.test.cjs) that eslint flagged in/alongside the touched files. Closes #2007 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a62079b2da |
fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR (#2024)
* fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR The gsd_run preamble resolved the Claude global install only at $HOME/.claude/gsd-core/bin/, but the installer honors CLAUDE_CONFIG_DIR — so a global install redirected via CLAUDE_CONFIG_DIR was invisible to every gsd_run call (every command failed with 'gsd-tools.cjs not found'). The Claude resolver arm now uses ${CLAUDE_CONFIG_DIR:-$HOME/.claude}, matching the installer + the other runtimes' ${VAR:-default} pattern. Default $HOME/.claude behavior is unchanged. - _runtime-launcher.snippet.sh: Claude arm honors CLAUDE_CONFIG_DIR. - sync-runtime-launcher.cjs re-run: 95 workflows/agents re-synced. - review.md / discuss-phase.md: trimmed to stay under their byte budgets. - runtime-launcher-parity.test.cjs: (A) substring updated for the new form + explicit #1865 assertion that the snippet honors CLAUDE_CONFIG_DIR. - goldens + size baselines recaptured. Closes #1865 * docs(#1865): backfill changeset pr 2024 |
||
|
|
7bef6a6496 |
fix(#1863): use named flags for state.* calls in executor + workflows (#1873)
* fix(#1863): use named flags for state.* calls in executor + workflows The named-only state-command router (parseNamedArgs) silently drops positional args, so state.cjs threw its required-arg error and metrics/decisions/blockers/session continuity were never recorded. Convert record-metric / add-decision / add-blocker / record-session in agents/gsd-executor.md to the named-flag form (mirroring execute-plan.md), and fix the two remaining positional record-session calls in gsd-core/workflows/milestone-summary.md and forensics.md. Recapture the golden-install-parity fixtures and size baselines for the edited files. Also fix a pre-existing detached-rebuild handle leak in tests/graphify-auto-update.slow.test.cjs: three dispatch tests returned after observing only the synchronous "running" status without awaiting the detached rebuild's terminal state. That leak was latent until the new #1863 regression block's added runtime shifted --test-force-exit timing and surfaced it as a non-zero chunk exit. The three tests now await terminal status via the file's existing waitForBuildStatus helper. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1863): add changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
3c13903dcd |
feat(#1866): agent-side self-load of configured agent_skills
Each of the 22 consumer agents now self-loads its configured agent_skills in its mandatory init step, so .planning/config.json agent_skills.<type> reaches the agent on every runtime — including Cursor and /gsd-autonomous, where Skill()-delegated workflow bash init did not reliably execute. - gsd-core/references/agent-skills-bootstrap.md: shared contract (query + Read + dedup guard that skips when <agent_skills> is already in the prompt, so Claude's orchestrator-side injection never doubles) - 22 agents/gsd-*.md: one self-load line naming the agent's own type - gsd-core/workflows/autonomous.md: note that delegated agents self-load - tests/agent-skills-bootstrap.test.cjs: regression + parity (CONSUMER_AGENTS bijection + fast-check property) — Generative-Fix-Divergence guard - docs: ADR-1866, CONFIGURATION dual-injection How It Works, INVENTORY row, Changed changeset Closes #1866 |
||
|
|
1bd04e1565 |
docs(#1847): add Claude Sonnet 5 changeset for next-line changelog + refresh stale model examples
The 1.6.1 forward-port (#1851) used the no-changelog opt-out instead of carrying a changeset, so Sonnet 5 — unlike every other 1.6.1 fix (#1580/#1591/#1693 whose fragments live on next) — had no fragment and would be MISSING from the 1.7.0 changelog. Add the fragment so the release render reflects current shipping code. Also note the bold-checklist form in the #1591 fragment, and refresh two stale claude-sonnet-4-6 illustrative examples to current IDs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
18995380ce |
feat(#1154): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1738)
* feat(verify-phase): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1154) Carry the edge-probe's existing `backstop` (non-inferable) tier through the plan-phase projection as a structured flat-scalar marker instead of a prose parenthetical, and make verify-phase abstain -> human_needed (never silent-pass) on a backstop truth it cannot confirm with explicit evidence. Truth-axis mirror of #644's prohibition judgment-tier (ADR-550 D4). Engine (deterministic, CI-tested per ADR-550 D5 — never the LLM verdict): - src/probe-core.cts: truthStatement/truthVerification normalizers, projectTruths (conservative serializer), dispositionForUnverifiableTruth (backstop+no-evidence -> unverified/flagged/insufficient_spec; backstop+evidence -> green; inferable -> green, the over-abstention guard). - src/roadmap.cts: coerceTruthToString now reads `statement` first so an object-form backstop truth is surfaced, not dropped (Hyrum backward-compat for truth-readers). Workflow/agent/docs: plan-phase emits the structured marker (flat scalar, ADR-550 #1278); verify-phase + gsd-verifier add the abstain arm; new references/honest-verifier.md; FEATURES/COMMANDS document insufficient_spec; ADR-550 amended (truth-axis D4 mirror). Decisions adopted (trek-e review): insufficient_spec feeds existing human_needed with a distinguishable reason (no new VERIFIER_STATUS); changeset Changed; round-trip parity test; abstain-on-unconfirmed-backstop regression test red-first. Implementation notes (deviations from the issue's proposed file list, verified live): - frontmatter.cts needs no change — its flat parser already round-trips object-form truths. - verify.cts needs no change — it grades artifacts/key_links structurally; truths are LLM-graded at the workflow layer, so consumption lives there + the deterministic helper. - No CJS<->SDK hand-sync — the SDK seam was retired (ADR-0174); src/*.cts is sole source. Regenerated artifacts: golden-install-parity fixtures, INVENTORY-MANIFEST, size baselines. * chore(#1154): add changeset (Changed) for honest verifier User-facing changelog fragment for #1738. Typed `Changed` (not `Added`) per trek-e review condition 3 — the verify behavior shifts for backstop-bearing specs (a confident silent `passed` becomes `human_needed`), which is user-visible even though the schema marker is additive. * docs(#1154): score-formula also excludes abstained insufficient_spec truths (review nit-1) trek-e review nit: the verify-phase score sentence said PRESENT_BEHAVIOR_UNVERIFIED truths were "the only ones excluded" from verified_truths. Post-#1154 an abstained `insufficient_spec` backstop truth is also excluded (it is not ✓ VERIFIED and routes to human_needed). Behavior was already correct; this tightens the wording. Regenerated golden-install-parity fixtures + workflow-size baseline for the touched verify-phase.md. (Nit-2 — a dedicated insufficient_spec_items frontmatter list — is intentionally not taken: the current design is ADR-550-D4-conformant, the abstain cause rides as a distinguishable report reason, and adding it would exceed the approved scope.) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
1a46109b97 |
enhance(#1579): deterministic gsd-tools query eval.score verb (#1583)
* feat(#1579): deterministic gsd-tools query eval.score verb Split C of #1573 (pure code, lowest risk). Adds an eval.score query verb (coverage*0.6 + infra*0.4; bands 80/60/40) mirroring the verify.* chain; gsd-eval-auditor consumes it instead of doing weighted arithmetic in-prompt. Non-breaking — additive only. arXiv: 2601.15130 (Plausibility Trap/DPDM), 2507.10281 (Table Agent), 2508.15754 (TIR). * fix(#1579): address review — domain guard, property test, glossary, SKIP_ROOT, inventory/baseline - C3 input-domain: reject out-of-domain eval.score (require 0<=covered<=total; was emitting overall_score>100 / negatives) - C1 property test: add tests/eval.property.test.cjs (fast-check) — determinism, band monotonicity, [0,100] bounds, never-throws - C2 glossary: CONTEXT.md "Eval Scoring Module" entry (source-of-truth path + interface) - C4: add `eval` to SKIP_ROOT_RESOLUTION (pure arithmetic; no .planning/ access) - inventory: register generated eval.cjs/eval-command-router.cjs (INVENTORY-MANIFEST.json + INVENTORY.md rows) - size: regen agent-size baseline for gsd-eval-auditor (reused gsd_run shim + eval.score step) - eslint: ignore generated eval*.cjs (ADR-457 bin/lib migration coverage) * fix(#1579): register eval family in alias-drift gates Add EVAL_COMMAND_ALIASES/EVAL_SUBCOMMANDS to scripts/check-alias-drift.cjs families and to familyArrayKeys in the manifest-coverage test, so the eval family lands under the same drift guard as every sibling family (state/verify/init/phase/phases/validate/roadmap). Addresses trek-e review. check:alias-drift ok; feat-3251 coverage 9/9; eval suites 10/10. * docs(#1579): use half-open verdict band ranges in CLI-TOOLS overall_score is fractional and thresholds are >=80/>=60/>=40, so a score in [79,80) is correctly NEEDS WORK despite the old '60-79' label. Relabel bands as 60-<80 / 40-<60 / 0-<40 to match the code. Addresses trek-e nit. * fix(#1579): validate eval.score CLI inputs Reject unknown infra tokens and fractional counts, and pin the 80-point verdict boundary including rounding-before-banding behavior. |
||
|
|
a63684c222 |
enhance(#1577): WebFetch/WebSearch injection isolation + opt-in blocking (#1585)
* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the largest untrusted channel) in gsd-read-injection-scanner; shared untrusted-input-boundary reference @-included by the 8 ingest agents (randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring); opt-in security.injection_blocking (default advisory — non-breaking). arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth). * fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized - A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already in the transcript. The prompt-level data/instruction boundary is the primary control. - A2: registered security.injection_blocking in the config schema + defaults manifests (default false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads. - A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention). - A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale). - A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input. - Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) + drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as the noted pre-existing follow-up. * fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate The new reference quotes injection phrases ('ignore previous instructions', 'you are now…') as examples agents must NOT comply with, tripping the repo's own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red on HEAD). Allowlist it alongside the other security docs (security-model.md, TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS scanner test doesn't scan references/, so only the shell gate needed it. Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15. * fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so no named web-ingress agent is uncovered, keeping the two justified additions (gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10. - gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset. - gsd-assumptions-analyzer reads 5-15 codebase source files (external/source- document ingress per the boundary), though it has no web tools. INGEST_AGENTS in the isolation test now asserts all 10; size baselines regenerated (+60 bytes each, both well under the DEFAULT cap); changeset reworded 8 -> 10. Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39. * docs(#1577): document security.injection_blocking + boundary seam trek-e Major 2 + Minor: - docs/CONFIGURATION.md: add the top-level security.injection_blocking key to the Full Schema and a Security Settings subsection, distinguishing it from the workflow.security_* namespace; honest circuit-breaker-not-redactor framing matching ADR-1577 / security-model. - CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry. Verified: lint:docs ok; config-field-docs + contributor-standards green. * test(#1577): make read-injection property test git-text, not binary trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as degenerate-edge inputs. The NUL is what actually made git classify it binary (git binary = NUL in first 8K). Replace both with text-safe escapes that keep the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File now diffs/blames line-by-line. Verified: property test 2/2; no NUL/raw-noncharacter bytes remain. * docs(#1577): align untrusted boundary docs Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner. * docs(#1577): align ADR ingest agent count Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
207d8f1697 |
fix(#1626): make the security gate severity-aware via per-threat severity (#1635)
workflow.security_block_on was documented as the minimum threat severity that blocks advancement, but threats carried no severity and the auditor's threats_open count (the SECURITY.md gate field) counted every open threat regardless of severity — so the threshold had no effect, and the auditor's block_on vocabulary (open/unregistered/none) did not even match the config enum (critical/high/medium/low/none). - planner: add a Severity column to the STRIDE threat register; assign severity per threat. - auditor: read severity; reconcile the <config> block_on domain to the severity enum; redefine threats_open as the count of OPEN threats whose severity is at or above block_on (none => 0). Below-threshold opens are reported as non-blocking and excluded from threats_open. - SECURITY.md template + planning-config.md reconciled. No gate-check site changed: threats_open == 0 stays the gate everywhere; only its computation is now severity-filtered. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f9d9dfb4bc |
fix(#1627): scale security rigor by ASVS level (planner disposition + auditor depth) (#1636)
workflow.security_asvs_level was display-only — the planner hardcoded 'mitigate if ASVS L1 requires it' and the auditor only echoed the level, so L2/L3 behaved identically to L1. - New reference gsd-core/references/security-asvs-levels.md defines L1 (opportunistic), L2 (standard), L3 (comprehensive) for both planner threat disposition and auditor verification depth (higher = superset). - planner: disposition now scales with the configured ASVS level (no hardcoded L1) + @-pointer to the reference. - auditor: verification depth scales with asvs_level (L1 grep-presence, L2 boundary/vector check, L3 end-to-end trace + bypass check). - planning-config.md + INVENTORY updated; planner kept under its 48K cap by extracting the goal-backward worked example to planner-guidance.md. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
2dedbdd11c |
fix(#1455): resolve project-code-prefixed roadmap headings
* fix: remove hardcoded phase project-code prefix cap
Centralize project-code prefix stripping/matching and replace fixed {1,6} caps so long codes (for example MANIFOLD-117) resolve across phase, roadmap parser, roadmap upgrade, and validate flows.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: address review feedback on prefix parsing
- allow project_code prefixes with digits and underscores while preserving milestone parsing
- use shared optional project-code prefix source in phase dir parsing
- extend regression coverage for APP1/APP_1 prefixed phases
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: resolve review nit in phase-id test comment
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix: resolve project-code-prefixed roadmap headings
Make getRoadmapPhaseInternal recover from drifted project-code-prefixed ROADMAP headings while preserving canonical bare-heading preference. Add init.phase-op and parser regressions for #1455 and guard gsd-roadmapper against emitting project_code in headings.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: add changeset for project-code phase fix
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* chore: update roadmapper agent size baseline
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(#1455): tighten project-code prefix regex to [A-Z] start; add boundary test; document source order
Resolves blockers from review:
- Regex changed from [A-Z_][A-Z0-9_]* to [A-Z][A-Z0-9_]* so leading
underscores (_FOO-7, _-7) are never misread as project-code prefixes;
adds boundary test asserting both do NOT strip.
- Adds comment to roadmapPhaseLookupSources explaining why 3 sources are
needed (order-dependent canonical-heading preference).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Solvely-Colin <211764741+Solvely-Colin@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
||
|
|
fe0f9f904d |
fix(#1478,#1479,#1480): prohibit ungrounded baselines, error-suppressing fallbacks, and stale-artifact authority in planner verify blocks (#1482)
* fix(#1478,#1479,#1480): prohibit ungrounded baselines, error-suppressing fallbacks, and stale-artifact authority in verify blocks Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore: add changeset for #1478/#1479/#1480 planner verify gate fix Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: correct changeset format Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1478,#1479,#1480): fix test contract violations from new planner dimensions gsd-planner.md exceeded both the planner-decomposition 48K char limit and the reachability-check 50K char limit after the new HARD RULE blocks were added inline. The full rule details already exist in planner-antipatterns.md (added in the same PR); replace the verbose inline blocks with a single @-reference pointer to the antipatterns file, reducing the file from 50981 to 49130 chars (under both limits). Also regenerate tests/agent-size-baseline.json to reflect the new sizes of gsd-planner.md and gsd-plan-checker.md. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
9540fe43b9 | fix: have executor self-report worktree metadata (#1349) | ||
|
|
2d264da661 |
enhance(#968): region-scoped negative-grep idiom + cross-task conflict warning (#1320)
Adds a region/function-scoped negative-grep idiom to the gsd-planner verification guidance plus a warn-only `validate_plan` check (`scanFileWideNegativeGateConflict`) that flags when a task's file-wide negative grep bans a construct a sibling task legitimately requires elsewhere in the same file. ReDoS-safe (linear, no RegExp on author patterns); region-scoped gates are exempt. Warn-only — never errors, never flips `valid`. Closes #968 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
b108f101b0 |
fix(#1284): grant mcp__perplexity__* to researcher agents + dispatch-table parity guard (#1288)
Adds mcp__perplexity__* to both researcher profiles (generated source-of-truth) and regenerates the agents; adds a generative dispatch-table↔tools parity guard so future provider drift fails CI. Regenerates the agent-size baseline for the +20-byte frontmatter growth. Fixes #1284 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0a856f06cd |
enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier (#1271)
* enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier Introduce a per-truth PRESENT_BEHAVIOR_UNVERIFIED state for must-haves that assert a state transition or a cancellation/cleanup/ordering invariant whose only evidence is symbol presence + wiring. Such truths are excluded from the verified_truths score, reported as a behavior_unverified count, recorded in an always-on behavior_unverified_items frontmatter list, and routed to the existing human_needed sink — so a clean N/N can no longer be reached on symbol presence alone. The overall-status vocabulary and the src/verification.cts seam are unchanged (the new state is per-truth only); gaps_found keeps decision-tree precedence; override-passed truths still count toward verified_truths. Mirrors the calibration into the shipped verify-phase.md workflow (with an infra/foundation carve-out), the VERIFICATION.md templates, and docs (planning-artifacts.md, AGENTS.md). gsd-verifier.md kept under its 48KB LARGE cap; size baselines regenerated. Closes #966 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#966): add changeset fragment for gsd-verifier behavior-unverified calibration Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
cf68841220 |
enh(#1243): consume Claude plugin-provided skills in agent_skills (epic #1258 Phase B) (#1261)
* feat(#1243): consume Claude plugin-provided skills via native Skill-tool directive + grant Skill to agent_skills-consumer agents - Relax global skill name validation to accept namespaced form `^[A-Za-z0-9_-]+(:[A-Za-z0-9_-]+)*$` - Namespaced names (containing colon) on claude runtime emit a Skill-tool load directive instead of a @-include line - Namespaced names on non-claude runtimes are skipped with a warning - Bare unresolved names retain existing warn-and-skip behavior (no promotion to directive) - Grant `Skill` tool to all 22 agent_skills consumer agents; 5 generated agents updated via research-profiles.cjs + regen, 17 hand-authored agents edited directly - Add 16 TDD tests in describe('bug #1243') covering happy/mixed/precedence/negative/cross-runtime/regression/grant cases Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * docs(#1243): document plugin-provided skills in agent_skills Update the Agent Skills Injection reference in CONFIGURATION.md with the three entry forms (project-relative, global:<name>, global:<plugin>:<skill>), the Claude-only runtime behaviour of the namespaced form and the warn-skip on other runtimes, the plugin pre-install prerequisite, and the consumer-agent Skill tool grant. Add docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md with a step-by-step guide for installing the plugin, locating the namespaced skill name, wiring it into agent_skills, and verifying injection. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1243): align agent_skills docs with emitted block format + mixed-block regression test (code-review) - Replace two-section mixed-block example (bogus "Load these plugin-provided skills using the Skill tool:" header) with the actual single-section inline format in CONFIGURATION.md and docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md - Fix quoted warning text in how-to doc to exactly match the emitted string: [agent-skills] WARNING: Plugin-namespaced skill "global:<name>" requires a Skill-tool-capable runtime (claude) — skipping on runtime "<runtime>" - Replace phantom agent slugs (gsd-checker, gsd-researcher, gsd-advisor, gsd-synthesizer) in CONFIGURATION.md Supported Agent Types with real agents/gsd-*.md examples (gsd-plan-checker, gsd-phase-researcher, gsd-code-reviewer, gsd-ui-auditor, gsd-research-synthesizer) - Add byte-identical mixed-block regression test: one path-resolvable global skill + one plugin-namespaced skill on claude runtime → asserts r.ir.block === single-section interleaved block, no secondary header Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore(#1243): regenerate agent-size baseline for the Skill-tool grant The 22 agent_skills-consumer agents each grew +7 bytes from adding `Skill` to their tools list; refresh the committed per-agent size baseline (#1074 guard). * chore(#1243): add Added changeset fragment * fix(#1243): traceable allow-test-rule ref + separator-agnostic byte-identical tests (CI) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
395fb519e7 |
feat(spec-phase): prohibition probe — surface "must-NOT" constraints (#644) (#1149)
Adds the spec-time prohibition probe (spec-phase Step 5.6) — the second adapter of the probe-core resolution model. Surfaces unwritten must-NOT constraints as negative SPEC acceptance criteria with test/judgment verification tiers; fail-closed at verify time. Per ADR-550. Closes #644. |
||
|
|
1a186013a4 |
fix(#1205): roadmapper applies phase_id_convention to generated phase IDs (#1215)
* fix(#1205): roadmapper applies phase_id_convention to generated phase IDs - Add Phase ID Convention section to <phase_identification> block: documents sequential (default) vs milestone-prefixed forms, and instructs the agent to read phase_id_convention from config.json - Update <output_formats> to show both header and checklist forms for sequential and milestone-prefixed conventions with examples (e.g. ### Phase 1-01: Name, - [ ] **Phase 1-01: Name**) - Add TDD regression test tests/bug-1205-roadmapper-convention.test.cjs (5 assertions, confirmed fail-first then pass after fix) - Update tests/agent-size-baseline.json to reflect legitimate growth - Add .changeset/brave-otters-leap.md (Fixed, pr:0 placeholder) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: backfill changeset pr: 1215 for fix/1205 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1205): move phase_id_convention regression into roadmapper-granularity.test.cjs lint-regression-test-names rejects new standalone bug-NNNN-*.test.cjs files; regression cases must live in the owning module's test file. Move the 5 phase_id_convention assertions (#1205 regression) from the removed tests/bug-1205-roadmapper-convention.test.cjs into tests/roadmapper-granularity.test.cjs as a new describe block, alongside the existing granularity calibration tests. Also update the allow-test-rule comment to cover both #163 and #1205 surface contracts. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1205): fix lint-allow-test-rule-refs for roadmapper-granularity - Add issue ref (see #1205) to allow-test-rule comment in tests/roadmapper-granularity.test.cjs so lint-allow-test-rule-refs passes (new exemptions require #NNN per ADR-456) - Prune stale 'source-text-is-the-product' entry from scripts/lint-allow-test-rule-refs.allowlist.json (ratchet-down; comment now compliant and no longer needs grandfathering) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a375c4b354 |
feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) (#1201)
* feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) Adds an opt-in, default-resilient ADR-857 feature capability that wires MemPalace (local-first memory: MCP server + CLI) into the GSD loop: deliberate recall before discuss/plan and verbatim + temporal-KG capture at phase boundaries. Three memory modes (augment default; kg_backend and replace forward-declared). Master gate mempalace.enabled (default off); every hook onError:skip, zero gates; absent/disabled MemPalace => loop unchanged. Transport is rendered-markdown only — MemPalace runs out-of-process, no third-party code in gsd-core (ADR-857 §7). Capability: capabilities/mempalace/ (manifest + 2 fragments), skills commands/gsd/mempalace-{recall,capture}.md, agent agents/gsd-mempalace-curator.md. Registration: ns-context router, utility cluster, KNOWN_SKILLS, help full.md, model-catalog, copilot install list, size baselines; regenerated capability-registry + inventory manifest. ship:post wired into ship.md (wire-on-demand). HELD on #1196: this capability also declares hooks at discuss:pre and discuss:post, which are structurally un-wireable until the host-loop conformance model covers the discuss phase (discuss-phase.md is not in HOST_LOOP_FILES). The phase6-capstone-conformance gate therefore fails on exactly those two orphaned points by design — see #1196. Once #1196 lands, rebase onto next and the gate goes green with no further change. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#956): backfill changeset PR number (#1201) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
7c07fce70f |
fix(#381): make gsd_run launcher reachable in fresh-shell-per-block runtimes (#1084)
* fix(#381): make gsd_run launcher reachable in fresh-shell-per-block runtimes On runtimes that execute each fenced bash block in a separate shell process (e.g. Claude Code — documented behavior: each Bash command is a separate process; inline shell functions and exported vars do not persist between calls), the once-per-file gsd_run() function was undefined in every block after the preamble block, and the call was swallowed by `2>/dev/null || echo "{}"` into silent empty state. Fix (budget-neutral session-level resolution): - Ship gsd-core/bin/gsd_run, a POSIX sh wrapper that symlink-resolves its own location and execs the co-located gsd-tools.cjs. Exposed on PATH via the npm `bin` field (global installs) and shipped to local installs via the recursive gsd-core/ copy. - The per-file launcher preamble now appends `export PATH='<bindir>':"$PATH"` to the file named by $CLAUDE_ENV_FILE (Claude Code's documented env-persistence mechanism) so later fresh-shell blocks resolve gsd_run from PATH. Guarded as a strict no-op when CLAUDE_ENV_FILE is unset; the inline gsd_run() definition remains the fallback for all other runtimes. The single-quoted dir neutralizes shell metacharacters at source time. - Propagated via scripts/sync-runtime-launcher.cjs to all launcher-using files. - XL workflow byte budget 93000 -> 93200 (the ~130B clause pushes plan-phase.md to 93135; legitimate content growth, ratchet-up per #717). Regression tests (I)/(J) in runtime-launcher-parity.test.cjs cover wrapper delegation and end-to-end PATH persistence (sourcing the env file with a space-bearing install path). Known limitation: an install path containing a literal single-quote yields a malformed env-file line and falls back to the status quo (no regression); rare on sanitized home directories. Closes #381 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#381): add changeset for gsd_run fresh-shell reachability fix Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#381): scope test (J) bare-PATH execution to POSIX (Windows Git Bash exec bit) Windows Git Bash (msys2) does not honor Node's chmod exec bit for PATH-executing extension-less scripts, so the bare `gsd_run` command lookup failed there even though the env-file PATH persistence was correct. The env-file content assertions (the fix's actual cross-platform logic) still run on every platform; only the final source-and-execute sub-step is gated to non-win32. Global installs on Windows are covered by npm's generated bin shim. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8813ee5f95 |
feat(#429): HARD GATE on negative-grep literals echoed in plan <action> bodies (#1062)
Convert the planner's soft comment-text guideline into a plan-write-time HARD GATE. When an acceptance criterion negative-greps for a literal (`grep -c 'LIT' file == 0`) and that same literal appears verbatim in an `<action>` body (JSDoc samples, head-comment references, "what NOT to do" snippets), the executor's commit-time verify gate later fails on the comment echo rather than a real regression — wasting cycles and training the executor to distrust the gate. `verify.plan-structure` (the `validate_plan` step) now scans for this: - confidently-extracted (quoted) negative-grep literal echoed in an <action> → error (valid:false), failing plan creation - unquoted/ambiguous grep target → warning (fallback policy) - `<!-- planner-discipline-allow: LIT -->` escape hatch skips a literal - positive-count gates (`== N`) and `!= 0`/`>= 0` are out of scope Adds the `<comment_text_discipline>` block to gsd-planner.md, the full rules + allowlist example to planner-antipatterns.md, and regression fixtures for downstream incidents 12-04, 11-04, 12-02 (plus a boundary case proving positive-count gate 11-02 is not flagged). Closes #429 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1fab2e10ba |
fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver (#1045)
* fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver Source agents/*.md (gsd-planner, gsd-executor, gsd-verifier, gsd-plan-checker, gsd-intel-updater, gsd-debugger, …) called bare "gsd-tools …" in shell blocks. On a shim-only install — where gsd-tools.cjs exists under the runtime home but gsd-tools is NOT on PATH — those calls fail with "command not found" and the agent silently skips init/state/validate/commit ceremony, deferring to the orchestrator or bypassing GSD bookkeeping entirely. were never migrated, so it persisted on Claude Code and every other runtime that consumes the source agents directly. Only gsd-phase-researcher.md carried a resolver — and a stale, claude-only truncated one. Fix (all runtimes): - Inject the canonical multi-runtime gsd_run preamble (byte-equal to _runtime-launcher.snippet.sh — claude/codex/cursor/gemini/copilot/windsurf/ augment/trae/qwen/cline/opencode/kilo/hermes/antigravity homes) at the top of the first gsd_run block of all 12 gsd-tools-calling agents, and rewrite every command-position bare gsd-tools to gsd_run. - Upgrade gsd-phase-researcher.md's stale resolver to the canonical one. - Extend scripts/sync-runtime-launcher.cjs to maintain agents/ in parity (the sync caught and corrected a mis-placed preamble during development). - Extend the bare-gsd-tools (#2851) and launcher-parity (#373) regression guards to agents/ so no runtime can silently regress. Closes #1041 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1041): backfill changeset PR number to 1045 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
12d6285b4f |
fix(#1000): align gsd-intel-updater output to canonical intel filenames (#1037)
* fix(#1000): align gsd-intel-updater output to canonical intel filenames The intel-updater agent was instructed to write short names (files.json, apis.json, deps.json) and a markdown arch.md, but the intel library + gsd-tools intel CLI read only the canonical long names from INTEL_FILES (file-roles.json, api-map.json, dependency-graph.json, arch-decisions.json as JSON). After /gsd:map-codebase --query refresh the agent output was orphaned — intel status and validate reported the canonical files missing and intel query returned nothing. Renames every short reference to its INTEL_FILES canonical name and converts the arch output from markdown to queryable arch-decisions.json. Adds a drift-proof regression test (derived from the exported INTEL_FILES map) in the owning module's test file tests/intel.test.cjs, reviving the maintainer-approved approach from closed PR #608. Closes #1000 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1000): backfill changeset PR number to 1037 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f61b97276e |
fix(#724): block convergence on actionable review findings (#728)
* fix(#724): block convergence on actionable review findings * merge: integrate clean next (#936 inline) onto author tip + re-apply cursor fixes (Mode field, REVIEWS.md extraction) and review hardening (#724) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
972a41a528 |
fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:) (#990)
* fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:) Closes #967 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore(#967): backfill changeset pr number (990) --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |