e4df05126deaf5ad1c29bf35b9dfe2193c80cb0b
178 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c5e0371775 |
feat(#1951): reversibility tagging — gate one-way-door decisions (#2471)
* test(#1951): add failing-first tests for reversibility tagging Red phase for issue #1951 (reversibility tagging: classify decisions by undo cost, gate one-way doors behind a checkpoint:decision). Tests assert, per the issue's acceptance criteria: - discuss-phase CONTEXT.md template records a **Reversibility:** field with a rationale on captured decisions, and states it is optional - gsd-planner @-references planner-reversibility.md and stays under the 49152-char agent cap (LARGE_CAP, tests/agent-size-budget.test.cjs) - a one-way rating inserts a checkpoint:decision before the dependent task; reversible inserts none; costly is flagged but never blocks - the taxonomy defaults to reversible when unsure (checkpoint-fatigue guard) and inserting a checkpoint implies autonomous: false - docs/reference/plan-md.md documents <reversibility> as optional with all three ratings - --no-reversibility-gates parses to REVERSIBILITY_GATES=false, is injected into the planner prompt, and is advertised in the command argument-hint and help full mode (argument-hint parity) - the override suppresses the gate but still persists the rating - cmdVerifyPlanStructure accepts every rating and the absent case (additive-validator guarantee, behavioral via runGsdTools) - parity: thinking-models-planning.md #4 adopts the canonical three-level taxonomy and the binary REVERSIBLE/IRREVERSIBLE vocabulary is gone - no content loss from the planner extraction made to fit under the cap Prose-contract assertions are Red until the implementation lands. The behavioral validator assertions pass immediately — regression guards proving the validator already accepts unknown optional tags. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1951): reversibility tagging — gate one-way-door decisions Classify planning decisions by what undoing them would cost, and give a one-way door a human beat before the agent walks through it (issue #1951, The Pragmatic Programmer Topic 15 'Reversibility'; Bezos's one-way/two-way door framing). Acceptance criteria met: - discuss-phase records an optional reversibility rating with a rationale on <decisions> entries in the phase CONTEXT.md template. Unrated decisions are treated as reversible, so existing phases are unaffected. - a one-way rating makes gsd-planner insert a checkpoint:decision before the task that implements the decision, reusing the existing checkpoint mechanism -- no new checkpoint machinery. - reversible ratings trigger no checkpoint; costly ratings are flagged in the plan but never block. - the rating persists on the task as the optional <reversibility rating=> element. cmdVerifyPlanStructure accepts every rating and the absent case; the structural validator does not reject unknown optional tags. - --no-reversibility-gates (REVERSIBILITY_GATES=false) suppresses checkpoint insertion for intentionally-unattended runs while still recording ratings -- the override changes what stops the run, not what the plan remembers. Single taxonomy, not two: references/thinking-models-planning.md #4 already shipped a binary REVERSIBLE/IRREVERSIBLE classification and is loaded by both gsd-planner and gsd-plan-checker. It is rewritten onto the canonical three-level vocabulary and now points at planner-reversibility.md as the taxonomy owner, with a parity test that fails if the surfaces diverge (DEFECT.GENERATIVE-FIX-DIVERGENCE). agents/gsd-planner.md sat 47 chars under the 49152 LARGE_CAP, so the checkpoint DO/DON'T guidance was relocated verbatim into planner-antipatterns.md -- already @-referenced from the same section for the same topic, so the planner still loads it and nothing was dropped. A test guards the relocation against content loss. Files: gsd-core/references/planner-reversibility.md (NEW, canonical taxonomy + emission rules + anti-patterns), gsd-planner.md, plan-phase workflow/command/help (flag wiring + parity), plan-md.md schema, discuss-phase context template, CONTEXT.md glossary, INVENTORY + manifest, size baselines, install goldens, plugin skills regen, changeset. Closes #1951 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1951): address orthogonal review findings Two isolated reviewers (correctness + security), neither of which authored the change. Every finding fixed: Security — the rationale is untrusted input (ADR-1577). It originates in conversation and flows CONTEXT.md -> planner -> PLAN.md -> executor, each hop an LLM reading the previous hop's output, with no validation on the path. planner-reversibility.md and the discuss-phase template now state it is data and never instructions, and name the </reversibility> early-termination hazard explicitly -- a rationale that closes its own element injects sibling structure the executor reads as real tasks. Four tests guard it. Correctness 1 — nothing machine-enforced the feature's own promise: a task rated one-way with no preceding checkpoint:decision validated as fully clean, so a planner error silently reopened the gap this feature exists to close. cmdVerifyPlanStructure now warns on an ungated one-way rating. A warning, not an error: <reversibility> stays additive and the plan stays valid. Four tests cover ungated (warns), gated (silent), still-valid, and reversible/costly never flagged. Correctness 2 — pass-always test. The --no-reversibility-gates parse test substring-matched the whole workflow file, and plan-phase.md prose mentions both tokens in one sentence, so it passed with the bash conditional deleted: it was testing the documentation, not the parser. Now scoped to the fenced bash blocks and matched as one physical line, with a negative control confirming prose alone cannot satisfy it. Correctness 3 — costly had no itemized emission rule, only one-way did, so two agents could diverge on whether to tag costly at all. Correctness 4 — template convention break: the example ratings were bare while every sibling field uses [...] to signal substitution, inviting an LLM to copy one-way/costly forward as boilerplate. Now bracketed. Correctness 5 — latent false-green: .includes('reversible') also matches inside irreversible/irreversibility, which appear in anti-pattern prose, so a surface that dropped the real taxonomy entry would still pass. Now word-boundary matched. ADR-857 phase-6 ceiling — the first gsd-test run caught plan-phase.md 1216 bytes over its frozen 94519 ceiling (it had 49 bytes of headroom on next). The ceiling may only rise for privileged host machinery, and reversibility gating is optional-feature logic, so the wiring was slimmed to its minimum and the explanatory prose moved to the reference files the planner already loads. plan-phase.md is now 94400 bytes -- 119 under the ceiling and 70 bytes SMALLER than on next, so the host loop shrank while gaining the feature, which is what phase 6 ratchets toward. The tracer contract (tests/tracer-bullet.test.cjs) is unchanged. Lint — fixed an unnecessary non-null assertion in verify.cts and a CRLF-fragile bare \n regex in the new test (DEFECT.WINDOWS-CRLF-TEST- PORTABILITY, the #1658/#1668/#2206/#2449/#2450 class). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture must carry the common task elements The gated-one-way fixture built a checkpoint:decision task from the abbreviated skeleton in gsd-planner.md, which shows only the checkpoint-specific elements (<decision>/<context>/<resume-signal>). cmdVerifyPlanStructure requires <name> and <action> on EVERY task regardless of type, so the fixture failed validation for reasons that had nothing to do with reversibility: errors: ["Task missing <name> element", "Task 'unnamed' missing <action>"] Caught by gsd-test on 14d14a39 (2 failures, both this fixture). The canonical shape is in tests/verify.test.cjs:266 — a checkpoint task carries <name>/<files>/<action>/<verify> like any other. Fixture corrected to match. Verified behaviorally against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Not a product defect: the validator's every-task contract is intentional and pre-existing, and docs/reference/plan-md.md scopes its required-element list to type=auto/tracer only because those are the elements a planner must author, not because checkpoints are exempt from <name>. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1951): backfill changeset pr number to 2471 * fix(#1951): CodeQL incomplete-sanitization + prompt-injection scan collision Both CI failures were real defects in code this PR added, not false positives. CodeQL js/incomplete-sanitization (high), reversibility-tagging.test.cjs:46 — the namesRating helper built its regex with `rating.replace(/[-]/g, '\\-')`, which escapes the hyphen but not backslash, so the escape was incomplete. It was also unnecessary: `-` carries no special meaning outside a character class. Replaced with a complete metacharacter escape (backslash included). Word-boundary behavior verified unchanged across all three ratings — notably that "irreversible" prose still does not satisfy a "reversible" match, which is the false-green this helper exists to prevent. Prompt injection scan — the checkpoint fixture used the human-verification child element inside <verify>. That tag name is a fake-instruction-boundary pattern in scripts/prompt-injection-scan.sh, and the scan runs over changed files, so copying the shape from tests/verify.test.cjs (unflagged only because it is not in this diff) tripped the gate. Switched to the documented plain-prose <verify> form. The first attempt at that fix failed the same gate a second time: the comment explaining the collision quoted the offending tag literally. The comment now names it in prose instead — the scanner does not care whether a match is code or commentary, which is the whole point of the DEFECT.PROMPT-INJECTION-SCAN-COLLISION note in CLAUDE.md. Verified locally before push: scan reports 0 findings across 57 changed files, eslint clean, and both fixtures still validate as designed (gated one-way silent, ungated one-way warns, neither errors). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): record measured cost and halve gsd-tools spawns The Windows shard 1/3 job timeout was traced to the sharding layer, not to this PR's assertions — see #2472. Two contributing factors were this file's own, and are fixed here. 1. tests/test-timings.json had no entry for reversibility-tagging.test.cjs, so scripts/run-tests.cjs weighted it at the table's median fallback (~315ms) for LPT chunk packing. It actually measures 5595ms — an 18x under-weight. Recorded the measured value from the green gsd-test run (max across the node22/node24 lanes, per gen-test-timings.cjs's convention). Only this one entry: a full regen churns 634 entries of run-to-run drift, and the table is explicitly advisory and un-gated, so a 637-line diff does not belong in a feature PR. 2. Each verifyPlan() spawns gsd-tools, which dominates this file's cost. Spawns cut from 9 to 6 with no coverage lost: - the ungated-one-way warning and its stays-valid assertion now share one plan instead of building the same plan twice; - the reversible/costly never-flagged-as-ungated test was strictly subsumed by the additive suite, which already runs those two ratings ungated and asserts no /reversibilit/ warning at all — and the gate warning's text contains both "reversibility" and "one-way", so the broader assertion catches it. It only re-spawned gsd-tools twice to prove the same thing. Both are symptom fixes. The shard imbalance itself (19/11/10 minutes against a 20-minute cap, from a cost-blind round-robin partition that also reshuffles downstream files whenever one is inserted) is tracked in #2472. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture adopts the #2444 type-branched contract Surfaced by rebasing onto next, which gained #2444 (branch plan-structure validation on task type=checkpoint:*) while this PR was in review. cmdVerifyPlanStructure no longer applies one required-element set to every task. A checkpoint:decision now requires <name> + <resume-signal> + <decision> + <options>, and is exempt from the <action>/<verify>/<done>/ <files> set that auto and tracer tasks carry. The gated-one-way fixture predated that split and failed on the new requirement: errors: ["Task 'Task 0: Confirm the on-disk format' missing <options>"] Fixture rewritten to mirror the checkpoint:decision contract exactly — real <options> with two <option> children — rather than padding it with fields checkpoints no longer need. That also drops the plain-prose <verify> the earlier revision carried purely to dodge the prompt-injection scan; a checkpoint task has no <verify> requirement at all, so the workaround is moot. Verified against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
455ad49ae3 |
feat(#2296): config-gated provider escalation on quota-exceeded (#2458)
* test(#2296): failing-first coverage for provider escalation on quota-exceeded
Covers the provider-escalation ladder layered onto EXEC.CLASSIFY: back-compat
(no escalation block without --failure-class), cap boundaries at
min(max_escalations, list length) at limit-1/limit/limit+1, opt-in gating,
malformed/hostile provider_escalation config, the --failure-class CLI negative
matrix, config-key registration, and a fast-check budget-limit property.
Red until the resolver, CLI flag, and manifest key land.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#2296): config-gated provider escalation on quota-exceeded
The dynamic_routing tier ladder escalates within one provider, which does not
help when that provider is what ran out of quota. Add an opt-in provider ladder
layered on the existing EXEC.CLASSIFY seam.
- model-resolver: resolveProviderEscalation walks dynamic_routing.provider_escalation
capped at min(max_escalations, list length), reporting from/to/attempted/exhausted.
Invalid entries are dropped (ADR 227 shape validation). Stays a leaf module —
the quota-class policy decision is the caller's, per the CONTEXT.md contract.
- agent-command-router: export a frozen AGENT_FAILURE_CLASSES so the new CLI
validator cannot drift from the classifier that produces the values.
- resolve-execution: --failure-class flag; emits an escalation block ONLY when
passed, so the existing JSON contract is byte-identical for every caller.
- config-schema.manifest: register dynamic_routing.provider_escalation.
- execute-phase step 7.1: auto-escalate, honor Retry-After, fail loudly naming
every model tried once the ladder is spent.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2296): extract quota recovery to a reference fragment; regen goldens
The step 7.1a addition pushed gsd-core/workflows/execute-phase.md from 93390 to
95111 LF bytes, past the frozen ADR-857 Phase 6 ceiling (hard <93600, margin
<=93400) asserted by tests/fix-2285-claude-orchestration-wiring.test.cjs. The
base sat 10 bytes under the margin, so no inline wording would have fit.
That gate's own rationale is that optional-feature detail belongs in a fragment,
not the host loop. Moved BOTH the new provider-escalation branch and the
pre-existing manual recovery prompt into
gsd-core/references/execute-phase-quota-recovery.md, leaving step 7.1 as a
one-line pointer. execute-phase.md is now 92880 bytes — 510 SMALLER than base.
Also regenerates the fixtures that legitimately moved because three shipped
files changed (gsd-tools.cjs, config-schema.manifest.json, execute-phase.md):
golden-install-parity + install-tree for all 16 runtimes, INVENTORY.md +
INVENTORY-MANIFEST.json for the new reference, and the workflow size baseline.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#2351): make the C1 orphan-reaping test load-independent
tests/run-with-timeout.test.cjs C1 asserted the child heartbeat file exists
after a 1s group-kill window, but the child only wrote it on the first 100ms
setInterval tick. Nothing synchronized the two: on a loaded container the group
is SIGKILLed before that tick lands, the file never appears, and the assertion
fails for a reason unrelated to reaping. Observed failing on both linux-node22
and linux-node24.
The behavior actually under test is the FREEZE assertion (heartbeat stops
advancing => descendant was reaped, not orphaned). That is unaffected by
sampling once more at t=0.
Child now writes its first heartbeat synchronously at startup before arming the
interval, and the kill window widens 1s -> 3s to cover child boot under load.
Both remove the timing dependency; neither weakens what the test proves.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#2296): backfill pr:2458 in .changeset/rapid-jays-bark.md
* chore(#2296): regenerate fixtures after rebase onto #2402
The rebase conflicted on the generated golden-install-parity fixtures and
workflow-size-baseline.json because #2402 (
|
||
|
|
b6e6a22fce |
fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer (#2457)
* fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer Replays the in-flight bot branch fix/2402-response-language-orchestrator-coverage (seven commits, never pushed) onto current origin/next as a single squashed commit. The original work was substantial and correct; this commit preserves its full scope, trimmed where rebase conflicts + workflow size budgets required it. Three independent layers where response_language was being dropped are closed: Layer 1 — orchestrator-facing directives across workflows. Adds the strong "All user-facing output in this workflow MUST be presented in {response_language}; technical terms, code, paths, and subagent prompts stay in English" directive to ~40 workflows that previously either lacked it entirely (verify-work, new-project, new-milestone, quick, manager, and ~35 more) or carried only the weak subagent-prompt-only form (plan-phase, execute-phase). The directive covers narration between tool calls and banner output, not just the AskUserQuestion prompts. Layer 2 — UAT checkpoint renderer (src/uat.cts). buildCheckpoint now accepts an optional responseLanguage parameter and renders the frame strings ("CHECKPOINT: Verification Required", "Type `pass` or describe what's wrong.") in any of 9 languages (English/Spanish/French/German/Portuguese/Japanese/ Chinese/Korean/Italian) with an alias table covering ~30 input variants (en, es, español, ja, 日本語, etc.). cmdRenderCheckpoint reads config.response_language via loadConfig(cwd) and passes it through, so the byte-for-byte block verify-work.md reprints verbatim is already localized when written — preserving the anti-injection hygiene rule at verify-work.md (the model is forbidden to translate after the fact). CJK display width is computed by East Asian Width property ranges (W/F) so the right ║ border of the banner stays aligned for full-width characters. English fallback is byte-identical to the pre-fix behavior when response_language is unset or unrecognized. Layer 3 — literal English report templates in execute-phase. The top-of- workflow directive covers all template sites (templates are a structural source, not literal output). Inline render-language notes that previously sat at each template site were removed during the squash because they pushed execute-phase.md over its frozen pre-phase-6 byte ceiling (93600 — ADR-857 Phase 6 capstone). The single top directive covers the same surface with fewer bytes. Also extends src/docs.cts and src/init.cts to propagate response_language into the init JSON bundle of the additional workflows so the directive can read it. Tests added: - tests/uat.test.cjs: buildCheckpoint with unset/unrecognized language falls back to English default; recognized language swaps only the two frame strings while structural lines stay untouched; CJK display-width regression (independent recomputation of East Asian Width W/F ranges). - tests/workspace.test.cjs, tests/docs-update.test.cjs: response_language wiring through docs.cts/init.cts. References: #2402; reporter's three-layer triage + Layer-4 follow-up; the byte-for-byte anti-injection hygiene rule at verify-work.md (the reason Layer 2 must be renderer-side, not model-translated). This is a squash of the in-flight bot branch — seven commits representing the original implementation plus its subsequent fix/CJK-padding/test/ changeset/regen cycles, none of which were ever pushed or PR'd. The squash captures the final coherent state. * chore(#2402): backfill pr:2457 in .changeset/2402-response-language-orchestrator-coverage.md * chore(#2402): regen golden + size baseline after rebase against #2315 (PR #2451) Rebase conflicts were entirely in generated artifacts (golden-install-parity fixtures + workflow-size-baseline.json). After taking theirs during rebase, regenerated cleanly against the merged source tree. |
||
|
|
d16a66479a |
feat(#1950): broken-windows ledger — cross-phase defect register gating ship (#2441)
* feat(#1950): broken-windows ledger — cross-phase defect register gating ship Adds a new capability (#1950) that operationalizes GSD's no-defer discipline as a tracked, enforced artifact: accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths across phases, and /gsd-ship blocks while any entry is open. Implementation: - src/broken-windows.cts → gsd-core/bin/lib/broken-windows.cjs: typed IR + I/O entry points (parseLedger/renderLedger/appendWindow/markWaived/markFixed + cmdWindowsStatus/Append/Waive/MarkFixed). Frozen REASON enum for typed error assertions. Windows-safe atomic rename with retry on transient EPERM/EBUSY/EACCES. - gsd-tools.cjs: new subcommand (status | append | waive | fixed), wired via routeWindows + HOST_COMMAND_ROUTERS.windows. - capabilities/broken-windows/capability.json: one ship:pre gate with artifact-frontmatter-equals predicate on WINDOWS.md open_count == 0. activationKey windows.enabled (default true) + sibling windows.enforce (default true, separate so tracking can precede enforcement). - gsd-core/workflows/ship.md: capId==broken-windows branch in preflight, sibling to security — reads gsd_run windows status --raw, fails closed on open_count > 0 or unreadable ledger. - agents/gsd-executor.md: extends the existing ## Known Stubs instruction to also append to WINDOWS.md via gsd_run windows append (best-effort, never blocks execution). - agents/gsd-verifier.md: new Step 8b — record unmet truths + human-verify items in WINDOWS.md. - gsd-core/workflows/progress.md: surfaces open + waived counts. - docs/COMMANDS.md + CONTEXT.md glossary entry + docs/INVENTORY.md: document the gate, waiver mechanism, and new module. - tests/broken-windows.test.cjs: pure + CLI behavioral coverage + fast-check roundtrip property; fail-closed on malformed ledger; security boundary on path traversal in --file. Backward-compatible: a project with no .planning/WINDOWS.md reports open_count: 0 and ships cleanly. Disable enforcement per-project with gsd config-set windows.enforce false (tracking continues, gate stays open). * chore(#1950): ratchet size baselines, defer verifier integration - Workflow size baseline: ship.md 25575→27928, progress.md 31789→32632 (broken-windows preflight branch + open-windows surface). - Agent size baseline: gsd-executor.md 46644→47951 (Known Stubs → also appends to WINDOWS.md). gsd-verifier.md unchanged. - LARGE_CAP (49152) preempted the planned verifier integration (gsd-verifier.md was at 49140 pre-PR — 12 bytes of headroom, not the documented 'real headroom'). Verifier integration deferred to a follow-up PR that extracts the VERIFICATION.md template (lines 739-859) to gsd-core/references/ — a pre-existing cap-tightness defect this PR exposed but does not expand scope to fix. Verifier integration is not in the issue's acceptance criteria (executor writes is; unmet-truths recording was an enhancement, not a gate). * fix(#1950): gate default-off, rename to workflow.windows_enforce, regen goldens Test-failure-driven fixes after first gsd-test run on db8733c8f failed 44 cases (pre-existing structural tests encoded 'ship:pre has 1 gate' / 'all caps off → empty hooks'): - capability manifest: rename windows.enabled+windows.enforce (default true) → single federated key workflow.windows_enforce (default FALSE, opt-in). Matches security's workflow.security_enforce convention and makes the adr857 all-caps-off test pass without modification (the test's buildAllFalseConfig handles workflow.* out of the box). Default-OFF keeps the gate out of the registry's default ship:pre resolution so existing loop-hooks-ship-pre-e2e structural assertions (exactly 1 gate, capId 'security') stay valid; users opt in via gsd config-set workflow.windows_enforce true. - drop activationKey (security doesn't have one either; workflow.* key doubles as the activation toggle). - regenerate docs/reference/capability-matrix.md to include broken-windows (capability-matrix-sync test). - regenerate tests/fixtures/golden-install-parity/*.json (18 runtimes) — installer now emits the new capability + lib file. - update CONTEXT.md, docs/COMMANDS.md, docs/FEATURES.md, ship.md, agents/gsd-executor.md to use the new key name and /gsd:colon slash syntax (slash-command-namespace test). - restore accidentally-regressed /gsd:capture in progress.md. Tracking-only by default; enforcement is opt-in. Acceptance criterion '/gsd-ship fails while any ledger entry is open' is met when workflow.windows_enforce=true (test fixture enables it). * test(#1950): update ship:pre structural invariants for 2-gate registry - loop-hooks-ship-pre-e2e: the registry now declares 2 gates at ship:pre (security + broken-windows), regardless of activation. Activation tests above still pin security-only or empty behavior via fixtures; these structural tests pin the REGISTRY shape, which has 2 gates as of #1950. - workflow-size-baseline: ship.md 27928→27945 (workflow.windows_enforce rename added 17 bytes). * fix(#1950): review H1+H2+M1+M2+M3 — fence-injection, EACCES fail-closed, cleanup, strict line, stryker Adversarial isolated review (Step 6.3) found 2 HIGH findings that block the PR and 3 mediums. All addressed: H1 (HIGH): description containing the markdown 3-backtick fence would terminate the ledger's JSON code block early inside JSON.stringify output (JSON doesn't escape backticks), corrupting the file and bricking the next parse. Fix: use a 4-backtick fence (json ... ) which JSON.stringify cannot produce on its own, AND validate that no entry text field contains a 4-backtick run (reject at append time with new WINDOWS_INVALID_TEXT reason code). Locked by a regression test. H2 (HIGH): readLedgerOrNull swallowed ALL fs errors as 'no ledger', silently returning open_count:0 on EACCES/EPERM/EIO. The ship gate would then pass on an unreadable ledger — the precise vector the workflow doc claims is impossible. Fix: only ENOENT returns null; every other fs error propagates as WINDOWS_LEDGER_MALFORMED so the gate blocks and the operator sees a real diagnostic. Locked by a regression test that chmod 000s a ledger with open_count=1 and asserts the result is never a false-green 0. M1: writeLedgerAtomic left an orphaned .tmp file on rename failure. Wrapped renameWithRetry in try/catch with best-effort unlink. M2: validateLine silently coerced 'abc' → NaN → null, hiding type drift. Removed the line === 0 special case (was undocumented) and made the error message match the strict check. Now any non-positive- integer line value throws, including strings. M3: tests/broken-windows.test.cjs (with its fast-check property test) was not in stryker.config.mjs DEFAULT_TEST_CMD — Stryker would mutate src/broken-windows.cts but no test would catch the mutations, producing false surviving-mutant scores. Added to the list. L1 (dead throw e after error()), L7 (line boundary tests, H1/H2 regression tests, 4-backtick CLI test) also addressed. * docs(#1950): inline concurrency + busy-wait notes (review L2+L3) * fix(#1950): regen goldens against latest gsd-tools; correct --line 0 boundary test gsd-test v4 caught two issues: - goldens I regenerated earlier (commit 526682084) predated the L1 routeWindows catch-block cleanup (commit dd844d565). Regenerated via 'npm run gen:golden' against current HEAD so the install parity hash for gsd-tools.cjs matches. - 'append --line boundary' test expected --line 0 to succeed with null entry.line, but the M2 fix correctly rejects 0 (lines are 1-indexed; 0 is not a valid source line). Updated the boundary test to assert --line 0 fails alongside -1 and 'abc'. * chore(#1950): regen goldens after rebase onto next * chore(#1950): quick.md baseline 50699→50993 (correct resolution from next rebase) * chore(changeset): backfill pr:2441 in .changeset/broken-windows-ledger.md * fix(#1950): renderTable escapes backslash before pipe (CodeQL incomplete-sanitization) CodeQL flagged the markdown-table cell escaper: String(s ?? '').replace(/\|/g, '\\|') — it escapes pipe but not backslash first. A description containing '\|' would render as '\\|' which markdown parses as 'literal backslash' + 'cell separator', splitting the column. Fix: escape backslash FIRST (each \ → \\), then pipe (each | → \|). Now a description with '\|' renders as '\\\\|' (literal '\\' + escaped pipe), which markdown renders as a single '\|' inside the cell. The JSON code block (the parse source-of-truth) was already correctly escaped via JSON.stringify; only the display-only table was affected. Locked by a regression test that: 1. Verifies the JSON block reparses with the description intact. 2. Walks the rendered table row counting unescaped pipes — must be exactly 11 (the row separators for 10 cells), proving no in-cell pipe added a split. |
||
|
|
d04e287fa9 |
fix(#2365): stop api-coverage detector false-positiving non-API phases (#2397)
* fix(#2365): stop the api-coverage detector false-positiving non-API phases detectApiIntegration fired on any integration verb co-occurring anywhere on a line with any API noun, treated / as a word boundary (so a first-party Next.js src/app/api/... route path matched the noun "api"), and read any capitalized word before API/SDK/REST/GraphQL as a service name behind a fixed stopword denylist (so threat-model prose like "Resolver-only API" fired). Because the verify:pre seal gate is BLOCKING, a phase touching no external API could not reach UAT without fabricating a coverage matrix. The compound rule now requires the verb and noun to share one clause (sentence punctuation and table-cell walls end a clause) within a bounded word gap. Non-prose spans are excluded before matching: fenced code (already), inline code spans (new stripInlineCode in the markdown-sectionizer seam), and path-shaped tokens. The <Service> API surface rule requires proper-noun position — a clause-initial capitalized word is ordinary English and needs dependency evidence (URL / package reference) on the same line — and rejects compound modifiers ("Resolver-only", lowercase after the hyphen). A phase that integrates no external API now has a first-class, reasoned way to say so: a COVERAGE.md containing "No external API integration: <reason>" satisfies the gate (declaration + rows is contradictory and blocks). The true-positive path is pinned by regression tests: every default-vocabulary positive still fires, including the widest word-gap pairing and the surface-rule-only shape. Fixes #2365 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2365): tighten api-coverage detector per Codex review (round 2) Applies the Codex review findings on the initial #2365 fix: - S-1: a COVERAGE.md "no external API integration" declaration is the human override for a fallible detector, so it must PASS even when detection still fires — but the contradiction is now SURFACED in the gate output (overridden signal count + terms) instead of passing silently. - S-2: verb/noun pairing is now a term-group nearest-pair merge walk over precomputed word ordinals (computeWordStarts / minWordGap), not a match×match cross product — a hostile line repeating one pair thousands of times stays linear instead of going quadratic. - FN-4: package-shaped inline-code spans (`stripe-sdk`, `@stripe/stripe-js`) are kept as noun/dependency evidence rather than being fully masked, so a genuine dependency reference inside code ticks still corroborates. - C-1: the <Service> API surface rule now scans every candidate in every clause; a rejected first candidate no longer shadows a later genuine service. - Cross-clause binding: a verb may bind a noun in the immediately following clause only when its own clause names a service object, within a tight gap — admits "Integrate Stripe, exposing its endpoints …" without re-admitting the unrelated-clauses false-positive class. - Internal-descriptor negative evidence ("internal Payments API", "the internal endpoint") never pairs; URL/scheme matching generalized beyond http(s). All 5 acceptance criteria still hold: the three reported false positives are clean and "integrate the Stripe API" still fires. Built .cjs committed alongside the .cts. tsc + eslint (incl. no-adhoc-markdown-parsing) + lint:regression-names clean; affected suites 256/256 green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#2365): retune api-coverage detector fail-closed per Codex review (round 3) Codex's second-round review found the round-2 tightening had over-corrected into FAIL-OPEN false negatives — realistic external-API prose that the BLOCKING seal gate silently let through (the catastrophic class, since a missed API surface is worse than a dismissable false positive). Retuned the detector to be explicitly fail-closed: lean toward detecting, and let the one-line COVERAGE.md "no external API integration" declaration dismiss the residual false positives. Fail-open false negatives fixed (all now detect): - F1 clause-initial `<Service> API` with a plain follower ("Stripe API for payment processing") — dropped the follower-allowlist / corroboration gate on clause-initial surfaces; a service that is not a stopword, descriptor, or compound modifier is a real name from any clause position. - F2 scheme-less external host ("api.stripe.com/v1") — a dotted host with an alphabetic final label now contributes its API nouns; a first-party route path (no dotted host) still does not. - F3 vendor's first-party SDK ("Integrate Shopify's first-party SDK") — the compound path no longer filters nouns on "internal"/"first-party" (Codex: the qualifier can describe the vendor's own API, not the consuming project's). - F4 long single integration clause — removed the word-gap cap entirely: it could not separate a 21-word genuine clause from an 18-word internal one, so the clause boundary is now the whole relationship test. - F5 lowercase cross-clause service — cross-clause binding no longer requires a capitalized "service object". New false positives fixed (all now clean): - F6 a URL token that swallowed a trailing clause comma, merging two clauses — trailing clause punctuation is kept literal so the split survives. - F7 a capitalized internal component authorizing cross-clause binding — the new gate requires a dependent elaboration, not a new coordinate clause opened by a conjunction ("…, then document…"). - F8 a protocol name read as a service ("REST API", "GraphQL API") — protocol and locality descriptors are rejected in the `<Service>` position. - Finding 9: the inline-code-span scanner was O(n^2) on pathological backtick runs; rewritten to linear via a per-length run cursor (2 MB: 4.15 s -> ~6 ms), semantics preserved (148 sectionizer tests unchanged). Net simplification: the fail-closed model removed the round-2 minWordGap / groupByTerm / follower / corroboration machinery (350 insertions vs 445 deletions across the touched files). Under fail-closed, three round-2 negative tests now correctly detect (integration verb + "internal"-qualified noun, and the distant-same-clause case); none were trek-e acceptance FPs. Verified: 1491/1491 unit tests pass; tsc + eslint (incl. no-adhoc-markdown- parsing) + lint:regression-names clean; all 8 review findings reproduced as regression tests, both directions. Built .cjs committed alongside the .cts. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#2365): resolve round-3 Codex review findings (fail-closed, round 4) Codex's round-3 adversarial review found the fail-closed retune had introduced new holes in both directions. Resolved: Fail-open false negatives (now detect): - External host addressing a PATH ("graph.microsoft.com/v1.0/me") is itself an integration surface and contributes an endpoint noun even when the host names no vocabulary word. A bare domain link with no path ("https://example.com") stays a non-signal, so "Integrate … from example.com, document …" is still clean. - Locality qualification ("internal", "private") no longer leaks across a sentence or clause boundary: only plain spaces may separate the descriptor from the service, so "The cache is private. Stripe API …" now detects. - Cross-clause binding: the fragile head-word cap (which could not tell a genuine "Connect … to Stripe payments, exposing its endpoints" from an unrelated "Integrate … from URL, document …" — both 4 words after the verb) is replaced by a participial-continuation rule: a verb binds a noun in the next clause only when that clause begins with an "-ing" elaboration. This fixes the 4-word-head false negative AND the false positive below at once. False positives (now clean): - Cross-clause no longer binds a finite continuation regardless of separator: "Wire the settings form. Document endpoint props." / "…; document …" / "…, document …" are separate actions, not elaborations. Perf (quadratic → linear): - The trailing-punctuation peel is a backward char scan instead of an unanchored `[…]+$` regex (16k chars: 156 ms → ~1 ms). - SERVICE_SURFACE_API_RE bounds the service-name length {1,40} so a hostile "A-A-…-x" run cannot drive O(n^2) backtracking (16k: 385 ms → ~3 ms). Consumer fail-open (blocking gate): - readPhaseScope now distinguishes "no plans" from a plan that EXISTS but is unreadable. On a read error the gate BLOCKS ("could not read the phase scope …") instead of silently certifying no-integration from partial scope — an unreadable plan could be the one describing the integration. Documented fail-closed tradeoffs, now pinned with tests so they are not "fixed" back into a fail-open: a clause-initial capitalized common word before "API" ("Payment API", "Search API") reads as a service name; a long clause pairs a verb with a distant noun; and a CommonMark inline code span that wraps a newline is matched within-line only. Codex judged these acceptable because the COVERAGE.md declaration is a cheap override. One documented limitation remains out of scope: "Integrate Stripe, and authenticate requests with its API" (a coordinate finite clause whose noun refers back by pronoun) needs coreference resolution, beyond a lexical detector. Verified: 379/379 affected + command-router tests pass (+14 new regression tests covering every round-3 finding, both directions); tsc + eslint (no-adhoc-markdown-parsing) + lint:regression-names clean. Built .cjs committed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#2365): simplify to robust core — remove whack-a-mole heuristics (round 5) Round-4 review confirmed the detector's two most complex features generate findings in both directions no matter how they are tuned, because they need a vendor dictionary + coreference the issue rules out in principle. Per the operator's "ship the robust core" decision, both are removed and their gaps are documented rather than chased further: - Cross-clause binding DELETED (allowsCrossClause / participle rule). It caused a fail-open on finite continuations ("Integrate Stripe; use its OAuth endpoints" — missed) and a false positive on "-ing"-SPELLED nouns ("…, billing endpoint terminology…" — wrongly fired). Detection is now same-clause only. - URL-path-as-evidence REVERTED. Treating every path-bearing URL as an endpoint fired on ordinary asset/link URLs ("…/theme.css", "…?next=/x", a docs/repo link) and recreated routine UI-phase false positives. An external URL is evidence only when it NAMES an API vocabulary word ("api.stripe.com/v1"). Two fail-open cases are now DOCUMENTED limitations, pinned by tests so a future maintainer does not re-add the heuristics that caused the false positives above: a service named only in a clause separate from its API noun, and a bare external host that names no vocabulary word. Both are cheaply covered by the COVERAGE.md declaration and rare in real phase prose ("integrate the X API"). Also fixed from the round-4 review: - Qualification now survives markdown emphasis ("The **internal** Payments API" stays clean) while still not crossing a sentence/clause boundary. - readPhaseScope fail-closes on a REAL read failure (EACCES/EIO) enumerating the phase directory or reading the roadmap fallback — not only per-plan-file failures; a missing directory/section remains a legitimate no-op. The declaration-override path surfaces scope_read_error so an incomplete-scope override stays visible. - SERVICE_SURFACE_API_RE length-bound comment no longer overclaims. Net: the detector is same-clause verb+noun + `<Service> API` surface, with path/code/inline masking and a fail-closed posture. All five acceptance criteria hold. 1573/1573 unit tests pass; tsc + eslint (no-adhoc-markdown-parsing) + lint:regression-names clean. Built .cjs committed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#2365): close roadmap-fallback fail-open + stale JSDoc (round-5 review) The round-5 sanity review confirmed the detector simplification is sound (all acceptance positives fire, all required negatives clean) and flagged one real blocker plus a nit: - Blocker: readPhaseScope's roadmap fallback could still silently pass an UNREADABLE roadmap. getRoadmapPhaseWithFallback gated on fs.existsSync(), which returns false on EACCES/EIO too — so an unreadable ROADMAP.md read as "absent", no exception reached isRealReadFailure, and the blocking gate certified empty scope. Fixed at the source: read the roadmap directly and honor the function's OWN documented contract — null only on ENOENT (genuinely absent), otherwise throw. Both existing callers already wrap it in try/catch expecting that throw, and readPhaseScope now fail-closes (blocks) via its roadmap catch. Verified by a new e2e test (unreadable roadmap fallback → block). - Nit: the detectApiIntegration JSDoc still described the removed cross-clause participial binding and "every external hostname counts" — corrected to the actual same-clause-only behavior and the names-a-vocab-word URL rule. Verified: full unit suite green; tsc + eslint + lint:regression-names clean. Built .cjs committed (roadmap.cjs is gitignored/rebuilt, per repo convention). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#2365): backfill changeset PR number (#2397) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#2365): sync generated capability-registry + recapture install goldens CI surfaced two generated-artifact staleness issues (all failing test shards + lint-tests traced to these, not to a logic defect): - gsd-core/bin/lib/capability-registry.cjs was stale: the initial fix edited the ai-integration `api-coverage-plan-pre.md` fragment (added the "No external API integration" declaration section) but did not regenerate the registry, which embeds an inline copy of that fragment. Regenerated via `gen-capability-registry.cjs --write` — the diff is exactly the fragment text sync. Fixes `lint:generated-sync` and the "committed registry is in sync" + "registry integration" tests. - The 18 golden-install-parity fixtures were stale by exactly one hash line each — `gsd-core/references/api-coverage.md`, which this PR edits and which is a hashed installed artifact. Recaptured with `UPDATE_GOLDEN=1`; the diff is that single hash per runtime and nothing else. Fixes the `golden parity — *` tests. No source or behavior change — generated artifacts only. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#2365): flip representative-corpus manifest to assert the fixed behavior The #2371 representative corpus (merged into next after this branch was cut) is a known-bug tripwire: it asserts each fixture's currentBuggyOutput so the test fails loudly the moment #2365 is fixed, at which point — per its own contract in representative-corpus.test.cjs — the fixer removes currentBuggyOutput so the assertion checks expectedDetected instead. This is that moment. Removed currentBuggyOutput from the three detector fixtures (nextjs-route-path, unrelated-verb-noun, threat-model-prose); the corpus now asserts detected:false, which the fail-closed same-clause detector satisfies. Notes updated to describe the fix rather than the bug. The #2366 matrix corpus is left untouched — that tripwire belongs to its own PR (#2374). Corpus test: 7/7 pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#2365): skip chmod-000 fail-closed e2e tests on Windows The three fail-closed gate tests induce an unreadable plan / directory / roadmap with chmod 000, but Windows does not enforce POSIX mode bits — readFileSync still succeeds, so the gate never reaches the read-error path and the assertion fails on the windows-latest CI leg. The fail-closed LOGIC is platform- independent (readError → block) and is fully exercised on the macOS/Linux legs; only the method of inducing EACCES is POSIX-specific. Guard the three tests to skip on win32 as well as root, mirroring golden-install-parity's win32 skip. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#2365): address trek-e review — glossary, clock-seam, IO injection, bounds Review response to PR #2397 (trek-e, CHANGES_REQUESTED). Fix logic unchanged; this closes the test/process-hygiene findings. Major: - CONTEXT.md "Markdown Sectionizer" glossary now lists the two exports this fix relies on, `stripInlineCode` and `scanInlineCodeSpans` (glossary is a PR gate). - Replaced the banned wall-clock assertion in the "hostile repeated-term line" test (Clock Seams rule — no elapsed-time asserts) with a deterministic signal-count assertion, which also directly verifies the term-dedup that keeps pairing linear (one signal for a 10k-pair line, not thousands). - Rewrote the three fail-closed read-failure tests: instead of chmod 0o000 (a no-op under root / on Windows, the pattern the repo's IO-failure convention avoids) they now exercise the newly-exported `readPhaseScope` in-process and inject the failure by monkeypatching fs.readFileSync/readdirSync to throw, restoring in finally. Deterministic and platform-independent (no skip needed), and they add the ENOENT-is-absence case that the chmod tests couldn't express. Minor: - Added limit / limit+1 boundary tests for SERVICE_SURFACE_API_RE's {1,40} service-name bound, QUALIFIER_LOOKBACK's 24-char window, and REASON_MAX_LEN (200) on the declaration reason. - Added a fast-check property that fuzzes the tokenizer / clause splitter / masking (scanLineTokens, splitClauses, collectTermMatches) with adversarial tokens (slashes, backticks, URLs, clause punctuation) and asserts the detector is total (never throws), shape-stable, holds detected <=> signals, and is deterministic. readPhaseScope is exported for the in-process tests. Verified: 125 detector + 19 gate tests pass; tsc + eslint + generated-sync (glossary/registry) + lint-regression-test-names + lint-test-file-count clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1720aacf0c |
feat(#1949): <precondition> task element — Design by Contract (#2422)
* test(#1949): add failing-first tests for <precondition> element Red phase for issue #1949 (Design by Contract: <precondition> element asserted before task execution). Tests assert: - docs/reference/plan-md.md documents the new <precondition> element - agents/gsd-planner.md @-references planner-preconditions.md and stays under the 49152-char cap (progressive-disclosure requirement) - gsd-core/references/planner-preconditions.md exists and documents the three emission cases mandated by the issue (user_setup / prior-phase artifact / env-var) and the contract triad mapping - agents/gsd-executor.md asserts <precondition> before task execution and routes unmet preconditions through existing checkpoint machinery - cmdVerifyPlanStructure (behavioral via runGsdTools) accepts plans both with and without <precondition> — the additive-validation guarantee - Parity assertion: plan-md.md and planner-preconditions.md agree on the canonical tag spelling (DEFECT.GENERATIVE-FIX-DIVERGENCE guard) Most prose-contract assertions are Red until the implementation lands. The behavioral validator assertions pass immediately (regression guards proving the validator already accepts unknown optional tags). * feat(#1949): <precondition> task element — Design by Contract Add an optional <precondition> element to <task> in PLAN.md (issue #1949, The Pragmatic Programmer Topic 23). The front-of-task side of the plan contract — preconditions (before) ↔ postconditions (<verify>/<done>/ <acceptance_criteria>, after) ↔ invariants (must_haves.truths, across the whole plan). Together with the tracer-bullet proposal (#1945), this closes both ends of the 'outrunning your headlights' failure mode for an autonomous AI executor. Acceptance criteria met: - <precondition> is an optional element on <task>; plans that omit it validate unchanged (cmdVerifyPlanStructure checks for presence of required tags, does not reject unknown optional tags). - gsd-executor evaluates the precondition before any other task work. Unmet halts execution with a checkpoint:human-verify and no partial commit; met or absent produces no visible change to execution flow. Unmet is never auto-approved under AUTO_CFG=true — a missing prerequisite is a fact the executor cannot establish on its own. - gsd-planner emits <precondition> in exactly the three cases the issue mandates: user_setup consumption, prior-phase artifact dependency, and env-var/runtime-config dependency. - Tests cover met, unmet, and absent preconditions plus the additive- validator guarantee. Files: - gsd-core/references/planner-preconditions.md (NEW): full emission rules, the three cases with worked examples, format guidance, anti-patterns, the contract triad mapping, and the executor assertion contract. Progressive disclosure. - agents/gsd-planner.md: slim <precondition> note in Task Anatomy with @-reference to the new file. To stay under the 49152-char agent-file cap (27-char headroom before this change), the inline <comment_text_discipline> and <region_scoped_negative_gate> summaries are compressed to one-line pointers — their full rules already live in planner-antipatterns.md, so no content is lost. - agents/gsd-executor.md: new step 0 'Precondition check' in the execute_tasks loop, before the type dispatch, routing unmet through checkpoint_return_format. - docs/reference/plan-md.md: new Preconditions section in the schema reference, with the canonical example and the three emission cases. - CONTEXT.md: Precondition glossary entry as a sibling of Tracer Bullet. - docs/INVENTORY.md + INVENTORY-MANIFEST.json: row for the new references/planner-preconditions.md (regen via gen-inventory-manifest). - tests/precondition-element.test.cjs: failing-first tests covering schema docs, planner emission contract, executor assertion contract, reference-file presence + the three cases, behavioral additive- validator guarantee, and a parity assertion (DEFECT.GENERATIVE-FIX- DIVERGENCE guard). - .changeset/quick-hawks-bark.md: Added fragment. Companion to #1945 (tracer bullets). * chore(#1949): regen agent-size baseline + install-tree goldens Documented baseline regenerations required by the feat(#1949) prose changes (RULESET.AGENT_SIZE_BUDGET + golden-install-parity): - npm run size:baseline — locks in the new gsd-executor.md size (+1050 bytes: the precondition-check step 0 block). gsd-planner.md is net smaller (-142 bytes: compressed two inline summary blocks whose full rules already lived in planner-antipatterns.md to make room for the slim <precondition> pointer). No hard-cap breach. - npm run gen:golden — pick up the new references/planner-preconditions.md + the two changed agent files across all 18 runtime install trees. Both regens are CI-mandated after intentional agent/reference changes; see CLAUDE.md 'RULESET.AGENT_SIZE_BUDGET' and the comments in tests/golden-install-parity.test.cjs. * fix(#1949): bound <precondition> checks to read-only (security review) Apply the security-review finding (LOW, isolated /security-review subagent): the executor's 'run the cheapest check' phrasing for a plan-author-controlled prose line was broader than ideal — a hostile plan author could craft a <precondition> whose 'cheapest check' is side-effecting (curl to an attacker host under the guise of verification, rm -rf before checking, secret emission). The risk is inherited from GSD's existing plan-trust model (<verify>, <action>, <done> already direct the executor to run arbitrary shell), so <precondition> does not materially expand it. But the new prose actively directs execution ('run the check') rather than passively consuming the element, so the bound is worth making explicit. Tightened across all four surfaces that describe the check shape: - agents/gsd-executor.md step 0: 'Verify with read-only checks only — file existence, env var presence (no value output), idempotent GET /health-style pings. Do NOT run commands with side effects (writes, network POSTs, secret emission) as the check; if a side-effecting check seems required, halt and surface via checkpoint instead.' - gsd-core/references/planner-preconditions.md Format section: same bound, plus the halt-and-surface escape hatch. - docs/reference/plan-md.md Preconditions section: mirrored. - CONTEXT.md Precondition glossary entry: mirrored. Regenerated agent-size baseline (executor grew 46186 -> 46440; still under the 49152 cap) and install-tree goldens. * chore(#1949): backfill changeset pr number 2422 Per CONTRIBUTING.md changeset workflow + feature-builder directive Step 8.7: backfill the placeholder pr:0 with the real PR number immediately after gh pr create returns. Avoids the fail_invalid_fragment gate. * fix(#1949): cite [#1949] on allow-test-rule exemption (ADR-456) CI's lint:ci runs lint-allow-test-rule-refs which per ADR-456 requires every // allow-test-rule: exemption on a NEW test file to carry an issue reference (#NNN or URL). My earlier push omitted it. Local 'npm run lint' (eslint) does NOT run this check — only 'npm run lint:ci' does. CLAUDE.md explicitly warns: 'lint:ci ≠ lint — CI runs lint:ci; a local pass is not the gate.' I should have run lint:ci before pushing; correcting now. Pattern matches the companion feature's test file: tests/tracer-bullet.test.cjs:1 // allow-test-rule: source-text-is-the-product [#1945] |
||
|
|
8d2f8bcb23 |
fix(#2388): gate shared requirement completion on sibling plans, revert on gaps (#2424)
* fix(#2388): gate shared-ID requirement marking and revert on gaps_found Adds requirements.ready-ids (execute-plan.md's update_requirements step) so a requirement ID declared by multiple plans in a phase only marks Complete once every declaring plan has produced a SUMMARY.md, and requirements.revert-phase (execute-phase.md's gaps_found branch) so a gaps_found verdict reverts the phase's own prematurely-Complete IDs before the gap report renders. Single-plan IDs still mark immediately. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2388): regenerate fixtures + lint gate-prep * fix(#2388): repair failing tests after gate verification * chore(#2388): add changeset (#2424) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
dd5a2211c9 |
enhance(#1964): semantic knowledge-base recall via MemPalace (keyword fallback) (#2416)
* test(#1964): add failing-first semantic-recall contract tests Epic #1957 Phase 3C (final). Source-text-is-the-product contract tests: semantic recall via MemPalace (top-k meaning-similar prior resolutions, catches same-root-cause/different-wording cases), indexing resolved sessions at archive, graceful degradation to keyword matching when MemPalace is absent, knowledge-base.md stays the durable plain-text source of truth, agent Phase 0 / Matching Logic is semantic-first (the stale 'keyword overlap, not semantic similarity' claim must go), and no new embedding/vector infra (reuse MemPalace). Failing-first: reference, the Matching Logic reframe, the Phase 0 consolidation, and the archive indexing step do not yet exist. * feat(#1964): semantic knowledge-base recall via MemPalace (keyword fallback) Epic #1957 Phase 3C (FINAL). Replaces keyword-overlap matching with semantic recall: at Phase 0 the debugger queries MemPalace with the current symptoms and surfaces the top-k meaning-similar prior resolutions, catching the same-root-cause/different-wording cases keyword overlap missed (the self-noted 'keyword overlap, not semantic similarity' limitation). Resolved sessions are indexed into MemPalace at archive (symptoms + root_cause(s) + fix + recurrence guard). knowledge-base.md remains the durable plain-text source of truth; when MemPalace is absent the debugger falls back to keyword-overlap matching (logged, never a silent skip). No new embedding/vector infrastructure — MemPalace is reused. Size-neutral agent edits: the Matching Logic section reframed (keyword-only -> semantic-first + keyword-fallback + @-include); Phase 0's three keyword bullets consolidated into one semantic-first bullet; one MemPalace-indexing step added at archive. Agent at 57222 B (122 B headroom — final phase). Full rules in gsd-core/references/debugger-semantic-recall.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1964): address orthogonal review (invocation mechanism, index Resolution-not-symptoms + redaction, fallback detail) - HIGH: the 'query MemPalace' instruction was WHAT-level only; the agent has no MCP tools. Added an Invocation section naming the Bash CLI (mempalace search --wing <wing>) + MCP-when-registered + wing resolution (config.mempalace.wing -> project_code -> project dir), matching every other MemPalace integration. Without this the feature silently degraded to keyword matching even when MemPalace was present. - MEDIUM (security x2): index the agent-authored Resolution summary (root_cause + fix + recurrence_guard), NOT raw user-supplied Symptoms — excludes attacker-controlled prose from the cross-session index AND reduces secret/PII leakage. Redact secret-shaped values before indexing. Stated the write order (KB append + commit MUST succeed before indexing). - LOW: restored 'identifiers' + 'case-insensitive' to the keyword fallback; added a test asserting the fallback mechanics survived the Phase 0 consolidation (Error patterns field, 2+ token overlap, identifiers, case-insensitive). * chore(#1964): ratchet agent-size baseline downward (leaner archive bullet shrank gsd-debugger.md 57222->57197) * chore(#1964): backfill changeset pr number (PR #2416) |
||
|
|
c67f301867 |
feat(#1963): emit blameless-postmortem Prevention block at resolution (#2410)
* test(#1963): add failing-first prevention/postmortem contract tests Epic #1957 Phase 3B. Source-text-is-the-product contract tests: blameless 5-Whys that BRANCHES per Phase 2A RCA (not a single-cause chain; treats agent error as 'why was that possible?'), the 'why wasn't this caught?' question, the recurrence-guard taxonomy (regression test / assertion / lint rule / KB pattern), the KB-entry why_not_caught + recurrence_guard fields with backward compat, the session-manager prevention summary line, and the Zawinski scope-boundary (a block, not a subsystem). Failing-first: reference, archive_session edit, KB schema extension, and session-manager summary do not yet exist. * feat(#1963): emit blameless-postmortem Prevention block at resolution Epic #1957 Phase 3B. At archive_session the debugger now produces a Prevention block with three blame-free components: a branching 5-Whys causal chain (branches per Phase 2A RCA, not a single chain; 'agent error' prompts 'why was that possible?', never blame), a 'why wasn't this caught?' answer naming the missed gate (test/typecheck/lint/review/verify), and a concrete recurrence guard (regression test / assertion / lint rule / KB pattern). The knowledge-base entry gains two structured fields (why_not_caught + recurrence_guard) so future Phase-0 recall surfaces the prior prevention, not just the prior fix. Additive: old entries without the fields still load. The session-manager compact summary surfaces a one-line prevention summary. Full rules extracted to gsd-core/references/debugger-prevention.md (slim archive_session step + 2 KB fields kept in the agent). INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1963): address orthogonal review (CRITICAL append-template drift + Phase-0 consumption + parity test) - CRITICAL: the archive_session KB append template omitted Why not caught + Recurrence guard (only the Entry Format had them) — the feature's core deliverable silently did not happen. Added both fields to the append template the agent actually follows (nearest-instruction wins). - HIGH: Phase 0 (KB read) only surfaced root_cause + fix; the new fields were dead data. Extended the Phase 0 Evidence line to consume why_not_caught + recurrence_guard when present (absent on old entries — backward compat holds). - MEDIUM: added a cross-section parity test (every Entry-Format field must also appear in the append template — the guard that would have caught the Critical) + a Phase-0-consumption assertion. - MEDIUM: the 'branches per Phase 2A' claim is now wired — reuses reasoning_checkpoint.candidate_causes across the four categories. - MEDIUM: recurrence-guard taxonomy gains type refinement + config-default change; LOW: added 'build' gate to both surfaces for parity. - NIT: compact-summary fallback shape ('no gate existed'); verify the guard artifact exists before recording it. * test(#1963): anchor Phase-0 consumption test on the specific heading The regex /Phase 0[\s\S]{0,1200}/ matched the first 'Phase 0' in the file (in knowledge_base_protocol prose), not the Phase 0 block in investigation_loop. Anchor on '**Phase 0: Check knowledge base**' and widen to 1500 chars. * chore(#1963): backfill changeset pr number (PR #2410) |
||
|
|
36a311c5bb |
enhance(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries) (#2409)
* test(#1962): add failing-first repro-hardening contract tests Epic #1957 Phase 3A. Source-text-is-the-product contract tests: PBT shrinking (fast-check/Hypothesis, minimized seed, manual-minimization degradation), the four oracle types (specified/derived/metamorphic/implicit with implicit flagged weakest), boundary neighbors (off-by-one/min-max/empty-singleton tied to the equivalence class), oracle_type in DEBUG Resolution, and the Phase 1A tie-in (minimized seed + real oracle => the mutation guardrail bites). Failing-first: reference, agent cross-refs, and template field do not yet exist. * feat(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries) Epic #1957 Phase 3A. Extends Minimal Reproduction (shrinking) and Test-First Debugging (oracle classification + boundary neighbors): - Shrinking: wrap an input-space failing input in a property (fast-check JS/TS, Hypothesis Python) and store the MINIMIZED counterexample as the regression seed; degrade to manual minimization when no PBT framework is present. - Oracle classification: state specified / derived (contract/model) / metamorphic / implicit (crash, weakest) before writing the assertion; record under Resolution.oracle_type; never default to implicit silently. - Boundary neighbors: off-by-one, min/max, empty/singleton around the fixed defect's equivalence class. Together they turn the regression test into a root-cause check — what the Phase 1A mutation guardrail needs to bite. Full rules extracted to gsd-core/references/ debugger-repro-hardening.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md + DEBUG template updated. * fix(#1962): address orthogonal review (bounding, provenance, oracle scope, sufficient-triple) - HIGH: added a 'Bound the property/shrink run' section (60s timeout, degrade- to-manual on timeout, do-not-raise-default-run-limits, argv-not-shell) — the gauntlet violation the sibling references already honored. - Medium: test-provenance caveat (the failing input often comes from the bug report — author the generator from a sanitized description, cross-ref debugger-fix-acceptance.md). - Medium: oracle scope note — the 4 types cover deterministic bugs; non- deterministic failures re-route to stability-stress per bug-taxonomy. - Medium: Phase 1A tie-in corrected — seed+oracle is necessary not sufficient; boundary neighbors close the adjacent-input escape; the sufficient triple is seed+oracle+neighbors. - Low: preserve the original noisy repro as a secondary reference; operationalize 'equivalence class' (the predicate the fix draws). Nit: degradation reworded. * chore(#1962): backfill changeset pr number (PR #2409) --------- Co-authored-by: sim <sim@local> |
||
|
|
6baa2a8182 |
feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger (#2407)
* test(#1961): add failing-first bug-taxonomy routing contract tests Epic #1957 Phase 2B. Source-text-is-the-product contract tests (3 taxonomy classes, explicit class->technique routing table, Bohrbug->repro+SBFL+bisect, Heisenbug->record-replay/stability+SKIP-SBFL, Concurrency->atomicity/order/ deadlock checklist, bug_class in DEBUG Current Focus, supersede-not-append) plus a routing-table specification object pinning the documented decisions (SBFL forbidden on Heisenbug is the load-bearing 1B/2B seam). Failing-first: reference, Phase 1.75, and routing-table reframe do not yet exist. * feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger Epic #1957 Phase 2B (reliability-critical). Adds Phase 1.75: classify the failure as Bohrbug / Heisenbug-Mandelbug / Concurrency, then route the investigation technique via an explicit class->technique table (Kernighan: no opaque heuristic). Bohrbug -> reproduction + SBFL (Phase 1.25) + git bisect; Heisenbug/Mandelbug -> record-replay (rr) + stability-stress + statistical sampling, with SBFL explicitly SKIPPED (a flaky spectrum poisons the Ochiai ranking — the load-bearing 1B/2B seam); Concurrency -> the atomicity/order/deadlock checklist first. Reframes (supersedes, not appends — Zawinski) the flat 'Technique Selection by situation' table into a class-routed table; the 11 techniques remain as routed targets. bug_class recorded in Current Focus (DEBUG template); common-bug- patterns catalog cross-referenced to the taxonomy. Full rules extracted to gsd-core/references/debugger-bug-taxonomy.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1961): address orthogonal review (phase-name drift, General lane, revoke framing, row-scoped tests, bounding) - HIGH: reference said 'Phase 1B' (epic shorthand); corrected to the deployed 'Phase 1.25' (matches the agent + SBFL reference). - HIGH: 6 of 11 techniques (Rubber duck, Delta, Working backwards, Differential, Comment-out, Follow-the-indirection) were orphaned by the situation-table reframe. Added a 'General (any class, situation-cued)' lane to BOTH the reference routing table and the agent's Technique Selection table that re-homes them — supersede-not-append now holds. - MEDIUM: the SBFL-skip is structurally retroactive (Phase 1.25 runs before Phase 1.75 classification), so reframed the table column from 'Do NOT use' to 'Revoke if already run' + an explicit 'retroactive revocation, not proactive skip' note stating the ordering honestly. - MEDIUM: contract tests are now row-scoped (parse the table by class, assert per-row) instead of presence-only; added a guard that the previously- orphaned techniques now have a General-lane route. - LOW: pinned the canonical bug_class value form (lowercase-kebab: bohrbug|heisenbug-mandelbug|concurrency; prose may use title-case). - NIT: added a 'Bound the Heisenbug-chase runs' note (rr/stability/sampling timeouts) per the unbounded-subprocess gauntlet. * chore(#1961): backfill changeset pr number (PR #2407) |
||
|
|
f8b16d1874 |
enhance(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger (#2405)
* test(#1960): add failing-first RCA-branching contract + schema-invariant tests Epic #1957 Phase 2A. Source-text-is-the-product contract tests (fishbone >=2 categories, AND-gate, multi-cause root_cause, backward compat, reasoning checkpoint candidate_causes+and_gate fields, debugger-philosophy single-cause note, DEBUG template) plus behavioral schema-invariant checks on two fixtures: two contributing causes (AND-gate yes) -> both recorded; single-cause (AND-gate no) -> one root_cause, identical to today. Failing-first: reference, agent edits, and template note do not yet exist. * feat(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger Epic #1957 Phase 2A. Guards against 5-Whys single-cause bias: before committing root_cause, the debugger enumerates candidate causes across >=2 Ishikawa categories (code/config/environment/data) and explicitly answers an AND-gate question. When the AND-gate fires, every contributing cause is recorded, so a multi-cause fix no longer recurs via the unaddressed second cause. Resolution.root_cause may hold one OR a small set (additive; single-cause sessions are byte-identical to today). The Structured Reasoning Checkpoint gains candidate_causes + and_gate fields; debugger-philosophy.md adds the single-cause-bias trap. Full rules extracted to gsd-core/references/debugger-rca-branching.md (slim Phase 2 routing + 2 checkpoint fields kept in the agent). INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md + DEBUG template updated. * fix(#1960): address orthogonal review (AND-gate self-consistency, parity guard, narrowed claim, ripples) - Reference: the collapse rule now enforces AND-gate self-consistency — and_gate=yes with a single confirmed cause is flagged as incomplete (return to Phase 3); a race/timing note clarifies such bugs bridge categories; the 'byte-identical' backward-compat claim narrowed to 'root_cause shape unchanged; reasoning_checkpoint gains 2 fields in every session'. - DEBUG.md: stale 'five-field' mirror prose -> seven-field (parallel-surface drift the reviewer flagged); new debug-session-management parity test pins the field-count claim to the gsd-debugger.md YAML keys (CRLF-safe). - Scalar-assuming consumers of set-valued root_cause updated: session-manager compact summaries (319/332), diagnose-only return (1062), archive entry (1216), ROOT CAUSE FOUND return (1322). - Test: added the AND-gate-yes/single-cause invariant + fixture; rephrased the fixture describe block honestly as a schema-invariant specification. - Phase 2 bullet phrasing clarified ('at hypothesis formation, before the Phase 4 commit'). * test(#1960): parity regex accepts word-form count ('seven-field' or '7-field') * test(#1960): parity regex counts array-valued YAML keys (no inline value) * chore(#1960): backfill changeset pr number (PR #2405) |
||
|
|
56a5c6404c |
feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger (#2403)
* test(#1959): add failing-first SBFL contract + Ochiai correctness tests Epic #1957 Phase 1B. Source-text-is-the-product contract tests (Ochiai formula documented, Tarantula fallback, top-N seeding, no-coverage skip logged, ranking->Evidence, Bohrbug gating) plus a behavioral Ochiai formula-correctness section: bound [0,1], max-score invariant, a known-fault fixture proving the fault ranks #1 (criterion 2), clean degradation on zero failing tests, and two fast-check properties. Failing-first: reference file and agent routing do not yet exist. * feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger Epic #1957 Phase 1B. When a runnable test suite with per-test coverage exists (>=1 failing AND >=1 passing test), the debugger computes an Ochiai suspiciousness ranking over the coverage spectrum and seeds the top-N suspicious locations into Evidence as first-class hypothesis candidates, narrowing the search space deterministically before LLM reasoning. Tarantula documented as fallback. Degrades cleanly (logged, never silent) when there is no test suite, no failing tests, or no per-test coverage, and is explicitly not trusted on flaky/Heisenbug spectra (pairs with Phase 2B bug-taxonomy). Full rules extracted to gsd-core/references/debugger-sbfl.md (slim Phase 1.25 routing kept in the agent to respect the size cap). No new coverage framework — reuses the project's existing test/coverage runner. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * test(#1959): bound property generators to valid coverage counts The [0,1] property generated failedExec independently of totalFailed, but Ochiai's score is only bounded by 1 under the coverage invariant failedExec <= totalFailed (a failing test that executed s is one of the totalFailed failing tests). Out-of-domain inputs (failedExec=100, totalFailed=5) make the formula correctly return >1. Bound failedExec by totalFailed via fc.chain so the property tests the real domain. Also cleaned up the ranking property (removed dead code). * fix(#1959): address orthogonal review (monotonicity property, degradation row, coverage bounding) - Replace vacuous ranking property (true-by-sort-construction) with a non-trivial monotonicity property: holding totalFailed + passedExec fixed, ochiai is non-decreasing in failedExec. An inverted formula would fail it. - Add the missing 'no passing tests' degradation row (preconditions require >=1 passing test; Tarantula would divide by totalPassed=0). - Bound the coverage subprocess (CLAUDE.md gauntlet): cap the coverage run, degrade-to-skip on timeout, never hang the debug session. - Reword 'discard the ranking' -> 'mark the Evidence entry as revoked (do not delete)' per Kernighan auditability. * test(#1959): bound monotonicity-property generator to valid coverage (failedExecA <= totalFailed) * chore(#1959): backfill changeset pr number (PR #2403) |
||
|
|
5e52350736 |
feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger (#2396)
* test(#1958): add failing-first guardrail contract tests Epic #1957 Phase 1A. Adds source-text-is-the-product tests asserting the 5-signal fix-acceptance guardrail contract (target test, mutation check, no-op/deletion detector, adjacent tests, revert-and-reconfirm), graceful degradation, FIX REJECTED BY GUARDRAIL return path, per-signal debug-file recording, and subprocess bounding. Failing-first: reference file and agent sections do not yet exist. * feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger Epic #1957 Phase 1A. Prevents accepting a fix that merely greens the test (Goodhart defense / APR overfitting). Adds a 5-signal gate run before fix acceptance: target test, mutation check (Stryker), no-op/behavior-deleting detector, adjacent/held-out tests, revert-and-reconfirm. Degrades gracefully when Stryker or a test suite is absent (each skip logged, never a silent pass), records per-signal results under Resolution.verification, and returns a FIX REJECTED BY GUARDRAIL outcome the session-manager surfaces for revise / accept-as-debt / abandon. Full rules extracted to gsd-core/references/debugger-fix-acceptance.md (slim routing kept in the agent to respect the agent-size cap). Debug template + INVENTORY + manifest + agent-size baseline + AGENTS.md updated. * test(#1958): correct newline-tolerant assertion + regen install-parity goldens The revert-and-reconfirm assertion collapsed whitespace before matching so markdown line-wrapping does not break it. Regenerated the golden-install-parity and install-tree fixtures (npm run gen:golden) to absorb the intentional gsd-debugger.md / gsd-debug-session-manager.md / DEBUG.md / new reference-file changes to the installed artifact tree. * fix(#1958): tighten guardrail per orthogonal review Addresses the isolated reviewer's findings: - signal 5 now states its recorded-repro dependency and routes the no-repro case to the degradation row; revert mechanism specified (git stash / git revert -n); minimality flag tied to diff structure, not revert-ability. - bounded-subprocesses section now bounds the git subprocess (5-30s) too, requires argv-array argument passing, and scopes Stryker to the driving regression test (a mutant killed only by a non-driving test is a finding). - new test-provenance (security) clause: the driving test must be agent-authored; bug-report repro scripts are DATA, never executed verbatim. - tightened 3 contract assertions to bind to specific clauses (guardrail_verdict field, deletion-reject-unless-RCA, 60s+git bounding). - Goodhart framing softened to 'partially-independent'; DEBUG.md template verification field notes the nested map shape. * chore(#1958): backfill changeset pr number (PR #2396) * fix(#1958): add issue ref to allow-test-rule annotation (ADR-456) CI lint-allow-test-rule-refs requires every allow-test-rule exemption to carry a 'see #NNN' issue ref per ADR-456. The new test file's annotation lacked it; this adds (see #1958). |
||
|
|
b302f53ee6 |
refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2) (#2370)
* refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2) Behavior-preserving relocation of the 706-line case 'capability': arm from gsd-tools.cjs into a new hand-authored bin/lib/capability-command-router.cjs (sibling of ensure-runtime-build.cjs). The 15 bin/-relative require paths are rewritten to sibling-relative (correct for bin/lib/). dispatchHostCommand is now async (capability's install/upgrade/consent ops await the lifecycle); sync routers (state/phase/…) pass through await unchanged. case 'capability': removed; capability dispatches via HOST_COMMAND_ROUTERS. Validated by the existing capability-lifecycle / -consent / -trust / -loader test suites (no logic changed). Golden install-parity fixtures regenerated. Closes #2368 (Slice 1 — relocation). Probe consolidation (capHostVersion→ readHostVersion, capReadStrict dedup) deferred to a follow-up slice. * fix(#2368): add capabilityState/capabilityWriter requires + INVENTORY row The relocated capability arm references capabilityState (cmdCapabilityState, resolveCapabilityRuntimeState) and capabilityWriter (cmdCapabilitySet) — both module-scope requires in gsd-tools.cjs (L288/289) that the initial closure-dep scan missed. Added as sibling requires to capability-command-router.cjs. Also adds the new cli module to docs/INVENTORY.md + regenerates the manifest. * fix(#2368): correct capHostVersion __dirname depth for bin/lib/ relocation capHostVersion's VERSION/package.json paths were bin/-relative ('..' and '..','..'); on relocation to bin/lib/ they resolved one level too deep, so capHostVersion returned 0.0.0 and capability install failed the engines.gsd gate (#1920). Added one more '..' to each (now resolves gsd-core/VERSION and repo-root package.json correctly). * test(#2368): drop capability from the invocation loop (async/FS vs /fake/cwd) capability is async and does FS/config reads, so invoking it against the unit test's /fake/cwd is fragile. The 6 sync Tier-1 routers stay in the invocation loop; capability is covered by the non-invoking registry- ownership assertion + the dedicated capability-* test suites. * chore: retrigger CI (no-changelog label now present) |
||
|
|
2cbf186420 |
chore(#2143): fail-loud Result + per-surface write-set contract — Phase 3 (#2251)
* chore(#2143): fail-loud Result + per-surface write-set contract — Phase 3 Phase 3 of epic #2143 (ADR-2143 §5/§6). The three target bugs (#2140, #2112, #2118) were already fixed tactically on next; this introduces the reusable structural contracts and rewires the primary #2140 site onto them. - src/write-set.cts (new): the parse `Result<T> = {ok,value|reason}` (§5) and the per-surface write-set (`WriteOutcome {surface, applied, requirement?}`, `WriteSet`, `writeSetComplete`) (§6). markdown-table.cts now imports + re-exports `Result` from here (single source; distinct from command-routing-hub's Result). - requirements mark-complete (src/milestone.cts): returns a PER-REQUIREMENT, per-surface write-set; `write_set_complete` is true only if every surface of every requirement applied — structurally forbidding the #2140 OR-into-one-flag masking, including across a multi-ID batch (adversarial-review regression). Pre-existing output fields unchanged (behaviour-preserving; #2140 already fixed). - deriveProgressFromRoadmap (src/phase-lifecycle.cts): removed the vestigial null-swallowing try/catch (findTableWithColumns never throws) — ADR §5 no-swallow; RoadmapProgress return contract unchanged. - commit --files (#2112) and milestone complete --dry-run (#2118) left as-is (single-surface commit / pre-mutation preview — not genuine multi-surface writes). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2244): backfill changeset PR number (#2251) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
efd04716da |
chore(#2143): withSection bounded-mutation seam + phase.cts migration — Phase 2 (#2250)
* chore(#2143): bounded-mutation seam (withSection/withPhaseSection) + phase.cts migration — Phase 2 Phase 2 of epic #2143 (ADR-2143 §4): add a bounded-mutation primitive so a per-phase ROADMAP edit is structurally confined to that phase's own section, and migrate the phase-scoped mutation sites in `phase.cts` onto it. - `src/markdown-sectionizer.cts`: `withSection(content, target, edit, opts?)` — resolves a section via `collectSection` and applies `edit` to ONLY that section's body, re-serialising via `replaceSection`. The edit callback sees only the section body, so any regex it runs is physically confined. - `src/roadmap-parser.cts`: `withPhaseSection(content, phaseId, edit)` — resolves a phase's `### Phase N` detail-section heading via the #2121 phase-id source and delegates to `withSection`. Heading match is anchored to the heading start (a sibling phase whose title mentions the number is not hijacked) and bounds at the next ATX heading of any level (`levelBounded:false`). - `src/phase.cts`: `mutateMilestonePhase`'s plan-count and per-plan-checkbox writes now route through `withPhaseSection` — structurally retiring the #2130 / #2067 / #2080 boundary-crossing class for these sites. The phase-LIST checkbox is intentionally left milestone-slice-scoped (it lives outside any `### Phase N` detail section). Cross-phase renumbering is untouched. - Property test (fast-check): editing phase k leaves every sibling section byte-identical; regression tests for title-collision + mixed heading depth. Behaviour-preserving (verified by old-vs-new differential runs on real fixtures). Extend-never-mutate (ADR-2143 §2). Registration: CONTEXT.md + docs/INVENTORY.md export lists. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2243): backfill changeset PR number (#2250) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d49ac81306 |
chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1 (#2248)
* chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1 Phase 1 of epic #2143 (ADR-2143): consolidate markdown table parsing onto a canonical seam and migrate the pilot reader. - Add src/markdown-table.cts: parseMarkdownTable (GFM tables -> typed {columns, rows} addressed by column NAME; ragged rows are typed parse errors, not silent), a single-source TABLE_SCHEMAS registry (RoadmapProgress / RequirementsTraceability / QuickTasks / Security, with variants under one id), matchTableSchema, and findTableBySchema. Result<T> is scoped to this seam (distinct from the dispatch Result). - Migrate deriveProgressFromRoadmap (src/phase-lifecycle.cts) off the position-anchored regex to name-based resolution via the seam — fixes #2137 (the 5-column milestone-grouped Progress table previously returned all-null). - Add a schema-backed `gsd-tools quick-tasks-append` subcommand and route fast.md's log_to_state through it, retiring the inline `awk NF-2` column arithmetic — fixes #2133 (addresses #2012, #2119). Cell values are escaped (| and newlines) and the STATE.md read-modify-write is atomic under readModifyWriteStateMd (lost-update race, cf. #500/#905/#1230). - Writer/reader/template parity test guards TABLE_SCHEMAS against drift (ADR-2143 §3 Generative-Fix-Divergence). Registration: .gitignore, eslint.config.mjs, docs/INVENTORY.md + INVENTORY-MANIFEST.json, CONTEXT.md glossary, docs/CLI-TOOLS.md. Behaviour-preserving for the canonical 4-column Progress table; the named bugs are driven fail-first. Extend-never-mutate (ADR-2143 §2). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2242): backfill changeset PR number (#2248) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2242): escape backslash before pipe in markdown-table cell escaping CodeQL js/incomplete-sanitization (high): escapeCell escaped | -> \| but not the backslash itself. Now escapes \ -> \\ before | -> \|, and splitTableRow unescapes both \\ -> \ and \| -> | symmetrically so cell values (incl. literal backslashes) round-trip exactly. Added backslash round-trip tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2242): read ROADMAP Progress table by column name — supersede #2168 ad-hoc scan Rebase reconciliation with #2168 (the tactical #2137 fix that marked itself "pending #2143"). deriveProgressFromRoadmap now resolves the Progress table via a new seam helper findTableWithColumns (first table whose header is a superset of Phase/Plans Complete/Status/Completed, any order, extra columns ignored) and reads cells by NAME — order/injection-invariant per ADR-2143 §3 — instead of the exact TABLE_SCHEMAS match. This satisfies #2168's column-invariance property test while staying seam-based and preserving its `## Progress` scoping (#2012/#1445). Ragged Progress tables now resolve to null (ADR-2143 fail-loud); updated the stale state.test.cjs assertion that predated the Phase-1 migration. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
bd613566cb |
feat(#2100): drive Windsurf through the EoS descriptor + wire Cascade's blocking hook bus (ADR-1239)
Fold all 10 residual isWindsurf branches in bin/install.js onto descriptor-driven hostBehaviors (byte-parity — no fold changes any install output): - 2 dead destructures dropped (uninstall, finishInstall); the dead `else if (isWindsurf)` legacy agent-loop arm removed (windsurf ∈ _DESCRIPTOR_AGENTS_RUNTIMES → unreachable). - skipSharedHooksInstall:true folds the two `!isWindsurf` shared-hooks exclusions. - legacyDevinSkillsCleanup:true folds the `.devin`→`.windsurf` one-time cleanup gate. - installsCommandBodiesForWorkflowDelegation:true folds the #1629 command-body copy (workflow-delegation target — load-bearing; local-install verified intact). - verificationStyle:"windsurf-workflows" folds the workflow-count report. - Corrected stale _LEGACY_SCAN_SUBDIR_NAMES + hooks-json manifest comments (cursor + windsurf). Zero live runtime==='windsurf'/isWindsurf branches remain across bin/install.js, install-engine.cts, surface.cts, runtime-artifact-conversion.cts (AC2 guard scans all four). UPGRADE (Cascade hook bus): wire GSD's write/command safety guards into Windsurf's native hook bus. New hooksSurface 'windsurf-hooks-json' (VALID_HOOKS_SURFACES 7→8, GATE A profile-marker-only allowlist, the HooksSurface union) + writeWindsurfHooksJson (Cursor-templated, Cascade's flat {hooks:{<event>:[{command}]}} shape) writing .windsurf/hooks.json with two BLOCKING pre-hooks: - pre_write_code → gsd-windsurf-pre-write.js: blocks writes to a file outside the active git worktree / into .git internals. - pre_run_command → gsd-windsurf-pre-command.js: conservative destructive-command deny-list (rm -rf of root/home incl. sudo/env/path-prefixed forms; fork bombs; force-push refspec forms — HEAD:main, +main, --force/-f — to main/master/next). Both use Cascade's protocol (stdin JSON, exit 2 + stderr to block, exit 0 to allow, fail-open on error/timeout). Tokenize-based classifier (no catastrophic-backtracking regex; 4096-char cap) with the fail-closed false-positives fixed post-review. The 4 advisory GSD guards + pre_mcp_tool_use + 5 post_* logging events are deliberately NOT wired: Cascade has no context-injection channel for advisory hooks and GSD has no MCP guard — porting them would be non-functional padding (documented; codebuddy #2098 / copilot #2099 faithful-subset precedent). extendedHookEvents stays []. Golden: the 2 guard scripts ship in the shared hook bundle (HOOKS_TO_COPY + the shared managed-hooks-registry), exactly like cursor's 6 gsd-cursor-*.js scripts — so the 8 shared-bundle runtimes' fixtures gain the 2 inert windsurf scripts + the registry hash (functionally inert for non-windsurf; the established cursor pattern). No install-output change beyond that (the folds are byte-parity; skip-bundle runtimes untouched). New scripts registered in managed-hooks-registry + build-hooks + INVENTORY. Tests: declarative-reference- windsurf (adapter/axes/fail-closed + AC2 guard) + windsurf-hooks-bridge (live exit-2 blocking + allow/fail-open + ReDoS-bound + writer/reconcile/remove idempotency); VALID_HOOKS_SURFACES pin updated to 8. Matrix hookBus delta + changeset (Changed). capability-registry regenerated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c1756d0cd5 |
chore(#1867): register ui-consideration probe + regen install cascade (SHIP-01)
Ship-safe registration + regenerated snapshots for the #1867 UI-consideration probe (Phase 3, SHIP-01): - CONTEXT.md: PROBE.ui.{verification,axis,seam} predicates + ui-consideration -probe added to PROBE.family (machine-canon for the 3rd adapter, MIXED axis). - agents/gsd-ui-{researcher,checker}.md: one @-include of references/ui-consideration-probe.md each (both under the 24576 agent cap). - docs/INVENTORY.md + INVENTORY-MANIFEST.json: register the reference doc and the compiled ui-consideration-probe.cjs (inventory-manifest-sync green). - tests/fixtures/golden-install-parity/*.json (16 runtimes): recaptured against a clean full build — folds in the deferred Phase-1 (ref doc, plan-phase lift) and Phase-2 (ui-phase step, UI-SPEC section) install-surface changes. - tests/agent-size-baseline.json: ratcheted the two grown UI agents. - .changeset/vivid-orcas-chatter.md: type Added (pr updated at PR-open). Inventory/golden/size gates green; lint:ci + lint:docs + lint:changeset green. The plan-phase.md PRE_PHASE6 ceiling stays RED pending #1852 (unchanged). Claude-Session: https://claude.ai/code/session_01BKt4hgNZwXSeJYJtYAQUSS |
||
|
|
303a796579 | docs(changeset): #2089 cursor host-integration migration + golden fixture | ||
|
|
5aa7771dd3 | Merge branch 'next' into fix/2009-load-failed-capability-injects-blocking- | ||
|
|
015c3a7fda |
fix(#2071): extract install-time effort resolvers so effort sync stops requiring the un-shipped bin/install.js
`gsd-tools effort sync` crashed in every installed runtime (e.g. ~/.claude/gsd-core/) with `Cannot find module '../../../bin/install.js'`: cmdEffortSync (src/commands.cts) required the package-root bin/install.js for its install-time effort resolvers, but the installer only copies the gsd-core/ subtree into a runtime home — bin/install.js is never present there. So `effort` config changes silently never reached installed agents without a full reinstall (exactly the gap #488 was meant to close). 4th instance of the recurring "runtime code under gsd-core/ requires a file outside the shipped subtree via ../../../" anti-pattern (#1223/#1920/#1383 were the prior three, all already mitigated). Fix (ADR-457 direction — extract, single source): move readGsdEffectiveEffortConfig + resolveInstallTimeEffort (with their _getGsdEffortCatalog + _readGsdConfigFile helpers) out of the hand-authored bin/install.js into a new src/install-effort-resolver.cts that compiles into the shipped gsd-core/bin/lib/install-effort-resolver.cjs. commands.cts now requires it as a sibling (`./install-effort-resolver.cjs`) — always present in the installed tree — instead of `../../../bin/install.js`. bin/install.js imports the same four symbols back from the new module (it still calls them + re-exports them), so there is one source of truth and no duplication/drift. The lazy manifest read is repointed from the package-root layout (`.., gsd-core, bin, shared`) to the bin/lib layout (`.., shared`). Scope note: this is one of four instances of the anti-pattern; the other three are already shipped/guarded. A build-time guard rejecting new cross-boundary requires whose target isn't in the installer copy manifest (to prevent instance #5) is recommended on the issue but kept out of this fix. Tests: tests/effort-sync-installed-runtime.test.cjs does a real minimal install into a temp home (the golden-parity helper) and runs the issue's exact repro (`gsd-tools effort sync --config-dir <temp>`), asserting no MODULE_NOT_FOUND for bin/install.js. Fail-first verified: against pristine next the same test throws `Cannot find module '../../../bin/install.js'` at cmdEffortSync; post-fix it syncs cleanly. New module registered in .gitignore (ADR-457), eslint ignores, docs/INVENTORY.md + INVENTORY-MANIFEST.json. bin/install.js is not shipped and the new module is under bin/lib (excluded from golden parity), so no golden fixtures change. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
46f8d3814c |
fix(#2009): load-failed capability gates fail open with a loud warning
Previously a capability that failed to LOAD (e.g. incompatible engines.gsd) but declared a gate-kind loop hook caused the loop resolver to inject a BLOCKING synthetic gate (blocking:true, onError:halt) at every declared point, halting every ship:pre / verify:post project-wide over an unrelated load error, with no remediation surfaced. Per maintainer decision (#2009) it now fails OPEN: no gate is injected (the loop proceeds; --active-cap correctly reports the failed cap inactive) and a loud warning is emitted — to stderr (the channel host workflows/agents actually see) and in the envelope 'warnings' array — naming the load reason and the exact 'gsd capability remove <id>' remediation. The loader still records blockedGates; only the consequence changes from block to warn. Security (review): capId and reason originate from a third-party manifest / directory name. capId is validated against the canonical kebab-case id shape before it is placed in the runnable remediation command (withheld otherwise); reason is stripped of control chars and backticks. This closes an argument/ prompt-injection vector in the surfaced message. Docs updated to the fail-open-warning posture (ARCHITECTURE, INVENTORY, CONFIGURATION, README, capability-overlay-model). Also removes a dead 'before' import surfaced by lint in the issue-2045 test. |
||
|
|
6addeccd19 |
feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates an external API/SDK/service can no longer seal without a decided coverage matrix. - src/api-coverage.cts: deterministic detector (compound verb+noun signal + <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix parse/validate/render with field-length caps. - check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as a token under .planning/phases/ only (traversal-neutralized); validates COVERAGE.md or blocks iff a strong integration signal is detected and no matrix exists; fail-closed when phases tree exists but phase unresolvable. - capabilities/ai-integration: workflow.api_coverage_gate config key (default true), plan:pre contribution, blocking verify:pre gate. Data-driven. - gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch. - Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e. Code+security review findings fixed (stopword FP, scope containment, pipe/cap rejection, prompt-injection message hygiene). - Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs. Closes #1562 |
||
|
|
603593d41d |
fix(#1857): test gates normalize to one-shot + bounded timeout (no watch-mode hang)
A GSD verification gate resolves a project's test command and runs it. vitest defaults to WATCH mode in an interactive TTY — exactly where a user runs `gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never exited and the orchestrator waited indefinitely. Recovery needed the user to manually prompt "something blocking?". Fix — one shared helper + a bounded, surfacing timeout on the test-command gates: - New pure module src/normalize-test-command.cts + `gsd-tools query normalize-test-command` verb: rewrites a resolved command to a best-effort one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`; a package-manager `test` script whose package.json runner is watch-vitest → `CI=true` prefix; handles `--dir`; already-one-shot commands unchanged — never double-flagged). Named `normalize-test-command` (not `test-*`) so the file does not match node --test's default `test-*` discovery glob. - The three gates that HUNG or silently-continued route through that ONE helper and bound execution with `timeout $(config-get workflow.test_gate_timeout)` (new config key, default 600s): the regression gate (extracted to execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen — it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124. - verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint on 124, staying under its frozen 40960-byte tier cap. Security hardening (review): the normalizer only rewrites a runner named as a standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never mangled), is length-capped and uses only linear-time split-based scanning (no super-linear backtracking on an adversarial `workflow.test_command`), and reads package.json only when it is a regular file (never blocks on a FIFO via `--dir`). Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores, inventory manifest/index. All 16 golden-install-parity fixtures + workflow size baseline regenerated for the changed shipped files; bin/lib is excluded from the parity manifest. Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route through the shared helper + configured timeout + exit-124 watch-mode hint; verify-phase asserted as normalize-only/already-bounded). tests/execute-phase-active-flags.test.cjs repointed at the extracted step; tests/planner-language-regression.test.cjs allowlist comment updated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ea4378063f | Merge branch 'next' into feat/1820-specless-predicate-rail | ||
|
|
7ef834cabc | feat(#1820): spec-optional predicate rail — author probe predicates into must_haves when SPEC omits them | ||
|
|
896c2740d3 | feat(#1990): add onboard command for brownfield setup | ||
|
|
e5ef323b15 |
feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command
Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.
* feat: add /gsd-start smart-entry command
State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.
- src/smart-entry.cts: deterministic situation classifier (no-project,
paused, blocked, verify-failed, needs-first-phase, planning, executing,
verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
dispatcher presenting an AskUserQuestion menu (with --text fallback for
non-Claude runtimes) and dispatching to existing commands. Falls back
to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
(markdown-layer invariants + every emitted command resolves to a real
slash command).
Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.
* refactor: rename smart-entry command to /gsd:next
Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.
All affected tests (188) pass; lint:ci clean.
* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)
Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.
- detectSignals now reads total_phases + percent from nested progress{}
first, then scalar fm, then body; current_phase falls back to the
body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
body Phase field) covering verify-pending + executing situations.
Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.
* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)
Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.
Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
layer is down — read .planning/STATE.md directly with the Read tool
and synthesize a minimal situation + actions menu so /gsd:next stays
useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
/gsd:progress as before.
Matches the direct-read resilience the live agent already did by hand.
* docs: add gsd-next skill surface
* chore: trigger no-mistakes validation
* no-mistakes(review): Fix smart-entry phase ordering
* no-mistakes(review): Fix decimal smart-entry phase ordering
* no-mistakes(test): Fix smart-entry next test contracts
* no-mistakes(document): Docs synced for smart entry
* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: shorten next.md description and update golden install parity fixtures
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: trigger no-mistakes validation
* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files
Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.
* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: regenerate golden install parity fixtures for /gsd:next
Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.
* Fix smart-entry verify-failed phase scoping and empty resolve shim step
Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.
* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: recapture all 16 golden fixtures with updated smart-entry.md hash
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: regenerate fixtures + inventory manifest after rebase onto next
Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next
Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).
Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)
The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs
Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo
Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
recommended action, 1-4 unique-id /gsd:* actions (previously the
one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout
Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:
- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
consolidation file doing dozens of real installs — 41s even on a fast Mac
(much worse on the slow Windows I/O path), plus an install-heavy cluster.
Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.
Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
3c13903dcd |
feat(#1866): agent-side self-load of configured agent_skills
Each of the 22 consumer agents now self-loads its configured agent_skills in its mandatory init step, so .planning/config.json agent_skills.<type> reaches the agent on every runtime — including Cursor and /gsd-autonomous, where Skill()-delegated workflow bash init did not reliably execute. - gsd-core/references/agent-skills-bootstrap.md: shared contract (query + Read + dedup guard that skips when <agent_skills> is already in the prompt, so Claude's orchestrator-side injection never doubles) - 22 agents/gsd-*.md: one self-load line naming the agent's own type - gsd-core/workflows/autonomous.md: note that delegated agents self-load - tests/agent-skills-bootstrap.test.cjs: regression + parity (CONSUMER_AGENTS bijection + fast-check property) — Generative-Fix-Divergence guard - docs: ADR-1866, CONFIGURATION dual-injection How It Works, INVENTORY row, Changed changeset Closes #1866 |
||
|
|
18995380ce |
feat(#1154): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1738)
* feat(verify-phase): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1154) Carry the edge-probe's existing `backstop` (non-inferable) tier through the plan-phase projection as a structured flat-scalar marker instead of a prose parenthetical, and make verify-phase abstain -> human_needed (never silent-pass) on a backstop truth it cannot confirm with explicit evidence. Truth-axis mirror of #644's prohibition judgment-tier (ADR-550 D4). Engine (deterministic, CI-tested per ADR-550 D5 — never the LLM verdict): - src/probe-core.cts: truthStatement/truthVerification normalizers, projectTruths (conservative serializer), dispositionForUnverifiableTruth (backstop+no-evidence -> unverified/flagged/insufficient_spec; backstop+evidence -> green; inferable -> green, the over-abstention guard). - src/roadmap.cts: coerceTruthToString now reads `statement` first so an object-form backstop truth is surfaced, not dropped (Hyrum backward-compat for truth-readers). Workflow/agent/docs: plan-phase emits the structured marker (flat scalar, ADR-550 #1278); verify-phase + gsd-verifier add the abstain arm; new references/honest-verifier.md; FEATURES/COMMANDS document insufficient_spec; ADR-550 amended (truth-axis D4 mirror). Decisions adopted (trek-e review): insufficient_spec feeds existing human_needed with a distinguishable reason (no new VERIFIER_STATUS); changeset Changed; round-trip parity test; abstain-on-unconfirmed-backstop regression test red-first. Implementation notes (deviations from the issue's proposed file list, verified live): - frontmatter.cts needs no change — its flat parser already round-trips object-form truths. - verify.cts needs no change — it grades artifacts/key_links structurally; truths are LLM-graded at the workflow layer, so consumption lives there + the deterministic helper. - No CJS<->SDK hand-sync — the SDK seam was retired (ADR-0174); src/*.cts is sole source. Regenerated artifacts: golden-install-parity fixtures, INVENTORY-MANIFEST, size baselines. * chore(#1154): add changeset (Changed) for honest verifier User-facing changelog fragment for #1738. Typed `Changed` (not `Added`) per trek-e review condition 3 — the verify behavior shifts for backstop-bearing specs (a confident silent `passed` becomes `human_needed`), which is user-visible even though the schema marker is additive. * docs(#1154): score-formula also excludes abstained insufficient_spec truths (review nit-1) trek-e review nit: the verify-phase score sentence said PRESENT_BEHAVIOR_UNVERIFIED truths were "the only ones excluded" from verified_truths. Post-#1154 an abstained `insufficient_spec` backstop truth is also excluded (it is not ✓ VERIFIED and routes to human_needed). Behavior was already correct; this tightens the wording. Regenerated golden-install-parity fixtures + workflow-size baseline for the touched verify-phase.md. (Nit-2 — a dedicated insufficient_spec_items frontmatter list — is intentionally not taken: the current design is ADR-550-D4-conformant, the abstain cause rides as a distinguishable report reason, and adding it would exceed the approved scope.) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
b0d5ca3379 |
feat(#1517): support custom reviewer instances for /gsd:review (#1766)
* feat(#1517): support custom reviewer instances for /gsd:review Add a bounded review.reviewer_instances config surface so one model-capable adapter (e.g. opencode) can run as several independent reviewer identities in a single /gsd:review pass. Instances participate only via review.default_reviewers, expand before built-in slugs, are available iff their cli is detected, and a non-matching entry is a hard error (typo must be loud). >=2 same-cli instances emit a shared-adapter caveat in REVIEWS.md. Default path with no instances is byte-for-byte unchanged. Single-source instance->cli resolution lives in resolveReviewerSelection / normalizeReviewerInstances (parity-locked in tests/review-reviewer-instances.test.cjs). cli validated against KNOWN_REVIEWER_SLUGS only (never arbitrary shell); model/agent opaque, never shell-interpolated. Closes #1517 * chore(#1517): backfill changeset pr:1766 --------- Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
b307c4cfde |
refactor(#1734): extract install engine from bin/install.js (ADR-1239 Phase B deep move) (#1735)
* refactor(#1734): extract install engine from bin/install.js (ADR-1239 Phase B deep move) Relocate the runtime-artifact install cluster out of the 12,490-line bin/install.js into a dedicated src/install-engine.cts -> install-engine.cjs: installRuntimeArtifacts, uninstallRuntimeArtifacts, installOpencodeFamilySkills, and their cluster helpers (_copyStaged, snapshot/restore, legacy migration, GSD-entry pruning, preserve/restoreUserArtifacts, OpenCode-family converters, USER_OWNED_ARTIFACTS). - bin/install.js imports the engine and re-exports the moved symbols for back-compat; getCommitAttribution STAYS in install.js (impure config I/O + argv explicitConfigDir global) and is injected via a resolveAttribution param. - 17 test files migrated to import the moved symbols from the engine. - Bookkeeping: eslint built-artifact ignore, .gitignore, INVENTORY manifest+row, CONTEXT.md Install Engine Module glossary seam. Behaviour-preserving: install output is byte-identical for all 16 runtimes (golden-parity harness #1730) — the only delta is the new install-engine.cjs file shipping in the installed gsd-core/bin/lib/ tree. Closes #1734 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1734): backfill changeset PR number (#1735) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
30d4b85de5 |
feat(#1684): negotiated host-integration interface (ADR-1239 Phase A) (#1690)
* feat(#1684): add negotiated host-integration interface module ADR-1239 Phase A: a pure, additive, no-I/O module exposing PROTOCOL_VERSION, the 8-axis HOST_INTEGRATION_AXES closed vocabulary, the UNDOCUMENTED fail-closed sentinel, negotiateHostCapabilities (effective subset of host-declared and engine-known), a typed degradation ladder, and host-capability profiles. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1684): validate and document host-integration axes (16 runtimes) Extend validateRuntimeBody to validate the 8 hostIntegration axes (closed enums + undocumented sentinel + dispatch struct + reserved-key guards) and the widened runtime vocabulary; author a documentation-sourced hostIntegration block in all 16 runtime descriptors; regenerate the registry. Every per-CLI value is documented (cited) or the explicit undocumented sentinel. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add host-integration capability matrix and adr amendment New per-CLI, per-axis citation reference (value/source/evidence for all 16 CLIs); ADR-1239 Phase-A-implemented amendment; CONTEXT.md glossary seam entry. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1684): harden dispatch negotiation edge cases Code-review hardening: treat NaN/Infinity maxDepth as missing (fail-closed, +warning); reset nested/background when namedDispatch collapses to false (struct consistency); SAFE_DEFAULTS dispatch floor to read-only; warn on non-finite protocolVersion; symmetric undocumented warnings for dispatch fields. Pure module — no consumers; behaviour fail-closed throughout. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1684): register host-integration.cjs in lint-ignore and inventory New tsc-generated bin/lib artifact: add to the eslint ignore list (ADR-457 — lint the .cts source), regenerate docs/INVENTORY-MANIFEST.json, and add the docs/INVENTORY.md CLI-modules row. Fixes the 3 gsd-test failures (551-eslint-bin-lib-coverage x2 + inventory-manifest-sync). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add changeset fragment for host-integration interface Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add how-to for sourcing a host's integration axes Diataxis how-to guide for adding/updating a host's runtime.hostIntegration axes from authoritative docs, the undocumented-sentinel rule, validation, and extending the closed vocabulary. Completes the Step-5 doc quadrants (reference + explanation + how-to). Indexed in docs/README.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1a46109b97 |
enhance(#1579): deterministic gsd-tools query eval.score verb (#1583)
* feat(#1579): deterministic gsd-tools query eval.score verb Split C of #1573 (pure code, lowest risk). Adds an eval.score query verb (coverage*0.6 + infra*0.4; bands 80/60/40) mirroring the verify.* chain; gsd-eval-auditor consumes it instead of doing weighted arithmetic in-prompt. Non-breaking — additive only. arXiv: 2601.15130 (Plausibility Trap/DPDM), 2507.10281 (Table Agent), 2508.15754 (TIR). * fix(#1579): address review — domain guard, property test, glossary, SKIP_ROOT, inventory/baseline - C3 input-domain: reject out-of-domain eval.score (require 0<=covered<=total; was emitting overall_score>100 / negatives) - C1 property test: add tests/eval.property.test.cjs (fast-check) — determinism, band monotonicity, [0,100] bounds, never-throws - C2 glossary: CONTEXT.md "Eval Scoring Module" entry (source-of-truth path + interface) - C4: add `eval` to SKIP_ROOT_RESOLUTION (pure arithmetic; no .planning/ access) - inventory: register generated eval.cjs/eval-command-router.cjs (INVENTORY-MANIFEST.json + INVENTORY.md rows) - size: regen agent-size baseline for gsd-eval-auditor (reused gsd_run shim + eval.score step) - eslint: ignore generated eval*.cjs (ADR-457 bin/lib migration coverage) * fix(#1579): register eval family in alias-drift gates Add EVAL_COMMAND_ALIASES/EVAL_SUBCOMMANDS to scripts/check-alias-drift.cjs families and to familyArrayKeys in the manifest-coverage test, so the eval family lands under the same drift guard as every sibling family (state/verify/init/phase/phases/validate/roadmap). Addresses trek-e review. check:alias-drift ok; feat-3251 coverage 9/9; eval suites 10/10. * docs(#1579): use half-open verdict band ranges in CLI-TOOLS overall_score is fractional and thresholds are >=80/>=60/>=40, so a score in [79,80) is correctly NEEDS WORK despite the old '60-79' label. Relabel bands as 60-<80 / 40-<60 / 0-<40 to match the code. Addresses trek-e nit. * fix(#1579): validate eval.score CLI inputs Reject unknown infra tokens and fractional counts, and pin the 80-point verdict boundary including rounding-before-banding behavior. |
||
|
|
a63684c222 |
enhance(#1577): WebFetch/WebSearch injection isolation + opt-in blocking (#1585)
* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the largest untrusted channel) in gsd-read-injection-scanner; shared untrusted-input-boundary reference @-included by the 8 ingest agents (randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring); opt-in security.injection_blocking (default advisory — non-breaking). arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth). * fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized - A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already in the transcript. The prompt-level data/instruction boundary is the primary control. - A2: registered security.injection_blocking in the config schema + defaults manifests (default false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads. - A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention). - A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale). - A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input. - Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) + drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as the noted pre-existing follow-up. * fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate The new reference quotes injection phrases ('ignore previous instructions', 'you are now…') as examples agents must NOT comply with, tripping the repo's own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red on HEAD). Allowlist it alongside the other security docs (security-model.md, TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS scanner test doesn't scan references/, so only the shell gate needed it. Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15. * fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so no named web-ingress agent is uncovered, keeping the two justified additions (gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10. - gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset. - gsd-assumptions-analyzer reads 5-15 codebase source files (external/source- document ingress per the boundary), though it has no web tools. INGEST_AGENTS in the isolation test now asserts all 10; size baselines regenerated (+60 bytes each, both well under the DEFAULT cap); changeset reworded 8 -> 10. Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39. * docs(#1577): document security.injection_blocking + boundary seam trek-e Major 2 + Minor: - docs/CONFIGURATION.md: add the top-level security.injection_blocking key to the Full Schema and a Security Settings subsection, distinguishing it from the workflow.security_* namespace; honest circuit-breaker-not-redactor framing matching ADR-1577 / security-model. - CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry. Verified: lint:docs ok; config-field-docs + contributor-standards green. * test(#1577): make read-injection property test git-text, not binary trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as degenerate-edge inputs. The NUL is what actually made git classify it binary (git binary = NUL in first 8K). Replace both with text-safe escapes that keep the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File now diffs/blames line-by-line. Verified: property test 2/2; no NUL/raw-noncharacter bytes remain. * docs(#1577): align untrusted boundary docs Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner. * docs(#1577): align ADR ingest agent count Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
f9d9dfb4bc |
fix(#1627): scale security rigor by ASVS level (planner disposition + auditor depth) (#1636)
workflow.security_asvs_level was display-only — the planner hardcoded 'mitigate if ASVS L1 requires it' and the auditor only echoed the level, so L2/L3 behaved identically to L1. - New reference gsd-core/references/security-asvs-levels.md defines L1 (opportunistic), L2 (standard), L3 (comprehensive) for both planner threat disposition and auditor verification depth (higher = superset). - planner: disposition now scales with the configured ASVS level (no hardcoded L1) + @-pointer to the reference. - auditor: verification depth scales with asvs_level (L1 grep-presence, L2 boundary/vector check, L3 end-to-end trace + bypass check). - planning-config.md + INVENTORY updated; planner kept under its 48K cap by extracting the goal-backward worked example to planner-guidance.md. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ba96c70b14 |
feat(#1602): deterministic coverage-metadata UAT routing for verify-work
Add an optional structured `coverage:` block to SUMMARY.md frontmatter and a deterministic classifier that `verify-work` consumes to route deliverables to auto-pass vs human-UAT — replacing the rejected #1598/#1599 post-hoc heuristic. - New `src/coverage.cts` (→ bin/lib/coverage.cjs) parses the nested coverage block (extractFrontmatter can't — its `-` items are scalars-only; this is a focused parser, sibling of parseMustHavesBlock), validates each entry, and classifies into auto_passed vs present. Frozen MODE/PRESENT_REASON/ERROR_CODE typed-IR surface. Exposed via `uat classify-coverage --summary <f>`. - Auto-pass is the narrow proven case only: strict-boolean human_judgment:false AND non-empty all-`pass` verification AND zero validation errors. Everything else — judgment, empty/failing verification, malformed entry — routes to the human (fail-safe). A malformed block falls back to legacy prose extraction and surfaces an error; an absent block is byte-identical to pre-#1602. - execute-plan create_summary populates the block (fail-safe default human_judgment:true); verify-work extract_tests consumes it; create_uat_file marks auto-passed entries `source: automated`. - Templates (summary + 3 variants), CONTEXT.md predicate + glossary, INVENTORY, eslint/gitignore registration, and Diataxis docs (COMMANDS reference + USER-GUIDE explanation) updated. - Behavioral tests via the CLI (no source-grep); parser-robustness regressions for the null-entry/comment-header/mis-indent cases found in adversarial review. Closes #1602 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e12a2abfd8 |
feat(#441): add /gsd-capture --list-seeds for seed listing and audit (#722)
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed), enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way to browse or audit parked seeds on demand. This adds a read-only listing, following the established --list → workflow pattern (per the approved scope on - gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the seeds dir, returns { count, seeds[], summary } JSON with each seed's id, slug, status, scope, trigger_when, planted, title. Optional case-insensitive status filter. User-controlled content is sanitized (sanitizeForDisplay) and every path validated (requireSafePath); read-only. Independent of audit.scanSeeds, which only returns unimplemented seeds for the milestone surface. - /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that renders the seed table. Closes #441 * chore(#441): point changeset fragment at PR #722 * test(#441): allowlist list-seeds test in prompt-injection scan The test asserts that list-seeds neutralizes injection payloads (<system>, [INST]) embedded in seed content, so the fixtures legitimately contain those patterns — same as the sibling security tests already on the allowlist. * fix(#441): use canonical /gsd:capture colon form in list-seeds workflow Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new list-seeds workflow used the hyphen form. * docs(#441): sync help full.md + INVENTORY for --list-seeds Adds the --list-seeds entry to the help reference (help/modes/full.md, per bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json. * docs(#441): add --list-seeds how-to + drop phantom statuses Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers): - USER-GUIDE.md Seeds section (how-to): extend the task to cover auditing parked seeds on demand via --list-seeds, including the status filter — kept task-oriented per Diataxis how-to mode. - CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected from the list-seeds filter vocabulary; the system only produces dormant|active|triggered (src/audit.cts scanSeeds). Reference must be factually accurate and complete. * fix(#441): guard non-scalar status frontmatter in cmdListSeeds A seed with a bare `status:` line (extractFrontmatter yields {}) or a `status: [a, b]` value (yields an array) crashed the whole audit list: `(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string. Coerce every frontmatter read through a `fmStr` helper (mirrors the existing `typeof fm.id === 'string'` guard), so a non-scalar status falls back to dormant and non-scalar scope/trigger_when/title can no longer leak a raw array/object into the JSON contract. Title is now capped symmetrically. Adds regression coverage for empty and array `status:` and non-scalar fields. Refs #441 * docs(#441): align list-seeds workflow status vocabulary The load_seeds step listed `implemented` as an example status filter, but the real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds); `implemented` has no producer. Matches the earlier CLI-TOOLS.md correction. Refs #441 * refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested in-process (review minor #1). No behavior change. Filter comparison now matches the raw lowercased status (both sides already normalized) instead of sanitizeForDisplay(status); sanitization is for output, not matching (review nit #3). * test(#441): add fast-check property coverage and count=1 boundary for list-seeds Adds tests/list-seeds.property.test.cjs with four fast-check properties over deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing (review minor #1). Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2). * chore(#441): sync runtime launcher snippet into list-seeds workflow Propagate the current _runtime-launcher.snippet.sh (with non-Claude runtime home probes) into the new list-seeds.md workflow via scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation. * test(#441): record list-seeds.md in workflow size baseline (#1074) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
405ae9b3b7 | refactor(#1557): add runtime artifact install plan module (#1560) | ||
|
|
b63a500fad |
fix(#1505): extract context_guard step to reference file; fix allow-test-rule see ref
- Extract execute-phase.md context_guard step prose to gsd-core/references/execute-phase-context-guard.md (@-ref lazy load), bringing execute-phase.md back under the ADR-857 phase-6 size ceiling (92914 < 93166 bytes) - Fix allow-test-rule comment: add `see #1452` per ADR-456 lint rule - Update feat-1452 tests to check the reference file for extracted content - Register execute-phase-context-guard.md in INVENTORY-MANIFEST.json and INVENTORY.md Workflow References section - Regenerate workflow-size-baseline.json after file shrinkage Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
3a3b2135c2 |
chore(#1073): purge phantom pre-migration issue refs from source, tests, docs (#1471)
#2551/#3182/#2361 are pre-migration get-shit-done-redux issue numbers with no equivalent in open-gsd/gsd-core; they mislead triage and manufacture phantom blockers. Repoint to real successors (#717 byte-budget rework, #720) or rewrite as prose referencing the discuss-phase/modes progressive-disclosure split. Correct co-located 'line budget'/'<500 lines' framing to the byte-based reality (#717). Add a CI guard (tests/no-phantom-issue-refs.test.cjs) that fails if a phantom ref is reintroduced. SSH-key patterns (id_ed25519) left untouched. No user-facing runtime behavior change. Closes #1073 |
||
|
|
e7855bc217 | fix(#1459): user-owned consent store gates third-party capability activation; env/cwd in disclosure; loader validator parity (#1473) | ||
|
|
0d56f544d2 |
feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation (#1458)
* feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation ADR-1244 Phase 6. Adds the capability matrix reference, generated FROM the committed registry so it can never drift from the actual capability set: - scripts/gen-capability-matrix.cjs (--write / --check); --check is a CI drift guard. - tests/capability-matrix-sync.test.cjs (4 tests): drift guard, buildMatrix==committed, every cap present, no placeholders. - docs/reference/capability-matrix.md regenerated from the registry (release-stable: shows engines.gsd, omits the lockstep per-cap version that would churn the file every release). - Consolidated the duplicate trust-model doc: deleted docs/explanation/the-capability-trust-model.md, merged its content into capability-trust-model.md, redirected ~10 references; no stale links remain. - Diataxis verification (now that gsd capability is a real command): corrected the matrix's third-party section — the matrix is the first-party catalogue; the overlay-aware view of installed third-party capabilities is 'gsd capability list', not this generated file. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1435): Added changeset for the capability matrix reference Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1435): address code-review — non-vacuous matrix test + generator polish - capability-matrix-sync.test.cjs: assert the 'security registers a ship:pre gate' precondition unconditionally so the extension-point check can never degrade to a vacuous pass on registry drift. - gen-capability-matrix.cjs: warn (stderr) on an unknown loop point at generation time; rename enginesOf -> fmtEngines for consistency with the other fmt* helpers (output unchanged). - capability-trust-model.md: point the two how-to links at the real files (import-a-capability-from-a-url.md, version-a-capability.md) instead of the bare directory. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1435): backfill changeset PR number → #1458 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9219af3360 |
feat(#1433): capability trust gate + upgrade/compat (ADR-1244 Phase 4) (#1449)
ADR-1244 Phase 4 (D5 trust + D6 upgrade/compat). capability-trust.cjs (disclosure/consent, strict_known_registries, engines+compatVersions, reserved namespace) + capability-lifecycle.cjs (install/upgrade/remove/reconcile; ledger-as-commit-point _pending intent; atomic stage-then-swap; surgical marker-isolated shared-edit strip; owner-token lock) + capability-source promote/skipEnginesGate seams + loader pending-skip + config keys. No sandbox re-derived (consent+integrity+reversibility). 6 Codex adversarial rounds + /security-review (no HIGH) + /code-review; gsd-test green both platforms; CI green. Phase 5 (#1434) wires the CLI dispatch. Closes #1433. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1abb0d3427 |
feat(#1432): capability source resolver + ledger (ADR-1244 Phase 3) (#1443)
* feat(#1432): capability source resolver + ledger (ADR-1244 Phase 3) ADR-1244 D3 + D4 — additive, testable-in-isolation modules (the /gsd:capability install command + consent gate are Phase 4). NEW modules only; not yet wired into bin/install.js. - src/capability-source.cts → resolveCapabilitySource(spec, opts): one seam, one adapter per source kind — local (fs copy), git (execGit clone+checkout, https/ssh/ git transports only), npm (execNpm pack --ignore-scripts + tar, NEVER npm install), tarball (https download + sha512 integrity-before-extraction + tar), registry (stub). SECURITY: install never executes capability code (copy/extract only, --ignore-scripts); integrity verified before staging; symlink members rejected at interior, source-root, AND tar-member (verbose-listing) layers; tar-slip member paths rejected pre-extraction; npm specs with shell metacharacters (incl %) rejected (execNpm uses a Windows shell); git ext::/file:// transports + leading-dash/metachar refs rejected; atomic staging (.staging→rename, restore-on-failure); full Phase 1/2 validator suite + engines.gsd pre-check on the fetched manifest. Test seam _setCapabilitySourceHttpGet. - src/capability-ledger.cts → per-runtime .gsd-capabilities.json install manifest: readLedger/writeLedger (atomic via platformWriteSync)/recordInstall (idempotent, prototype-pollution-guarded)/removeEntry/reconcile (orphan report, hardened against hostile/non-string/'..' files[] members — never throws, never oracles outside runtimeDir). - 48 new tests (34 source incl. the full security matrix, 14 ledger). Two Codex adversarial rounds; all 12 findings fixed + regression-tested. - .gitignore + eslint.config.mjs: built capability-source/ledger.cjs git+eslint-ignored (ADR-457 #551 migration coverage); CONTEXT glossary + INVENTORY rows + manifest regen. Closes #1432 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1432): add changeset for capability source resolver + ledger Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
353f63d170 |
feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) (#1440)
* feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) Promote the registry from a frozen data file to loadRegistry({includeInstalled}), composing the first-party registry with a validated installed overlay (ADR-1244 D2): - Extract the conformance validator to a shared runtime-callable module (gsd-core/bin/lib/capability-validator.cjs); the generator re-exports it verbatim, guarded by a generative-parity test (no build-time/runtime drift). - capability-loader.cts: loadRegistry({includeInstalled}) composes first-party ∪ validated overlay from $GSD_HOME/.gsd/capabilities (global) and <root>/.gsd/capabilities (project) via the canonical buildRegistry. First-party always wins (id/skill/agent/config/command-family + reserved gsd-/anthropic- prefixes); full merged-set cross-capability validation; engines.gsd load-time re-gate (skip-with-warning); gate-kind capabilities FAIL CLOSED; fragment-path escapes rejected. - semverSatisfies (hand-written, no dep) for the engines.gsd gate, fail-closed. - Wire surface/state + loop to the overlay; loop injects a blocking gate for each skipped gate-kind overlay (fail-closed). - cwd-aware overlay config-key federation: config-loader _federatedConfigSchema(cwd) + config-schema isValidConfigKey(key, cwd) compose the overlay per loadConfig/ config-set call (never eager at module load, never wrong-cwd); first-party path unchanged with no cwd. - run-tests.cjs sandboxes GSD_HOME (idempotent — nested spawns reuse it) for test hermeticity; capability-loader.cjs git+eslint-ignored (tsc artifact); capability-validator.cjs stays linted (#551 migration coverage). Closes #1431 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1431): add changeset for runtime capability registry overlay Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1431): kill config-schema cwd-aware federation mutants (Stryker ≥52) The cwd-aware overlay config-key federation added to config-schema.cts (_capabilityConfigSchema(cwd) + isCapabilityConfigKey/isValidConfigKey cwd threading) introduced mutable surface uncovered by config-schema's mutation test set, dropping its score to 39.58% (below the 52 break threshold). Add a real-overlay-fixture describe block exercising every branch (cwd guard, overlay loadRegistry, found-branch, first-party fallback, cwd threading); local Stryker score 39.58% -> 77.08%. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
7ca8011cd9 |
refactor(#1373): add canonical markdown-sectionizer seam (epic #1372 T0) (#1381)
* refactor(#1373): add markdown-sectionizer seam (ADR-1372 T0) Establishes the canonical markdown-structure parsing seam per ADR-1372. No existing parsers are modified; this is the foundational T0 tier only. - docs/adr/1372-markdown-sectionizer-seam.md: Accepted ADR defining the seam interface, the tiered migration plan (T0-T7), and the prohibition enforcement approach (no-adhoc-markdown-parsing ESLint rule in T7). - src/markdown-sectionizer.cts: Pure module, Node built-ins only. Exports: stripFencedCode (CommonMark-correct state machine ported from uat-predicate.cts _stripFencedBlocks, CRLF-safe, unterminatedFence signal), tokenizeHeadings (ATX headings outside fenced blocks), collectSections (line-by-line predicate-driven section collection), collectSection (single named section, levelBounded stop, optional stripFences), iterateBullets (dash/checkbox/numbered + continuation). - tests/markdown-sectionizer.test.cjs: 54-test behavioral suite covering the parser QA matrix (LF/CRLF, Unicode headings, headings-inside-fences, unterminated fences, nested levels, all bullet markers, continuation lines, empty/non-string input) plus 4 fast-check property tests (idempotence, output shape, never-throws, length monotonicity). - CONTEXT.md: Markdown Sectionizer glossary entry added (PR review gate). Tests: 54 pass, 0 fail. Existing adr-parser + uat-passed tests: 22 pass. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(#1373): add extractTaggedBlocks + replaceSection to seam; register inventory - src/markdown-sectionizer.cts: extend Section type with bodyStart/bodyEnd offsets; add extractTaggedBlocks(content, tagName) (inner text of <tag>…</tag> blocks, tagName regex-escaped, caller decides fence-stripping) and replaceSection(content, section, newBody) (pure character-offset splice for read-modify-write callers); update collectSections/collectSection to populate bodyStart/bodyEnd. - tests/markdown-sectionizer.test.cjs: add 33 new behavioral tests for extractTaggedBlocks, replaceSection, and a DEFECT.GENERATIVE-FIX parity guard that asserts stripFencedCode and uat-predicate's _stripFencedBlocks agree on a shared 9-item corpus; documents the known 4-space-indent divergence. - docs/adr/1372-markdown-sectionizer-seam.md: list extractTaggedBlocks and replaceSection in §"The seam". - CONTEXT.md: update ### Markdown Sectionizer glossary entry with the two new exports. - docs/INVENTORY.md: add markdown-sectionizer.cjs row (alphabetically between loop-resolver and milestone). - docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest --write. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1373): clear no-unsafe-assignment + unused-var lint in markdown-sectionizer Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1373): correct section offset/round-trip + CommonMark heading/fence edges; register eslint coverage FIX 1 (CRITICAL): Enforce content.slice(bodyStart,bodyEnd) === body invariant in both collectSection and collectSections. bodyEnd is now bodyStart + body.length instead of the raw stop-line offset, eliminating the trailing-newline overcounting that caused replaceSection to drop separator newlines (## A\nbody## B gluing bug). FIX 2 (MED): tokenizeHeadings now accepts ≤3-space indent (CommonMark §4.5) and empty ATX headings (## / ## ), text=''. 4-space indent correctly excluded. FIX 3 (MED): collectSection gains stopAtLevel option — stops at the next heading whose level ≤ stopAtLevel, independent of the opener's level. Enables state.cts ## sections that also stop at ### without abusing levelBounded. FIX 4 (MED): Backtick fence opener info string must not contain a backtick (CommonMark). Applied in both stripFencedCode and tokenizeHeadings fence state machines. Tilde fences unaffected. FIX 5 (LOW): "byte offset" → "character (string-index) offset" in HeadingToken / Section doc comments. FIX 6 (LOW): extractTaggedBlocks doc comment documents nested-tag non-support; test locks the non-greedy close-at-first-</tag> behavior. FIX 7: Add gsd-core/bin/lib/markdown-sectionizer.cjs to eslint.config.mjs ignores so tests/551-eslint-bin-lib-coverage.test.cjs passes (3/3). Tests: 107 pass / 0 fail (was 87; +20 new tests for FIX 1–4, 6). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1373): gitignore tsc-built markdown-sectionizer.cjs (ADR-457 build-at-publish) The seam's compiled artifact must be a build-at-publish output like every other src/*.cts->bin/lib/*.cjs module (decisions, core, state, ...), not a committed file. Add it to the ADR-457 ignore list and untrack it; build:lib/CI regenerate it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
8f428bdd41 |
feat(1259-01): deterministic prohibition-enforcement producer + check route
- Author src/prohibition-enforcement.cts: the test-tier prohibition PRODUCER/gate (ADR-550 D5d heavy half) locate wired check (node-test|lint-rule) -> confirm fail-first -> run -> build enforcementEvidence -> dispositionForProhibition pure/deterministic with injectable runCheck; missing/failing/non-fail-first -> hard-gate (both modes); passing -> green - Route check prohibition-enforcement in src/check-command-router.cts (same family as ui-plan-gate / tdd-review-checkpoint) - Add tests/prohibition-enforcement.test.cjs (behavioral, typed-field, injected runner) — new module within <=2 budget - Register the built bin/lib surface: docs/INVENTORY.md row + regenerated docs/INVENTORY-MANIFEST.json - .gitignore: add the emitted gsd-core/bin/lib/prohibition-enforcement.cjs (build artifact, ADR-457) - No src/probe-core.cts edit — the green/fail-closed policy seam already exists |