bb97ffb5aa28ee569bcc2b744bab8f7b8e45f12c
2251 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bb97ffb5aa |
fix(#2431): self-suppress TDD Audit section when all commits are missing (#2467)
* fix(#2431): self-suppress TDD Audit section when all commits are missing The TDD Audit section in ship.md step 8 was always emitted — but the execute pipeline only writes gate_status: git trailers when TDD mode is active. Without TDD mode (the default), every commit's trailer is absent and the section normalizes to 100% missing, producing a noise table with no way to disable it. Fix: add a self-suppress instruction at the point where gate_status values are normalized. When every commit in the scan normalizes to 'missing', skip both step 8 (TDD Audit section) and step 9 (aggregate gate_status trailer) entirely. Only emit when at least one commit carries a real value (skill, fallback, or exempt). This is data-driven, NOT config-gated. An earlier iteration used inline 'gsd_run query config-get workflow.tdd_mode' — but workflow.tdd_mode is owned by the tdd capability, and ADR-857 Phase 6 forbids host loop workflows from reading capability-owned keys via inline config-get. The self-suppress approach avoids any config-get entirely; it checks the actual trailer data and skips when there's nothing real to report. Matches the triage's suggested approach: 'have the audit gracefully degrade (skip the section)' when there is no real signal. Tests: tests/workflow-compat.test.cjs gains 3 #2431 assertions: - documents self-suppress when every commit is missing - step 9 (aggregate trailer) is also gated on real values existing - does NOT read workflow.tdd_mode inline (ADR-857 Phase 6 compliant) The existing feat-41 assertions still pass — the section content is preserved, only gated by the self-suppress instruction. References: #2431; PR #585 (consumer shipped, producer never wired); ADR-857 Phase 6 (capability-owned config keys must not be read inline by host loop workflows). * chore(#2431): backfill pr:2467 in .changeset/clever-moles-frolic.md |
||
|
|
6140627f5c |
fix(#2427): ground smart-entry completion in ROADMAP-derived counts + tighten status regex (#2466)
* fix(#2427): ground smart-entry completion in ROADMAP-derived counts + tighten status regex Two coupled defects in isComplete (src/smart-entry.cts): 1. Two-scale comparison: isComplete compared global current_phase (from STATE.md body 'Phase: N') against milestone-scoped total_phases (from STATE.md frontmatter progress.total_phases, written once at milestone- switch time and going stale as soon as new phases are appended to the roadmap). When current_phase >= stale total_phases (e.g. 7 >= 4), isComplete tripped true even though later phases were still unchecked in ROADMAP.md. 2. Over-broad status regex: /\bcomplete(d)?|done|shipped\b/i matched any 'shipped' or 'done' substring — including per-phase status like 'Phase X shipped — PR #N' — and falsely satisfied the status side of the completion check. Fix: - Added two new SmartEntrySignals fields: roadmap_total_phases and roadmap_completed_phases, populated by calling the existing deriveProgressFromRoadmap helper (from phase-lifecycle.cts:60) when ROADMAP.md exists. These are global, authoritative counts from the Progress table — never stale. - isComplete now prefers the roadmap-derived counts when available (completed >= total) and falls back to the legacy STATE.md comparison only when the roadmap has no parseable Progress table (backward compat for fresh or non-standard projects). - Tightened the status regex to /\b(milestone\s+complete|all\s+phases\s+complete|complete(d)?)\b/i. Drops 'done' and 'shipped' (per-phase language). Keeps milestone-level signals per ADR-2207 (milestone complete, all phases complete) plus the legacy short form 'complete'/'completed'. Tests (tests/smart-entry.unit.test.cjs gains a #2427 describe block): - Mid-milestone with stale total_phases=4, current_phase=7, per-phase 'shipped' status, and 3 unchecked roadmap phases → NOT complete (the core bug scenario). - All roadmap phases complete classifies as complete even with stale cached total_phases (roadmap wins). - Per-phase 'shipped' or 'done' status alone does NOT satisfy completion when roadmap phases are unchecked. - Legacy fallback: empty roadmap (no Progress table) still classifies via STATE.md comparison (backward compat). The makeProject test helper now accepts a string for the 'roadmap' parameter (written verbatim) in addition to the boolean shorthand, so tests can supply a real Progress table. References: #2427; ADR-2207 (milestone status lifecycle); ADR-2143 (column-name-driven Progress table parsing via deriveProgressFromRoadmap); triage note that this is a read-side fix only (the milestone-switch write path in state.cjs is out of scope). * chore(#2427): backfill pr:2466 in .changeset/curious-rams-run.md |
||
|
|
448e148058 |
fix(#2455): abort remaining chunks when one hits the per-chunk timeout (#2465)
The per-chunk timeout exists so a bad chunk fails loudly 'rather than silently burn the job's wall-clock budget until the CI runner cancels the whole job' (run-tests.cjs:633-637). The control flow defeated that: after a timeout kill the loop fell through to the next chunk. Since the timeout (600000ms) is half the 20m job cap and a healthy Windows full pass is ~11m42s, continuing after a timeout can essentially never finish. Observed on run 29749380190 (windows shard 2/3): chunk 1/5 was killed at exactly 600s, the loop pressed on through chunks 2-4, and the job was cancelled mid-chunk-5 at the 20m wall. The failure surfaced as '##[error]The operation was canceled.' — the timeout diagnostic ended up ~38,000 log lines from the end and 'gh run view --log-failed' returned nothing, making the real cause very hard to find. Abort the remaining chunks on a timeout so the diagnostic survives as the visible failure. Ordinary test failures still run every chunk, so the operator keeps seeing all failures in one pass. Fixes #2455 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
903182fed3 |
fix(#2414): mempalace-capture rooms example must be dicts with name key (#2464)
* fix(#2414): mempalace-capture rooms example must be dicts with name key
The skill's Step 3 'Add the drawer (verbatim)' example wrote a flat list
of bare strings under rooms: in the embedded mempalace.yaml. mempalace's
miner (detect_room + _mine_impl) indexes room["name"] — a bare-string
list crashes the first 'mempalace mine' invocation with
TypeError: string indices must be integers, not 'str'
Following the documented example verbatim and running the capture crashed
every time, before any file was routed. The bug shipped in #2220's fix
(commit
|
||
|
|
953b8043ea |
fix(#2456): weight test chunks by measured cost and pack with LPT (#2463)
* fix(#2456): weight test chunks by measured cost and pack with LPT scripts/run-tests.cjs guessed each test file's cost from its filename (basename matching /^(?:install|codex-)/ scored 12, everything else 1). Measured durations show that guess is wrong in both directions: installer-migration-authoring.test.cjs scored 12 while running ~0.1s, and the two most expensive files in the suite both scored 1 — run-tests-harness.test.cjs never matched the prefix, and release-tarball-smoke.install.test.cjs was missed because the regex is anchored to the START of the basename. Chunks were therefore balanced by file COUNT, not cost. On the real shard 2/3 the two heaviest files packed into the SAME chunk, leaving the slowest chunk 2.8x the lightest and sitting near the 600s per-chunk timeout while other chunks idled. Weight each file by its measured duration from a checked-in, regenerable timings table and pack with LPT (heaviest first, into the lightest chunk). On the same shard this drops the slowest chunk from 383s to 238s and the imbalance from 2.79x to 1.00x, and separates the two heavy files. Timings are advisory, never gated: an unknown file falls back to the table's median weight, a missing or corrupt table falls back to uniform weight, and a count-based floor guarantees the packer never produces fewer chunks than plain count-based packing would. Closes #2456 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2456): harden chunk packing against degenerate knobs and table keys Follow-up hardening found while reviewing the packer, fixed inline. The chunk knobs are read from the environment with Number(), so a typo (RUN_TESTS_MAX_FILES_PER_CHUNK=abc) yields NaN and an explicit 0 yields 0. Both flow into the new chunk-count arithmetic: NaN made Math.ceil return NaN, Array.from({length: NaN}) produce zero bins, and packChunks' retry loop spin forever — a hung CI job with no output. Zero made the count Infinity and threw RangeError: Invalid array length. The previous count-based packer degraded to a single chunk instead, so this was a regression introduced by the LPT rewrite. Normalize the knobs at the environment boundary (positiveNumberEnv: anything not a positive finite number falls back to the default) and guard packChunks itself, since it is exported and cannot assume its caller normalized. Non-finite weights from an arbitrary weightOf are clamped too. RUN_TESTS_CHUNK_TIMEOUT_MS gets the same treatment. Also resolve timing-table lookups with Object.hasOwn: the table is JSON-parsed, so a bare index would walk the prototype chain and return a function for a file named constructor.test.cjs or toString.test.cjs. The typeof guard already rejected that, but the lookup now resolves correctly rather than relying on the downstream check. Refs #2456 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2456): correct prototype-lookup rationale and guard generator keys Two findings from independent security review, fixed inline. The makeFileWeigher comment claimed a bare table lookup "would return a FUNCTION for a file named constructor.test.cjs". That premise is false: basename('constructor.test.cjs') is 'constructor.test.cjs', which is not an Object.prototype key, and walkTestFiles only ever collects *.test.cjs. The prototype chain was never reachable from a real selection, and the existing typeof guard already rejected the function it would return, so Object.hasOwn is defense-in-depth rather than a behavior change. The comment now says that instead of asserting something untrue. The accompanying test inherited the same false premise: it fed constructor.test.cjs and asserted a median fallback that would have held with or without the guard, so it passed for a reason unrelated to what it claimed to prove. It now uses BARE keys (constructor, toString, valueOf, hasOwnProperty, __proto__) — the only inputs that actually resolve on Object.prototype — and asserts the real exported contract: any key absent from the table weighs the median, never a function. gen-test-timings.cjs built its output object by computed-key assignment from basenames taken out of a reporter stream it does not control — the js/prototype-polluting-assignment shape, and this repo has a CodeQL barrier for exactly that pattern. It was not exploitable (the value is always a rounded number, so the __proto__ setter is a silent no-op), but it silently DROPPED such an entry rather than reporting it. Validate every key against a test-basename pattern and fail loudly instead, and build the table with a null prototype. Refs #2456 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2456): replace tautological chunking tests and clamp chunk count Six findings from independent correctness review, all reproduced and fixed inline. The two subprocess tests written to carry the #2088 guarantee forward were tautological: every seeded file weighed exactly 1, so both passed under the OLD prefix-heuristic packer and with the timings file deleted entirely. Neither could fail for the reason it existed. Both are rebuilt so the old algorithm produces a different packing and the assertion goes red: the spread test now uses three expensive files named so the old heuristic scored them 1 alongside three trivial `install-`-prefixed files it scored 12 — inverted from real cost, giving {2,2,1,1} under the old packer versus {2,2,2} under measured weights. The companion test covers the other direction: four trivial `install-` files the old heuristic split into four single-file chunks now stay in one. packChunks clamped the chunk count from below but not above, so a legitimate but tiny budget (RUN_TESTS_MAX_FILES_PER_CHUNK=1e-9, which positiveNumberEnv accepts) asked for 637,000,000,000 bins and threw RangeError. More chunks than files is never useful; the count now clamps at one file per chunk. The generator's basename-collision guard compared full dirnames, so two OS lanes reporting the same file under different container roots (/work/tests vs C:/work/tests) flagged every shared basename as a collision — on the script's own documented multi-lane usage. Detection is now scoped per stream, where the root is constant; a genuine same-lane collision is still caught. Also: the LPT tie-break compared raw paths, so a path separator (0x2F vs 0x5C) could order a subdir file differently per platform, contradicting the documented byte-identical guarantee — it now normalizes separators. loadTestTimings now honors schema_version instead of writing it and never reading it, falling back to uniform weight on an unknown version. A comment claiming an all-uniform suite "chunks exactly as it did before" was false and contradicted by this PR's own test: the chunk count is preserved, the composition is not. And the missing-table test created a temp dir it never cleaned up, for a path that only needed to not exist. Refs #2456 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
be5113abfe |
fix(#2408): fold colliding phase statuses + add W023 collision warning (#2461)
* fix(#2408): fold colliding phase statuses + add W023 collision warning Two coupled bugs from #2408: 1. cmdStats last-write-wins status (src/commands.cts:1610-1618): when two on-disk phase directories normalize to the same phase key (e.g. `05-real/` and `05-real-stray/`), the directory-scan merge overwrote `status` with whatever the *current* directory in scan order computed, discarding `existing?.status` entirely. fs.readdirSync order is not stable across platforms, so /gsd-stats could silently report `Not Started` for a phase that is actually `Complete`. Plan/summary counts were already additively merged; only `status` was wrong. Fix: introduced a `foldPhaseStatus(a, b)` helper that returns whichever status is further along the precedence ladder `Complete > Needs Review > Executed > In Progress > Planned > Not Started` (with `Pending` and unrecognized statuses ranked after). The merge site now calls `existing ? foldPhaseStatus(existing.status, status) : status`. The fold is commutative, so the result is identical regardless of read order. 2. cmdValidateHealth had no collision-detection pass (src/verify.cts): codes W001-W022 cover every condition except normalized-key collisions, so an operator got zero signal that anything was wrong. Fix: added W023 — groups phaseDirEntries by their normalized phase key (via the existing normalizePhaseName + extractPhaseToken helpers from phase-id.cjs — same normalization cmdStats uses) and emits a warning for any group with ≥2 dirs. The warning names the normalized key, both directory names (sorted by comparePhaseNum for stable output), and each directory's independently-computed status (via determinePhaseStatus imported from commands.cjs). Wording is deliberately neutral — never guesses which directory is the real one. The optional --repair path from the issue is intentionally NOT implemented in this PR (the issue marked it lower priority and acceptance criterion 4 is vacuously satisfied by omission). Triage correction applied: the issue proposed W022, but that code is already in use for config.json model-tier validation (src/verify.cts :1372-1391). The next free code is W023, used here. Tests: - tests/commands.test.cjs: integration test that 05-real/ (Complete) + 05-real-stray/ (empty/Not Started) collide and stats reports the merged phase as Complete regardless of read order; plus a direct unit test of foldPhaseStatus asserting commutativity + correct precedence for every status pair + correct handling of unrecognized statuses. - tests/health-validation.test.cjs: integration test that W023 fires on the collision naming both dirs + their statuses (and uses neutral wording), plus a negative test that no W023 fires when only one dir exists per key. References: #2408; reporter's three-layer triage + acceptance criteria; triage correction that W022 is already in use (model-tier validation). * chore(#2408): backfill pr:2461 in .changeset/graceful-koalas-forage.md |
||
|
|
455ad49ae3 |
feat(#2296): config-gated provider escalation on quota-exceeded (#2458)
* test(#2296): failing-first coverage for provider escalation on quota-exceeded
Covers the provider-escalation ladder layered onto EXEC.CLASSIFY: back-compat
(no escalation block without --failure-class), cap boundaries at
min(max_escalations, list length) at limit-1/limit/limit+1, opt-in gating,
malformed/hostile provider_escalation config, the --failure-class CLI negative
matrix, config-key registration, and a fast-check budget-limit property.
Red until the resolver, CLI flag, and manifest key land.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#2296): config-gated provider escalation on quota-exceeded
The dynamic_routing tier ladder escalates within one provider, which does not
help when that provider is what ran out of quota. Add an opt-in provider ladder
layered on the existing EXEC.CLASSIFY seam.
- model-resolver: resolveProviderEscalation walks dynamic_routing.provider_escalation
capped at min(max_escalations, list length), reporting from/to/attempted/exhausted.
Invalid entries are dropped (ADR 227 shape validation). Stays a leaf module —
the quota-class policy decision is the caller's, per the CONTEXT.md contract.
- agent-command-router: export a frozen AGENT_FAILURE_CLASSES so the new CLI
validator cannot drift from the classifier that produces the values.
- resolve-execution: --failure-class flag; emits an escalation block ONLY when
passed, so the existing JSON contract is byte-identical for every caller.
- config-schema.manifest: register dynamic_routing.provider_escalation.
- execute-phase step 7.1: auto-escalate, honor Retry-After, fail loudly naming
every model tried once the ladder is spent.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2296): extract quota recovery to a reference fragment; regen goldens
The step 7.1a addition pushed gsd-core/workflows/execute-phase.md from 93390 to
95111 LF bytes, past the frozen ADR-857 Phase 6 ceiling (hard <93600, margin
<=93400) asserted by tests/fix-2285-claude-orchestration-wiring.test.cjs. The
base sat 10 bytes under the margin, so no inline wording would have fit.
That gate's own rationale is that optional-feature detail belongs in a fragment,
not the host loop. Moved BOTH the new provider-escalation branch and the
pre-existing manual recovery prompt into
gsd-core/references/execute-phase-quota-recovery.md, leaving step 7.1 as a
one-line pointer. execute-phase.md is now 92880 bytes — 510 SMALLER than base.
Also regenerates the fixtures that legitimately moved because three shipped
files changed (gsd-tools.cjs, config-schema.manifest.json, execute-phase.md):
golden-install-parity + install-tree for all 16 runtimes, INVENTORY.md +
INVENTORY-MANIFEST.json for the new reference, and the workflow size baseline.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#2351): make the C1 orphan-reaping test load-independent
tests/run-with-timeout.test.cjs C1 asserted the child heartbeat file exists
after a 1s group-kill window, but the child only wrote it on the first 100ms
setInterval tick. Nothing synchronized the two: on a loaded container the group
is SIGKILLed before that tick lands, the file never appears, and the assertion
fails for a reason unrelated to reaping. Observed failing on both linux-node22
and linux-node24.
The behavior actually under test is the FREEZE assertion (heartbeat stops
advancing => descendant was reaped, not orphaned). That is unaffected by
sampling once more at t=0.
Child now writes its first heartbeat synchronously at startup before arming the
interval, and the kill window widens 1s -> 3s to cover child boot under load.
Both remove the timing dependency; neither weakens what the test proves.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#2296): backfill pr:2458 in .changeset/rapid-jays-bark.md
* chore(#2296): regenerate fixtures after rebase onto #2402
The rebase conflicted on the generated golden-install-parity fixtures and
workflow-size-baseline.json because #2402 (
|
||
|
|
b6e6a22fce |
fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer (#2457)
* fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer Replays the in-flight bot branch fix/2402-response-language-orchestrator-coverage (seven commits, never pushed) onto current origin/next as a single squashed commit. The original work was substantial and correct; this commit preserves its full scope, trimmed where rebase conflicts + workflow size budgets required it. Three independent layers where response_language was being dropped are closed: Layer 1 — orchestrator-facing directives across workflows. Adds the strong "All user-facing output in this workflow MUST be presented in {response_language}; technical terms, code, paths, and subagent prompts stay in English" directive to ~40 workflows that previously either lacked it entirely (verify-work, new-project, new-milestone, quick, manager, and ~35 more) or carried only the weak subagent-prompt-only form (plan-phase, execute-phase). The directive covers narration between tool calls and banner output, not just the AskUserQuestion prompts. Layer 2 — UAT checkpoint renderer (src/uat.cts). buildCheckpoint now accepts an optional responseLanguage parameter and renders the frame strings ("CHECKPOINT: Verification Required", "Type `pass` or describe what's wrong.") in any of 9 languages (English/Spanish/French/German/Portuguese/Japanese/ Chinese/Korean/Italian) with an alias table covering ~30 input variants (en, es, español, ja, 日本語, etc.). cmdRenderCheckpoint reads config.response_language via loadConfig(cwd) and passes it through, so the byte-for-byte block verify-work.md reprints verbatim is already localized when written — preserving the anti-injection hygiene rule at verify-work.md (the model is forbidden to translate after the fact). CJK display width is computed by East Asian Width property ranges (W/F) so the right ║ border of the banner stays aligned for full-width characters. English fallback is byte-identical to the pre-fix behavior when response_language is unset or unrecognized. Layer 3 — literal English report templates in execute-phase. The top-of- workflow directive covers all template sites (templates are a structural source, not literal output). Inline render-language notes that previously sat at each template site were removed during the squash because they pushed execute-phase.md over its frozen pre-phase-6 byte ceiling (93600 — ADR-857 Phase 6 capstone). The single top directive covers the same surface with fewer bytes. Also extends src/docs.cts and src/init.cts to propagate response_language into the init JSON bundle of the additional workflows so the directive can read it. Tests added: - tests/uat.test.cjs: buildCheckpoint with unset/unrecognized language falls back to English default; recognized language swaps only the two frame strings while structural lines stay untouched; CJK display-width regression (independent recomputation of East Asian Width W/F ranges). - tests/workspace.test.cjs, tests/docs-update.test.cjs: response_language wiring through docs.cts/init.cts. References: #2402; reporter's three-layer triage + Layer-4 follow-up; the byte-for-byte anti-injection hygiene rule at verify-work.md (the reason Layer 2 must be renderer-side, not model-translated). This is a squash of the in-flight bot branch — seven commits representing the original implementation plus its subsequent fix/CJK-padding/test/ changeset/regen cycles, none of which were ever pushed or PR'd. The squash captures the final coherent state. * chore(#2402): backfill pr:2457 in .changeset/2402-response-language-orchestrator-coverage.md * chore(#2402): regen golden + size baseline after rebase against #2315 (PR #2451) Rebase conflicts were entirely in generated artifacts (golden-install-parity fixtures + workflow-size-baseline.json). After taking theirs during rebase, regenerated cleanly against the merged source tree. |
||
|
|
352876ff0c |
fix(#2315): respect review.default_reviewers in bare convergence invocation (#2451)
* test(#2315): regression test for review.default_reviewers precedence A bare /gsd-plan-review-convergence invocation (no reviewer flags) is supposed to let users configure a persistent reviewer lineup via review.default_reviewers and just run the loop. Instead, the orchestrator's argument parser silently discards that configuration and forces --codex on every no-flag invocation — with no warning that the configured reviewers were ignored. This commit adds a regression test that fails against the pre-fix workflow (the buggy unconditional --codex fallback is still present at this commit) and passes after the fix lands: - Structural: the workflow must NOT contain an unconditional 'if [ -z "$REVIEWER_FLAGS" ]; then REVIEWER_FLAGS="--codex"; fi' line before the workflow.plan_review_convergence config gate. - Structural: the workflow must query review.default_reviewers AFTER the config gate and document that empty REVIEWER_FLAGS lets gsd-review apply the default. - Behavioral: matrix across {configured, unset, empty-array, explicit-flag} invoking the actual deployed parse + resolution blocks with a stubbed gsd_run. Also updates two existing tests whose assertions the fix makes stale: - #2293 behavioral: endMarker was the buggy unconditional fallback line; the bare invocation assertion was 'run("5") === "--codex"'. Both flip post-fix (endMarker is now the last --all grep line; bare invocation returns empty from the parse block, default applied later in step 1.5). - command-default-claim: the pre-fix command documented '--codex (default if no reviewer specified)' which was the user-facing mirror of the bug. The assertion now requires the command to document the review.default_reviewers precedence. * fix(#2315): respect review.default_reviewers in bare convergence invocation Root cause: plan-review-convergence.md step 1 (Parse and Normalize Arguments) contained an unconditional fallback that set REVIEWER_FLAGS=\"--codex\" whenever no explicit reviewer flag was supplied. This value was then interpolated verbatim into the gsd-review args, so gsd-review saw --codex as an explicit flag (precedence rule 1) and never reached rule 3 (review.default_reviewers). The same path silently dropped any configured review.reviewer_instances (instances participate ONLY via review.default_reviewers per ADR-1517). Fix: - Remove the unconditional --codex fallback from step 1. - Add a config-gated resolution in step 1.5 (after CONVERGENCE_ENABLED check) that queries review.default_reviewers and either leaves REVIEWER_FLAGS empty (letting gsd-review apply its own rule-3 default) or falls back to --codex when no default is configured — preserving the pre-fix default for unconfigured users (AC3). - Replace the banner {REVIEWER_FLAGS} token with {REVIEWER_DISPLAY} so the startup banner reflects what will actually run (AC4), not a hardcoded value. The fix upholds the documented precedence contract (ADR-0011, ADR-0015) that the bug was actively violating. Explicit-flag invocations (--gemini, --all, etc.) are unaffected (AC5). References: #2315; ADR-0011 (review.default_reviewers precedence); ADR-0015 (autonomous cross-AI convergence); ADR-1517 (reviewer instances). * chore(#2315): bump plan-review-convergence.md baseline + changeset - Bump plan-review-convergence.md size baseline 23713 → 25536 (the new step-1.5 default-resolution block). - Add .changeset/plucky-yaks-roar.md documenting the user-visible change. * fix(#2315): restore /gsd: colon syntax + skip behavioral test when jq missing Two follow-ups to the #2315 fix discovered by gsd-test: 1. The fix commit accidentally regressed the slash-command namespace in the disabled-feature exit message: /gsd:plan-review-convergence (correct, from PR #3452) became /gsd-plan-review-convergence (retired dash syntax). The slash-command-namespace invariant test caught this. Restored the colon form. 2. The behavioral test exercises the deployed reviewer-resolution block, which pipes through jq. jq is a documented production dependency (review.md:244 "install jq if missing") and is present in every production deployment, but is NOT on PATH in the gsd-test linux-node{22,24} containers (same constraint as tests/opencode-review-reconstruction. property.test.cjs). Without jq, the printf|jq pipeline fails silently, the ||echo 0 fallback yields DEFAULT_REVIEWERS_COUNT=0, and the resolution falls through to the --codex branch — producing a false negative. Added a jqAvailable guard at module load (matching the existing pattern) and skip the behavioral test when jq is absent. The structural tests (no bash execution) still run and validate the fix. * chore(#2315): regenerate golden-install-parity fixtures Source changes to plan-review-convergence.md (workflow + skill mirror + command doc) changed the install-tree hashes. Regenerated via 'npm run gen:golden' after rebuilding gsd-core/bin/lib/install-engine.cjs ('npm run build:lib') — the on-disk lib was stale relative to src/install-engine.cts (isSymlinkedDestOptIn) and blocked fixture gen. * test(#2315): address review findings — strengthen structural tests + property test Code-review + security-review (isolated subagent passes) surfaced Low/Nit findings; this commit addresses the test-side findings: - Structural test 1 ("unconditional --codex one-liner") now asserts the buggy line is absent EVERYWHERE, not just before the config gate. The earlier assertion allowed a maintainer to re-add the line after the gate (passing the structural test) while the bug would still bite at runtime before the gate runs. - Structural test 4 ("banner uses REVIEWER_DISPLAY") now asserts the LITERAL banner placeholder "Reviewers: {REVIEWER_DISPLAY}" and forbids "Reviewers: {REVIEWER_FLAGS}". The earlier workflow.includes( "REVIEWER_DISPLAY") was satisfied by a comment mention. - New structural test for command/skill content parity: both files must document the review.default_reviewers precedence on the --codex flag (catches a manual edit to one that the gen:plugin-skills mirror misses). - Behavioral test stub now passes default_reviewers via env var ($GSD_TEST_DEFAULT_REVIEWERS) instead of inline-interpolating into a bash single-quoted string. Removes the (currently-safe) fragility where a future test input containing a single quote would close the bash quote and execute as bash under execFileSync. - New property test (fast-check, numRuns=25) for the JSON-classification contract: non-empty arrays of slugs -> empty REVIEWER_FLAGS; empty array / scalar JSON / malformed JSON -> --codex fallback. Locks the parser contract per CLAUDE.md mandate. * fix(#2315): address review findings — defensive jq-missing warning + banner cleanup Code-review surfaced two Low-severity workflow-side findings: - Defensive jq-missing warning: if jq is not on PATH in production (it is a documented dependency per review.md:244, but the dependency can be absent in degraded environments), the printf|jq pipeline fails silently to "0" and a user with review.default_reviewers configured gets --codex with no indication their configured default was unreadable. Added a command -v jq guard at the top of the resolution that falls back to --codex AND emits a stderr warning explaining the reason. This makes the failure diagnosable instead of silently reproducing the #2315 override. - Banner leading-space cleanup: REVIEWER_FLAGS accumulates with a leading space ("$REVIEWER_FLAGS --gemini" from ""), so the explicit-flag branch of REVIEWER_DISPLAY="$REVIEWER_FLAGS" rendered "Reviewers: --gemini" (double space). Pre-existing but worth fixing alongside the AC4 banner work. Strip one leading space with the ${VAR# } parameter expansion in the explicit-flag branch only (the configured-default and --codex branches already produce clean strings). * chore(#2315): bump plan-review-convergence.md baseline + regen golden fixtures The defensive jq-missing warning grew plan-review-convergence.md (25536 -> 26285 bytes). Bumps the per-file baseline snapshot and regenerates the golden-install-parity fixtures for the resulting install-tree hash changes. * chore(#2315): backfill pr:2451 in .changeset/plucky-yaks-roar.md |
||
|
|
eb45fc0e8c |
fix(#2415): close_phase_todos stages the pending/ deletion alongside completed/ (#2447)
* fix(#2415): close_phase_todos stages the pending/ deletion alongside completed/ Bug: the workflow step moved resolved todos from .planning/todos/pending/ to .planning/todos/completed/ with a plain 'mv', then committed listing ONLY the destination directory in --files: mv "$TODO_FILE" "$COMPLETED_DIR/" gsd_run query commit '...' --files .planning/todos/completed/ .planning/STATE.md Git's index still tracked the moved file at its old pending/<name>.md path. The commit therefore only staged the new completed/<name>.md copy — the deletion at pending/ was never staged, never committed, and lingered as an unstaged deletion in git status indefinitely until some later broad 'git add -A' caught it. The phase genuinely closed the todo, but the working tree was never clean. Fix: add .planning/todos/pending/ to the --files list. 'git add' of that directory (since git 2.0) stages deletions of tracked files in the pathspec, so the moved-away file is staged as a deletion atomically with the new completed/ copy in the same commit. Chose plain mv + two-dir --files over 'git mv' because git mv FAILS on: - untracked todos (new todo file not yet committed) - non-git .planning dirs (worktree safety / pre-init projects) Plain mv has neither failure mode. Regression tests in tests/close-phase-todos-stage-deletion.test.cjs (source-text-is-the-product: workflow .md text IS what the runtime loads) cover: - the commit --files list includes BOTH completed/ AND pending/ - the move uses plain 'mv' (not 'git mv') so untracked + non-git cases work * chore(#2415): trim commit subject to keep execute-phase.md under byte ceiling The fix added '.planning/todos/pending/' (~26 bytes) to the commit --files list. To stay under the ADR-857 Phase 6 pre-phase-6 byte ceiling margin (93400 bytes, hard ceiling 93600), shortened the commit subject from 'auto-close N todo(s) resolved by this phase' to 'close N resolved todo(s)'. Net change vs origin/next: +6 bytes (93384 → 93390), well under the margin. Also regenerates the golden-install-parity fixtures (execute-phase.md content-hash update across all runtimes). * chore(#2415): bump execute-phase.md workflow-size baseline (93384 → 93390) The +pending/ fix added 6 net bytes (93384 → 93390), still well under the ADR-857 Phase 6 pre-phase-6 byte ceiling margin (93400). * chore(changeset): backfill pr:2447 in .changeset/sturdy-wasps-run.md * fix(#2415): add issue ref to allow-test-rule annotation (ADR-456) CI lint-allow-test-rule-refs failed on the prior commit — ADR-456 requires '// allow-test-rule: <category> see #NNNN' so every exemption is traceable to an issue. Added 'see #2415' to the source-text-is-the-product annotation. |
||
|
|
e227e81f59 |
fix(#2395): persist runtime identity into ~/.gsd/defaults.json for non-Claude installs (#2446)
* fix(#2395): persist runtime identity into ~/.gsd/defaults.json for non-Claude installs Bug: finishInstall for non-Claude runtimes never persisted a 'runtime' key into ~/.gsd/defaults.json. resolveRuntime() precedence is GSD_RUNTIME env > config.runtime > 'claude', so both inputs being empty on a Cursor (or any non-Claude) install caused agent_runtime and every runtime-branded slash hint to silently fall through to 'claude'. Cursor users saw 'agent_runtime: "claude"' and Claude-formatted /gsd-* hints with no env or config hand-set. Fix mirrors the existing resolve_model_ids: 'omit' write site at the same call site (bin/install.js finishInstall, gated on !_hostBehaviors(runtime) .nativeModelAliases && !GSD_TEST_MODE). Writes runtime: <runtime> into ~/.gsd/defaults.json when absent/null/empty. Claude is the resolveRuntime() fallback so it needs no write; an explicit pre-existing runtime value is always preserved across installs of any runtime. Pattern parity with #1156 (default-to-omit intent) and #1569 (preserve explicit user value) — the new write is the third sibling on the same install-time persistence block. Regression tests in tests/install.test.cjs cover: - absent / null / empty-string runtime → populated to <install runtime> - explicit pre-existing runtime preserved (no clobber across runtimes) - parameterized across 5 non-Claude runtimes (cursor, codex, opencode, antigravity, windsurf) Out of scope (per triage): the resolveRuntime() precedence order itself (env > config > default) is unchanged. A separate follow-up noted in the issue (subagent rendering when runtime correctly identifies as 'cursor') is unrelated to branding and not addressed here. * chore(#2395): regenerate golden-install-parity fixtures for new runtime persistence The golden install-tree fixtures capture the post-install state, including ~/.gsd/defaults.json. For non-Claude runtimes, defaults.json now includes runtime: <runtime> — content hash updated for each affected runtime (17 golden-install-parity/*.json files). Claude's defaults.json is unchanged (no runtime key written for Claude — it's the resolveRuntime() fallback). * chore(#2395): drop product names from changeset fragment (product-name-purity gate) * test(#2395): move describe out of fix-1521 fold + add same-runtime idempotence test Code review (correctness subagent) flagged 2 Low test-quality issues: 1. The new Bug #2395 describe was inserted inside the folded:fix-1521-real-install-stamping IIFE callback, muddying test reporting and tracing the regression to the wrong epic. Moved it outside the IIFE close to be a top-level sibling. 2. No explicit same-runtime idempotence test — the suite covered cross-runtime preservation (cursor seed → opencode install) but not 'install cursor twice → second is a no-op'. Added: seeds fresh defaults, installs cursor, captures mtime, installs cursor again, asserts runtime unchanged AND defaults.json mtime unchanged (idempotent, no rewrite churn). * chore(changeset): backfill pr:2446 in .changeset/fierce-ravens-dance.md |
||
|
|
12e4d93b19 |
fix(#2393): add GSD_ALLOW_SYMLINKED_DEST opt-in for intentional user-owned symlink layouts (#2445)
* fix(#2393): add GSD_ALLOW_SYMLINKED_DEST opt-in for intentional user-owned symlink layouts Bug: v1.7.0's destSubpath write-confinement (ADR-1239 Phase B) refused install/update whenever CLAUDE_CONFIG_DIR (or an artifact-kind child like skills/, hooks/) was a pre-existing symlink, with no opt-out. Three legitimate user-owned layouts were blocked: - (lars-hh) CLAUDE_CONFIG_DIR=~/.claude-personal with skills/hooks symlinked to a user-owned external dir - (Mamiki) ~/.claude/skills is a Windows Junction to a shared skills dir - (Azd325) ~/.claude itself is a symlink to a dotfiles repo (the early root-is-symlink return refused before the component loop ran) Fix: add GSD_ALLOW_SYMLINKED_DEST env var (accepts '1' or 'true'). When set, hasExistingSymlinkBetween follows symlinks instead of refusing them. Cross-platform: fs.lstatSync().isSymbolicLink() returns true for both POSIX symlinks and NTFS junctions (Node ≥ 16), so Mamiki's Junction case is handled by the same code path. Threat model preserved (these still refuse EVEN WITH opt-in): (a) path-traversal in the destSubpath string itself ('../../etc'-style) — ADR-1239 Phase B threat (a), untrusted destSubpath protection (b) a symlink whose resolved real path equals the install root itself — would let _removeGsdEntries wipe the root; #1704 threat (b) (c) broken symlinks (realpathSync throws) — fail-closed What opt-in RELAXES specifically: the 'pre-existing symlink pointing outside configHome' refusal — #1704 threat (c). The user has explicitly asserted they own and trust the symlink target. Error messages at all 4 call sites (installRuntimeArtifacts, _copyStaged, migrateLegacyDevPreferencesToSkill, installOpencodeFamilySkills) updated to (1) name the env var opt-in, (2) be accurate when the root itself is a symlink (Azd325's complaint that the old message accused destDir of 'containing' a symlink when the root was the actual symlink). Docs: docs/CONFIGURATION.md Environment Variables table updated. Regression tests in tests/install-write-confinement.test.cjs cover: - child-symlink layout (lars-hh / Mamiki): default refuses, opt-in allows - root-is-symlink layout (Azd325): default refuses, opt-in follows - path-traversal '../../etc' refused EVEN WITH opt-in (threat a preserved) - resolved-target-equals-install-root refused EVEN WITH opt-in (threat b) - broken symlink refused EVEN WITH opt-in (fail-closed) * test(#2393): import beforeEach/afterEach in install-write-confinement suite The original file imported only { describe, test } from node:test. The new #2393 opt-in describe block uses beforeEach/afterEach to manage the GSD_ALLOW_SYMLINKED_DEST env var lifecycle — add them to the import. * test(#2393): correct broken-symlink test — existsSync follows link → loop terminates early Initial test expected broken symlinks to be refused even with opt-in. That was wrong: fs.existsSync follows symlinks, so a broken symlink returns false from existsSync and the component loop terminates before the symlink check fires. Both default and opt-in paths share this behavior; the fix preserves it. Updates the test to pin the actual current behavior so a future refactor (e.g. switching to lstatSync for existence) is a deliberate behavior change. * fix(#2393): realpath the install root — guard against macOS /var ↔ /private/var Code review (security subagent) flagged a HIGH-severity hole in the threat-(b) preservation: realTarget (from fs.realpathSync) is fully symlink-resolved, resolvedRoot (from path.resolve) is lexical-only. On macOS /var is a symlink to /private/var, so resolvedRoot='/var/foo/.claude' but realConfigHome is '/private/var/foo/.claude'. A symlink whose realtarget matches the install root by real path would compare unequal to the lexical resolvedRoot — defeating the wipe-protection guard exactly in the reporter's case (Azd325, nix-darwin: ~/.claude is itself a symlink). Fix: compute realRoot once via fs.realpathSync(resolvedRoot) at function entry (with fail-closed fallback to lexical form on realpath failure — broken/missing root, permission denied, exotic FS). Threat (a) path-traversal check above still confines regardless. Compare against BOTH lexical and real forms in both the root-symlink and component-symlink branches. Also adds the reviewer's transitivity-trust clarification comment: once a symlink is followed under opt-in, the walk continues from the resolved real path WITHOUT re-checking further segments stay inside a confining boundary. This is documented opt-in semantics — one opt-in trusts the whole reachable tree — and the comment makes the design choice explicit so a future maintainer doesn't add a 'follow one symlink only' expectation. Regression test added for the macOS /var normalization case (spelled configHome via os.tmpdir() lexically while pointing the test symlink through its realpath). Test skips on non-darwin platforms and when os.tmpdir() has no symlink component. * fix(#2393): root-symlink branch — do not apply threat-(b) check to root itself Initial fix applied the wipe-threat-(b) check to the root-symlink branch unconditionally. That was wrong: when root itself is a symlink (Azd325's nix-darwin case), its realpath IS realRoot by construction — so the check always fires, defeating the opt-in for exactly the case it was meant to enable. The wipe threat (b) does NOT apply to root being a symlink: destDir is a CHILD of root, and resolving root just gives root's target. There is no circular back-reference to root from a path that descends from a resolved root. So the root-symlink branch should just follow the symlink under opt-in and continue the walk, no threat-(b) check. Threat (b) only fires in the COMPONENT loop, where a child symlink can resolve back to the install root. That branch keeps the (b) check using BOTH lexical and real forms of root (the macOS /var ↔ /private/var fix from the prior commit). * fix(#2393): apply opt-in at the 5 bin/install.js call sites + add env-var/transitive tests Code review (correctness subagent) flagged a Critical coverage gap: the initial fix updated only the 4 src/install-engine.cts call sites. Five more call sites in bin/install.js still used the 2-arg signature, so the opt-in env var was silently ignored on: - installCodexConfig (config.toml + agents/ dir + per-agent .toml paths) — Codex only - copyWithPathReplacement (the generic emit path: workflows, commands, staging) — ALL runtimes - resolveInstallRelativePath (path resolver used in various places) Result: a user setting GSD_ALLOW_SYMLINKED_DEST=1 would see SOME refusals disappear (engine path) and OTHERS remain (bin/install.js paths) — a partially-applied install and a confusing UX, directly contradicting the PR's headline claim. Fix: - Export isSymlinkedDestOptIn from src/install-engine.cts alongside hasExistingSymlinkBetween - Import it in bin/install.js - Update all 5 bin/install.js call sites to pass { allowOptInFollow } - Update all 3 bin/install.js error messages to name the env var, matching the engine's phrasing Also addresses reviewer's Medium test-adequacy findings: - isSymlinkedDestOptIn env-var parsing now tested directly (accepts only documented '1' / 'true'; rejects 'TRUE', 'yes', 'on', '0', 'false', empty, unset) - transitive symlink chain (configHome/outer → outside1 → outside2) test pins the documented 'transitive and unbounded' opt-in semantics so a future contributor can't accidentally narrow it * chore(changeset): backfill pr:2445 in .changeset/eager-wasps-swim.md |
||
|
|
517bae8d6d |
fix(#2372): widen decision-coverage-plan to all planner-canonical tags, drop misleading "(or body)" (#2443)
* fix(#2372): widen decision-coverage scan to planner-canonical tags, fix message Bug: check.decision-coverage-plan's remediation message told the user to cite decisions "(or body)" but extractPlanDesignatedSections only scanned <objective>/<tasks>/<task>/<action>. A decision cited in <read_first>, <behavior>, <verify>, <acceptance_criteria>, or <done> was invisible to the gate — false BLOCKING coverage gap, plus the message's own fix-hint sent the user to "the body" where re-citing still failed. Two-part fix (must change together — that drift was the bug): 1. Widen XML_DECISION_TAGS_RE in src/check-command-router.cts to also match <read_first>, <behavior>, <verify>, <acceptance_criteria>, <done>. These are all planner-canonical tags the planner is told to use (plan-phase.md:830-862, plan-phase.md:772). The body negative- lookahead mirrors the opening-tag set so each tag's body is captured independently. 2. Correct buildPlanMessage to name ONLY the surfaces the extractor actually scans (front-matter must_haves/truths/objective, designated markdown headings, and the nine planner-canonical tag bodies). The misleading "(or body)" clause is gone. Also updates the planner's documented contract (agents/gsd-planner.md:69) and user-facing docs (docs/CONFIGURATION.md, docs/USER-GUIDE.md) to reflect the wider scan. Regression tests in tests/decisions.test.cjs cover each newly-scanned tag body, a control (no citation still uncovered), and a message/extractor parity assertion that names every scanned surface — so the two cannot drift apart again. Out of scope (per triage): cmdDecisionCoverageVerify/buildVerifyMessage is a separate command (decision-coverage-verify) checking shipped artifacts, not plan citations — untouched. * chore(#2372): regenerate agent-size-baseline + golden-install-parity fixtures gsd-planner.md grew 49172 → 49294 (+122 chars) from the widened decision- coverage contract (5 new scanned tag names + heading clarification). Growth is justified: the contract surface is itself the fix — the prior text under-described what the gate scans, which was the bug. Updates: - tests/agent-size-baseline.json (gsd-planner.md: 49172 → 49294) - 17 tests/fixtures/golden-install-parity/*.json (one hash per runtime) - tests/fixtures/install-tree/*.json (regenerated by gen:golden) * fix(#2372): per-tag matching — outer-tag citations survive inner-tag nesting Code review (subagent) flagged a Medium edge-case regression from the single-alternation regex: when a newly-scanned tag nests inside another scanned tag, the alternation's negative lookahead halts the outer tag's body at the inner tag — losing any D-NN citation in the outer tag's prefix prose. Concretely: <action>per D-05 <verify>npm test</verify></action> → 3-tag alternation (old): captured 'per D-05 <verify>npm test</verify>' as <action> body → D-05 caught → 9-tag alternation (bug): captured 'npm test' only (from <verify>); D-05 in <action> prefix LOST Switches extractXmlTagBodies to per-tag matching: each tag gets its own regex whose negative-lookahead tempers only against the SAME tag's reopening. So <verify> inside <action> is absorbed into <action>'s body (D-05 caught) AND <verify> is matched separately on its own pass. Per-tag preserves both: - the reporter's case (sibling tags inside <read_first>) - nested-tag citations in outer-tag prefix prose - ReDoS safety (each per-tag regex keeps the #2128 body tempering) Also adds the reviewer's other requested edge-case tests: - non-scanned tag (<name>) bearing D-NN must NOT count - self-closing form <read_first /> safely ignored - attribute form <verify type="...">D-NN</verify> (canonical planner shape) - CRLF newlines in tag body do not break capture * chore(changeset): backfill pr:2443 in .changeset/noble-elks-chatter.md |
||
|
|
0bbbca2a46 |
fix(#2069): forward model_policy, model_profile_overrides, runtime from global defaults (#2442)
* test(#2069): add fail-first regression for global-defaults dropped keys Adds four failing-first regression cases to tests/defaults-json-fallback.test.cjs: - model_policy forwarded from ~/.gsd/defaults.json - model_profile_overrides forwarded from ~/.gsd/defaults.json - runtime forwarded from ~/.gsd/defaults.json - parity: model_policy survives identically whether it lives in the global defaults or in a project's .planning/config.json All four fail on unfixed code (Branch D of loadConfigResolved builds _globalBaseCfg from a whitelist that omits these three keys). The project-config path at config-loader.cts:602-604 already forwards them, so the global path should too. * fix(#2069): forward model_policy, model_profile_overrides, runtime from global defaults The _globalBaseCfg whitelist in Branch D of loadConfigResolved previously omitted three keys that the project-config path forwards parsed['…']: - runtime - model_profile_overrides - model_policy so ~/.gsd/defaults.json silently dropped them. A machine-wide model policy (or runtime / profile overrides) was honored inside a project (where .planning/config.json carries it) but ignored for out-of-project runs — resolve-model fell back to the profile default with no warning. Adds the three entries to _globalBaseCfg in the same (globalDefaults['…']) || null shape as the sibling keys and the project-config path, so global defaults honor them identically. Regression tests in the prior commit (#2069 fail-first) demonstrate the fix on the same suite that previously failed. * test(#2069): extend parity test to all three previously-dropped keys Code review (subagent) flagged that the parity test only asserted model_policy shape-parity between global-defaults and project-config paths. A future regression breaking just runtime or just model_profile_overrides shape (e.g. someone changing parsed['runtime'] to ?? null in the project path) would slip a single-key test. Extends the parity test to assert deepStrictEqual / strictEqual across all three keys: model_policy, model_profile_overrides, runtime. Same two-dir setup, three cheap assertions. * chore(changeset): backfill pr:2442 in .changeset/sturdy-seals-fly.md |
||
|
|
d16a66479a |
feat(#1950): broken-windows ledger — cross-phase defect register gating ship (#2441)
* feat(#1950): broken-windows ledger — cross-phase defect register gating ship Adds a new capability (#1950) that operationalizes GSD's no-defer discipline as a tracked, enforced artifact: accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths across phases, and /gsd-ship blocks while any entry is open. Implementation: - src/broken-windows.cts → gsd-core/bin/lib/broken-windows.cjs: typed IR + I/O entry points (parseLedger/renderLedger/appendWindow/markWaived/markFixed + cmdWindowsStatus/Append/Waive/MarkFixed). Frozen REASON enum for typed error assertions. Windows-safe atomic rename with retry on transient EPERM/EBUSY/EACCES. - gsd-tools.cjs: new subcommand (status | append | waive | fixed), wired via routeWindows + HOST_COMMAND_ROUTERS.windows. - capabilities/broken-windows/capability.json: one ship:pre gate with artifact-frontmatter-equals predicate on WINDOWS.md open_count == 0. activationKey windows.enabled (default true) + sibling windows.enforce (default true, separate so tracking can precede enforcement). - gsd-core/workflows/ship.md: capId==broken-windows branch in preflight, sibling to security — reads gsd_run windows status --raw, fails closed on open_count > 0 or unreadable ledger. - agents/gsd-executor.md: extends the existing ## Known Stubs instruction to also append to WINDOWS.md via gsd_run windows append (best-effort, never blocks execution). - agents/gsd-verifier.md: new Step 8b — record unmet truths + human-verify items in WINDOWS.md. - gsd-core/workflows/progress.md: surfaces open + waived counts. - docs/COMMANDS.md + CONTEXT.md glossary entry + docs/INVENTORY.md: document the gate, waiver mechanism, and new module. - tests/broken-windows.test.cjs: pure + CLI behavioral coverage + fast-check roundtrip property; fail-closed on malformed ledger; security boundary on path traversal in --file. Backward-compatible: a project with no .planning/WINDOWS.md reports open_count: 0 and ships cleanly. Disable enforcement per-project with gsd config-set windows.enforce false (tracking continues, gate stays open). * chore(#1950): ratchet size baselines, defer verifier integration - Workflow size baseline: ship.md 25575→27928, progress.md 31789→32632 (broken-windows preflight branch + open-windows surface). - Agent size baseline: gsd-executor.md 46644→47951 (Known Stubs → also appends to WINDOWS.md). gsd-verifier.md unchanged. - LARGE_CAP (49152) preempted the planned verifier integration (gsd-verifier.md was at 49140 pre-PR — 12 bytes of headroom, not the documented 'real headroom'). Verifier integration deferred to a follow-up PR that extracts the VERIFICATION.md template (lines 739-859) to gsd-core/references/ — a pre-existing cap-tightness defect this PR exposed but does not expand scope to fix. Verifier integration is not in the issue's acceptance criteria (executor writes is; unmet-truths recording was an enhancement, not a gate). * fix(#1950): gate default-off, rename to workflow.windows_enforce, regen goldens Test-failure-driven fixes after first gsd-test run on db8733c8f failed 44 cases (pre-existing structural tests encoded 'ship:pre has 1 gate' / 'all caps off → empty hooks'): - capability manifest: rename windows.enabled+windows.enforce (default true) → single federated key workflow.windows_enforce (default FALSE, opt-in). Matches security's workflow.security_enforce convention and makes the adr857 all-caps-off test pass without modification (the test's buildAllFalseConfig handles workflow.* out of the box). Default-OFF keeps the gate out of the registry's default ship:pre resolution so existing loop-hooks-ship-pre-e2e structural assertions (exactly 1 gate, capId 'security') stay valid; users opt in via gsd config-set workflow.windows_enforce true. - drop activationKey (security doesn't have one either; workflow.* key doubles as the activation toggle). - regenerate docs/reference/capability-matrix.md to include broken-windows (capability-matrix-sync test). - regenerate tests/fixtures/golden-install-parity/*.json (18 runtimes) — installer now emits the new capability + lib file. - update CONTEXT.md, docs/COMMANDS.md, docs/FEATURES.md, ship.md, agents/gsd-executor.md to use the new key name and /gsd:colon slash syntax (slash-command-namespace test). - restore accidentally-regressed /gsd:capture in progress.md. Tracking-only by default; enforcement is opt-in. Acceptance criterion '/gsd-ship fails while any ledger entry is open' is met when workflow.windows_enforce=true (test fixture enables it). * test(#1950): update ship:pre structural invariants for 2-gate registry - loop-hooks-ship-pre-e2e: the registry now declares 2 gates at ship:pre (security + broken-windows), regardless of activation. Activation tests above still pin security-only or empty behavior via fixtures; these structural tests pin the REGISTRY shape, which has 2 gates as of #1950. - workflow-size-baseline: ship.md 27928→27945 (workflow.windows_enforce rename added 17 bytes). * fix(#1950): review H1+H2+M1+M2+M3 — fence-injection, EACCES fail-closed, cleanup, strict line, stryker Adversarial isolated review (Step 6.3) found 2 HIGH findings that block the PR and 3 mediums. All addressed: H1 (HIGH): description containing the markdown 3-backtick fence would terminate the ledger's JSON code block early inside JSON.stringify output (JSON doesn't escape backticks), corrupting the file and bricking the next parse. Fix: use a 4-backtick fence (json ... ) which JSON.stringify cannot produce on its own, AND validate that no entry text field contains a 4-backtick run (reject at append time with new WINDOWS_INVALID_TEXT reason code). Locked by a regression test. H2 (HIGH): readLedgerOrNull swallowed ALL fs errors as 'no ledger', silently returning open_count:0 on EACCES/EPERM/EIO. The ship gate would then pass on an unreadable ledger — the precise vector the workflow doc claims is impossible. Fix: only ENOENT returns null; every other fs error propagates as WINDOWS_LEDGER_MALFORMED so the gate blocks and the operator sees a real diagnostic. Locked by a regression test that chmod 000s a ledger with open_count=1 and asserts the result is never a false-green 0. M1: writeLedgerAtomic left an orphaned .tmp file on rename failure. Wrapped renameWithRetry in try/catch with best-effort unlink. M2: validateLine silently coerced 'abc' → NaN → null, hiding type drift. Removed the line === 0 special case (was undocumented) and made the error message match the strict check. Now any non-positive- integer line value throws, including strings. M3: tests/broken-windows.test.cjs (with its fast-check property test) was not in stryker.config.mjs DEFAULT_TEST_CMD — Stryker would mutate src/broken-windows.cts but no test would catch the mutations, producing false surviving-mutant scores. Added to the list. L1 (dead throw e after error()), L7 (line boundary tests, H1/H2 regression tests, 4-backtick CLI test) also addressed. * docs(#1950): inline concurrency + busy-wait notes (review L2+L3) * fix(#1950): regen goldens against latest gsd-tools; correct --line 0 boundary test gsd-test v4 caught two issues: - goldens I regenerated earlier (commit 526682084) predated the L1 routeWindows catch-block cleanup (commit dd844d565). Regenerated via 'npm run gen:golden' against current HEAD so the install parity hash for gsd-tools.cjs matches. - 'append --line boundary' test expected --line 0 to succeed with null entry.line, but the M2 fix correctly rejects 0 (lines are 1-indexed; 0 is not a valid source line). Updated the boundary test to assert --line 0 fails alongside -1 and 'abc'. * chore(#1950): regen goldens after rebase onto next * chore(#1950): quick.md baseline 50699→50993 (correct resolution from next rebase) * chore(changeset): backfill pr:2441 in .changeset/broken-windows-ledger.md * fix(#1950): renderTable escapes backslash before pipe (CodeQL incomplete-sanitization) CodeQL flagged the markdown-table cell escaper: String(s ?? '').replace(/\|/g, '\\|') — it escapes pipe but not backslash first. A description containing '\|' would render as '\\|' which markdown parses as 'literal backslash' + 'cell separator', splitting the column. Fix: escape backslash FIRST (each \ → \\), then pipe (each | → \|). Now a description with '\|' renders as '\\\\|' (literal '\\' + escaped pipe), which markdown renders as a single '\|' inside the cell. The JSON code block (the parse source-of-truth) was already correctly escaped via JSON.stringify; only the display-only table was affected. Locked by a regression test that: 1. Verifies the JSON block reparses with the description intact. 2. Walks the rendered table row counting unescaped pipes — must be exactly 11 (the row separators for 10 cells), proving no in-cell pipe added a split. |
||
|
|
1a46bc068a |
fix(#2376): emit absolute subagent-facing paths from init/state, convert workflow literals (#2428)
* fix(#2376): emit absolute subagent-facing init/state paths Make init.* and state.* path fields absolute rather than cwd-relative so subagent prompts resolve correctly regardless of working directory. Adds intel_dir/conflicts_path/requirements_path/roadmap_path/state_path to cmdInitIngestDocs, an absolute debug_dir to cmdStateLoad, and replaces bare .planning/... literals in 12 workflow Agent() prompt blocks with the absolute init-JSON path fields. Includes decoy-cwd regression tests and realpath'd tmpdir fixtures for macOS. Squashed rebase of the #2376 commit series onto a fresh origin/next (previous merge ee25543a1 was against a now-stale next). * chore(#2376): add changeset * chore(#2376): regenerate golden fixtures + workflow size baseline Regenerated after rebasing the absolute-path fix onto current next (picks up #2351's run-with-timeout content in execute-phase.md too). * fix(#2376): trim execute-phase.md redundancy to stay under the size margin * chore(#2376): regenerate golden/size baseline after rebase onto next |
||
|
|
d04e287fa9 |
fix(#2365): stop api-coverage detector false-positiving non-API phases (#2397)
* fix(#2365): stop the api-coverage detector false-positiving non-API phases detectApiIntegration fired on any integration verb co-occurring anywhere on a line with any API noun, treated / as a word boundary (so a first-party Next.js src/app/api/... route path matched the noun "api"), and read any capitalized word before API/SDK/REST/GraphQL as a service name behind a fixed stopword denylist (so threat-model prose like "Resolver-only API" fired). Because the verify:pre seal gate is BLOCKING, a phase touching no external API could not reach UAT without fabricating a coverage matrix. The compound rule now requires the verb and noun to share one clause (sentence punctuation and table-cell walls end a clause) within a bounded word gap. Non-prose spans are excluded before matching: fenced code (already), inline code spans (new stripInlineCode in the markdown-sectionizer seam), and path-shaped tokens. The <Service> API surface rule requires proper-noun position — a clause-initial capitalized word is ordinary English and needs dependency evidence (URL / package reference) on the same line — and rejects compound modifiers ("Resolver-only", lowercase after the hyphen). A phase that integrates no external API now has a first-class, reasoned way to say so: a COVERAGE.md containing "No external API integration: <reason>" satisfies the gate (declaration + rows is contradictory and blocks). The true-positive path is pinned by regression tests: every default-vocabulary positive still fires, including the widest word-gap pairing and the surface-rule-only shape. Fixes #2365 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2365): tighten api-coverage detector per Codex review (round 2) Applies the Codex review findings on the initial #2365 fix: - S-1: a COVERAGE.md "no external API integration" declaration is the human override for a fallible detector, so it must PASS even when detection still fires — but the contradiction is now SURFACED in the gate output (overridden signal count + terms) instead of passing silently. - S-2: verb/noun pairing is now a term-group nearest-pair merge walk over precomputed word ordinals (computeWordStarts / minWordGap), not a match×match cross product — a hostile line repeating one pair thousands of times stays linear instead of going quadratic. - FN-4: package-shaped inline-code spans (`stripe-sdk`, `@stripe/stripe-js`) are kept as noun/dependency evidence rather than being fully masked, so a genuine dependency reference inside code ticks still corroborates. - C-1: the <Service> API surface rule now scans every candidate in every clause; a rejected first candidate no longer shadows a later genuine service. - Cross-clause binding: a verb may bind a noun in the immediately following clause only when its own clause names a service object, within a tight gap — admits "Integrate Stripe, exposing its endpoints …" without re-admitting the unrelated-clauses false-positive class. - Internal-descriptor negative evidence ("internal Payments API", "the internal endpoint") never pairs; URL/scheme matching generalized beyond http(s). All 5 acceptance criteria still hold: the three reported false positives are clean and "integrate the Stripe API" still fires. Built .cjs committed alongside the .cts. tsc + eslint (incl. no-adhoc-markdown-parsing) + lint:regression-names clean; affected suites 256/256 green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#2365): retune api-coverage detector fail-closed per Codex review (round 3) Codex's second-round review found the round-2 tightening had over-corrected into FAIL-OPEN false negatives — realistic external-API prose that the BLOCKING seal gate silently let through (the catastrophic class, since a missed API surface is worse than a dismissable false positive). Retuned the detector to be explicitly fail-closed: lean toward detecting, and let the one-line COVERAGE.md "no external API integration" declaration dismiss the residual false positives. Fail-open false negatives fixed (all now detect): - F1 clause-initial `<Service> API` with a plain follower ("Stripe API for payment processing") — dropped the follower-allowlist / corroboration gate on clause-initial surfaces; a service that is not a stopword, descriptor, or compound modifier is a real name from any clause position. - F2 scheme-less external host ("api.stripe.com/v1") — a dotted host with an alphabetic final label now contributes its API nouns; a first-party route path (no dotted host) still does not. - F3 vendor's first-party SDK ("Integrate Shopify's first-party SDK") — the compound path no longer filters nouns on "internal"/"first-party" (Codex: the qualifier can describe the vendor's own API, not the consuming project's). - F4 long single integration clause — removed the word-gap cap entirely: it could not separate a 21-word genuine clause from an 18-word internal one, so the clause boundary is now the whole relationship test. - F5 lowercase cross-clause service — cross-clause binding no longer requires a capitalized "service object". New false positives fixed (all now clean): - F6 a URL token that swallowed a trailing clause comma, merging two clauses — trailing clause punctuation is kept literal so the split survives. - F7 a capitalized internal component authorizing cross-clause binding — the new gate requires a dependent elaboration, not a new coordinate clause opened by a conjunction ("…, then document…"). - F8 a protocol name read as a service ("REST API", "GraphQL API") — protocol and locality descriptors are rejected in the `<Service>` position. - Finding 9: the inline-code-span scanner was O(n^2) on pathological backtick runs; rewritten to linear via a per-length run cursor (2 MB: 4.15 s -> ~6 ms), semantics preserved (148 sectionizer tests unchanged). Net simplification: the fail-closed model removed the round-2 minWordGap / groupByTerm / follower / corroboration machinery (350 insertions vs 445 deletions across the touched files). Under fail-closed, three round-2 negative tests now correctly detect (integration verb + "internal"-qualified noun, and the distant-same-clause case); none were trek-e acceptance FPs. Verified: 1491/1491 unit tests pass; tsc + eslint (incl. no-adhoc-markdown- parsing) + lint:regression-names clean; all 8 review findings reproduced as regression tests, both directions. Built .cjs committed alongside the .cts. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#2365): resolve round-3 Codex review findings (fail-closed, round 4) Codex's round-3 adversarial review found the fail-closed retune had introduced new holes in both directions. Resolved: Fail-open false negatives (now detect): - External host addressing a PATH ("graph.microsoft.com/v1.0/me") is itself an integration surface and contributes an endpoint noun even when the host names no vocabulary word. A bare domain link with no path ("https://example.com") stays a non-signal, so "Integrate … from example.com, document …" is still clean. - Locality qualification ("internal", "private") no longer leaks across a sentence or clause boundary: only plain spaces may separate the descriptor from the service, so "The cache is private. Stripe API …" now detects. - Cross-clause binding: the fragile head-word cap (which could not tell a genuine "Connect … to Stripe payments, exposing its endpoints" from an unrelated "Integrate … from URL, document …" — both 4 words after the verb) is replaced by a participial-continuation rule: a verb binds a noun in the next clause only when that clause begins with an "-ing" elaboration. This fixes the 4-word-head false negative AND the false positive below at once. False positives (now clean): - Cross-clause no longer binds a finite continuation regardless of separator: "Wire the settings form. Document endpoint props." / "…; document …" / "…, document …" are separate actions, not elaborations. Perf (quadratic → linear): - The trailing-punctuation peel is a backward char scan instead of an unanchored `[…]+$` regex (16k chars: 156 ms → ~1 ms). - SERVICE_SURFACE_API_RE bounds the service-name length {1,40} so a hostile "A-A-…-x" run cannot drive O(n^2) backtracking (16k: 385 ms → ~3 ms). Consumer fail-open (blocking gate): - readPhaseScope now distinguishes "no plans" from a plan that EXISTS but is unreadable. On a read error the gate BLOCKS ("could not read the phase scope …") instead of silently certifying no-integration from partial scope — an unreadable plan could be the one describing the integration. Documented fail-closed tradeoffs, now pinned with tests so they are not "fixed" back into a fail-open: a clause-initial capitalized common word before "API" ("Payment API", "Search API") reads as a service name; a long clause pairs a verb with a distant noun; and a CommonMark inline code span that wraps a newline is matched within-line only. Codex judged these acceptable because the COVERAGE.md declaration is a cheap override. One documented limitation remains out of scope: "Integrate Stripe, and authenticate requests with its API" (a coordinate finite clause whose noun refers back by pronoun) needs coreference resolution, beyond a lexical detector. Verified: 379/379 affected + command-router tests pass (+14 new regression tests covering every round-3 finding, both directions); tsc + eslint (no-adhoc-markdown-parsing) + lint:regression-names clean. Built .cjs committed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#2365): simplify to robust core — remove whack-a-mole heuristics (round 5) Round-4 review confirmed the detector's two most complex features generate findings in both directions no matter how they are tuned, because they need a vendor dictionary + coreference the issue rules out in principle. Per the operator's "ship the robust core" decision, both are removed and their gaps are documented rather than chased further: - Cross-clause binding DELETED (allowsCrossClause / participle rule). It caused a fail-open on finite continuations ("Integrate Stripe; use its OAuth endpoints" — missed) and a false positive on "-ing"-SPELLED nouns ("…, billing endpoint terminology…" — wrongly fired). Detection is now same-clause only. - URL-path-as-evidence REVERTED. Treating every path-bearing URL as an endpoint fired on ordinary asset/link URLs ("…/theme.css", "…?next=/x", a docs/repo link) and recreated routine UI-phase false positives. An external URL is evidence only when it NAMES an API vocabulary word ("api.stripe.com/v1"). Two fail-open cases are now DOCUMENTED limitations, pinned by tests so a future maintainer does not re-add the heuristics that caused the false positives above: a service named only in a clause separate from its API noun, and a bare external host that names no vocabulary word. Both are cheaply covered by the COVERAGE.md declaration and rare in real phase prose ("integrate the X API"). Also fixed from the round-4 review: - Qualification now survives markdown emphasis ("The **internal** Payments API" stays clean) while still not crossing a sentence/clause boundary. - readPhaseScope fail-closes on a REAL read failure (EACCES/EIO) enumerating the phase directory or reading the roadmap fallback — not only per-plan-file failures; a missing directory/section remains a legitimate no-op. The declaration-override path surfaces scope_read_error so an incomplete-scope override stays visible. - SERVICE_SURFACE_API_RE length-bound comment no longer overclaims. Net: the detector is same-clause verb+noun + `<Service> API` surface, with path/code/inline masking and a fail-closed posture. All five acceptance criteria hold. 1573/1573 unit tests pass; tsc + eslint (no-adhoc-markdown-parsing) + lint:regression-names clean. Built .cjs committed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#2365): close roadmap-fallback fail-open + stale JSDoc (round-5 review) The round-5 sanity review confirmed the detector simplification is sound (all acceptance positives fire, all required negatives clean) and flagged one real blocker plus a nit: - Blocker: readPhaseScope's roadmap fallback could still silently pass an UNREADABLE roadmap. getRoadmapPhaseWithFallback gated on fs.existsSync(), which returns false on EACCES/EIO too — so an unreadable ROADMAP.md read as "absent", no exception reached isRealReadFailure, and the blocking gate certified empty scope. Fixed at the source: read the roadmap directly and honor the function's OWN documented contract — null only on ENOENT (genuinely absent), otherwise throw. Both existing callers already wrap it in try/catch expecting that throw, and readPhaseScope now fail-closes (blocks) via its roadmap catch. Verified by a new e2e test (unreadable roadmap fallback → block). - Nit: the detectApiIntegration JSDoc still described the removed cross-clause participial binding and "every external hostname counts" — corrected to the actual same-clause-only behavior and the names-a-vocab-word URL rule. Verified: full unit suite green; tsc + eslint + lint:regression-names clean. Built .cjs committed (roadmap.cjs is gitignored/rebuilt, per repo convention). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#2365): backfill changeset PR number (#2397) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#2365): sync generated capability-registry + recapture install goldens CI surfaced two generated-artifact staleness issues (all failing test shards + lint-tests traced to these, not to a logic defect): - gsd-core/bin/lib/capability-registry.cjs was stale: the initial fix edited the ai-integration `api-coverage-plan-pre.md` fragment (added the "No external API integration" declaration section) but did not regenerate the registry, which embeds an inline copy of that fragment. Regenerated via `gen-capability-registry.cjs --write` — the diff is exactly the fragment text sync. Fixes `lint:generated-sync` and the "committed registry is in sync" + "registry integration" tests. - The 18 golden-install-parity fixtures were stale by exactly one hash line each — `gsd-core/references/api-coverage.md`, which this PR edits and which is a hashed installed artifact. Recaptured with `UPDATE_GOLDEN=1`; the diff is that single hash per runtime and nothing else. Fixes the `golden parity — *` tests. No source or behavior change — generated artifacts only. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#2365): flip representative-corpus manifest to assert the fixed behavior The #2371 representative corpus (merged into next after this branch was cut) is a known-bug tripwire: it asserts each fixture's currentBuggyOutput so the test fails loudly the moment #2365 is fixed, at which point — per its own contract in representative-corpus.test.cjs — the fixer removes currentBuggyOutput so the assertion checks expectedDetected instead. This is that moment. Removed currentBuggyOutput from the three detector fixtures (nextjs-route-path, unrelated-verb-noun, threat-model-prose); the corpus now asserts detected:false, which the fail-closed same-clause detector satisfies. Notes updated to describe the fix rather than the bug. The #2366 matrix corpus is left untouched — that tripwire belongs to its own PR (#2374). Corpus test: 7/7 pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#2365): skip chmod-000 fail-closed e2e tests on Windows The three fail-closed gate tests induce an unreadable plan / directory / roadmap with chmod 000, but Windows does not enforce POSIX mode bits — readFileSync still succeeds, so the gate never reaches the read-error path and the assertion fails on the windows-latest CI leg. The fail-closed LOGIC is platform- independent (readError → block) and is fully exercised on the macOS/Linux legs; only the method of inducing EACCES is POSIX-specific. Guard the three tests to skip on win32 as well as root, mirroring golden-install-parity's win32 skip. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#2365): address trek-e review — glossary, clock-seam, IO injection, bounds Review response to PR #2397 (trek-e, CHANGES_REQUESTED). Fix logic unchanged; this closes the test/process-hygiene findings. Major: - CONTEXT.md "Markdown Sectionizer" glossary now lists the two exports this fix relies on, `stripInlineCode` and `scanInlineCodeSpans` (glossary is a PR gate). - Replaced the banned wall-clock assertion in the "hostile repeated-term line" test (Clock Seams rule — no elapsed-time asserts) with a deterministic signal-count assertion, which also directly verifies the term-dedup that keeps pairing linear (one signal for a 10k-pair line, not thousands). - Rewrote the three fail-closed read-failure tests: instead of chmod 0o000 (a no-op under root / on Windows, the pattern the repo's IO-failure convention avoids) they now exercise the newly-exported `readPhaseScope` in-process and inject the failure by monkeypatching fs.readFileSync/readdirSync to throw, restoring in finally. Deterministic and platform-independent (no skip needed), and they add the ENOENT-is-absence case that the chmod tests couldn't express. Minor: - Added limit / limit+1 boundary tests for SERVICE_SURFACE_API_RE's {1,40} service-name bound, QUALIFIER_LOOKBACK's 24-char window, and REASON_MAX_LEN (200) on the declaration reason. - Added a fast-check property that fuzzes the tokenizer / clause splitter / masking (scanLineTokens, splitClauses, collectTermMatches) with adversarial tokens (slashes, backticks, URLs, clause punctuation) and asserts the detector is total (never throws), shape-stable, holds detected <=> signals, and is deterministic. readPhaseScope is exported for the in-process tests. Verified: 125 detector + 19 gate tests pass; tsc + eslint + generated-sync (glossary/registry) + lint-regression-test-names + lint-test-file-count clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d0bacc2517 |
fix(#2351): replace hardcoded timeout with portable run-with-timeout (#2426)
* fix(#2351): replace hardcoded gnu timeout with portable run-with-timeout Stock macOS ships neither `timeout` nor `gtimeout` (GNU coreutils). The 10 hardcoded `timeout <n> <cmd>` calls across the workflow/agent/reference gates exited 127 ("command not found") on such hosts, and the gates — which only distinguish 0/124/other — misreported a passing build or test as a FAILURE. Fix: a single Node-based `gsd_run run-with-timeout <secs> [--] <cmd…>` verb in gsd-tools.cjs. Coreutils-independent (stock macOS AND Windows), keeps GNU `timeout`'s exit-code contract (124 timeout, passthrough, 127/126 ENOENT/EACCES, 128+signum on signal), inherits stdio so pipes/redirects work, and reaps the whole process group so a watch-mode runner cannot outlive its budget. Runs before gsd-tools' flag parsing so the wrapped argv stays opaque. Hardened per adversarial review: - On timeout, SIGKILL the group SYNCHRONOUSLY before resolving — a descendant that traps SIGTERM was otherwise orphaned holding stdout, hanging captured gates (the exact watch-mode hang the feature prevents). - Forward SIGINT/SIGTERM to the child tree instead of dying and orphaning it. - Reject blank/whitespace <seconds> (was a silent unbounded run); clamp the timer to the 32-bit setTimeout ceiling (was a spurious immediate timeout). - Lint detector: catch GNU long options / `-k5` / `$((...))`; anchor to command position so prose "timeout 30 seconds" no longer false-positives. Resolution lives once in the CLI; all 10 sites call the shared verb. A parity guard (scripts/lint-portable-timeout.cjs, wired into lint:ci) fails the build if a bare `timeout`/`gtimeout` execution reappears (the portable `command -v timeout` probe form is intentionally allowed). Also fixes the identical bug in the zh-CN checkpoints translation, updates the tests that asserted the old strings, trims a redundant phrase in gsd-verifier.md to keep it under its size hard cap, and refreshes the size baselines + golden install-parity fixtures. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2351): add changeset (#2426) * chore: regenerate golden/size baseline after rebase onto next --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6b196ef638 |
fix(#2423): finalize job syncs next package.json after final release (#2437)
* fix(release): finalize job calls sync-next-version.cjs to bump next after final release (#2423) The release pipeline's 'finalize' job shipped X.Y.0 to npm 'latest' and merged to main, but never bumped 'next' to match. 'scripts/sync-next-version.cjs' exists exactly for this — its docstring promises to run 'for every release type (rc / hotfix / final)' — but it was wired only into the 'rc' job (release.yml:479), not 'finalize'. As a result, after 1.7.0 shipped on 2026-07-15, 'next' stayed at 1.7.0-rc.6 and every npm script banner on 'next' (and feature branches cut from it) reported the stale rc version. Regression of #1104 — closed incomplete (covered rc only, not final). This patch: - adds a 'Sync next branch to the published release' step to the 'finalize' job, mirroring the rc job's pattern at line 479 (uses VERSION from inputs.version rather than PRE_VERSION from steps.prerelease, since finalize does not run the prerelease step); - gates it on !inputs.dry_run + continue-on-error:true (matches rc); - adds tests/release-finalize-syncs-next-version.test.cjs — a structural YAML assertion that fails against the pre-fix workflow and passes after. Failing-first demonstrated during development; the test parses job blocks by indentation rather than grep so it stays valid as the file grows. Root-cause diagnosis: scripts/sync-next-version.cjs:14 docstring admits 'used by release.yml's rc job, which has no next-targeting PR of its own'. release.yml:506-685 (finalize job) had no sync-next-version step before this patch. Every prior rc.N release has a matching 'chore: sync next package version to 1.7.0-rc.N' commit; there is no such commit for 1.7.0. * chore(release): sync next package version to 1.7.0 (#2423) Replays the canonical 'chore: sync next package version to <v>' commit that the release pipeline's rc job auto-produces via scripts/sync-next-version.cjs, for the 1.7.0 final release that shipped on 2026-07-15 (commit |
||
|
|
40ce95f882 |
fix(#2358): scope review.md and ship.md temp files to a per-run mktemp directory (#2433)
* fix(#2358): scope review workflow temp files to a per-run mktemp dir /gsd-review wrote every prompt/section/output temp file to a hardcoded /tmp path keyed only on the phase number, so two GSD projects sharing a small phase number collide on the exact same path and a crashed run's leftover file becomes bait a later, unrelated run can silently read. ship.md's external peer-review stderr capture was strictly worse — one shared, unqualified path across every project/phase/run. Thread a single mktemp -d "${TMPDIR:-/tmp}/gsd-review.XXXXXX" run directory through every review.md temp path (67 sites) via a new {run_dir}/$RUN_DIR placeholder, mirroring the existing {phase} substitution mechanism, and clean it up at the end of the run. Route ship.md's stderr capture through a per-run mktemp file the same way. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2358): regenerate fixtures + lint gate-prep * fix(#2358): repair failing tests after gate verification * fix(#2358): thread RUN_DIR scoping into reviewer-instances.md (#1517) review.md's own invoke_reviewers step lazily loads gsd-core/references/reviewer-instances.md for the review.reviewer_instances codepath, but that doc was missed when review.md and ship.md were moved to the run-scoped {run_dir} temp directory. It still read the combined prompt from the old /tmp/gsd-review-prompt-{phase}.md (which build_prompt no longer writes, breaking reviewer-instances functionality outright) and wrote each instance's output to the old unscoped /tmp/gsd-review-${INSTANCE_NAME}-{phase}.md, leaving the exact cross-project temp-file collision bug open for that code path. Both paths now thread through {run_dir}, matching every other reviewer block in review.md. Extends the existing #2358 regression test with assertions pinning reviewer-instances.md's prompt read and output write to {run_dir}, and adds the Fixed changeset fragment. Regenerated the golden-install-parity content hashes for reviewer-instances.md via `npm run gen:golden` (paths unchanged; only the modified file's hash moved). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2358): backfill changeset pr (#2433) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
cd6665d73b |
fix(#2406): stop Codex config.toml from double-registering agent roles (#2432)
* fix(#2406): stop Codex config.toml from double-registering agent roles generateCodexConfigBlock emitted an [agents.<name>] role table per agent pointing config_file back at the standalone agents/<name>.toml Codex already auto-discovers, so every install declared each role twice in one config layer and Codex logged a duplicate-role warning per agent. Remove the redundant role-table loop; the standalone per-agent TOML is now the sole canonical registration source. The existing marker-truncate and leaked-section stripping in mergeCodexConfig already clean up legacy [agents.gsd-*] tables from prior installs, so updates converge to zero duplicates without any new migration path. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2406): regenerate fixtures + lint gate-prep * fix(#2406): repair failing tests after gate verification * docs(#2406): add changeset for Codex duplicate agent-role fix Adds the missing .changeset/*.md fragment for the Codex config.toml double-registration fix, closing the PR-gate finding from review. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2406): backfill changeset pr (#2432) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
88d6b392af |
fix(#2362): materialize third-party capability skills on opencode/kilo install (#2434)
* fix(#2362): materialize third-party capability skills on opencode/kilo install installOpencodeFamilyArtifacts/installOpencodeFamilySkills never received the capability registry, so a registered+surfaced+active third-party capability skill was silently dropped by the OpenCode/Kilo combined-family INSTALL path (registry said surfaced:true, disk had nothing). Thread capabilityRegistry through and reuse the existing #2322 seam's exported helpers (readInstalledCapabilitySkill, capabilityClusterStems, CAPABILITY_SKILL_MARKER) to fill in third-party skills after the first-party loop, with the same guarantees: registry-bound ownership, first-party-wins, full/'*' sentinel support, and graceful degradation on a missing/corrupt capability skill. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#2362): cover the tiered (non-'*') profile candidateStems branch Review found the new third-party capability-skill fill-in for OpenCode/Kilo had zero coverage of the tiered-profile path — all 12 regression tests only exercised the '*' full-profile sentinel. Adds a case that resolves a `standard`-tier profile through a synthetic registry (mirroring the seam's own __registryFor pattern) and asserts the resulting concrete Set still materializes the registered capability skill, with its prune-parity marker, alongside the tier's first-party skills. Also adds the .changeset/ fragment this fix was still missing. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2362): backfill changeset pr (#2434) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a7d83dc234 |
fix(#2390): warn on goal-shaped phase.add titles, correct auto-detect docs (#2425)
* fix(#2390): phase.add title warning + auto-detect doc fix phase.add now returns a `warning` field when a description reads as goal-shaped (>80 chars and/or multi-sentence) rather than title-shaped, instead of silently writing the whole paragraph verbatim as the `### Phase N:` header. The CLI still creates the phase as-is (the strict two-layer slash-vs-CLI interface is unchanged); the warning just surfaces the gap. Also clarifies six doc sites (command argument hints, workflow detection steps, and how-to/reference docs) that described the phase-number argument as "auto-detecting" the next unplanned phase -- that detection is an orchestrating-workflow/LLM step reading ROADMAP.md (concretely: `query roadmap.analyze`'s `next_phase` field), not a `gsd-tools.cjs` CLI feature. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2390): regenerate fixtures + lint gate-prep * fix(#2390): repair failing tests after gate verification * chore(#2390): add changeset (#2425) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1720aacf0c |
feat(#1949): <precondition> task element — Design by Contract (#2422)
* test(#1949): add failing-first tests for <precondition> element Red phase for issue #1949 (Design by Contract: <precondition> element asserted before task execution). Tests assert: - docs/reference/plan-md.md documents the new <precondition> element - agents/gsd-planner.md @-references planner-preconditions.md and stays under the 49152-char cap (progressive-disclosure requirement) - gsd-core/references/planner-preconditions.md exists and documents the three emission cases mandated by the issue (user_setup / prior-phase artifact / env-var) and the contract triad mapping - agents/gsd-executor.md asserts <precondition> before task execution and routes unmet preconditions through existing checkpoint machinery - cmdVerifyPlanStructure (behavioral via runGsdTools) accepts plans both with and without <precondition> — the additive-validation guarantee - Parity assertion: plan-md.md and planner-preconditions.md agree on the canonical tag spelling (DEFECT.GENERATIVE-FIX-DIVERGENCE guard) Most prose-contract assertions are Red until the implementation lands. The behavioral validator assertions pass immediately (regression guards proving the validator already accepts unknown optional tags). * feat(#1949): <precondition> task element — Design by Contract Add an optional <precondition> element to <task> in PLAN.md (issue #1949, The Pragmatic Programmer Topic 23). The front-of-task side of the plan contract — preconditions (before) ↔ postconditions (<verify>/<done>/ <acceptance_criteria>, after) ↔ invariants (must_haves.truths, across the whole plan). Together with the tracer-bullet proposal (#1945), this closes both ends of the 'outrunning your headlights' failure mode for an autonomous AI executor. Acceptance criteria met: - <precondition> is an optional element on <task>; plans that omit it validate unchanged (cmdVerifyPlanStructure checks for presence of required tags, does not reject unknown optional tags). - gsd-executor evaluates the precondition before any other task work. Unmet halts execution with a checkpoint:human-verify and no partial commit; met or absent produces no visible change to execution flow. Unmet is never auto-approved under AUTO_CFG=true — a missing prerequisite is a fact the executor cannot establish on its own. - gsd-planner emits <precondition> in exactly the three cases the issue mandates: user_setup consumption, prior-phase artifact dependency, and env-var/runtime-config dependency. - Tests cover met, unmet, and absent preconditions plus the additive- validator guarantee. Files: - gsd-core/references/planner-preconditions.md (NEW): full emission rules, the three cases with worked examples, format guidance, anti-patterns, the contract triad mapping, and the executor assertion contract. Progressive disclosure. - agents/gsd-planner.md: slim <precondition> note in Task Anatomy with @-reference to the new file. To stay under the 49152-char agent-file cap (27-char headroom before this change), the inline <comment_text_discipline> and <region_scoped_negative_gate> summaries are compressed to one-line pointers — their full rules already live in planner-antipatterns.md, so no content is lost. - agents/gsd-executor.md: new step 0 'Precondition check' in the execute_tasks loop, before the type dispatch, routing unmet through checkpoint_return_format. - docs/reference/plan-md.md: new Preconditions section in the schema reference, with the canonical example and the three emission cases. - CONTEXT.md: Precondition glossary entry as a sibling of Tracer Bullet. - docs/INVENTORY.md + INVENTORY-MANIFEST.json: row for the new references/planner-preconditions.md (regen via gen-inventory-manifest). - tests/precondition-element.test.cjs: failing-first tests covering schema docs, planner emission contract, executor assertion contract, reference-file presence + the three cases, behavioral additive- validator guarantee, and a parity assertion (DEFECT.GENERATIVE-FIX- DIVERGENCE guard). - .changeset/quick-hawks-bark.md: Added fragment. Companion to #1945 (tracer bullets). * chore(#1949): regen agent-size baseline + install-tree goldens Documented baseline regenerations required by the feat(#1949) prose changes (RULESET.AGENT_SIZE_BUDGET + golden-install-parity): - npm run size:baseline — locks in the new gsd-executor.md size (+1050 bytes: the precondition-check step 0 block). gsd-planner.md is net smaller (-142 bytes: compressed two inline summary blocks whose full rules already lived in planner-antipatterns.md to make room for the slim <precondition> pointer). No hard-cap breach. - npm run gen:golden — pick up the new references/planner-preconditions.md + the two changed agent files across all 18 runtime install trees. Both regens are CI-mandated after intentional agent/reference changes; see CLAUDE.md 'RULESET.AGENT_SIZE_BUDGET' and the comments in tests/golden-install-parity.test.cjs. * fix(#1949): bound <precondition> checks to read-only (security review) Apply the security-review finding (LOW, isolated /security-review subagent): the executor's 'run the cheapest check' phrasing for a plan-author-controlled prose line was broader than ideal — a hostile plan author could craft a <precondition> whose 'cheapest check' is side-effecting (curl to an attacker host under the guise of verification, rm -rf before checking, secret emission). The risk is inherited from GSD's existing plan-trust model (<verify>, <action>, <done> already direct the executor to run arbitrary shell), so <precondition> does not materially expand it. But the new prose actively directs execution ('run the check') rather than passively consuming the element, so the bound is worth making explicit. Tightened across all four surfaces that describe the check shape: - agents/gsd-executor.md step 0: 'Verify with read-only checks only — file existence, env var presence (no value output), idempotent GET /health-style pings. Do NOT run commands with side effects (writes, network POSTs, secret emission) as the check; if a side-effecting check seems required, halt and surface via checkpoint instead.' - gsd-core/references/planner-preconditions.md Format section: same bound, plus the halt-and-surface escape hatch. - docs/reference/plan-md.md Preconditions section: mirrored. - CONTEXT.md Precondition glossary entry: mirrored. Regenerated agent-size baseline (executor grew 46186 -> 46440; still under the 49152 cap) and install-tree goldens. * chore(#1949): backfill changeset pr number 2422 Per CONTRIBUTING.md changeset workflow + feature-builder directive Step 8.7: backfill the placeholder pr:0 with the real PR number immediately after gh pr create returns. Avoids the fail_invalid_fragment gate. * fix(#1949): cite [#1949] on allow-test-rule exemption (ADR-456) CI's lint:ci runs lint-allow-test-rule-refs which per ADR-456 requires every // allow-test-rule: exemption on a NEW test file to carry an issue reference (#NNN or URL). My earlier push omitted it. Local 'npm run lint' (eslint) does NOT run this check — only 'npm run lint:ci' does. CLAUDE.md explicitly warns: 'lint:ci ≠ lint — CI runs lint:ci; a local pass is not the gate.' I should have run lint:ci before pushing; correcting now. Pattern matches the companion feature's test file: tests/tracer-bullet.test.cjs:1 // allow-test-rule: source-text-is-the-product [#1945] |
||
|
|
8d2f8bcb23 |
fix(#2388): gate shared requirement completion on sibling plans, revert on gaps (#2424)
* fix(#2388): gate shared-ID requirement marking and revert on gaps_found Adds requirements.ready-ids (execute-plan.md's update_requirements step) so a requirement ID declared by multiple plans in a phase only marks Complete once every declaring plan has produced a SUMMARY.md, and requirements.revert-phase (execute-phase.md's gaps_found branch) so a gaps_found verdict reverts the phase's own prematurely-Complete IDs before the gap report renders. Single-plan IDs still mark immediately. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2388): regenerate fixtures + lint gate-prep * fix(#2388): repair failing tests after gate verification * chore(#2388): add changeset (#2424) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
873bdf51e5 |
fix(#2352): expand tilde paths in review scope before the deleted-file filter (#2419)
* fix(#2352): tilde-expand SUMMARY.md key-files paths before deleted-file filter compute_file_scope's "Filter deleted files" step tested the literal `~/...` value from SUMMARY.md key-files entries with `[ -f "$file" ]`, which bash never tilde-expands (only a literal `~` in source text expands, not one arriving as an already-expanded variable value). Real files recorded with a `~/...` path were silently misclassified as deleted and dropped from REVIEW_FILES, and a phase whose every recorded file used a tilde path hit the empty-scope skip as a false negative. Adds a tilde-normalization loop as step 1 of post-processing (all tiers), before the deleted-file filter, rewriting a leading `~/` to `${HOME}/...` so downstream existence checks, the empty-scope short-circuit, and the FILES_TO_READ/CONFIG_FILES construction all see a real, openable path. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2352): regenerate fixtures + lint gate-prep * chore(#2352): add Fixed changeset fragment (pr 2419) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
dd5a2211c9 |
enhance(#1964): semantic knowledge-base recall via MemPalace (keyword fallback) (#2416)
* test(#1964): add failing-first semantic-recall contract tests Epic #1957 Phase 3C (final). Source-text-is-the-product contract tests: semantic recall via MemPalace (top-k meaning-similar prior resolutions, catches same-root-cause/different-wording cases), indexing resolved sessions at archive, graceful degradation to keyword matching when MemPalace is absent, knowledge-base.md stays the durable plain-text source of truth, agent Phase 0 / Matching Logic is semantic-first (the stale 'keyword overlap, not semantic similarity' claim must go), and no new embedding/vector infra (reuse MemPalace). Failing-first: reference, the Matching Logic reframe, the Phase 0 consolidation, and the archive indexing step do not yet exist. * feat(#1964): semantic knowledge-base recall via MemPalace (keyword fallback) Epic #1957 Phase 3C (FINAL). Replaces keyword-overlap matching with semantic recall: at Phase 0 the debugger queries MemPalace with the current symptoms and surfaces the top-k meaning-similar prior resolutions, catching the same-root-cause/different-wording cases keyword overlap missed (the self-noted 'keyword overlap, not semantic similarity' limitation). Resolved sessions are indexed into MemPalace at archive (symptoms + root_cause(s) + fix + recurrence guard). knowledge-base.md remains the durable plain-text source of truth; when MemPalace is absent the debugger falls back to keyword-overlap matching (logged, never a silent skip). No new embedding/vector infrastructure — MemPalace is reused. Size-neutral agent edits: the Matching Logic section reframed (keyword-only -> semantic-first + keyword-fallback + @-include); Phase 0's three keyword bullets consolidated into one semantic-first bullet; one MemPalace-indexing step added at archive. Agent at 57222 B (122 B headroom — final phase). Full rules in gsd-core/references/debugger-semantic-recall.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1964): address orthogonal review (invocation mechanism, index Resolution-not-symptoms + redaction, fallback detail) - HIGH: the 'query MemPalace' instruction was WHAT-level only; the agent has no MCP tools. Added an Invocation section naming the Bash CLI (mempalace search --wing <wing>) + MCP-when-registered + wing resolution (config.mempalace.wing -> project_code -> project dir), matching every other MemPalace integration. Without this the feature silently degraded to keyword matching even when MemPalace was present. - MEDIUM (security x2): index the agent-authored Resolution summary (root_cause + fix + recurrence_guard), NOT raw user-supplied Symptoms — excludes attacker-controlled prose from the cross-session index AND reduces secret/PII leakage. Redact secret-shaped values before indexing. Stated the write order (KB append + commit MUST succeed before indexing). - LOW: restored 'identifiers' + 'case-insensitive' to the keyword fallback; added a test asserting the fallback mechanics survived the Phase 0 consolidation (Error patterns field, 2+ token overlap, identifiers, case-insensitive). * chore(#1964): ratchet agent-size baseline downward (leaner archive bullet shrank gsd-debugger.md 57222->57197) * chore(#1964): backfill changeset pr number (PR #2416) |
||
|
|
c67f301867 |
feat(#1963): emit blameless-postmortem Prevention block at resolution (#2410)
* test(#1963): add failing-first prevention/postmortem contract tests Epic #1957 Phase 3B. Source-text-is-the-product contract tests: blameless 5-Whys that BRANCHES per Phase 2A RCA (not a single-cause chain; treats agent error as 'why was that possible?'), the 'why wasn't this caught?' question, the recurrence-guard taxonomy (regression test / assertion / lint rule / KB pattern), the KB-entry why_not_caught + recurrence_guard fields with backward compat, the session-manager prevention summary line, and the Zawinski scope-boundary (a block, not a subsystem). Failing-first: reference, archive_session edit, KB schema extension, and session-manager summary do not yet exist. * feat(#1963): emit blameless-postmortem Prevention block at resolution Epic #1957 Phase 3B. At archive_session the debugger now produces a Prevention block with three blame-free components: a branching 5-Whys causal chain (branches per Phase 2A RCA, not a single chain; 'agent error' prompts 'why was that possible?', never blame), a 'why wasn't this caught?' answer naming the missed gate (test/typecheck/lint/review/verify), and a concrete recurrence guard (regression test / assertion / lint rule / KB pattern). The knowledge-base entry gains two structured fields (why_not_caught + recurrence_guard) so future Phase-0 recall surfaces the prior prevention, not just the prior fix. Additive: old entries without the fields still load. The session-manager compact summary surfaces a one-line prevention summary. Full rules extracted to gsd-core/references/debugger-prevention.md (slim archive_session step + 2 KB fields kept in the agent). INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1963): address orthogonal review (CRITICAL append-template drift + Phase-0 consumption + parity test) - CRITICAL: the archive_session KB append template omitted Why not caught + Recurrence guard (only the Entry Format had them) — the feature's core deliverable silently did not happen. Added both fields to the append template the agent actually follows (nearest-instruction wins). - HIGH: Phase 0 (KB read) only surfaced root_cause + fix; the new fields were dead data. Extended the Phase 0 Evidence line to consume why_not_caught + recurrence_guard when present (absent on old entries — backward compat holds). - MEDIUM: added a cross-section parity test (every Entry-Format field must also appear in the append template — the guard that would have caught the Critical) + a Phase-0-consumption assertion. - MEDIUM: the 'branches per Phase 2A' claim is now wired — reuses reasoning_checkpoint.candidate_causes across the four categories. - MEDIUM: recurrence-guard taxonomy gains type refinement + config-default change; LOW: added 'build' gate to both surfaces for parity. - NIT: compact-summary fallback shape ('no gate existed'); verify the guard artifact exists before recording it. * test(#1963): anchor Phase-0 consumption test on the specific heading The regex /Phase 0[\s\S]{0,1200}/ matched the first 'Phase 0' in the file (in knowledge_base_protocol prose), not the Phase 0 block in investigation_loop. Anchor on '**Phase 0: Check knowledge base**' and widen to 1500 chars. * chore(#1963): backfill changeset pr number (PR #2410) |
||
|
|
36a311c5bb |
enhance(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries) (#2409)
* test(#1962): add failing-first repro-hardening contract tests Epic #1957 Phase 3A. Source-text-is-the-product contract tests: PBT shrinking (fast-check/Hypothesis, minimized seed, manual-minimization degradation), the four oracle types (specified/derived/metamorphic/implicit with implicit flagged weakest), boundary neighbors (off-by-one/min-max/empty-singleton tied to the equivalence class), oracle_type in DEBUG Resolution, and the Phase 1A tie-in (minimized seed + real oracle => the mutation guardrail bites). Failing-first: reference, agent cross-refs, and template field do not yet exist. * feat(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries) Epic #1957 Phase 3A. Extends Minimal Reproduction (shrinking) and Test-First Debugging (oracle classification + boundary neighbors): - Shrinking: wrap an input-space failing input in a property (fast-check JS/TS, Hypothesis Python) and store the MINIMIZED counterexample as the regression seed; degrade to manual minimization when no PBT framework is present. - Oracle classification: state specified / derived (contract/model) / metamorphic / implicit (crash, weakest) before writing the assertion; record under Resolution.oracle_type; never default to implicit silently. - Boundary neighbors: off-by-one, min/max, empty/singleton around the fixed defect's equivalence class. Together they turn the regression test into a root-cause check — what the Phase 1A mutation guardrail needs to bite. Full rules extracted to gsd-core/references/ debugger-repro-hardening.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md + DEBUG template updated. * fix(#1962): address orthogonal review (bounding, provenance, oracle scope, sufficient-triple) - HIGH: added a 'Bound the property/shrink run' section (60s timeout, degrade- to-manual on timeout, do-not-raise-default-run-limits, argv-not-shell) — the gauntlet violation the sibling references already honored. - Medium: test-provenance caveat (the failing input often comes from the bug report — author the generator from a sanitized description, cross-ref debugger-fix-acceptance.md). - Medium: oracle scope note — the 4 types cover deterministic bugs; non- deterministic failures re-route to stability-stress per bug-taxonomy. - Medium: Phase 1A tie-in corrected — seed+oracle is necessary not sufficient; boundary neighbors close the adjacent-input escape; the sufficient triple is seed+oracle+neighbors. - Low: preserve the original noisy repro as a secondary reference; operationalize 'equivalence class' (the predicate the fix draws). Nit: degradation reworded. * chore(#1962): backfill changeset pr number (PR #2409) --------- Co-authored-by: sim <sim@local> |
||
|
|
6baa2a8182 |
feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger (#2407)
* test(#1961): add failing-first bug-taxonomy routing contract tests Epic #1957 Phase 2B. Source-text-is-the-product contract tests (3 taxonomy classes, explicit class->technique routing table, Bohrbug->repro+SBFL+bisect, Heisenbug->record-replay/stability+SKIP-SBFL, Concurrency->atomicity/order/ deadlock checklist, bug_class in DEBUG Current Focus, supersede-not-append) plus a routing-table specification object pinning the documented decisions (SBFL forbidden on Heisenbug is the load-bearing 1B/2B seam). Failing-first: reference, Phase 1.75, and routing-table reframe do not yet exist. * feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger Epic #1957 Phase 2B (reliability-critical). Adds Phase 1.75: classify the failure as Bohrbug / Heisenbug-Mandelbug / Concurrency, then route the investigation technique via an explicit class->technique table (Kernighan: no opaque heuristic). Bohrbug -> reproduction + SBFL (Phase 1.25) + git bisect; Heisenbug/Mandelbug -> record-replay (rr) + stability-stress + statistical sampling, with SBFL explicitly SKIPPED (a flaky spectrum poisons the Ochiai ranking — the load-bearing 1B/2B seam); Concurrency -> the atomicity/order/deadlock checklist first. Reframes (supersedes, not appends — Zawinski) the flat 'Technique Selection by situation' table into a class-routed table; the 11 techniques remain as routed targets. bug_class recorded in Current Focus (DEBUG template); common-bug- patterns catalog cross-referenced to the taxonomy. Full rules extracted to gsd-core/references/debugger-bug-taxonomy.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1961): address orthogonal review (phase-name drift, General lane, revoke framing, row-scoped tests, bounding) - HIGH: reference said 'Phase 1B' (epic shorthand); corrected to the deployed 'Phase 1.25' (matches the agent + SBFL reference). - HIGH: 6 of 11 techniques (Rubber duck, Delta, Working backwards, Differential, Comment-out, Follow-the-indirection) were orphaned by the situation-table reframe. Added a 'General (any class, situation-cued)' lane to BOTH the reference routing table and the agent's Technique Selection table that re-homes them — supersede-not-append now holds. - MEDIUM: the SBFL-skip is structurally retroactive (Phase 1.25 runs before Phase 1.75 classification), so reframed the table column from 'Do NOT use' to 'Revoke if already run' + an explicit 'retroactive revocation, not proactive skip' note stating the ordering honestly. - MEDIUM: contract tests are now row-scoped (parse the table by class, assert per-row) instead of presence-only; added a guard that the previously- orphaned techniques now have a General-lane route. - LOW: pinned the canonical bug_class value form (lowercase-kebab: bohrbug|heisenbug-mandelbug|concurrency; prose may use title-case). - NIT: added a 'Bound the Heisenbug-chase runs' note (rr/stability/sampling timeouts) per the unbounded-subprocess gauntlet. * chore(#1961): backfill changeset pr number (PR #2407) |
||
|
|
f8b16d1874 |
enhance(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger (#2405)
* test(#1960): add failing-first RCA-branching contract + schema-invariant tests Epic #1957 Phase 2A. Source-text-is-the-product contract tests (fishbone >=2 categories, AND-gate, multi-cause root_cause, backward compat, reasoning checkpoint candidate_causes+and_gate fields, debugger-philosophy single-cause note, DEBUG template) plus behavioral schema-invariant checks on two fixtures: two contributing causes (AND-gate yes) -> both recorded; single-cause (AND-gate no) -> one root_cause, identical to today. Failing-first: reference, agent edits, and template note do not yet exist. * feat(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger Epic #1957 Phase 2A. Guards against 5-Whys single-cause bias: before committing root_cause, the debugger enumerates candidate causes across >=2 Ishikawa categories (code/config/environment/data) and explicitly answers an AND-gate question. When the AND-gate fires, every contributing cause is recorded, so a multi-cause fix no longer recurs via the unaddressed second cause. Resolution.root_cause may hold one OR a small set (additive; single-cause sessions are byte-identical to today). The Structured Reasoning Checkpoint gains candidate_causes + and_gate fields; debugger-philosophy.md adds the single-cause-bias trap. Full rules extracted to gsd-core/references/debugger-rca-branching.md (slim Phase 2 routing + 2 checkpoint fields kept in the agent). INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md + DEBUG template updated. * fix(#1960): address orthogonal review (AND-gate self-consistency, parity guard, narrowed claim, ripples) - Reference: the collapse rule now enforces AND-gate self-consistency — and_gate=yes with a single confirmed cause is flagged as incomplete (return to Phase 3); a race/timing note clarifies such bugs bridge categories; the 'byte-identical' backward-compat claim narrowed to 'root_cause shape unchanged; reasoning_checkpoint gains 2 fields in every session'. - DEBUG.md: stale 'five-field' mirror prose -> seven-field (parallel-surface drift the reviewer flagged); new debug-session-management parity test pins the field-count claim to the gsd-debugger.md YAML keys (CRLF-safe). - Scalar-assuming consumers of set-valued root_cause updated: session-manager compact summaries (319/332), diagnose-only return (1062), archive entry (1216), ROOT CAUSE FOUND return (1322). - Test: added the AND-gate-yes/single-cause invariant + fixture; rephrased the fixture describe block honestly as a schema-invariant specification. - Phase 2 bullet phrasing clarified ('at hypothesis formation, before the Phase 4 commit'). * test(#1960): parity regex accepts word-form count ('seven-field' or '7-field') * test(#1960): parity regex counts array-valued YAML keys (no inline value) * chore(#1960): backfill changeset pr number (PR #2405) |
||
|
|
50efae13ce |
fix(#2305): stage the shared guard hooks Kilo's native plugin spawns (#2327)
* fix(#2305): stage shared guard hooks for Kilo — drop skipSharedHooksInstall Kilo's capability descriptor declared BOTH hostBehaviors.nativePlugin (a plugin that spawns the shared PreToolUse guard scripts as subprocesses) AND hostBehaviors.skipSharedHooksInstall:true, which suppresses staging of hooks/*.js into the Kilo config dir. The plugin's runHook treats an absent hook script as a silent allow, so every guard it spawned (gsd-prompt-guard, gsd-read-guard, gsd-worktree-path-guard) no-opped on every Kilo install. OpenCode uses the byte-identical plugin with hook staging on and is unaffected — it is the reference shape. The skip flag predates Kilo's plugin surface: it dates to #1821 (hooks were dead weight for a runtime with no hook consumer), and #2093 added the hooks-dependent nativePlugin without revisiting it. - capabilities/kilo/capability.json: remove skipSharedHooksInstall (regenerated gsd-core/bin/lib/capability-registry.cjs accordingly) - bin/install.js: correct the stale #1821 comments claiming Kilo has no plugin surface - tests/kilo-upgrades.test.cjs: install-fixture tests (global + local) asserting the guard scripts land where the plugin's walk-up resolves them; an end-to-end test driving a disallowed out-of-worktree write through the REAL installed Kilo tree and asserting the guard rejects it; a cross-runtime descriptor invariant (nativePlugin and skipSharedHooksInstall:true must never coexist) - tests/kilo-imperative-reference.test.cjs: flip the pinned assertion - golden fixtures regenerated (kilo now stages the 24 hook files, same set as OpenCode) Fixes #2305 * fix(#2305): warn loudly when a guard hook script is missing (runHook) runHook's absent-file branch returned a silent exit-0 allow — the mechanism that let #2305 ship undetected: with the hooks bundle never staged on Kilo, every PreToolUse guard the plugin spawned resolved to "file not found → allow" with zero signal anywhere. Keep the adapter's design contract (a missing hook must never break the tool call — pinned by the existing adapter test) but make the absence loud: console.error once per hook file, naming the unresolved path and the remediation. Applied identically to .kilo/ and .opencode/ plugin copies (byte-parity guard). Golden parity fixtures regenerated (the installed plugin file's hash changed). Fixes #2305 * chore(#2305): add changeset fragment * test(#2305): include gsd-workflow-guard.js in the staged-guards regression list The native plugin spawns four guards on write-like tool calls — the regression test's PLUGIN_GUARD_HOOKS list covered three. Staging itself was already asserted via the golden fixtures (the full bundle), but the named per-guard assertion should cover every guard the plugin actually dispatches. Surfaced by cross-AI review of PR #2327. * test(#2305): update the #1821 tests that encoded Kilo's false no-plugin premise The #1821 hook-copy test asserted Kilo must receive no staged hooks — the exact behavior this PR reverses (and the cause of all 8 CI failures). Kilo moves from the ZCode "no dead hooks" loop to the OpenCode group, with positive assertions on the new contract: the three guard hooks the plugin spawns, hooks/lib/git-cmd.js, and plugins/gsd-core.js all staged. The integration runtime contract flips kilo packageJson to true (the CommonJS marker ships with the bundle), and the pi contract comment no longer cites Kilo as a no-plugin runtime. * chore(#2305): scope the queued #1821 changeset fragment to ZCode only The fragment still claimed the installer skips hooks for Kilo — rendering both it and this PR's fragment into the same release would ship two contradictory statements about Kilo's install behavior. It now claims ZCode only and notes that #2327 reverses the Kilo half. * chore(#2305): rename changeset fragment to the generator naming convention 2305-kilo-stage-guard-hooks.md -> loud-guard-hooks.md, matching the <adjective>-<noun>-<noun> shape npm run changeset generates (review nit). --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
13d181aedf |
fix(#2349): exclude status: superseded plans from phase completion counts (#2404)
Adds a status: superseded plan-frontmatter marker that scanPhasePlans excludes from both plan and summary counts, so a phase with a deliberately-unexecuted plan no longer reads incomplete forever (the plan-level analogue of #1514). Includes all-superseded completion handling and a bounded, symlink-safe frontmatter read. Fixes #2349. |
||
|
|
56a5c6404c |
feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger (#2403)
* test(#1959): add failing-first SBFL contract + Ochiai correctness tests Epic #1957 Phase 1B. Source-text-is-the-product contract tests (Ochiai formula documented, Tarantula fallback, top-N seeding, no-coverage skip logged, ranking->Evidence, Bohrbug gating) plus a behavioral Ochiai formula-correctness section: bound [0,1], max-score invariant, a known-fault fixture proving the fault ranks #1 (criterion 2), clean degradation on zero failing tests, and two fast-check properties. Failing-first: reference file and agent routing do not yet exist. * feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger Epic #1957 Phase 1B. When a runnable test suite with per-test coverage exists (>=1 failing AND >=1 passing test), the debugger computes an Ochiai suspiciousness ranking over the coverage spectrum and seeds the top-N suspicious locations into Evidence as first-class hypothesis candidates, narrowing the search space deterministically before LLM reasoning. Tarantula documented as fallback. Degrades cleanly (logged, never silent) when there is no test suite, no failing tests, or no per-test coverage, and is explicitly not trusted on flaky/Heisenbug spectra (pairs with Phase 2B bug-taxonomy). Full rules extracted to gsd-core/references/debugger-sbfl.md (slim Phase 1.25 routing kept in the agent to respect the size cap). No new coverage framework — reuses the project's existing test/coverage runner. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * test(#1959): bound property generators to valid coverage counts The [0,1] property generated failedExec independently of totalFailed, but Ochiai's score is only bounded by 1 under the coverage invariant failedExec <= totalFailed (a failing test that executed s is one of the totalFailed failing tests). Out-of-domain inputs (failedExec=100, totalFailed=5) make the formula correctly return >1. Bound failedExec by totalFailed via fc.chain so the property tests the real domain. Also cleaned up the ranking property (removed dead code). * fix(#1959): address orthogonal review (monotonicity property, degradation row, coverage bounding) - Replace vacuous ranking property (true-by-sort-construction) with a non-trivial monotonicity property: holding totalFailed + passedExec fixed, ochiai is non-decreasing in failedExec. An inverted formula would fail it. - Add the missing 'no passing tests' degradation row (preconditions require >=1 passing test; Tarantula would divide by totalPassed=0). - Bound the coverage subprocess (CLAUDE.md gauntlet): cap the coverage run, degrade-to-skip on timeout, never hang the debug session. - Reword 'discard the ranking' -> 'mark the Evidence entry as revoked (do not delete)' per Kernighan auditability. * test(#1959): bound monotonicity-property generator to valid coverage (failedExecA <= totalFailed) * chore(#1959): backfill changeset pr number (PR #2403) |
||
|
|
863a54ec82 |
fix(#2350): pass --raw to config-get in every build/test gate (#2399)
Adds --raw to config-get workflow.build_command|test_command reads in the post-merge, regression, verify-phase, and audit-fix gates so an unset key is a genuinely empty string, not the literal "" — restoring the auto-detect cascade and graceful skip instead of a false exit-127 failure. Regression guard sweeps all four gate files. Fixes #2350. |
||
|
|
5e52350736 |
feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger (#2396)
* test(#1958): add failing-first guardrail contract tests Epic #1957 Phase 1A. Adds source-text-is-the-product tests asserting the 5-signal fix-acceptance guardrail contract (target test, mutation check, no-op/deletion detector, adjacent tests, revert-and-reconfirm), graceful degradation, FIX REJECTED BY GUARDRAIL return path, per-signal debug-file recording, and subprocess bounding. Failing-first: reference file and agent sections do not yet exist. * feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger Epic #1957 Phase 1A. Prevents accepting a fix that merely greens the test (Goodhart defense / APR overfitting). Adds a 5-signal gate run before fix acceptance: target test, mutation check (Stryker), no-op/behavior-deleting detector, adjacent/held-out tests, revert-and-reconfirm. Degrades gracefully when Stryker or a test suite is absent (each skip logged, never a silent pass), records per-signal results under Resolution.verification, and returns a FIX REJECTED BY GUARDRAIL outcome the session-manager surfaces for revise / accept-as-debt / abandon. Full rules extracted to gsd-core/references/debugger-fix-acceptance.md (slim routing kept in the agent to respect the agent-size cap). Debug template + INVENTORY + manifest + agent-size baseline + AGENTS.md updated. * test(#1958): correct newline-tolerant assertion + regen install-parity goldens The revert-and-reconfirm assertion collapsed whitespace before matching so markdown line-wrapping does not break it. Regenerated the golden-install-parity and install-tree fixtures (npm run gen:golden) to absorb the intentional gsd-debugger.md / gsd-debug-session-manager.md / DEBUG.md / new reference-file changes to the installed artifact tree. * fix(#1958): tighten guardrail per orthogonal review Addresses the isolated reviewer's findings: - signal 5 now states its recorded-repro dependency and routes the no-repro case to the degradation row; revert mechanism specified (git stash / git revert -n); minimality flag tied to diff structure, not revert-ability. - bounded-subprocesses section now bounds the git subprocess (5-30s) too, requires argv-array argument passing, and scopes Stryker to the driving regression test (a mutant killed only by a non-driving test is a finding). - new test-provenance (security) clause: the driving test must be agent-authored; bug-report repro scripts are DATA, never executed verbatim. - tightened 3 contract assertions to bind to specific clauses (guardrail_verdict field, deletion-reject-unless-RCA, 60s+git bounding). - Goodhart framing softened to 'partially-independent'; DEBUG.md template verification field notes the nested map shape. * chore(#1958): backfill changeset pr number (PR #2396) * fix(#1958): add issue ref to allow-test-rule annotation (ADR-456) CI lint-allow-test-rule-refs requires every allow-test-rule exemption to carry a 'see #NNN' issue ref per ADR-456. The new test file's annotation lacked it; this adds (see #1958). |
||
|
|
2c54f219c9 |
fix(#2348): derive verification staleness from git commit time, not mtime (#2394)
readVerificationStatus() decided a phase's verification was `stale` (a *-SUMMARY.md newer than the *-VERIFICATION.md) by comparing filesystem mtimes. mtimes are assigned at checkout time and are not preserved by `git clone` / `cp -R`, and any unrelated `touch` / reformat / editor-save re-stales a valid report — so a committed phase declaring `status: passed` could silently read `stale` on a fresh clone purely from checkout order, falsely rewriting a ROADMAP row and blocking milestone close (#2022 gate). Each file's effective "last changed" time is now its git commit time when the file is committed AND clean, and its mtime otherwise (uncommitted or working-tree-dirty). Both are real wall-clock change times, so a summary committed after — or edited after — the verification reads stale, while a clean fresh clone stays passed. Git commit time is content-tied and clone- stable; mtime is retained only where it is the true last-changed signal. Implementation: - Two bounded git calls per phase (never one-per-file): `git log --first-parent --format=%ct --name-only` for commit times, and `git diff --name-only HEAD` to drop dirty files. readVerificationStatus runs per-phase in the init/roadmap listing loops, so per-file spawning would fan out to P×(S+1) git processes ("Unbounded Subprocesses"). - `--first-parent` so merge commits report their file lists (plain `--name-only` omits merge diffs and would under-date merge-landed content). - The dirty-check fails SAFE: if `git diff` is inconclusive (errors / exits non-zero) the commit times are discarded so every file falls back to mtime, never trusting a possibly-stale commit time (no false "not stale"). - Paths matched back by `/`-bounded suffix (root vs nested `plans/` can't collide) and passed after `--` (dash-named files can't be read as flags). - A phase with no summaries skips git entirely; the scan short-circuits on the first stale summary. A `phaseCleanCommitTimesMs` seam keeps the unit tests hermetic (no git spawn); the resolver's two-call error handling is unit-tested via an injected execGit; two real-git integration tests lock the end-to-end path, the committed-then-edited (dirty) regression, and the `--` argv guard. |
||
|
|
f2c077df38 |
chore(#2387): refactor CONTEXT.md legacy content + add glossary drift gate (#2391)
* chore(#2387): refactor CONTEXT.md legacy content + add glossary drift gate Apply the audit-and-enforce concept from the ADR index (#2356) to CONTEXT.md: correct stale facts, and add a CI gate so the machine-verifiable claims can't silently re-rot. CONTEXT.md was entirely hand-maintained with nothing checking its claims against the shipped tree, so it had rotted. An audit against live code (Memtrace + filesystem + gh), each finding adversarially re-verified, drove 38 factual corrections + 1 surfaced by the new gate: - Dead references: Package Identity named @opengsd/get-shit-done-redux (package is @opengsd/gsd-core); Shell Command Projection named run-git/run-npm/run-tool (real exports execGit/execNpm/execTool); a partial docs/adr/1606 ref; retired sdk/ framing. - Superseded facts: allRuntimes 15 -> 17 (pi #2102, zcode); "seven nested-loader runtimes" -> five (claude reverted flat #924, antigravity flat); stacked-PR examples rebasing onto main -> next; QUOTA_SENTINELS precedence corrected to match src/agent-command-router.cts. - Drifted CONTRIBUTING.md line citations refreshed. Per CONTRIBUTING.md:179, only stale FACTS were corrected -- no maintainer intent, lesson, or opinion was rewritten, and the append-only session log is untouched except one dated in-place superseding note. The three tests that assert on CONTEXT.md content (phase6-capstone-conformance, tracer-bullet, external-job-waiting) keep all their anchors. New scripts/check-glossary-refs.cjs (--check, wired into lint:generated-sync): - Check A: every backticked file reference under a TRACKED_PREFIXES allowlist resolves on disk. Generated gsd-core/bin/lib/*.cjs (77 refs, gitignored), ~/-paths, .planning/, and bare filenames are deliberately skipped so a clean CI checkout never false-fails. - Check B: the allRuntimes count + member set in the glossary prose match bin/install.js's allRuntimes literal (drifts on every runtime addition). tests/check-glossary-refs.test.cjs covers both, including the false-positive guard that a missing bin/lib/*.cjs ref does NOT trip the gate. Closes #2387 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2387): confine glossary-gate file refs to ROOT (no `..` traversal) Pre-PR security review finding (low): extractTrackedRefs fed tokens straight to fs.existsSync(path.join(ROOT, token)), and PATH_TOKEN_RE admits `.` in a segment, so a CONTEXT.md token like `src/../../../etc/passwd` passed the `src/` prefix check and normalized to an out-of-tree absolute path — turning the doc lint into a filesystem-existence oracle on the CI host (existsSync only; CONTEXT.md is a trusted committed file, hence low severity, but a defense-in-depth gap). Add isWithinRoot() confinement in extractTrackedRefs: a token is dropped unless path.resolve(ROOT, token) stays within ROOT. A CONTEXT.md reference is always a plain in-repo path, so a `..` escape is never legitimate. Regression test asserts a `..`-bearing token is skipped and never named in output. Refs #2387 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2387): drop legacy `get-shit-done` name from a CONTEXT.md defect entry CI lint-legacy-dir-name failed: the line-928 upstream-issue re-point I applied wrote the historical provenance as "gsd-build/get-shit-done#3545", and scripts/lint-legacy-dir-name.cjs forbids the legacy `get-shit-done` name. Reword to "moved from #3545 in the predecessor repo" — same provenance, no legacy name. Caught by `npm run lint:ci` (the CI lint chain), which I had not run locally — lint:generated-sync + eslint do not include lint-legacy-dir-name. Refs #2387 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
81f7ab4df1 |
refactor(#2384): leaf dispatch table + runCommand collapse (ADR-2346 P4) (#2392)
* refactor(#2384): leaf dispatch table + runCommand collapse (ADR-2346 P4) Cutover all 55 remaining case arms from runCommand's switch to HOST_COMMAND_ROUTERS. runCommand now contains only its default case (~40 lines): the three-layer dispatch (capability → overlay → host table) plus the unknown-command diagnostic. The 73-case switch is dissolved. Each case body was relocated verbatim to a module-scope route*Command function via a brace-matching extractor; inner break; statements (from _dispatchNonFamily early-exit patterns) were converted to return; (5 arms affected); loop break; statements preserved. Closes #2384 * chore: retrigger CI |
||
|
|
f15eb5f5c9 |
fix(#2347): make the decision-shape evidence test format-agnostic (#2389)
#1365's fail-loud guard reused the parser's own D- grammar as its evidence test, so a populated <decisions> block using any other ID prefix (e.g. D5-01) was invisible to both parser and guard, collapsing could-not-parse into a clean none-present pass. Add an ID-shaped bold-lead-in probe as format-agnostic evidence on both parse paths; empty/prose scaffolds stay none-present. Graduates the #2371 d5-prefix representative fixture to its expected* assertion. Closes #2347. Admin-merged (self-review bypass) with full green CI. |
||
|
|
062f3fda90 |
chore(#2371): representative gate-fixture corpus + document-shaped property test (#2380)
* chore(#2371): representative gate-fixture corpus + document-shaped property test Adds tests/fixtures/representative/ — a permanent corpus of verbatim, incident-sourced fixtures (never author-invented) from #2286, #2347, #2365, #2366, each labeled with its expected gate verdict in a MANIFEST.json and driven through the real CLI gate entrypoint via tests/representative-corpus.test.cjs. Adds a document-shaped fast-check property test alongside the existing writer-seeded bijection test in tests/api-coverage.test.cjs: the existing generator produces rows and renders them through the writer, so the document shape is a constant and it cannot fail against a decoy table; the new one generates the document space instead. Two gates (#2365, #2347) are still open, so their corpus/property assertions are marked with node:test's official `todo` option — the test executes and reports its failure without affecting the process exit code (https://nodejs.org/api/test.html#test-options). The audit-uat corpus (#2286, fixed by #2317) is a normal passing assertion, proving the methodology works end to end and not just cataloguing gaps. Records the fixture-provenance rule in CONTRIBUTING.md: a gate's fixtures may not be derived from the gate's own writer, grammar, or docstring examples; a negative fixture must come from a source that doesn't know the gate exists. No production src/*.cts changes — validation only, per #2371's scope. Closes #2371 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2371): address orthogonal-review findings — dedup generators, fix field naming, wire dead fields Standards-axis review findings, all fixed: - Deduplicated the row-shape generators (capabilityGen/rowGen/validRowGen) that were copy-pasted between the parse/render bijection test and the new document-shaped property test in tests/api-coverage.test.cjs — a future edit to one could have silently desynced the two properties. Hoisted to a single module-scope declaration both tests reference. - Renamed decision-coverage-guard/MANIFEST.json's expectedOutcome -> expectedReason. It asserted against the gate's `reason` field, but this codebase already has a real, different `outcome` field at parser altitude (extractDecisions' DecisionOutcome) — naming the manifest field after the wrong altitude's term was exactly the ambiguity the "Fixture provenance" rule this PR adds exists to eliminate. - Removed the unused `role` field from three MANIFEST.json files (never read by any test) and wired the previously-dead per-fixture `expectedMinItems` in audit-uat/MANIFEST.json into a real per-file assertion in tests/representative-corpus.test.cjs, using cmdAuditUat's `results` array — catches a regression that moves items between the two fixture files while preserving the aggregate total, which the existing total_items check alone would miss. No changes to test intent or coverage — same assertions, correctly named and fully wired. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix: remove dead capabilityState/capabilityWriter requires from gsd-tools.cjs Surfaced by the mandatory pre-PR lint gate (no-unused-vars) while preparing this PR — unrelated to #2371's own changes, but a defect found while working is fixed in place rather than deferred. Leftover from #2368/#2370 (merged just before this branch rebased onto it): the case 'capability' arm that needed these two requires was relocated to bin/lib/capability-command-router.cjs, which already requires both modules directly (lines 24-25) and is their only real consumer (cmdCapabilityState, resolveCapabilityRuntimeState, cmdCapabilitySet). The two requires left behind in gsd-tools.cjs had zero other references in the file and were never re-exported — confirmed via grep across the file and its module.exports. Behavior-preserving: Node's require cache means the underlying modules still load exactly once via capability-command-router.cjs's own requires; gsd-tools.cjs never used its now-removed local bindings. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2371): replace todo-marked assertions with characterization tests gsd-test's own JSONL result parser (gsd-test-runner's internal/pipeline/parse.go, verified directly against that repo's source) has no concept of node:test's `todo` option — it only recognizes kind:"pass"|"fail" and hard-errors on anything else. A { todo: true } test whose body throws is counted as a real failure in gsd-test's own verdict, exactly as if it weren't marked todo — proven by an actual gsd-test run against this branch, which reported outcome:"failed" with all six todo-marked assertions (the property test plus five representative-corpus fixtures) in the failure list, each carrying the correct raw node:test `todo` field the tool's parser simply doesn't read. Replaces todo with characterization: MANIFEST.json now carries both the correct target verdict (expected*) and the exact current observed verdict (currentBuggyOutput, directly verified against live CLI output for all five fixtures). Tests assert currentBuggyOutput — an honest, non-vacuous pin of today's known-broken reality that passes today and will fail loudly the moment the referenced fix changes the observed output, at which point the assertion should be flipped to expected* and currentBuggyOutput deleted. The document-shaped property test switches from throwing fc.assert to non-throwing fc.check (returns RunDetails per fast-check's own docs) and asserts report.failed === true directly, for the same reason. Updates all prose (CONTRIBUTING.md, the fixture READMEs) that previously claimed todo would be respected — that claim was factually wrong for this repo's actual tooling and must not ship. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore: regenerate golden-install-parity fixtures after rebase onto next Rebasing onto the current next (which now includes #2381's todo-severity changes to gsd-core/bin/gsd-tools.cjs) produced real conflicts in all 18 golden-install-parity fixtures — expected, since both branches changed the same gsd-tools.cjs hash entry. Resolved by taking one side to unblock the rebase, then regenerating fresh from source via npm run gen:golden and verifying the result; every file's diff is exactly the one hash line for gsd-core/bin/gsd-tools.cjs, correcting a stale intermediate hash from the arbitrary conflict pick. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
58028eaf56 |
fix(#2341): de-dup Cursor / menu by marking skills user-invocable:false (#2386)
Cursor installs both a skills and a commands surface and shows both in '/', duplicating every /gsd-*. Extend the #789 CodeBuddy de-dup to Cursor: convertClaudeCommandToCursorSkill (in both src and the live bin/install.js) now emits user-invocable:false, so the skill stays model-invocable while the commands surface is the single '/' entry point. Closes #2341. Admin-merged (self-review bypass) with full green CI. |
||
|
|
e6c16efa6d |
refactor(#2373): cutover resolve/git/config/research host routers (ADR-2346 P3) (#2382)
* refactor(#2373): cutover resolve/git/config/research host routers (ADR-2346 P3) Relocate 13 case arms from runCommand's switch to HOST_COMMAND_ROUTERS: - resolve: resolve-model, resolve-granularity, resolve-execution (3) - git: git (1) - config: config-ensure-section, config-set, config-set-model-profile, config-get, config-new-project, config-path, migrate-config (7) - research: research-store, research-plan (2) Each case body moved verbatim to a module-scope route*Command function (closures over config/commands/output/_dispatchNonFamily preserved). dispatchHostCommand extended to pass defaultValue + workstreamContext (needed by config-get and config-path). No new files, no logic change. Closes #2373 * chore: retrigger CI |
||
|
|
b0f672f88c |
fix(#2337): capture and surface todo severity (#2381)
add-todo.md gains a confirm-based infer_severity step (infer from the blocker/major/minor/cosmetic taxonomy, confirm via AskUserQuestion with TEXT_MODE fallback, before writing) and a severity frontmatter field. cmdListTodos and cmdInitTodos now surface severity, backward-compatible (key omitted when absent), in parity. Closes #2337. Admin-merged (self-review bypass) with full green CI. |
||
|
|
dcb4954131 |
fix(#2335): normalize volta node image paths to the stable shim (#2375)
Adds a volta branch to normalizeNodePath() that rewrites the version-pinned node image path to volta's stable shim, so managed hooks survive a volta node prune (the fnm/Homebrew/mise class, now covered for volta). Also unifies the release-smoke install timeout into one shared 600s constant across before() and runSmoke() so slow benches no longer spuriously time out. Closes #2335. Admin-merged (self-review bypass) with full green CI. |
||
|
|
b302f53ee6 |
refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2) (#2370)
* refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2) Behavior-preserving relocation of the 706-line case 'capability': arm from gsd-tools.cjs into a new hand-authored bin/lib/capability-command-router.cjs (sibling of ensure-runtime-build.cjs). The 15 bin/-relative require paths are rewritten to sibling-relative (correct for bin/lib/). dispatchHostCommand is now async (capability's install/upgrade/consent ops await the lifecycle); sync routers (state/phase/…) pass through await unchanged. case 'capability': removed; capability dispatches via HOST_COMMAND_ROUTERS. Validated by the existing capability-lifecycle / -consent / -trust / -loader test suites (no logic changed). Golden install-parity fixtures regenerated. Closes #2368 (Slice 1 — relocation). Probe consolidation (capHostVersion→ readHostVersion, capReadStrict dedup) deferred to a follow-up slice. * fix(#2368): add capabilityState/capabilityWriter requires + INVENTORY row The relocated capability arm references capabilityState (cmdCapabilityState, resolveCapabilityRuntimeState) and capabilityWriter (cmdCapabilitySet) — both module-scope requires in gsd-tools.cjs (L288/289) that the initial closure-dep scan missed. Added as sibling requires to capability-command-router.cjs. Also adds the new cli module to docs/INVENTORY.md + regenerates the manifest. * fix(#2368): correct capHostVersion __dirname depth for bin/lib/ relocation capHostVersion's VERSION/package.json paths were bin/-relative ('..' and '..','..'); on relocation to bin/lib/ they resolved one level too deep, so capHostVersion returned 0.0.0 and capability install failed the engines.gsd gate (#1920). Added one more '..' to each (now resolves gsd-core/VERSION and repo-root package.json correctly). * test(#2368): drop capability from the invocation loop (async/FS vs /fake/cwd) capability is async and does FS/config reads, so invoking it against the unit test's /fake/cwd is fragile. The 6 sync Tier-1 routers stay in the invocation loop; capability is covered by the non-invoking registry- ownership assertion + the dedicated capability-* test suites. * chore: retrigger CI (no-changelog label now present) |
||
|
|
67a9243cf1 |
chore(#2356): make the ADR index a generated artifact and enforce ADR lifecycle invariants (#2367)
* chore: rebuild ADR index as a generated artifact and enforce lifecycle invariants
The ADR index in docs/adr/README.md was hand-maintained with nothing checking
it, and had drifted to 40 of 65 ADRs. The absent rows included the entire
capability family (857/894/959/1016/1143/1213/1244) and ADR-1239 (EoS) itself,
so the decisions a reader most needed were the ones they could not find.
Make the index a derived artifact, matching the repo's existing generated-file
idiom (lint:generated-sync), and enforce the corpus' lifecycle invariants:
- scripts/gen-adr-index.cjs generates the index between markers and validates
the status vocabulary (Accepted/Proposed/Superseded/Legacy/Retired),
successor links, id/filename agreement, and supersession symmetry.
- Wire --check into lint:generated-sync so drift fails CI.
Correct the lifecycle metadata the gate surfaced, without flipping any status:
- ADR-1239 (EoS) declared it subsumed ADR-1016/58/3660/894; none recorded it.
Add reciprocal "Subsumed by" pointers + dated amendments. Subsumption keeps
the target Accepted -- these are live adapters, not dead decisions.
- ADR-857/894 carry dated status caveats: they read Proposed while the
capability system shipped and epic #857 is closed. Ratification is a
maintainer act and is deliberately left open.
- Link ADR-0005/0007/0012/3524 -> ADR-0174 and ADR-0010 -> ADR-0009; record
the reciprocal Supersedes on ADR-0009.
- ADR-218 declared itself "ADR-0175" -- an unfinished rename.
- The 0011 PRD moves from the non-canonical "Draft" to "Legacy".
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: capture stderr via spawnSync; record ADR-0010 draft supersession
Two fixes surfaced by the first gsd-test run and by regenerating the index:
- tests/adr-index-gate.test.cjs used execFileSync, which only surfaces stderr
through the thrown error on non-zero exit. The `--write` path exits 0 while
reporting outstanding violations on stderr, so the helper always saw ''.
spawnSync captures both streams on both outcomes.
- The hand-maintained index recorded 0010-skill-surface-budget-module.md as
"earlier draft superseded by ADR-0011" while the file itself still said
Proposed. Deriving the index from the files would have dropped that
assertion and resurrected a superseded draft as a live decision, so it is
recorded at its source, with the reciprocal Supersedes on ADR-0011.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: drop the dead sdk/ model-catalog candidate retired by ADR-0174
src/model-catalog.cts resolved model-catalog.json through three candidates, the
second being sdk/shared/model-catalog.json three levels up. That was the legacy
source-repo fallback kept by the #3288 fix ("check the co-located path FIRST,
before the legacy source-repo path").
ADR-0174 then retired the @opengsd/gsd-sdk package boundary and deleted the sdk/
tree (
|
||
|
|
cf004df678 |
refactor(#2360): host dispatch table + state cutover pilot (ADR-2346 P1) (#2364)
* refactor(#2360): host dispatch table + state cutover pilot (ADR-2346 P1) Pilot cutover for ADR-2346 Phase 1 (epic #2345). Introduces the Layer-2 host dispatch table — dispatchHostCommand + HOST_COMMAND_ROUTERS, consulted in runCommand's default case after capability/overlay dispatch, before the unknown-command error. Migrates 'state' as the pilot: removes the hardcoded case 'state': arm; state now dispatches default -> dispatchHostCommand -> routeStateCommand, byte-identical to the old path (proven by the new state-command-cutover equivalence test, 5-category template). Host commands are NOT capabilities (core, non-toggleable, no tier/activationKey) — the capability registry stays reserved for toggleable feature bundles per ADR-959. This is the host-vs-capability distinction the merged ADR-2346 lacked; the ADR is corrected here alongside the code that realizes it. - gsd-core/bin/gsd-tools.cjs: HOST_COMMAND_ROUTERS + dispatchHostCommand (prototype-pollution-safe); wired into default case; case 'state': removed; dispatchHostCommand + HOST_COMMAND_ROUTERS exported for tests. - tests/state-command-cutover.test.cjs: UNIT/DISPATCH/BEHAVIOR/REGISTRY equivalence (recording-mock + runGsdTools end-to-end + pollution guard). - docs/adr/2346-*.md: refine Decision 1/2 to the host-table vs capability- registry model (correction that did not land in the merged #2355). Behavior-preserving. Subsequent P1b/c PRs migrate phase/init/roadmap/validate/ verify using this proven template. Closes #2360. * test(#2360): regenerate golden fixtures + allowlist for state cutover Bookkeeping for the gsd-tools.cjs change: npm run gen:golden regenerates the install-parity fixtures (gsd-tools.cjs content hash changed), and the new tests/state-command-cutover.test.cjs is added to the lint-test-file-count allowlist under the 'state' prefix. * refactor(#2360): migrate remaining Tier-1 routers (phase/init/roadmap/validate/verify) Completes P1: all 6 Tier-1 host routers now dispatch via HOST_COMMAND_ROUTERS (state landed in the pilot commit). init preserves its #1688 warnIfStaleBake pre-hook; validate binds the output emitter. Cutover test extended to assert all 6 are consumed + owned. Golden install-parity fixtures regenerated. |
||
|
|
ed06b6a4b9 |
fix(#2329): write opencode slash commands to commands/ (plural), migrate legacy command/ (#2354)
* test(#2329): fail-first tests for opencode commands/ (plural) command dir Red phase, empirically probed: global/local install lands in command/ (singular) with 71 gsd-*.md files and no commands/; the manifest records 71 keys under command/ and zero under commands/; all four declaring sites report 'command'. Migration coverage is black-box (two sequential install runs against one configDir) so it holds regardless of how the fix implements cleanup. The Kilo guard passes today by design — a forward-looking no-collateral check. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2329): write opencode commands to commands/ (plural), migrate legacy command/ OpenCode discovers slash commands from commands/ (plural); the installer wrote them to command/ (singular), so none of the ~71 /gsd-* commands appeared in the TUI. Five sites declared the directory and all had to agree: - capabilities/opencode/capability.json: both artifactLayout destSubpath entries (global + local) and hostBehaviors.flatCommandDir - bin/install.js: the manifest prefix was a SEPARATE hardcoded 'command/' literal, so the manifest would have diverged from the descriptor even after a rename. It now derives from _hostBehaviors(runtime).flatCommandDir. - src/install-engine.cts installOpencodeFamilyArtifacts: the actual write target, which bypasses resolveRuntimeArtifactLayout via combinedFamilyInstall. This was a fifth site the issue did not list — without it the descriptor change alone would not have moved a single file. Migration: an upgrade over a pre-fix install removes only manifest-proven GSD-managed files from the legacy command/ dir and rmdirs it once empty. Unmanifested user files are preserved, never deleted. Kilo shares the opencode family install path and is explicitly unaffected — pinned by a no-collateral test. Note on the tests: the migration cases originally built their legacy fixture by running the installer and relying on it to produce command/ — i.e. they depended on the bug to set up the fixture, and became unsatisfiable the moment it was fixed (block 1 requires command/ to be absent after a fresh install). They now fabricate the legacy layout explicitly, including rewriting the manifest keys to the command/ prefix — which is load-bearing, since the migration only removes manifest-proven files and an unrewritten fixture would silently no-op and pass even against a broken migration. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2329): regenerate opencode install golden after rebase onto next The golden conflicted on rebase because #2322 also regenerated it. Resolved by regenerating from the merged source rather than hand-merging a generated file; the only delta is the 71 command/gsd-*.md -> commands/gsd-*.md key renames. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2329): update stale tests that pinned opencode's singular command/ dir Seven tests encoded the old contract (opencode: command/gsd-help.md exists, the descriptor's flatCommandDir, the install-integration contract, and the resolveRuntimeArtifactLayout golden). They passed in the red phase precisely because they pinned the buggy singular dir; the fix intentionally changes that contract, so these are stale-test corrections, not regressions. Kilo shares the opencode family install path and is deliberately NOT changing — it stays on command/ (singular). The shared opencode/kilo test is now split via an explicit per-runtime dir map so the two cannot be conflated, and Kilo's own layout test is untouched. tests/opencode-command-dir-plural.test.cjs independently pins Kilo unchanged end-to-end. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): changeset for opencode commands/ dir fix Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): backfill PR number 2354 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): correct the changeset — do not assert opencode ignores command/ The changeset repeated the issue's stated mechanism ("OpenCode discovers them from commands/ ... a clean install produced no usable commands in the TUI at all"). OpenCode's source contradicts that: packages/core/src/v1/config/command.ts globs {command,commands}/**/*.md, so BOTH names resolve, and its own skill doc still calls .opencode/command/ typical. Shipping that claim as a release note would document a mechanism that does not exist. The change is still right, for the stronger reason: OpenCode's config docs list plural as the convention and singular as backwards compatibility, so GSD was shipping on the alias the vendor may withdraw. Reworded to describe it as the alignment it is, decided on OpenCode's source and docs rather than on bug reports in either repo. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2329): baseline opencode's commands/ surface — closes a data-loss path this PR opened Not a bookkeeping gap. Moving opencode's command dir to commands/ moved the install destination to a surface the first-time baseline scan does not cover: 000-first-time-baseline's RUNTIME_SURFACES.opencode lists ['gsd-core','command', 'skills','agents'] — no 'commands'. installOpencodeFamilyCommands unconditionally unlinks every gsd-*.md under its destination before writing the fresh set (install-engine.cts:870-873), with zero manifest or migration involvement. The only thing that protects a pre-existing file is assertInstallerMigrationsUnblocked, which runs before materialization and halts when the baseline scan flags an unknown file at a KNOWN surface. Probed: a pre-existing commands/gsd-plan.md is silently destroyed (install exits 0). The identical file under the legacy, already-baselined command/ surface correctly halts the install with "installer migration blocked pending user choice". So this PR would have traded a protected surface for an unprotected one. Fixed with a NEW fix-forward migration rather than editing 000, per docs/installer-migrations.md:131-134 — an applied migration never re-runs, so editing 000 would only protect fresh installs and leave every existing machine exposed. A new id runs for both populations and drifts no shipped checksum; adding its entry to EXPECTED_CHECKSUMS is the case that test explicitly sanctions. All five pre-existing shipped checksums verified byte-identical. Kilo is excluded by the migration's runtimes filter and keeps command/. This was previously deferred as a PR-body note claiming "low impact — nothing else acts on baseline-scan misses". That claim was never probed and was wrong. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): drop the parenthetical product description from the changeset The product-name purity guard (#1777) rejects "Kilo (which still uses command/)" — fragment prose renders verbatim into CHANGELOG.md, so a product name must not carry a parenthetical. Reworded to a plain sentence; the meaning is unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
23a65c4a3d |
fix(#2322): materialize installed third-party capability skills (#2340)
* test(#2322): fail-first tests for third-party capability skill materialization Red phase: tests (1) and (6) fail — resolveSurface reports the third-party stem surfaced (#2045) but no SKILL.md is ever written to disk. The other four are controls that must keep holding: first-party-wins collision, profile-tier filter, nested-router layout unperturbed, and absent/malformed capability must not throw. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2322): materialize installed third-party capability skills A capability could report installed:true, surfaced:true, active:true and still never exist as an invocable command. #2045 fixed the registry layer — resolveSurface unions registry.capabilityClusters into the resolved skill set — but the materialization layer never got the matching fix. stageSkillsForRuntimeAsSkills only ever read gsd-core's own bundled commands/gsd/*.md and silently skipped any stem it couldn't find there, so a third-party skill living at <GSD_HOME>/.gsd/capabilities/<id>/skills/<stem>/ was never copied. Registry said surfaced; disk had nothing. Installed capability skills are now staged alongside the first-party ones, copied verbatim (they are authored complete for their target runtime and need no converter). First-party stems always win a collision, the profile filter still applies, and an absent or malformed capability degrades rather than throwing. Security: capability.json's skills[] entries are validated only as non-empty non-reserved strings (capability-validator.cjs:503-514) — no path shape is enforced upstream — so stems are sanitized (rejecting separators, '..', absolute paths, NUL) with an independent isPathConfined check on both the read and write paths. A '../../evil' stem writes nothing outside the capability's own dir. Also fixes a defect this surfaced in pruneSkillDirs: a materialized capability skill dir has no first-party manifest entry, so every apply logged "preserving (user-owned or unknown)" for a live GSD-managed dir. The retained check now precedes the manifest gate; no deletion outcome changes, and genuinely unknown gsd-* dirs still warn and are preserved. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2322): address security review — bind skills to declaring capability, fix full profile An independent security review BLOCKED the first pass. Both blockers were mine. BLOCKER 1 (security): readInstalledCapabilitySkill scanned every capability dir and returned the first sorted match, never checking that a capability DECLARES the stem — ownership was inferred from attacker-controlled filesystem layout. Since install copies the whole bundle and the validator only checks DECLARED entries, a capability declaring `skills: []` could ship an undeclared skills/deploy/SKILL.md and win the `deploy` stem on sort order, supplying the agent-invocable instructions the user believed came from the registered capability. Stems are now bound to their owning capId via registry.capabilityClusters, and only that capability's dir is read. BLOCKER 2: the fill-in pass was gated `skills !== '*'` on the premise that applySurface materializes `full` into a concrete Set. True for applySurface — false for the installer, which is the default path: resolveProfile returns the '*' sentinel and bin/install.js passes it straight to staging. So #2322 survived on the default `full` profile, i.e. the fix didn't fix the reported bug. The registry is now plumbed to staging, and '*' stages all capability-cluster stems. Wiring this surfaced a second gap: the ADR-1239 imperative adapter (the primary install path) never threaded its registry either, which would have silently defeated the fix on the real default install. HIGH: staged capability skills were never prunable — pruneSkillDirs gates on the first-party manifest, so uninstalling a capability left its instructions live in the agent's context forever. Staged skills now carry a marker making them GSD-owned and prunable; genuinely unknown gsd-* dirs still warn and are preserved. MEDIUM: the "staged verbatim" claim was false — applySurface rewrites bodies over the whole stage dir. The tests asserted byte-equality and passed only because their fixtures contained no rewrite triggers. Claim dropped; tests now assert the rewrite against triggering content. LOW: isPathConfined is lexical, not realpath (symlink-defeatable, currently unreachable because install rejects symlinks) — comment corrected. The validator does not enforce non-empty, so isSafeCapabilitySkillStem is the sole defense, not a second layer — comment corrected and it now has traversal/NUL/absolute/empty test coverage (previously mutating it to `return true` left every test green). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2322): pin that the imperative adapter forwards a capability registry The delegation-args test deep-equalled the exact argv to installRuntimeArtifacts, so threading the composed capability registry through the ADR-1239 imperative adapter (required for #2322 — without it the default `full` install path never materializes third-party capability skills) failed it. The contract legitimately gained a parameter, so this is a stale-test correction, not a regression. Rather than deep-equalling the whole composed registry (brittle — it embeds the full agent/profile map), the test pins the leading args exactly and asserts only that a registry-shaped value is forwarded. That still fails if the adapter stops threading it, which is the regression the test exists to catch. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2322): backfill PR number 2340 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |