diff --git a/.changeset/nimble-orcas-tumble.md b/.changeset/nimble-orcas-tumble.md new file mode 100644 index 000000000..11f9996c6 --- /dev/null +++ b/.changeset/nimble-orcas-tumble.md @@ -0,0 +1,5 @@ +--- +type: Security +pr: 3516 +--- +**Capability installs no longer fetch from internal hosts, and unpinned installs say so in the consent prompt** — the URL importer refuses loopback/link-local/metadata hosts (including the cloud metadata addresses and localhost) before any bytes leave, and an http:// tarball URL fails with a clear https-only reason instead of a raw protocol error. Installs without an integrity pin now show a distinct 'NO PINNED HASH — staged unverified' line in the consent disclosure. (#3514) diff --git a/.pr-body-3514.md b/.pr-body-3514.md new file mode 100644 index 000000000..248e99da9 --- /dev/null +++ b/.pr-body-3514.md @@ -0,0 +1,73 @@ +## Fix PR + +> **Using the wrong template?** +> — Enhancement: use [enhancement.md](?template=enhancement.md) +> — Feature: use [feature.md](?template=feature.md) + +--- + +## Linked Issue + +> **Required.** This PR will be auto-closed if no valid issue link is found. + +Fixes #3514 + +> The linked issue must have the `confirmed-bug` label. If it doesn't, ask a maintainer to confirm the bug before continuing. + +--- + +## What was broken + +Three hardening gaps at the ADR-1244 D3/D5 edges (epic #1900, finding F21): the capability URL-import fetcher never validated the resolved host (a loopback or cloud-metadata **https** endpoint was fetched like any URL); an `http://` tarball spec failed with a raw `ERR_INVALID_PROTOCOL` instead of a named reason; and an install without an integrity pin produced a consent prompt that did not distinguish verified from unverified content. + +## What this fix does + +- **Pre-transport fetch gate** (`realHttpsGet` → `assertFetchableUrl`): loopback/link-local/metadata/unspecified hosts (`127/8`, `169.254/16` incl. `169.254.169.254`, `0/8`, `::1`, `fe80::/10`, `::`, IPv4-mapped spellings in both dotted and hex-normalized forms) and `localhost`/`*.localhost` names are denied with a named error **before any bytes leave** — the injected transport is provably never invoked (tests assert call-count 0). The WHATWG URL parser normalizes alternate IP spellings (decimal `2130706433`, hex `0x7f000001`, octal, short forms) to dotted-quad before the gate sees them — locked by tests. +- **Non-https, per the adopted split decision**: `parseSpec` still *classifies* `http://` tarball specs (internal-mirror flows are not broken at parse); the fetcher refuses with a clear https-only reason naming the mirror alternative. +- **Unverified-integrity disclosure**: `evaluateInstallTrust` accepts `integrityPinned` (a verified `--integrity` pin, or a git `#sha:` ref); the consent prompt renders `content: NO PINNED HASH — staged unverified` when no pin was supplied (and the pinned counterpart when one was). The status is **prompt-only** — deliberately excluded from `disclosureSignature` (consent's content binding is `bundleContentHash`, #1459; tests lock that it can never fire a spurious re-consent). Legacy callers see byte-identical output. + +## Root cause + +The D3 fetcher was built with a byte cap and a timeout but no host policy — the finding is an edge the pipeline's safe posture simply wasn't applied to. The integrity line was missing because the disclosure was built from the manifest alone; the pin fact lives on the resolve path, and nothing threaded it into the verdict. + +## Testing + +### How I verified the fix + +- Failing-first: `gsd-test` at the tests-only commit (holodeck, linux-node24) — verdict `failed` with exactly the 20 expected failures (12-case denylist matrix, https-reason, trust/lifecycle integrity suites), 34,343 green. +- GREEN: full `gsd-test` matrix at the final HEAD — verdict below. +- Controls lock the negative space: a public host and a private-range mirror still fetch through the gate; `parseSpec` still classifies `http://`; legacy `summarizeDisclosure` output is line-identical. + +### Regression test added? + +- [x] Yes — added a test that would have caught this bug + +### Platforms tested + +- [x] macOS +- [ ] Windows (including backslash path handling) +- [x] Linux + +### Runtimes tested + +- [ ] Claude Code +- [ ] Gemini CLI +- [ ] OpenCode +- [ ] Other: ___ +- [x] N/A (not runtime-specific — engine-internal source resolver + trust gate) + +--- + +## Checklist + +- [x] Issue linked above with `Fixes #3514` — **PR will be auto-closed if missing** +- [x] Linked issue has the `confirmed-bug` label +- [x] Fix is scoped to the reported bug — no unrelated changes included +- [x] Regression test added (or explained why not) +- [x] All existing tests pass (`npm test`) — full `gsd-test` matrix at final HEAD +- [x] `.changeset/` fragment added — `Security` type +- [x] No unnecessary dependencies added + +## Breaking changes + +None intended. `http://` tarball URLs already failed at the transport; they now fail with a better message. Denied hosts (loopback/link-local/metadata/localhost) are not legitimate install sources. **Known limits, documented** in `docs/explanation/capability-trust-model.md`: RFC1918 private ranges are deliberately *not* denied (internal mirrors); the check is on the URL's host literal — DNS rebinding is out of scope. Local/unpinned-git installs render the unverified line even though their trust basis is the path/commit (honest, slightly noisy). diff --git a/CONTEXT.md b/CONTEXT.md index df6a1c44a..5641adb4b 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -303,7 +303,7 @@ Human-facing discoverability catalog (`docs/registries/reviewer-registry.md`, ge Shared conformance validator (`gsd-core/bin/lib/capability-validator.cjs`, ADR-1244 D2) extracted from `scripts/gen-capability-registry.cjs` so the build-time generator and the runtime overlay loader share one validator implementation. Exports the same `validateCapability(manifest)` surface consumed by both the generator (build-time) and `capability-loader.cjs` (runtime). Generative-parity is CI-guarded: a drift between the generator's validation logic and the extracted module is a hard failure. Callers that previously inlined validation against the generator's internal helpers are migrated to import this module directly. Source of truth: `gsd-core/bin/lib/capability-validator.cjs`. ### Capability Source Resolver -ADR-1244 D3 fetch-and-stage seam (`gsd-core/bin/lib/capability-source.cjs`). Primary interface: `resolveCapabilitySource(spec, opts) → { id, version, stagedDir, integrity, source }`. Parses specs via `parseSpec` and dispatches to one adapter per source kind: `local` (fs copy from a `./`-prefixed path), `git` (clone `--depth 1` + checkout via `execGit`; https/ssh/git transports only — `ext::` and `file://` are rejected), `npm` (pack via `execNpm --ignore-scripts` + tar extract — NEVER `npm install`, no lifecycle scripts; shell-metacharacter spec rejection for Windows shell safety), `tarball` (HTTPS download + sha512 integrity verify BEFORE extraction + tar extract), and `registry` (explicit stub — no first-party endpoint yet). Security contract: install never executes capability code (copy/extract only); integrity is verified before staging; tar-slip member paths and symlinks are rejected. Staging is atomic: a per-pid/timestamp scratch directory under `.staging/` is renamed into `$GSD_HOME/.gsd/capabilities//` on success and removed on failure. The Phase 1/2 validator suite runs on the fetched manifest before finalizing; `engines.gsd` is pre-checked. Test seam: `_setCapabilitySourceHttpGet`. +ADR-1244 D3 fetch-and-stage seam (`gsd-core/bin/lib/capability-source.cjs`). Primary interface: `resolveCapabilitySource(spec, opts) → { id, version, stagedDir, integrity, source }`. Parses specs via `parseSpec` and dispatches to one adapter per source kind: `local` (fs copy from a `./`-prefixed path), `git` (clone `--depth 1` + checkout via `execGit`; https/ssh/git transports only — `ext::` and `file://` are rejected), `npm` (pack via `execNpm --ignore-scripts` + tar extract — NEVER `npm install`, no lifecycle scripts; shell-metacharacter spec rejection for Windows shell safety), `tarball` (HTTPS download + sha512 integrity verify BEFORE extraction + tar extract), and `registry` (explicit stub — no first-party endpoint yet). Security contract: install never executes capability code (copy/extract only); integrity is verified before staging; tar-slip member paths and symlinks are rejected. Staging is atomic: a per-pid/timestamp scratch directory under `.staging/` is renamed into `$GSD_HOME/.gsd/capabilities//` on success and removed on failure. The Phase 1/2 validator suite runs on the fetched manifest before finalizing; `engines.gsd` is pre-checked. #3514 (epic #1900 F21): `realHttpsGet` gates every fetch PRE-transport — loopback/link-local/metadata/unspecified hosts (127/8, 169.254/16 incl. cloud metadata, 0/8, ::1, fe80::/10, ::, v4-mapped spellings) and localhost-form names are denied with a named error, non-`https` URLs fail with a named https-only reason while `parseSpec` still CLASSIFIES `http://` tarball specs (no parse-time hard reject — internal-mirror flows; the transport never fetches plaintext), and RFC1918 ranges are deliberately allowed (internal mirrors; the denylist is not an allowlist; host-LITERAL check only — DNS rebinding out of scope). Test seams: `_setCapabilitySourceHttpGet` (whole fetch) and `_setHttpsGetImpl` (low-level transport, exercises the gate). ### Capability Ledger ADR-1244 D4 per-runtime install manifest (`gsd-core/bin/lib/capability-ledger.cjs`). Leaf module (only `node:fs`/`node:path` plus `shell-command-projection`'s `platformWriteSync`). Records `{ id, version, source, integrity, files[], sharedEdits[{file,marker}] }` per installed capability in `.gsd-capabilities.json` at the runtime config dir root. Exports: `readLedger` (structural-validated, never throws), `writeLedger` (atomic via `platformWriteSync`), `recordInstall` (idempotent, prototype-pollution-guarded), `removeEntry`, and `reconcile` (reports orphans whose `files[]` are missing on disk; hardened against non-string/`..` members; never mutates). Serves as the atomic commit point for Phase-4 upgrade/remove and the reconciliation basis for detecting stale entries after out-of-band deletions. @@ -315,8 +315,7 @@ Issue #1459 user-owned consent seam (`gsd-core/bin/lib/capability-consent.cjs`). Issue #1459 finding 4 shared cross-process lock primitive (`gsd-core/bin/lib/capability-lock.cjs`). Leaf module (`node:fs`/`node:path`/`node:os`/`node:crypto` + the ledger's bounded `readSmallRegularFile` + `shell-command-projection`'s `execTool` for the rare start-time shell-out). THE single hardened lockfile protocol shared by BOTH `capability-lifecycle` (the `.gsd/capabilities/.lock` mutation lock) and `capability-consent` (the consent-store `.consent.lock`) — extracted so the two locks cannot diverge (mirrors the shared-validator / shared bounded-reader lessons). Exports: `acquireLock(lockPath, opts?)` (O_EXCL create with a JSON `{token,pid,hostname,startTime,ts}` body; steal protocol binds age to the body's own `ts`, never stale-steals a VERIFIED-LIVE same-host holder — pid alive AND recorded start-time matches the pid's current start-time, defeating pid-reuse without ever stealing a live holder — and reclaims only a dead/unverifiable holder via the dead-pid fast path or the hard `LOCK_DEADMAN_MS` deadman; `opts.maxAttempts` raises the bounded retry budget and `opts.waitForFresh` makes a contended fresh/live holder be WAITED FOR rather than failed-fast so genuinely-racing consent writers serialize), `releaseLock(handle)` (token + inode owner-safe — never deletes a successor's lock), `getProcessStartTime`, and the `_setLockProbes`/`_resetLockProbes` test seams. Carries the #1462 lifecycle-lock invariants (process-start-time liveness, TOCTOU-safe pre-rename identity recheck, bounded iterative loop). ### Capability Trust Gate -ADR-1244 Phase 4 (D5) PURE policy module (`gsd-core/bin/lib/capability-trust.cjs`). Computes *what* a capability would do and *whether* policy permits it; performs no mutation and no I/O beyond existence-checking declared artifacts. Exports: `discloseExecutableSurfaces(manifest, stagedDir?, resolveHost?)` (enumerates the four executable surfaces — `hooks`, command modules, `mcpServers`, and reviewer lanes (ADR-2782 D5) — plus a fifth, non-executable class, instruction surfaces (`skills` stems only, ADR-2363 D5, #3248), returned as `instructionSurfaces`; flags `hasExecutable` from the four executable classes only — instruction surfaces deliberately do NOT contribute to it; a reviewer lane is the one class that *receives* data, so it discloses its binary + full args (spawn) or destination host + `hostConfigKey` (openai-http) together with the egress payload classes); `evaluateInstallTrust(args)` (composes source policy + reserved-namespace + engines gate + disclosure into `{ allowed, requiresConsent, disclosure, engines, blockReasons }`); `evaluateSourceAllowed(parsed, strictKnownRegistries)` enforcing `capabilities.strict_known_registries` (unset/null → permissive-with-consent; `[]` → block all external; non-empty → host-based allowlist, never substring); `checkEngines(manifest, hostVersion)` (engines.gsd hard gate via `semverSatisfies` + `compatVersions` graceful-downgrade picking the newest working version); `executableSetChanged(old, new)` (auto-update re-consent trigger); `checkReservedNamespace` (`gsd-`/`gsd-core-`/`anthropic-`); `collectInstructionSurfaces(manifest)` (the instruction-surface collector — `skills` stems only, independently testable, same total/`safeCollect` contract as the four executable collectors; ADR-2363 D3 classifies `agents` as an instruction surface too, but a third-party capability's declared `agents[]` are never staged into the agent's instruction context — `stageAgentsForRuntimeWithConverter` (`src/install-profiles.cts`) has no registry-aware third-party path the way `readInstalledCapabilitySkill` gives skills — so disclosing them would name a surface that does not exist; agents stay unimplemented pending a maintainer decision, and are NOT thereby safe or inert, only undisclosed); `summarizeInstructionSurfaces(disclosure)` (renders the instruction-surface section of the consent summary; called from BOTH branches of `summarizeDisclosure` because a skill-only capability has `hasExecutable === false` and takes the early return, so a section appended only at the end would never render for exactly the capabilities that need it). The MCP disclosure also captures each server's `env` (string→string, filtered) and `cwd` (#1459) — `disclosureSignature` folds them in as STABLE SORTED JSON so any env/cwd add/change forces re-consent while a key reorder does not; `signatureForManifest(manifest, stagedDir?)` is the single source of truth for that signature (consumed by the loader's consent check and the lifecycle's consent binding). #1459 finding 5: each MCP surface also carries `rawConfig` — the FULL declared server config the writer persists (`{...config}`), prototype-pollution-cleaned — folded into the signature as STABLE SORTED JSON so a change to ANY persisted field (not just the explicit whitelist — a future `envFile`/`workingDir`/launch option) forces re-consent, while a pure key reorder does not; the human summary stays readable via the key fields only. #3515 (epic #1900 F20): the MCP disclosure section additionally renders an INTENTIONALLY-NOT-CONFINED notice for every SPAWNED (stdio) server — command/args/cwd are written verbatim and may point anywhere on the machine, unlike the confined hook path (capability-lifecycle's MCP write documents the same asymmetry in code); remote-only (http/sse) servers render no notice (nothing local is spawned), and the line introduces NO new Disclosure field so disclosureSignature is untouched. Instruction surfaces are deliberately EXCLUDED from `disclosureSignature` (ADR-2363 D4, #3248) — a manifest gaining, losing, or changing `skills` produces a byte-identical signature and disturbs no stored consent record; any future binding arrives as a versioned v2, never an in-place re-encoding of v1. The barrier is consent + integrity + reversibility, NOT a sandbox — see `docs/explanation/capability-trust-model.md`. - +ADR-1244 Phase 4 (D5) PURE policy module (`gsd-core/bin/lib/capability-trust.cjs`). Computes *what* a capability would do and *whether* policy permits it; performs no mutation and no I/O beyond existence-checking declared artifacts. Exports: `discloseExecutableSurfaces(manifest, stagedDir?, resolveHost?)` (enumerates the four executable surfaces — `hooks`, command modules, `mcpServers`, and reviewer lanes (ADR-2782 D5) — plus a fifth, non-executable class, instruction surfaces (`skills` stems only, ADR-2363 D5, #3248), returned as `instructionSurfaces`; flags `hasExecutable` from the four executable classes only — instruction surfaces deliberately do NOT contribute to it; a reviewer lane is the one class that *receives* data, so it discloses its binary + full args (spawn) or destination host + `hostConfigKey` (openai-http) together with the egress payload classes); `evaluateInstallTrust(args)` (composes source policy + reserved-namespace + engines gate + disclosure into `{ allowed, requiresConsent, disclosure, engines, blockReasons }`); `evaluateSourceAllowed(parsed, strictKnownRegistries)` enforcing `capabilities.strict_known_registries` (unset/null → permissive-with-consent; `[]` → block all external; non-empty → host-based allowlist, never substring); `checkEngines(manifest, hostVersion)` (engines.gsd hard gate via `semverSatisfies` + `compatVersions` graceful-downgrade picking the newest working version); `executableSetChanged(old, new)` (auto-update re-consent trigger); `checkReservedNamespace` (`gsd-`/`gsd-core-`/`anthropic-`); `collectInstructionSurfaces(manifest)` (the instruction-surface collector — `skills` stems only, independently testable, same total/`safeCollect` contract as the four executable collectors; ADR-2363 D3 classifies `agents` as an instruction surface too, but a third-party capability's declared `agents[]` are never staged into the agent's instruction context — `stageAgentsForRuntimeWithConverter` (`src/install-profiles.cts`) has no registry-aware third-party path the way `readInstalledCapabilitySkill` gives skills — so disclosing them would name a surface that does not exist; agents stay unimplemented pending a maintainer decision, and are NOT thereby safe or inert, only undisclosed); `summarizeInstructionSurfaces(disclosure)` (renders the instruction-surface section of the consent summary; called from BOTH branches of `summarizeDisclosure` because a skill-only capability has `hasExecutable === false` and takes the early return, so a section appended only at the end would never render for exactly the capabilities that need it). The MCP disclosure also captures each server's `env` (string→string, filtered) and `cwd` (#1459) — `disclosureSignature` folds them in as STABLE SORTED JSON so any env/cwd add/change forces re-consent while a key reorder does not; `signatureForManifest(manifest, stagedDir?)` is the single source of truth for that signature (consumed by the loader's consent check and the lifecycle's consent binding). #3514 (epic #1900 F21c): `evaluateInstallTrust` accepts optional `integrityPin` ('sha512' = a verified `--integrity` pin | 'git-commit' = a git `#sha:<7-40-hex-commit>` ref, hex-validated so a mutable `#sha:main` is NOT a pin | 'none') and sets a PROMPT-ONLY `Disclosure.integrityStatus` ('pinned' | 'commit-pinned' | 'unverified') rendered by `summarizeDisclosure` as a `NO PINNED HASH — staged unverified` line when no pin was supplied — deliberately EXCLUDED from `disclosureSignature` (consent's content binding is `bundleContentHash`, #1459; a rendered line must never fire a spurious re-consent), mirroring the ADR-2363 D4 instruction-surface exclusion one field over. #1459 finding 5: each MCP surface also carries `rawConfig` — the FULL declared server config the writer persists (`{...config}`), prototype-pollution-cleaned — folded into the signature as STABLE SORTED JSON so a change to ANY persisted field (not just the explicit whitelist — a future `envFile`/`workingDir`/launch option) forces re-consent, while a pure key reorder does not; the human summary stays readable via the key fields only. #3515 (epic #1900 F20): the MCP disclosure section additionally renders an INTENTIONALLY-NOT-CONFINED notice for every SPAWNED (stdio) server — command/args/cwd are written verbatim and may point anywhere on the machine, unlike the confined hook path (capability-lifecycle's MCP write documents the same asymmetry in code); remote-only (http/sse) servers render no notice (nothing local is spawned), and the line introduces NO new Disclosure field so disclosureSignature is untouched. Instruction surfaces are deliberately EXCLUDED from `disclosureSignature` (ADR-2363 D4, #3248) — a manifest gaining, losing, or changing `skills` produces a byte-identical signature and disturbs no stored consent record; any future binding arrives as a versioned v2, never an in-place re-encoding of v1. The barrier is consent + integrity + reversibility, NOT a sandbox — see `docs/explanation/capability-trust-model.md`. ### Capability Lifecycle ADR-1244 Phase 4 (D5+D6) orchestration seam (`gsd-core/bin/lib/capability-lifecycle.cjs`) composing the source resolver, ledger, and trust gate into the mutating operations. Exports: `installCapability` (pre-fetch source gate → resolve copy-only with `promote:false` → trust verdict → promote + apply marker-stamped shared edits → **ledger commit**; nothing written on block/abort), `upgradeCapability` (atomic stage-then-swap: old set aside, new swapped in, shared edits re-derived, **ledger committed**, backup dropped; re-prompts when the executable set changed), `removeCapability` (strip only `_gsdCapability`-marked shared-config entries — user hand-edits preserved — delete exactly the ledger-recorded files, then drop the entry; `CAPABILITY_DATA` preserved unless `removeData`), `reconcileCapabilities` (crash recovery driven by the ledger's `_pending {kind,backupName,sharedFiles}` INTENT — not a version comparison: roll an uncommitted upgrade back by restoring the backup, an uncommitted fresh install away entirely, and re-sync shared config from the winning bundle, guaranteeing no half-state), plus `applyCapabilitySharedEdits`/`stripCapabilitySharedEdits` (marker-isolated JSON edits, prototype-pollution-guarded). All four mutating ops + reconcile take a cross-process lock (`.gsd/capabilities/.lock`, atomic stale-steal) so a concurrent reconcile can't clear a live intent. Capability code never executes during any operation. The source resolver's `promote:false`/`skipEnginesGate` options are the seams that let this module own the swap/commit ordering and the engines gate (with `compatVersions` downgrade hint). @@ -593,8 +592,6 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr `RULESET.CONTRIB.CLASSIFY.fix=requires confirmed-bug before implementation (legacy 'confirmed' label is back-compat only for duplicate-sweep exemption, not a valid implementation gate)` `RULESET.CONTRIB.CLASSIFY.enhancement=requires approved-enhancement before implementation` `RULESET.CONTRIB.CLASSIFY.feature=requires approved-feature before implementation` -`RULESET.CONTRIB.GATE.DELETED-MODULE-DISPOSITION=a PR deleting a module/package/lineage with a surviving counterpart carries a symbol-level disposition table for the deleted side — every declared name (exported or local) marked migrated (location named) | renamed → X (conventions normalized: verifyX → cmdVerifyX) | dropped because Y; SCREAMING_CASE constants ranked first (policy lives there; both confirmed ADR-0174 losses were bounds); the surviving tests passing is NOT evidence — the surviving tests belong to the surviving side (#3484, ADR-0174 behavior-carry-forward amendment; losses: #3427 shortFormToId tier + unresolved-deps warning, #3477 regexForKeyLinkPattern 512-char cap + nested-quantifier screen, epic #3473 B5 MAX_JSON_SEARCH_DEPTH=48). Method with mandatory positive control: docs/how-to/audit-a-retiring-lineage.md` -`RULESET.CONTRIB.GAP-TRACKING=a known gap is tracked only by an issue — a code comment, TODO, or changeset note is attached to code and dies with the refactor that deletes it, while the gap stays real (worked example: archived changeset clever-yaks-cheer.md recorded the shortFormToId backfill as 'tracked as a follow-up parity gap' + a // KNOWN GAP: comment marked the call site; the SDK retirement deleted both, no issue was opened, and it surfaced a year later as #3427). If you write 'tracked as a follow-up' anywhere, open the issue first and cite its number. Prose twin of #3473 criterion B6 (a guard needs a retirement condition)` ## Workspace seams (machine-oriented predicates) diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index bcc079b3d..1e62d1287 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -198,52 +198,6 @@ Contributor requirements (summary): - **CI must pass** — all configured matrix jobs must be green. Node 24 is the compatibility floor and primary target; Node 26 compatibility must be preserved for code and tests even when a Node 26 CI lane is not yet available. - **Scope matches the approved issue** — if your PR does more than what the issue describes, the extra changes will be asked to be removed or moved to a new issue -### Deleting a module that has a surviving counterpart — disposition required - -A consolidation PR that deletes a module, package, or whole lineage — where any part of -the deleted side has a surviving counterpart in the tree — must carry a **symbol-level -disposition table for the deleted side**: for every declared name (exported *or* local), -one of: - -- `migrated` — the behavior exists at a named location in the surviving tree, -- `renamed → ` — including convention drift (`verifyX` → `cmdVerifyX`, - `preserveExistingProgress` → `shouldPreserveExistingProgress`), -- `dropped because ` — a deliberate deletion, with the reason stated. - -Rank **SCREAMING_CASE constants first** — that is where policy lives; both confirmed -losses from the one retirement this repo has performed were bounds -([#3484](https://github.com/open-gsd/gsd-core/issues/3484)). Local-variable renames are -noise and may be summarized. - -**"The surviving tests pass" is not evidence.** The surviving tests belong to the -surviving side and cannot observe behavior only the deleted side had. ADR-0174 retired -the SDK package and deleted ADR-3524's parity apparatus in the same migration; three -invariants went into the bin with the package and nobody noticed for months (#3427, -#3477, epic #3473 B5) — see ADR-0174's behavior-carry-forward amendment (#3484) for the record. - -**The audit is a method, not a guess.** The calibrated procedure — pairing by basename, -comment-stripped declared-name diffing, rename normalization, SCREAMING_CASE-first -ranking, and a **mandatory positive control** (the audit must re-detect a known loss or -it is not calibrated; an export-only scan misses `shortFormToId`, which was a *local* -`const` in a surviving function) — is written up step by step in -[`docs/how-to/audit-a-retiring-lineage.md`](docs/how-to/audit-a-retiring-lineage.md). - -### A known gap gets an issue — a comment is not a tracking mechanism - -Recording a known-but-unfixed gap in a code comment, a changeset note, or a TODO, and -treating that as *tracking* it, is an anti-pattern: all three are attached to code, and -a refactor deletes the code they annotate while the gap stays real. The worked example -is `.changeset/archived/clever-yaks-cheer.md` (PR #3798): it accurately recorded the -`shortFormToId` backfill as "tracked as a follow-up parity gap", a `// KNOWN GAP:` -comment marked the call site, then the SDK retirement deleted both — and no issue was -ever opened. The gap surfaced a year later as production bug #3427. - -The rule: **a gap that matters gets an issue. A gap with no issue is not tracked.** If -you find yourself writing "tracked as a follow-up" in a comment or a changeset body, -open the issue first, then cite its number. This is the prose-level twin of epic #3473's -criterion B6 (a guard needs a retirement condition): the tracking mechanism must -survive the artifact it is attached to. - ## CHANGELOG Entries — Drop a Fragment **Do not edit `CHANGELOG.md` directly.** Two PRs that both append to a `### Fixed` block always conflict on merge — git can't pick a serialization order without a human. Instead, every PR with user-facing changes drops a fragment file in `.changeset/`. diff --git a/docs/CONTEXT-INDEX.json b/docs/CONTEXT-INDEX.json index 14a07cb3a..4dad4ea2f 100644 --- a/docs/CONTEXT-INDEX.json +++ b/docs/CONTEXT-INDEX.json @@ -1,6 +1,6 @@ { "schemaVersion": 1, - "count": 263, + "count": 261, "classes": { "ARCH": 1, "CI": 2, @@ -17,7 +17,7 @@ "PROC": 14, "PROHIB": 10, "RELEASE-NOTES": 31, - "RULESET": 60, + "RULESET": 58, "SESSION": 9, "WAVE": 5, "WORKSTREAM": 5, @@ -979,16 +979,6 @@ "klass": "RULESET", "value": "requires confirmed-bug before implementation (legacy 'confirmed' label is back-compat only for duplicate-sweep exemption, not a valid implementation gate)" }, - { - "id": "RULESET.CONTRIB.GAP-TRACKING", - "klass": "RULESET", - "value": "a known gap is tracked only by an issue — a code comment, TODO, or changeset note is attached to code and dies with the refactor that deletes it, while the gap stays real (worked example: archived changeset clever-yaks-cheer.md recorded the shortFormToId backfill as 'tracked as a follow-up parity gap' + a // KNOWN GAP: comment marked the call site; the SDK retirement deleted both, no issue was opened, and it surfaced a year later as #3427). If you write 'tracked as a follow-up' anywhere, open the issue first and cite its number. Prose twin of #3473 criterion B6 (a guard needs a retirement condition)" - }, - { - "id": "RULESET.CONTRIB.GATE.DELETED-MODULE-DISPOSITION", - "klass": "RULESET", - "value": "a PR deleting a module/package/lineage with a surviving counterpart carries a symbol-level disposition table for the deleted side — every declared name (exported or local) marked migrated (location named) | renamed → X (conventions normalized: verifyX → cmdVerifyX) | dropped because Y; SCREAMING_CASE constants ranked first (policy lives there; both confirmed ADR-0174 losses were bounds); the surviving tests passing is NOT evidence — the surviving tests belong to the surviving side (#3484, ADR-0174 behavior-carry-forward amendment; losses: #3427 shortFormToId tier + unresolved-deps warning, #3477 regexForKeyLinkPattern 512-char cap + nested-quantifier screen, epic #3473 B5 MAX_JSON_SEARCH_DEPTH=48). Method with mandatory positive control: docs/how-to/audit-a-retiring-lineage.md" - }, { "id": "RULESET.CONTRIB.GATE.ORDER", "klass": "RULESET", diff --git a/docs/README.md b/docs/README.md index f3694f4a3..8344bcffa 100644 --- a/docs/README.md +++ b/docs/README.md @@ -24,7 +24,6 @@ Language versions: [English](README.md) · [Português (pt-BR)](pt-BR/README.md) - [Resolve edge-coverage findings](how-to/resolve-edge-coverage-findings.md) — turn the spec phase's surfaced domain-boundary edges into covered, dismissed, or backstopped spec decisions - [Resolve prohibition findings](how-to/resolve-prohibition-findings.md) — turn the spec phase's surfaced must-NOT constraints into resolved, dismissed, or deferred spec decisions - [Resolve an ESLint glob-coverage finding](how-to/resolve-eslint-coverage-findings.md) — bring a source file that matches no lint rule under coverage, or record a reasoned exemption -- [Audit a retiring lineage](how-to/audit-a-retiring-lineage.md) — produce the symbol-level disposition table a module-deleting consolidation PR must carry, with a mandatory positive control - [Plan a phase](how-to/plan-a-phase.md) — run research, decompose work, and verify plan quality - [Execute a phase](how-to/execute-a-phase.md) — run plans in parallel waves with fresh-context subagents - [Verify and ship](how-to/verify-and-ship.md) — walk through completed work, diagnose failures, and create the PR diff --git a/docs/adr/0174-retire-gsd-sdk-package-boundary.md b/docs/adr/0174-retire-gsd-sdk-package-boundary.md index c91ce6249..6c4a7117c 100644 --- a/docs/adr/0174-retire-gsd-sdk-package-boundary.md +++ b/docs/adr/0174-retire-gsd-sdk-package-boundary.md @@ -1,6 +1,6 @@ # ADR-0174: Retire @opengsd/gsd-sdk package boundary — single-runtime collapse -- **Status:** Accepted (2026-05-23); amended #1642 (2026-06-23) — §5 reconciled to as-built Result type + `exitReason?` field added on `InvalidArgs`; amended #3484 (2026-08-14) — behavior-carry-forward amendment appended +- **Status:** Accepted (2026-05-23); amended #1642 (2026-06-23) — §5 reconciled to as-built Result type + `exitReason?` field added on `InvalidArgs` - **Date:** 2026-05-23 - **Tracking issue:** [#174](https://github.com/open-gsd/get-shit-done-redux/issues/174) — sub-issues #175–#197 @@ -161,46 +161,3 @@ Seven phases, ~15–18 PRs total. Each phase is a coherent slice that leaves the | 7 — Land this ADR's PR | The PR for this ADR closes the umbrella tracking issue. | 1 (this PR) | Implementation is tracked in [#174 — sub-issues #175–#197](https://github.com/open-gsd/get-shit-done-redux/issues/174). - -## Amendment (2026-08-14): Behavior carry-forward for lineage-retiring consolidations (#3484) - -This amendment was appended after the migration completed. It changes nothing about the -collapse decision above — one runtime still cannot drift from itself, and ADR-3524's -parity apparatus is still correctly gone. What this ADR lacked was a clause about the -*instant of collapse*: the two lineages were not equivalent at the moment `sdk/` was -deleted (drift had been accumulating for months — ADR-3524 itself names nine prior -one-sided fixes), and the collapse took the CJS lineage as canonical. Everything the SDK -side had and the CJS side did not was deleted with the package — and the parity tests -that would have caught exactly that were deleted by Migration Plan Phase 3 in the same -motion. The surviving tests were the CJS side's tests, and they passed throughout. - -**The requirement.** A consolidation that retires a lineage must enumerate the retiring -side's behavior and record, per item, whether it was **`migrated`**, **`renamed → X`**, -or **`dropped because Y`** — before the deletion merges. "The tests pass" is explicitly -not sufficient evidence, because the surviving tests belong to the surviving side and -cannot observe behavior only the deleted side had. - -**The evidence that produced this amendment — three confirmed losses through the gap:** - -| Lost in the collapse | Surviving site | Consequence | -|---|---|---| -| `shortFormToId` resolution tier + the `unresolved depends_on reference` warning | `src/phase.cts` | [#3427](https://github.com/open-gsd/gsd-core/issues/3427) — short-form `depends_on` edges silently dropped; dependency-ordered plans execute in parallel | -| `regexForKeyLinkPattern`'s 512-char cap + nested-quantifier screen | `src/verify.cts` | [#3477](https://github.com/open-gsd/gsd-core/issues/3477) — ReDoS; `verify-phase` hangs on an untrusted plan pattern | -| `MAX_JSON_SEARCH_DEPTH = 48` | `src/intel.cts` | unbounded recursion in `matchesInValue`; stack overflow on deeply nested intel JSON (carried by epic #3473, B5) | - -None was noticed for months; #3427 surfaced only when a user hit it in production. The -`shortFormToId` loss is the sharpest: the gap was *known* — recorded as "tracked as a -follow-up parity gap" in an archived changeset (`.changeset/archived/clever-yaks-cheer.md`) -and marked with a `// KNOWN GAP:` comment — then the SDK was retired, the comment died -with the lineage it annotated, and no issue was ever opened. A comment is attached to -code; it is not a tracking mechanism. A gap that matters gets an issue. - -A fourth-loss audit run on 2026-08-14 (method recorded in [#3484](https://github.com/open-gsd/gsd-core/issues/3484) -and in the contributor how-to) found no further losses — which calibrates the audit -method, not a claim that the class is closed. - -**Where the rule lives now.** The general form of this requirement — not scoped to this -ADR — is a merge gate in `CONTRIBUTING.md` ("Deleting a module that has a surviving -counterpart"): the deleting PR carries a symbol-level disposition table, SCREAMING_CASE -constants ranked first, with a mandatory positive control. Future lineage-retiring -consolidations cite this amendment and satisfy that gate. diff --git a/docs/explanation/capability-trust-model.md b/docs/explanation/capability-trust-model.md index cac32ed45..c5f974675 100644 --- a/docs/explanation/capability-trust-model.md +++ b/docs/explanation/capability-trust-model.md @@ -288,6 +288,37 @@ An `integrity` field in `capability.json` carries a `sha512-` digest of the capability bundle. When present, GSD verifies this digest before extracting any files. A mismatch aborts the install. +When NO pin is supplied, the consent prompt says so plainly: a +`content: NO PINNED HASH — staged unverified` line distinguishes an install +whose bytes were verified against a commitment from one that was not +([#3514](https://github.com/open-gsd/gsd-core/issues/3514)). A computed +sha512 of what was actually fetched is still recorded in the ledger at +install, so a later `trust` inspection shows exactly which bytes landed. +Prompt claims are exact per kind: a sha512 `--integrity` pin renders as +*supplied and verified*, a git source checked out at a `#sha:` ref +renders as *pinned to a git commit* (never as a sha512 pin — none was +supplied), and a mutable `#sha:` ref is not a pin at all. + +### Fetch-host denylist + +The URL importer's fetch transport refuses, before any bytes leave +([#3514](https://github.com/open-gsd/gsd-core/issues/3514)): + +- **loopback, link-local, and unspecified hosts** — `127.0.0.0/8`, + `169.254.0.0/16` (which contains the cloud metadata addresses), `0.0.0.0/8`, + `::1`, `fe80::/10`, `::`, their IPv4-mapped IPv6 spellings, and + `localhost`/`*.localhost` names. No legitimate capability install fetches + these. +- **plaintext `http://` URLs** — the transport is `https`-only; an `http://` + tarball spec still *classifies* (so an internal-mirror workflow fails with a + clear, named reason instead of a raw protocol error) but never fetches. + +Deliberate limits: RFC1918 private ranges (`10/8`, `172.16/12`, +`192.168/16`) are **not** denied — an internal https mirror is a legitimate +install source, and the denylist is not an allowlist. The check is on the +URL's host literal; a public hostname that *resolves* via DNS to a denied +range (rebinding) is out of scope. + What integrity pinning defends against: a capability hosted at a URL or in a registry that is later replaced with a different bundle (whether by an attacker who has compromised the hosting, or by an author publishing a silent breaking diff --git a/docs/how-to/audit-a-retiring-lineage.md b/docs/how-to/audit-a-retiring-lineage.md deleted file mode 100644 index 72c9063ea..000000000 --- a/docs/how-to/audit-a-retiring-lineage.md +++ /dev/null @@ -1,96 +0,0 @@ -# Audit a retiring lineage (behavior carry-forward) - -**You need this if** your PR deletes a module, package, or runtime lineage — and any part -of the deleted side has a surviving counterpart in the tree. The merge gate -([CONTRIBUTING.md → "Deleting a module that has a surviving counterpart"](../../CONTRIBUTING.md)) -requires the PR to carry a symbol-level disposition table. This page is the calibrated -method for producing that table. It was first run against the ADR-0174 SDK retirement -and found three confirmed losses after the fact (#3484); run it *before* the deletion -merges and those losses become review findings instead of production bugs. - -## Why "the tests pass" proves nothing here - -The surviving tests belong to the surviving side. They were written against the -surviving implementation and pass before, during, and after the deletion — including -when the deleted side carried an invariant the survivor never had. ADR-3524 existed -precisely because the two sides had already drifted in nine known places; deleting one -side without enumerating it is choosing the survivor's omissions by default. - -## The audit, step by step - -All commands run against `` — the last commit **before** the retirement (for the -SDK retirement this was `04b3be683`, the parent of the retiring commit). - -### 1. List the deleted source files - -```bash -git ls-tree -r --name-only -- sdk/src -``` - -172 files for the SDK run. Everything under the deleted root is a candidate; test files -are excluded from pairing (they are the deleted side's claims, not its behavior). - -### 2. Pair each file with a surviving counterpart - -Pair by **basename**: `sdk/src/query/phase.ts` ↔ `src/phase.cts`. For the SDK run this -yielded 22 paired modules. Files with **no** counterpart are whole surfaces deleted on -purpose — list them in the disposition as one `dropped because …` line each (or grouped -by surface), but they are out of scope for symbol diffing. - -### 3. Diff *declared* names, not raw tokens - -Strip comments from both sides, then diff the declared names (functions, constants, -types — exported or local). **Do not tokenize the raw text**: raw token diffing matches -prose in comments and produced ~500 false positives on the first SDK run. - -### 4. Normalize rename conventions before reporting - -The two lineages used different naming conventions. Normalize before diffing: - -| Deleted-side name | Surviving-side convention | -|---|---| -| `verifyX` | `cmdVerifyX` | -| `preserveExistingProgress` | `shouldPreserveExistingProgress` | - -Skipping this step is what makes the raw diff unreadable — every convention-renamed -symbol reads as a loss. - -### 5. Rank what remains by shape - -**SCREAMING_CASE constants are the high-signal class.** That is where policy lives -(bounds, limits, tiers, thresholds), and both real losses the SDK audit confirmed were -bounds (`regexForKeyLinkPattern`'s 512-char cap; `MAX_JSON_SEARCH_DEPTH = 48`). Local -variable renames are noise — summarize them. - -### 6. Positive control — mandatory - -Before trusting the audit, verify it **re-detects a known loss**. For the SDK lineage the -control is `shortFormToId` — a *local* `const` inside a surviving function at -`sdk/src/query/phase.ts:609`, not an export. An export-only scan misses it entirely and -reports a clean bill of health; a calibrated run finds both halves of #3427 (the tier -`shortFormToId` and the diagnostic `unresolvedDeps`). If your audit method cannot find -the control, it is not calibrated — fix the method, do not ship the table. - -### 7. Write the disposition table into the PR - -One row per name that survived steps 3–5: `migrated` (name the location), `renamed → X`, -or `dropped because Y`. Unpaired whole surfaces get their own grouped `dropped because` -rows. The table goes in the PR body where the reviewer of the deletion will see it. - -## Reading the result honestly - -- **No losses found** means the audit found none — with the method calibrated by the - positive control. It is evidence, not proof; say which control you ran. -- **A "migrated" claim you cannot point at** is a loss wearing a disposition. Every - `migrated` row names a file or symbol in the surviving tree. -- The audit is run **once per retirement**, against that retirement's parent commit — - it is not a standing CI job. Its value is the moment before the deletion merges. - -## Reason codes at a glance - -| Code | Meaning | -|---|---| -| `migrated` | Behavior exists at a named location in the surviving tree | -| `renamed → X` | Same behavior under the surviving convention's name | -| `dropped because Y` | Deliberate deletion, reason stated | -| *(nothing — no issue)* | Not tracked. A gap that matters gets an issue; a comment or changeset note dies with the code it annotates (`clever-yaks-cheer.md` → #3427) | diff --git a/examples/dynamic-context-management/CONTEXT-INDEX.json b/examples/dynamic-context-management/CONTEXT-INDEX.json index 8f288a67a..a7a8e1d4e 100644 --- a/examples/dynamic-context-management/CONTEXT-INDEX.json +++ b/examples/dynamic-context-management/CONTEXT-INDEX.json @@ -1,6 +1,6 @@ { "schemaVersion": 1, - "count": 263, + "count": 261, "classes": { "ARCH": 1, "CI": 2, @@ -17,7 +17,7 @@ "PROC": 14, "PROHIB": 10, "RELEASE-NOTES": 31, - "RULESET": 60, + "RULESET": 58, "SESSION": 9, "WAVE": 5, "WORKSTREAM": 5, @@ -28,1579 +28,1567 @@ "id": "ARCH.SKILL.improve-codebase.next-candidates", "klass": "ARCH", "value": "[Workstream Progress Projection Module]", - "line": 626 + "line": 603 }, { "id": "CI.GATE.changeset-lint", "klass": "CI", "value": "hard-fail for user-facing code diffs unless .changeset/* or PR has no-changelog label", - "line": 610 + "line": 587 }, { "id": "CI.GATE.issue-link-required", "klass": "CI", "value": "hard-fail if PR body lacks closes/fixes/resolves #", - "line": 609 + "line": 586 }, { "id": "CONFIG.LOCATION.SEAM.in-process-scrub", "klass": "CONFIG", "value": "TEST_ENV_BASE reaches CHILD env only; a test calling install() IN-PROCESS must additionally use helpers.scrubConfigLocationEnv() in beforeEach + its restorer in afterEach — HOME/USERPROFILE sandboxing is NOT sufficient because getGlobalConfigDir is env-FIRST", - "line": 644 + "line": 621 }, { "id": "CONFIG.LOCATION.SEAM.kimi-two-homes", "klass": "CONFIG", "value": "kimi declares TWO config-location vars: KIMI_CONFIG_DIR (registry, generic Agent-Skills root via resolveKimiGlobalDir) and KIMI_SHARE_DIR (KIMI_HOOKS_TOML_DESCRIPTOR, kimi's OWN native config.toml carrying GSD's [[hooks]] block via resolveKimiHooksTomlDir); a registry-only derivation covers the first and silently misses the second", - "line": 643 + "line": 620 }, { "id": "CONFIG.LOCATION.SEAM.scrub-set", "klass": "CONFIG", "value": "tests/helpers.cjs CONFIG_LOCATION_ENV_KEYS is DERIVED from five sources rather than maintained as one hand-written list (source 4 IS a literal residue list, for vars that fit no other rung — what is never hand-listed is the SET): capability-registry runtimes[].runtime.configHome.env AND [].configHome.skillsHome.env + runtime-homes NON_REGISTRY_CONFIG_HOME_DESCRIPTORS[].env AND [].skillsHome.env (a descriptor is a descriptor — BOTH descriptor rungs walk skillsHome, which resolves independently via resolveSkillsBaseFromDescriptor) + runtime-homes GSD_LOCATION_ENV_KEYS + a residue list (GROK_AGENTS_HOME, GSD_RUNTIME, GSD_PROJECT, GSD_WORKSTREAM) + WRITE_ESCAPE_PERMISSION_ENV_KEYS (GSD_ALLOW_SYMLINKED_DEST — a permission, not a location: it names no path but disarms the symlink-escape guard, so blanking it makes the guard STRICTER, never looser); adding a config-location var means making it ENUMERABLE at one of those sources, not appending a literal", - "line": 641 + "line": 618 }, { "id": "CONFIG.LOCATION.SEAM.two-families", "klass": "CONFIG", "value": "runtime configHomes (where a third-party runtime keeps config, registry- or descriptor-declared) and GSD's OWN location vars (GSD_HOME -> $GSD_HOME/.gsd store, GSD_AGENTS_DIR -> getAgentsDir priority 1) are DISTINCT families; no registry derivation reaches the second, and treating a miss there as a registry gap is what produced review round 2", - "line": 642 + "line": 619 }, { "id": "CONFIG.SEAM.loadConfig-context", "klass": "CONFIG", "value": "loadConfig(cwd,{workstream}) replaces env-mutation fallback; no temporary process.env GSD_WORKSTREAM rewrites", - "line": 640 + "line": 617 }, { "id": "EXEC.CLASSIFY.classes", "klass": "EXEC", "value": "{class:'quota-exceeded'|'classify-handoff-bug'|'unknown-failure', sentinel?, retryAfterSeconds?}", - "line": 857 + "line": 834 }, { "id": "EXEC.CLASSIFY.cross-runtime", "klass": "EXEC", "value": "Anthropic/CC: usage limit|rate limit|quota|429|retry-after; Copilot CLI: rate_limit (stem); Codex CLI: 429|usage_limit_reached|too many requests", - "line": 859 + "line": 836 }, { "id": "EXEC.CLASSIFY.handler", "klass": "EXEC", "value": "gsd-core/bin/lib/agent-command-router.cjs:classifyAgentFailure (registered via command-aliases.cjs; mutation:false outputMode:json)", - "line": 855 + "line": 832 }, { "id": "EXEC.CLASSIFY.precedence", "klass": "EXEC", "value": "quota sentinel wins over classifyHandoffIfNeeded bug when both appear", - "line": 860 + "line": 837 }, { "id": "EXEC.CLASSIFY.proactive-signal-not-usable", "klass": "EXEC", "value": "Anthropic exposes anthropic-ratelimit-* headers + Agent SDK RateLimitEvent; Claude Code subprocess does NOT forward to hooks/statusline today (upstream #33820, #22407, #32796)", - "line": 862 + "line": 839 }, { "id": "EXEC.CLASSIFY.retry-after-parser", "klass": "EXEC", "value": "\\bretry[-_ ]after[:\\s]+(\\d+)\\b avoids embedded-word false matches like noretry-after", - "line": 861 + "line": 838 }, { "id": "EXEC.CLASSIFY.sentinel-order", "klass": "EXEC", "value": "most specific first: 429 beats too-many-requests; resource_exhausted beats quota (array order in src/agent-command-router.cts QUOTA_SENTINELS checks resource_exhausted before quota); case-insensitive; canonical sentinel value is lower-cased form", - "line": 858 + "line": 835 }, { "id": "EXEC.CLASSIFY.workflow", "klass": "EXEC", "value": "gsd-core/workflows/execute-phase.md step 7; class-distinct prompts (quota-to-wait-for-reset; classify-handoff-bug-to-spot-check; unknown-to-continue/stop)", - "line": 856 + "line": 833 }, { "id": "GSD-RESEARCH.CONTEXT-DISCIPLINE", "klass": "GSD-RESEARCH", "value": "less-context levers: subagent isolation + compact provider output + fetches-to-disk + cache-returns-digest; API clear_tool_uses/memory tool are the conceptual model, not a Claude Code harness knob", - "line": 398 + "line": 377 }, { "id": "GSD-RESEARCH.INTEGRATION.L2-hybrid", "klass": "GSD-RESEARCH", "value": "code owns cache+legitimacy+confidence+provider-pick (gsd-tools query research-plan/research-store/package-legitimacy); MCP owns the fetch; agent returns RESEARCH.md path, never raw fetches", - "line": 396 + "line": 375 }, { "id": "GSD-RESEARCH.MODULE.package-legitimacy", "klass": "GSD-RESEARCH", "value": "registry-API verdicts (npm/PyPI/crates.io injectable adapters) computed from thresholds {minAgeDays:30,minWeeklyDownloads:1000,requireRepo:true}; verdict OK|SUS|SLOP per package; slopcheck=optional adapter that can only escalate, never the install-or-degrade gate", - "line": 395 + "line": 374 }, { "id": "GSD-RESEARCH.MODULE.research-provider", "klass": "GSD-RESEARCH", "value": "single source of truth PROVIDER_WATERFALL (docs Context7->Ref->Jina->websearch; web Exa->Tavily->Perplexity->Brave->websearch; scrape Firecrawl->Jina); planResearch returns cache-hits+fetch-plan; classifyConfidence stamps HIGH|MEDIUM|LOW by provider AUTHORITY + verification EVIDENCE (HIGH requires code-computed ground-truth corroboration e.g. legitimacyVerdict OK; provider authority alone caps at MEDIUM; SLOP caps at LOW); Firecrawl is scrape-only (not in docs/web discovery)", - "line": 394 + "line": 373 }, { "id": "GSD-RESEARCH.MODULE.research-store", "klass": "GSD-RESEARCH", "value": "content-addressed cache; key=sha256(ecosystem+library+version+query+kind); getResearch->{hit,stale} never throws (mirrors graphify staleness); ttlForSource curated HIGH 30d|MED 7d|web LOW 1d; tiers: curated-doc kinds -> ~/.gsd/research-cache (cross-project), web/synthesis -> project .planning/research/.cache", - "line": 393 + "line": 372 }, { "id": "GSD-RESEARCH.PROVIDER.availability", "klass": "GSD-RESEARCH", "value": "config flags brave_search/exa_search/firecrawl/tavily_search/ref_search/perplexity/jina (env _API_KEY or ~/.gsd/_api_key); context7/jina/websearch always available; planResearch falls through waterfall to websearch terminal", - "line": 397 + "line": 376 }, { "id": "LEARNING.prompt-budget.boundary-gap", "klass": "LEARNING", "value": "PR #3708 commit 2df566ed reserved NOTE_RESERVE_TOKENS in pressure-threshold AND in minSet pre-check; both buggy paths only fire when baseTokens ∈ (effectiveBudget - NOTE_RESERVE_TOKENS, effectiveBudget]; original test suite used budgets far from that band so neither path was exercised; fix bde1ae8f confines NOTE_RESERVE accounting to post-trim assembly path only; future budget/limit code MUST add boundary fixtures per RULESET.TESTS.boundary-coverage.fixtures", - "line": 555 + "line": 534 }, { "id": "LIVE-CONFIG.GUARD.SEAM.ci-blind", "klass": "LIVE-CONFIG", "value": "the AMBIENT-ENV half stays CI-blind — CI never has these vars set, so green CI is not evidence for it; what strict mode catches in CI is the suite's own default-root leaks (HOME/USERPROFILE-derived), the guard remains the only loud signal for ambient-var escapes", - "line": 650 + "line": 627 }, { "id": "LIVE-CONFIG.GUARD.SEAM.module", "klass": "LIVE-CONFIG", "value": "scripts/live-config-guard.cjs (deliberately NOT scripts/lib/, which the installer copies to users wholesale while uninstall removes only an allowlist; excluded from the npm tarball via package.json files[] together with its whole require chain run-tests.cjs/affected-tests-lib.cjs/run-affected-tests.cjs — a partial exclusion trips the #2858 shipped-requires-only-shipped gate); exports [resolveLiveConfigRoots, resolveExtraWatchTargets, snapshotLiveConfig, diffLiveConfig, formatViolations, newestMtime]; driven by scripts/run-tests.cjs pre/post suite", - "line": 645 + "line": 622 }, { "id": "LIVE-CONFIG.GUARD.SEAM.non-root-targets", "klass": "LIVE-CONFIG", "value": "resolveExtraWatchTargets covers THREE live write surfaces that are not runtime config ROOTS (skills bases are a DELIBERATE non-target — the config-root layout misfires beneath them, so they need their own layout): $GSD_HOME/.gsd watched WHOLESALE (exclusively GSD-owned, so the shared-root trap does not apply) plus ONE config.toml per NON_REGISTRY_CONFIG_HOME_DESCRIPTORS entry, each watched as a SINGLE FILE (those roots belong to their products) — today three targets, since #2755 split Kimi CLI (~/.kimi, KIMI_SHARE_DIR) from Kimi Code (~/.kimi-code, KIMI_CODE_HOME); the targets are DERIVED by iterating that array, never by calling a named resolver, so a further descriptor is picked up without editing the guard PROVIDED it owns the same NON_REGISTRY_OWNED_FILE ('config.toml') — one that owns a different filename needs a per-descriptor mapping, the named residual the guard states at its own definition. SECOND RESIDUAL: config.toml is not all GSD writes into those roots — installSharedHooksBundle also populates /hooks/, which is UNWATCHED; closing it is a layout decision, like skills bases; passed to snapshotLiveConfig explicitly so a fixture-root caller cannot pull the real ~/.gsd into its snapshot", - "line": 647 + "line": 624 }, { "id": "LIVE-CONFIG.GUARD.SEAM.scope", "klass": "LIVE-CONFIG", "value": "ownership-based, never whole-root: GSD_OWNED_ENTRIES top-level footprint + children whose name startsWith GSD_ARTIFACT_PREFIX ('gsd-') under GSD_PREFIXED_PARENTS (dirs shared with the host agent); watching a shared root wholesale false-positives on the host's own writes and a guard that cries wolf gets disabled", - "line": 646 + "line": 623 }, { "id": "LIVE-CONFIG.GUARD.SEAM.severity", "klass": "LIVE-CONFIG", "value": "reports by default locally; CI wires GSD_STRICT_LIVE_CONFIG_GUARD=1 on Linux/macOS lanes (test.yml, all three test jobs) so a suite-produced leak FAILS those runs; Windows lanes stay report-only pending the documented pre-existing USERPROFILE sweep (~190 test sites sandbox HOME alone) — promote once that lands; skipped by GSD_SKIP_LIVE_CONFIG_GUARD=1", - "line": 649 + "line": 626 }, { "id": "LIVE-CONFIG.GUARD.SEAM.truncation", "klass": "LIVE-CONFIG", "value": "MAX_ENTRIES/MAX_DEPTH bound the walk; a bound hit sets truncated and diffLiveConfig emits kind:'unverified' — a truncated scan MUST NOT read as clean; boundary covered at {limit-1,limit,limit+1} via newestMtime's injected budget plus fast-check monotonicity, per RULESET.TESTS.boundary-coverage + RULESET.TESTS.property-based-testing", - "line": 648 + "line": 625 }, { "id": "META.RULE.brief-must-cite-doc", "klass": "META", "value": "agent prompts MUST quote the canonical doc line being applied; paraphrasing from predicate memory drifts and produces violations", - "line": 701 + "line": 678 }, { "id": "META.RULE.brief-no-paraphrase", "klass": "META", "value": "writing \"k040 — never leave changelog box unchecked\" caused 5 of 8 agents to edit CHANGELOG.md in violation of CONTRIBUTING.md L110", - "line": 702 + "line": 679 }, { "id": "META.RULE.canonical-source-precedence", "klass": "META", "value": "CONTRIBUTING.md > docs/adr/* > CONTEXT.md > agent memory", - "line": 699 + "line": 676 }, { "id": "META.RULE.read-contributing-first", "klass": "META", "value": "read CONTRIBUTING.md sections \"Pull Request Guidelines\" + \"CHANGELOG Entries\" before EVERY agent dispatch", - "line": 700 + "line": 677 }, { "id": "PLANNING.PATH.PARITY.project-scope", "klass": "PLANNING", "value": ".planning/ (never .planning/projects/); mirror planning-workspace.cjs planningDir()", - "line": 635 + "line": 612 }, { "id": "PLANNING.PATH.SEAM.helpers", "klass": "PLANNING", "value": "helpers.planningPaths delegates to workspacePlanningPaths + resolveWorkspaceContext; precedence explicit-ws > env-ws > env-project > root", - "line": 636 + "line": 613 }, { "id": "PLANNING.PATH.SEAM.init-handlers", "klass": "PLANNING", "value": "[initExecutePhase, initPlanPhase, initPhaseOp, initMilestoneOp] consume helpers.planningPaths().planning (no direct relPlanningPath join)", - "line": 637 + "line": 614 }, { "id": "PR.3267.POSTMORTEM.recovery", "klass": "PR", "value": "[issue#3270 created, label approved-enhancement applied, PR reopened, body includes \"Closes #3270\", label no-changelog applied]", - "line": 614 + "line": 591 }, { "id": "PR.3267.POSTMORTEM.root-cause", "klass": "PR", "value": "[missing issue link, missing changeset/no-changelog]", - "line": 613 + "line": 590 }, { "id": "PRED.k320.canonical-source", "klass": "PRED", "value": "CONTRIBUTING.md L193-211", - "line": 705 + "line": 682 }, { "id": "PRED.k320.ci-enforcement", "klass": "PRED", "value": "scripts/changeset/lint.cjs", - "line": 711 + "line": 688 }, { "id": "PRED.k320.ci-paths-monitored", "klass": "PRED", "value": "bin/ gsd-core/ src/ agents/ commands/ hooks/ sdk/src/ sdk/prompts/", - "line": 712 + "line": 689 }, { "id": "PRED.k320.cure", "klass": "PRED", "value": "drop .changeset/--.md fragment ONLY", - "line": 707 + "line": 684 }, { "id": "PRED.k320.evidence", "klass": "PRED", "value": "PR #3302 merge-conflict against #3308 CHANGELOG.md row 2026-05-09", - "line": 714 + "line": 691 }, { "id": "PRED.k320.opt-out-label", "klass": "PRED", "value": "no-changelog", - "line": 710 + "line": 687 }, { "id": "PRED.k320.recovery", "klass": "PRED", "value": "open Removed-typed cleanup PR deleting only the redundant row", - "line": 713 + "line": 690 }, { "id": "PRED.k320.rule", "klass": "PRED", "value": "do not edit CHANGELOG.md in feature/fix/enhancement PRs", - "line": 706 + "line": 683 }, { "id": "PRED.k320.signal", "klass": "PRED", "value": "changelog-direct-edit-forbidden", - "line": 704 + "line": 681 }, { "id": "PRED.k320.tool", "klass": "PRED", "value": "npm run changeset -- --type --pr --body \"...\"", - "line": 708 + "line": 685 }, { "id": "PRED.k320.types", "klass": "PRED", "value": "Added|Changed|Deprecated|Removed|Fixed|Security", - "line": 709 + "line": 686 }, { "id": "PRED.k321.evidence", "klass": "PRED", "value": "PRs #3304/#3305 (2026-05-09): real Minor/Major findings in body, 0 threads", - "line": 720 + "line": 697 }, { "id": "PRED.k321.poll-shape", "klass": "PRED", "value": "parse pulls//reviews body AND graphql reviewThreads", - "line": 718 + "line": 695 }, { "id": "PRED.k321.resolution", "klass": "PRED", "value": "address in code; no GraphQL resolveReviewThread needed for body-only findings", - "line": 719 + "line": 696 }, { "id": "PRED.k321.shape", "klass": "PRED", "value": "CR posts \"[!CAUTION] outside the diff\" findings in review BODY, not in reviewThreads", - "line": 717 + "line": 694 }, { "id": "PRED.k321.signal", "klass": "PRED", "value": "cr-outside-diff-range-finding", - "line": 716 + "line": 693 }, { "id": "PRED.k322.cure-1", "klass": "PRED", "value": "2nd retrigger ~10min after first ack", - "line": 725 + "line": 702 }, { "id": "PRED.k322.cure-2", "klass": "PRED", "value": "if silent at 50min, treat as silent-pass with maintainer flag in merge-commit body", - "line": 726 + "line": 703 }, { "id": "PRED.k322.distinct-from", "klass": "PRED", "value": "k080", - "line": 723 + "line": 700 }, { "id": "PRED.k322.evidence", "klass": "PRED", "value": "PR #3306 (2026-05-09): 0 reviews after 50min + 2 retriggers", - "line": 728 + "line": 705 }, { "id": "PRED.k322.merge-gate-impact", "klass": "PRED", "value": "k070 real_coderabbit_review_present unsatisfied; requires maintainer judgment", - "line": 727 + "line": 704 }, { "id": "PRED.k322.shape", "klass": "PRED", "value": "ack posted, real review never lands within [5s, 410s] cooldown after burst of N PRs <15min", - "line": 724 + "line": 701 }, { "id": "PRED.k322.signal", "klass": "PRED", "value": "cr-sustained-throttle", - "line": 722 + "line": 699 }, { "id": "PRED.k323.cure-alt", "klass": "PRED", "value": "consolidate into single PR when 2+ issues share root cause", - "line": 733 + "line": 710 }, { "id": "PRED.k323.cure-pre-dispatch", "klass": "PRED", "value": "brief one agent canonical-owner; brief others to EXCLUDE shared site", - "line": 732 + "line": 709 }, { "id": "PRED.k323.evidence", "klass": "PRED", "value": "#3300 (#3297) overlapped #3306 (#3298) on add-backlog.md hunks 2026-05-09", - "line": 735 + "line": 712 }, { "id": "PRED.k323.recovery", "klass": "PRED", "value": "close smaller PR as \"subsumed by #N\" or rebase second to drop overlap hunk", - "line": 734 + "line": 711 }, { "id": "PRED.k323.shape", "klass": "PRED", "value": "2+ open issues touch same canonical bug site; each fix's sibling-audit produces overlapping diff", - "line": 731 + "line": 708 }, { "id": "PRED.k323.signal", "klass": "PRED", "value": "sibling-audit-cross-pr-overlap", - "line": 730 + "line": 707 }, { "id": "PRED.k324.cure", "klass": "PRED", "value": "verify via gh api on every agent-completion notification; never trust narrative", - "line": 739 + "line": 716 }, { "id": "PRED.k324.evidence", "klass": "PRED", "value": "2026-05-09 session: 5+ mid-monitor terminations across PRs #3232/#3271/#3251/#3255/#3262", - "line": 741 + "line": 718 }, { "id": "PRED.k324.k095-restatement", "klass": "PRED", "value": "k095 confirmed shape: agent reports \"waiting for monitor\" / \"tests still running\" then terminates", - "line": 738 + "line": 715 }, { "id": "PRED.k324.poll-shape", "klass": "PRED", "value": "gh pr view --json mergeStateStatus,statusCheckRollup + pulls//reviews + graphql reviewThreads + issues//comments tail", - "line": 740 + "line": 717 }, { "id": "PRED.k324.signal", "klass": "PRED", "value": "agent-terminates-mid-monitor", - "line": 737 + "line": 714 }, { "id": "PRED.k325.cleanup", "klass": "PRED", "value": "git worktree remove --force for aged agent worktrees", - "line": 746 + "line": 723 }, { "id": "PRED.k325.cure", "klass": "PRED", "value": "detached-HEAD: git checkout --detach $(git ls-remote origin ); modify; commit; git push --force-with-lease=: origin HEAD:refs/heads/", - "line": 745 + "line": 722 }, { "id": "PRED.k325.evidence", "klass": "PRED", "value": "2026-05-09 CHANGELOG.md strip on PRs #3300/#3302/#3304/#3305 required detached-HEAD", - "line": 747 + "line": 724 }, { "id": "PRED.k325.shape", "klass": "PRED", "value": "git checkout errors \"already used by worktree at \"", - "line": 744 + "line": 721 }, { "id": "PRED.k325.signal", "klass": "PRED", "value": "worktree-branch-lock-on-force-push", - "line": 743 + "line": 720 }, { "id": "PRED.k326.cure", "klass": "PRED", "value": "quote canonical doc verbatim in brief; mentally simulate \"if all N agents follow this brief literally, do they violate any rule?\"", - "line": 751 + "line": 728 }, { "id": "PRED.k326.evidence", "klass": "PRED", "value": "2026-05-09 brief \"k040 — update CHANGELOG.md\" → 5 of 8 agents violated CONTRIBUTING.md L110", - "line": 752 + "line": 729 }, { "id": "PRED.k326.shape", "klass": "PRED", "value": "N parallel agents amplify a single brief-vs-doc contradiction into N violations", - "line": 750 + "line": 727 }, { "id": "PRED.k326.signal", "klass": "PRED", "value": "brief-contradicts-canonical-doc", - "line": 749 + "line": 726 }, { "id": "PRED.k327.ack-shape", "klass": "PRED", "value": "body \"✅ Actions performed - Full review triggered\"", - "line": 755 + "line": 732 }, { "id": "PRED.k327.cooldown-normal", "klass": "PRED", "value": "[5s, 410s]", - "line": 758 + "line": 735 }, { "id": "PRED.k327.cooldown-throttled", "klass": "PRED", "value": "k322", - "line": 759 + "line": 736 }, { "id": "PRED.k327.distinguish-key", "klass": "PRED", "value": "len(pulls//reviews) — ack=0, real=≥1", - "line": 757 + "line": 734 }, { "id": "PRED.k327.real-review-shape", "klass": "PRED", "value": "body starts \"Actionable comments posted: N\" OR \"[!CAUTION] Some comments are outside the diff\"", - "line": 756 + "line": 733 }, { "id": "PRED.k327.signal", "klass": "PRED", "value": "cr-ack-vs-real-review", - "line": 754 + "line": 731 }, { "id": "PRED.k328.audit-list", "klass": "PRED", "value": "[heading-matches-class, closing-keyword-present, changeset-fragment-or-no-changelog-label]", - "line": 764 + "line": 741 }, { "id": "PRED.k328.canonical-source", "klass": "PRED", "value": "CONTRIBUTING.md L48,L64,L81 (template links) + .github/PULL_REQUEST_TEMPLATE/{fix,enhancement,feature}.md L1 (heading text)", - "line": 762 + "line": 739 }, { "id": "PRED.k328.k100-restatement", "klass": "PRED", "value": "heading must match issue class: bug→## Fix PR, enhancement→## Enhancement PR, feature→## Feature PR", - "line": 763 + "line": 740 }, { "id": "PRED.k328.signal", "klass": "PRED", "value": "pr-template-typed-heading-required", - "line": 761 + "line": 738 }, { "id": "PRED.k329.body", "klass": "PRED", "value": "**** — . (#)", - "line": 770 + "line": 747 }, { "id": "PRED.k329.canonical-source", "klass": "PRED", "value": "CONTRIBUTING.md L196-202 + .changeset/README.md", - "line": 767 + "line": 744 }, { "id": "PRED.k329.filename", "klass": "PRED", "value": ".changeset/--.md", - "line": 768 + "line": 745 }, { "id": "PRED.k329.frontmatter", "klass": "PRED", "value": "---\\\\ntype: \\\\npr: \\\\n---", - "line": 769 + "line": 746 }, { "id": "PRED.k329.observed-clean", "klass": "PRED", "value": "#3299 sunny-ibex-wave, #3301 sturdy-rams-caper, #3306 3298-phase-dir-prefix-drift-workflows", - "line": 771 + "line": 748 }, { "id": "PRED.k329.signal", "klass": "PRED", "value": "changeset-fragment-canonical-shape", - "line": 766 + "line": 743 }, { "id": "PRED.k330.fallback", "klass": "PRED", "value": "append predicate-format findings directly to CONTEXT.md", - "line": 775 + "line": 752 }, { "id": "PRED.k330.shape", "klass": "PRED", "value": "mempalace MCP tools require explicit user call; AI cannot trigger", - "line": 774 + "line": 751 }, { "id": "PRED.k330.signal", "klass": "PRED", "value": "mempalace-diary-not-callable-by-ai", - "line": 773 + "line": 750 }, { "id": "PRED.k331.cure", "klass": "PRED", "value": "gh pr close with NO --comment flag", - "line": 780 + "line": 757 }, { "id": "PRED.k331.evidence", "klass": "PRED", "value": "2026-05-09 wave-3: violation on #3300 close, deleted within 30s", - "line": 782 + "line": 759 }, { "id": "PRED.k331.k101-restatement", "klass": "PRED", "value": "k101 includes close-time --comment flag; rationale belongs in subsuming PR's squash-merge body", - "line": 779 + "line": 756 }, { "id": "PRED.k331.recovery", "klass": "PRED", "value": "if violation lands, gh api -X DELETE repos///issues/comments/", - "line": 781 + "line": 758 }, { "id": "PRED.k331.shape", "klass": "PRED", "value": "instruction \"close with no comment (rationale)\" — parenthetical is rationale, NOT comment body", - "line": 778 + "line": 755 }, { "id": "PRED.k331.signal", "klass": "PRED", "value": "close-with-no-comment-is-literal", - "line": 777 + "line": 754 }, { "id": "PROBE.ci.surface", "klass": "PROBE", "value": "the contract (parse/validate, projection round-trip, fail-closed guards), NEVER the LLM judgment (ADR-550 D5)", - "line": 524 + "line": 503 }, { "id": "PROBE.core.seam", "klass": "PROBE", "value": "analyzeCoverage(items,resolutions?,validators) ingests ALREADY-proposed items; does NOT assume deterministic propose (ADR-550 D7b)", - "line": 517 + "line": 496 }, { "id": "PROBE.edge.verification", "klass": "PROBE", "value": "explicit|backstop", - "line": 519 + "line": 498 }, { "id": "PROBE.family", "klass": "PROBE", "value": "edge-probe(shape-axis)+prohibition-probe(must-NOT-axis)+ui-consideration-probe(UI-state-axis), shared probe-core, run as spec-phase/ui-phase soft gates (ADR-550 D7; #1867)", - "line": 515 + "line": 494 }, { "id": "PROBE.item.axes", "klass": "PROBE", "value": "status{resolved|dismissed|unresolved} x verification{|null} — orthogonal; the lifecycle enum carries no verification fact (ADR-550 D7a)", - "line": 518 + "line": 497 }, { "id": "PROBE.principle", "klass": "PROBE", "value": "verifier-reach-equals-spec-reach (a goal-backward verifier only checks assertions that exist; probes make omitted assertions exist before code) — ADR-857 verification-substrate boundary; docs/design/verifier-reach.md", - "line": 514 + "line": 493 }, { "id": "PROBE.prohib.verification", "klass": "PROBE", "value": "test|judgment", - "line": 520 + "line": 499 }, { "id": "PROBE.protocol", "klass": "PROBE", "value": "recall(adversarial over-generate)->precision(drop routine-engineering); dismissals require a non-empty reason", - "line": 516 + "line": 495 }, { "id": "PROBE.ui.axis", "klass": "PROBE", "value": "MIXED — closed compiled shape-rooted 8 (empty/loading/error/populated/partial/overflow/zero-one-many/long-text) via ui-consideration-probe adapter; open UX (real-time/a11y/i18n-RTL) prose-owned in references/domain-probes.md, NOT compiled (#1867)", - "line": 522 + "line": 501 }, { "id": "PROBE.ui.seam", "klass": "PROBE", "value": "ui-phase Step 9.5 post-verification: element-cue classify -> propose-then-confirm (partial-cue mitigation, Goodhart) -> autoResolve --auto floor (never dismiss; unclassified stays unresolved #1110) -> ## UI Considerations write-back -> plan-phase `## UI Considerations` lift rule (#1867)", - "line": 523 + "line": 502 }, { "id": "PROBE.ui.verification", "klass": "PROBE", "value": "explicit|backstop", - "line": 521 + "line": 500 }, { "id": "PROC.AGENT-DISPATCH.completion-verify", "klass": "PROC", "value": "run k324.poll-shape on every agent-completion notification", - "line": 786 + "line": 763 }, { "id": "PROC.AGENT-DISPATCH.parallel-overlap-audit", "klass": "PROC", "value": "before dispatching N sibling-audit fixers, compute file-set union and assign canonical owners", - "line": 785 + "line": 762 }, { "id": "PROC.AGENT-DISPATCH.preflight", "klass": "PROC", "value": "[read-CONTRIBUTING.md-fresh, read-relevant-ADRs, cite-specific-line-in-brief, require-closing-keyword, require-changeset-fragment, forbid-CHANGELOG.md-edit, require-isolation-worktree, forbid-self-PR-comment, mandate-trust-but-verify]", - "line": 784 + "line": 761 }, { "id": "PROC.MERGE-WAVE.changelog-strip-pattern", "klass": "PROC", "value": "detached-HEAD per k325 + git checkout main -- CHANGELOG.md + commit + force-with-lease", - "line": 790 + "line": 767 }, { "id": "PROC.MERGE-WAVE.merge-tool", "klass": "PROC", "value": "gh pr merge --squash --delete-branch", - "line": 791 + "line": 768 }, { "id": "PROC.MERGE-WAVE.merge-tool-warning", "klass": "PROC", "value": "delete-branch may fail with \"used by worktree at\" — harmless; remote branch still deleted", - "line": 792 + "line": 769 }, { "id": "PROC.MERGE-WAVE.ordering", "klass": "PROC", "value": "[wave1: isolated-files, wave2: CHANGELOG-only-overlap (better: strip per k320), wave3: same-file-overlap with explicit decision]", - "line": 788 + "line": 765 }, { "id": "PROC.MERGE-WAVE.preflight", "klass": "PROC", "value": "gh pr view --json files for every PR; identify overlap pairs; surface to maintainer", - "line": 789 + "line": 766 }, { "id": "PROC.PARALLEL-FIX-DISPATCH.observed", "klass": "PROC", "value": "#3541 + #3542 dispatched simultaneously this session; PRs #3546 #3547 opened green; one syntax slip caught by AGENT-RETIRED-SLASH-SYNTAX-DRIFT and fixed before second PR opened", - "line": 866 + "line": 843 }, { "id": "PROC.PARALLEL-FIX-DISPATCH.pattern", "klass": "PROC", "value": "bot triage brief → worktree per branch → parallel sub-agents do rubber-duck/RCA/TDD implementation only → top-level orchestrator owns commit + gsd-test + push + PR + changeset-pr-backfill", - "line": 864 + "line": 841 }, { "id": "PROC.PARALLEL-FIX-DISPATCH.rationale", "klass": "PROC", "value": "long-running test runs need cross-turn notifications (orchestrator-only); CONTRIBUTING.md gh-templates-first hook requires session-scoped Read calls sub-agents wouldn't otherwise make; sequencing test runs avoids GSD-TEST-CONCURRENT-OUTPUT-COLLISION", - "line": 865 + "line": 842 }, { "id": "PROC.TRIAGE.comment-shape", "klass": "PROC", "value": "lead with \"duplicate of #NNNN, fixed by PR #MMMM, in v1.X.Y\"; show current code snippet proving bug-surface gone; give @latest and @next upgrade commands; close", - "line": 869 + "line": 846 }, { "id": "PROC.TRIAGE.no-duplicate-label", "klass": "PROC", "value": "this repo has no duplicate label; framing lives in comment text + closing the issue", - "line": 870 + "line": 847 }, { "id": "PROC.TRIAGE.routing-incoming", "klass": "PROC", "value": "stale-bug-already-fixed to close as duplicate of originating issue + cite fix PR + first stable tag; release-publish-or-backport to ready-for-human; reporter-can-self-test to awaiting-retest", - "line": 868 + "line": 845 }, { "id": "PROHIB.canon-referral", "klass": "PROHIB", "value": "OWASP/GDPR/fairness-canon are REFERRED to /gsd:secure-phase+eslint, never minted as prohibitions (ADR-550 D6)", - "line": 526 + "line": 505 }, { "id": "PROHIB.descriptor.shape", "klass": "PROHIB", "value": "5 FLAT scalars (check_kind,check_target,check_rule,check_violation_fixture,check_clean_fixture) — NEVER a nested check:{} (parseMustHavesBlock is a flat parser, src/frontmatter.cts)", - "line": 531 + "line": 510 }, { "id": "PROHIB.enforce.adr", "klass": "PROHIB", "value": "docs/adr/1606 (verify-time enforcement seam) + docs/adr/550 (spec-phase contract)", - "line": 534 + "line": 513 }, { "id": "PROHIB.enforce.causation", "klass": "PROHIB", "value": "clean-fixture control proves the red is content-caused not env-var-set; MANDATORY for node-test (#1906 supersedes #1346 opt-in) — absent clean-fixture ⇒ node-test un-provable/fail-closed; lint-rule needs none (its subject IS the linted file)", - "line": 530 + "line": 509 }, { "id": "PROHIB.enforce.failfirst", "klass": "PROHIB", "value": "MACHINE-PROVEN against an author-supplied violation fixture (#1279); caller failFirst attestation DEMOTED to a non-authoritative hint (FF-08)", - "line": 529 + "line": 508 }, { "id": "PROHIB.enforce.green-rule", "klass": "PROHIB", "value": "passed iff provenFailFirst===true && run.passed===true (runProhibitionEnforcement); every miss/fail/un-provable HARD-GATES both modes via dispositionForProhibition's fail-closed default", - "line": 527 + "line": 506 }, { "id": "PROHIB.enforce.kinds", "klass": "PROHIB", "value": "node-test (non-vacuous red via isNonVacuousNodeTestRed; pass-side vacuity via isNonVacuousNodeTestPass) | lint-rule (eslint --format json filtered by ruleId)", - "line": 528 + "line": 507 }, { "id": "PROHIB.judgment-tier", "klass": "PROHIB", "value": "never-silent / never-hard-halt soft gate; autonomous emits \"unverified-prohibition — human review recommended\" (exogenous grading, ADR-550 D4)", - "line": 533 + "line": 512 }, { "id": "PROHIB.rail", "klass": "PROHIB", "value": "core verify rail, non-toggleable (ADR-857 verification-substrate boundary / decision #6); the verifier<->predicate contract is NOT an off-by-default capability", - "line": 532 + "line": 511 }, { "id": "PROHIB.recall", "klass": "PROHIB", "value": "LLM-prose; no compiled prohibition-probe recall engine (only the schema/projection layer is code, ADR-550 D7b)", - "line": 525 + "line": 504 }, { "id": "RELEASE-NOTES.ANTI-PATTERN", "klass": "RELEASE-NOTES", "value": "raw \"What's Changed\" PR list as final body for hotfix or feature release; \"Full Changelog only\" body for tagged release with >0 user-facing fixes", - "line": 681 + "line": 658 }, { "id": "RELEASE-NOTES.ANTI-PATTERN.implementation-first", "klass": "RELEASE-NOTES", "value": "do not lead bullet with file path or function name; lead with symptom/user-visible behavior", - "line": 682 + "line": 659 }, { "id": "RELEASE-NOTES.ANTI-PATTERN.risk-commentary", "klass": "RELEASE-NOTES", "value": "do not include \"may break\", \"be careful\", \"test thoroughly\" - release notes state what changed, not hedges about what might go wrong", - "line": 683 + "line": 660 }, { "id": "RELEASE-NOTES.DEFAULT-STATE", "klass": "RELEASE-NOTES", "value": "auto-generated body is \"What's Changed\" PR list + Full Changelog link; treat as draft, not final", - "line": 657 + "line": 634 }, { "id": "RELEASE-NOTES.EXAMPLE.hotfix", "klass": "RELEASE-NOTES", "value": "v1.41.1 (https://github.com/open-gsd/gsd-core/releases/tag/v1.41.1) - 14 fixes grouped by 6 subgroups", - "line": 685 + "line": 662 }, { "id": "RELEASE-NOTES.EXAMPLE.minor-auto-acceptable", "klass": "RELEASE-NOTES", "value": "v1.41.0 - kept auto-generated body; many small fixes with clean conventional-commit titles", - "line": 687 + "line": 664 }, { "id": "RELEASE-NOTES.EXAMPLE.rc", "klass": "RELEASE-NOTES", "value": "v1.7.0-rc.1 (https://github.com/open-gsd/gsd-core/releases/tag/v1.7.0-rc.1) - intro + Added/Changed/Fixed/Documentation taxonomy", - "line": 686 + "line": 663 }, { "id": "RELEASE-NOTES.GATE.hotfix", "klass": "RELEASE-NOTES", "value": "manual edit required; auto-generated body for vX.Y.{Z>0} is \"Full Changelog only\" and must be replaced with structured body", - "line": 658 + "line": 635 }, { "id": "RELEASE-NOTES.GATE.minor", "klass": "RELEASE-NOTES", "value": "auto-generated body acceptable when PR titles are clean; promote to structured body when >20 PRs or contains feature+refactor+fix mix", - "line": 660 + "line": 637 }, { "id": "RELEASE-NOTES.GATE.rc", "klass": "RELEASE-NOTES", "value": "manual edit recommended; auto-generated PR list is acceptable for early RCs but final RC before vX.Y.0 should match standard", - "line": 659 + "line": 636 }, { "id": "RELEASE-NOTES.RELEASE-STREAM.main-branch", "klass": "RELEASE-NOTES", "value": "next (RCs) + latest (stable); install via @next or @latest", - "line": 692 + "line": 669 }, { "id": "RELEASE-NOTES.RELEASE-STREAM.rule", "klass": "RELEASE-NOTES", "value": "streams do not mix; do not document @next in hotfix/stable notes", - "line": 693 + "line": 670 }, { "id": "RELEASE-NOTES.SCOPE", "klass": "RELEASE-NOTES", "value": "GitHub Releases body for tags vX.Y.Z, vX.Y.Z-rc.N; not CHANGELOG.md (changeset workflow owns that)", - "line": 656 + "line": 633 }, { "id": "RELEASE-NOTES.SOURCE.changesets", "klass": "RELEASE-NOTES", "value": ".changeset/*.md (frontmatter pr: + body bullets)", - "line": 672 + "line": 649 }, { "id": "RELEASE-NOTES.SOURCE.commits", "klass": "RELEASE-NOTES", "value": "git log .. --pretty=format:'%s%n%n%b' --no-merges", - "line": 671 + "line": 648 }, { "id": "RELEASE-NOTES.SOURCE.pr-bodies", "klass": "RELEASE-NOTES", "value": "gh pr view --json title,body for fixes lacking a changeset", - "line": 673 + "line": 650 }, { "id": "RELEASE-NOTES.SOURCE.precedence", "klass": "RELEASE-NOTES", "value": "changeset body > commit body > PR body > commit subject (prefer authored content over auto-generated)", - "line": 674 + "line": 651 }, { "id": "RELEASE-NOTES.STANDARD.bullet-shape", "klass": "RELEASE-NOTES", "value": "**Bold user-visible change** — explanation of what was broken or what's new, leading with symptom not implementation. Trailing (#NNN) PR ref.", - "line": 664 + "line": 641 }, { "id": "RELEASE-NOTES.STANDARD.footer.full-changelog", "klass": "RELEASE-NOTES", "value": "**Full Changelog**: https://github.com/open-gsd/gsd-core/compare/...", - "line": 668 + "line": 645 }, { "id": "RELEASE-NOTES.STANDARD.footer.hotfix", "klass": "RELEASE-NOTES", "value": "Install/upgrade: \\`npx @opengsd/gsd-core@latest\\`", - "line": 666 + "line": 643 }, { "id": "RELEASE-NOTES.STANDARD.footer.rc", "klass": "RELEASE-NOTES", "value": "Install for testing: \\`npx @opengsd/gsd-core@next\\` (per branch->dist-tag policy)", - "line": 667 + "line": 644 }, { "id": "RELEASE-NOTES.STANDARD.heading-level", "klass": "RELEASE-NOTES", "value": "## for category, ### for subgroup (area), - for bullet", - "line": 663 + "line": 640 }, { "id": "RELEASE-NOTES.STANDARD.intro", "klass": "RELEASE-NOTES", "value": "optional one-paragraph framing for RC/feature releases; omit for pure-fix hotfixes", - "line": 669 + "line": 646 }, { "id": "RELEASE-NOTES.STANDARD.subgroups", "klass": "RELEASE-NOTES", "value": "phase-planning-state | workstream | query-dispatch-cli | code-review | install | capture | docs | architecture | security", - "line": 665 + "line": 642 }, { "id": "RELEASE-NOTES.STANDARD.taxonomy", "klass": "RELEASE-NOTES", "value": "Keep-a-Changelog 1.1.0: Added | Changed | Deprecated | Removed | Fixed | Security | Documentation", - "line": 662 + "line": 639 }, { "id": "RELEASE-NOTES.TEMPLATE.hotfix", "klass": "RELEASE-NOTES", "value": "## Fixed\\n\\n### \\n- **** — . (#)\\n\\n---\\n\\nInstall/upgrade: \\`npx @opengsd/gsd-core@latest\\`\\n\\n**Full Changelog**: ", - "line": 689 + "line": 666 }, { "id": "RELEASE-NOTES.TEMPLATE.rc", "klass": "RELEASE-NOTES", "value": "\\n\\n## Added\\n### \\n- **** — . (#)\\n\\n## Changed\\n### Architecture\\n- **** — . (#)\\n\\n## Fixed\\n### \\n- **** — . (#)\\n\\n## Documentation\\n- **** — . (#)\\n\\n---\\n\\nThis is a release candidate. Install for testing:\\n\\`\\`\\`bash\\nnpx @opengsd/gsd-core@next\\n\\`\\`\\`\\n\\n**Full Changelog**: ", - "line": 690 + "line": 667 }, { "id": "RELEASE-NOTES.WORKFLOW.edit", "klass": "RELEASE-NOTES", "value": "gh release edit --notes-file ", - "line": 676 + "line": 653 }, { "id": "RELEASE-NOTES.WORKFLOW.idempotency", "klass": "RELEASE-NOTES", "value": "gh release edit overwrites body wholesale; safe to re-run after refining", - "line": 679 + "line": 656 }, { "id": "RELEASE-NOTES.WORKFLOW.token", "klass": "RELEASE-NOTES", "value": "must use .envrc GITHUB_TOKEN per RULESET.GH.AUTH.DEFAULT (this doc); never ambient gh auth", - "line": 678 + "line": 655 }, { "id": "RELEASE-NOTES.WORKFLOW.view", "klass": "RELEASE-NOTES", "value": "gh release view --json body --jq .body", - "line": 677 + "line": 654 }, { "id": "RULESET.ADR-HEADER", "klass": "RULESET", "value": "every docs/adr/NNNN-*.md must open with - **Status:** Accepted|Proposed|Superseded (by [ADR-NNNN](file.md))|Legacy + - **Date:** YYYY-MM-DD immediately after title", - "line": 579 + "line": 558 }, { "id": "RULESET.AGENT_SIZE_BUDGET", "klass": "RULESET", "value": "agent-size-budget (#1074; sibling of WORKFLOW_SIZE_BUDGET; BYTES not lines per #717/#683, rebased from lines in PR 3/3) = differential attribution size ratchet (PRIMARY anti-creep since #2724/ADR-2719 §4, same mechanism and same ack fragments (tests/emitted-drift-acks/, #2914; legacy tests/emitted-drift-ack.json still honored) as WORKFLOW_SIZE_BUDGET, scoped to agents/gsd-*.md) + loose tier hard caps (red lines, never raised on approach: XL<=57344 / LARGE<=49152 / DEFAULT<=24576); net-new agents are DEFAULT-tier (no separate new-file cap). Sizes are measured via the shared scripts/workflow-size.cjs measureMdFiles(dir,predicate) counter (tests/helpers/emitted-runtime.cjs's currentSizes() and the guard's own tier-cap checks both import it). A grown agent fails the differential guard — ack + justify, or extract LAZILY to gsd-core/references/. DISTINCT from DEFECT.AGENT-FILE-SIZE-CAP-BREACH (a separate 45K-CHAR extraction-evidence threshold on gsd-planner via planner-decomposition/reachability tests): that guard proves mode-sections were extracted; this one bounds total agent bytes. Two guards, two units (chars vs bytes), two purposes. The prior per-file baseline (tests/agent-size-baseline.json, `npm run size:baseline`) is REMOVED by #2724", - "line": 568 + "line": 547 }, { "id": "RULESET.ALLOWED-TOOLS-FRONTMATTER", "klass": "RULESET", "value": "command's allowed-tools must cover every tool the workflow calls (including Write for file creation); thin-wrapper pattern makes this easy to miss", - "line": 575 + "line": 554 }, { "id": "RULESET.ARGUMENTS-SANITIZE", "klass": "RULESET", "value": "any workflow step constructing .planning/.../{SLUG}.md path from user input ($ARGUMENTS, parsed remainder) must sanitize inline ([a-z0-9-] only, reject ..//\\\\, max-length) — \"(already sanitized)\" must trace back to explicit guard; RESUME/fallback modes need own guards", - "line": 576 + "line": 555 }, { "id": "RULESET.AUDIT.search-source-not-generated", "klass": "RULESET", "value": "verify an invariant/validation EXISTS by searching the AUTHORED source (src/*.cts OR the scripts/gen-*.cjs generator), never the generated bin/lib/*.cjs (gitignored, ADR-457); gen-time checks live in gen-*.cjs not the .cts it consumes → search BOTH before declaring absent; read generated .cjs only for output drift. Repro: grep src/*.cts for VALID_CONVERTER_NAMES → false \"5e ConverterName unenforced\"; actually enforced in gen-capability-registry.cjs. cf RULESET.TESTS.no-source-grep", - "line": 564 + "line": 543 }, { "id": "RULESET.CAPABILITY.cutover-self-gating", "klass": "RULESET", "value": "a phase-6 per-feature cutover moves the host's phase-context detection + mode/flag logic INTO the skill (self-gating, per ADR-894); the loop hook is intentionally COARSE — \"invoke skill X at point Y when config Z\" — and carries no detection/mode. WORKED EXAMPLE: plan-phase.md §5.6 UI gate (frontend-detection via ui-safety-gate.cjs + --auto/manual branch + --skip-ui bypass) must move into gsd-ui-phase before its plan:pre hook can replace the inline call without behavior loss. Spike #1018 finding.", - "line": 348 + "line": 327 }, { "id": "RULESET.CAPABILITY.off-means-off", "klass": "RULESET", "value": "the host derives shared outputs from the ACTIVE hook set (via loop.render-hooks); a hook may ADD a labeled block or be COUNTED into a host-computed aggregate (e.g. a score denominator), but NEVER mutates host source — so a disabled capability yields the base output by construction, not by authoring discipline. Ratify in ADR-894; proven by spike #1018.", - "line": 346 + "line": 325 }, { "id": "RULESET.CAPABILITY.precedence-engine-single-owner", "klass": "RULESET", "value": "the config-key four-level precedence walk (loadConfig result → workstream config.json → root config.json → registry.configSchema default → absent) is owned solely by src/capability-activation.cts: raw-value primitive resolveConfigKey(dotKey, {config,cwd,registry}) and boolean wrapper _resolveActivationValue(dotKey,config,cwd,registry); loop-resolver.cts imports the engine (no duplicate); resolveConfigValues in loop-resolver.cts delegates to resolveConfigKey; resolveCapabilityRuntimeState does NOT return registry/config — callers import capability-registry.cjs and call loadConfig(cwd) directly.", - "line": 352 + "line": 331 }, { "id": "RULESET.CAPABILITY.step-additive-gate-blocks", "klass": "RULESET", "value": "a `step` hook is purely additive (invoke skill + produce artifacts, NEVER halts the host); host-blocking preconditions are `gate`s (blocking:true, onError:halt); runtime/mode context (auto/chain vs manual) self-gates IN THE SKILL, not via `when` (config-only). §5.6 = plan:pre step (ui-phase; skill self-gates on frontend+pipeline, auto-fires only in pipelines) + a NEW plan:pre gate (frontend-and-no-UI-SPEC → halt, when:workflow.ui_safety_gate); the loop.render-hooks dispatch template handles steps AND gates. Resolves #1022.", - "line": 350 + "line": 329 }, { "id": "RULESET.CODERABBIT.GUARD.COMPLETE", "klass": "RULESET", "value": "required_checks_green && coderabbit_check_pass && graphQL(reviewThreads.unresolved_count)==0", - "line": 603 + "line": 580 }, { "id": "RULESET.CODERABBIT.GUARD.GRAPHQL", "klass": "RULESET", "value": "reviewThreads(first:100){nodes{id isResolved comments{nodes{author body path line originalLine url}}}}; use unresolved threads as authoritative, not badge text alone", - "line": 604 + "line": 581 }, { "id": "RULESET.CODERABBIT.GUARD.OPEN_PRS", "klass": "RULESET", "value": "gh pr list --repo open-gsd/gsd-core --author @me --state open; repeat near end because open PR set can change mid-run", - "line": 602 + "line": 579 }, { "id": "RULESET.CODERABBIT.GUARD.RERUN", "klass": "RULESET", "value": "after every push wait for CodeRabbit completion, then re-query unresolved threads; CodeRabbit can add new findings after earlier threads were resolved", - "line": 605 + "line": 582 }, { "id": "RULESET.CODERABBIT.GUARD.RESOLVE", "klass": "RULESET", "value": "fix validated finding -> focused tests -> commit/push -> resolveReviewThread(threadId) -> wait CI/CodeRabbit -> final unresolved_count query", - "line": 606 + "line": 583 }, { "id": "RULESET.CODERABBIT.GUARD.SCOPE", "klass": "RULESET", "value": "if a new @me open PR appears during final list, include it in the same guard pass before declaring all-open-PRs complete", - "line": 607 + "line": 584 }, { "id": "RULESET.CONTENT-PATH-NORMALIZATION", "klass": "RULESET", "value": "filesystem paths substituted into markdown body text (@-references, workflow .md, agent .md, generated docs, command bodies) MUST be normalized to POSIX forward slashes via .replace(/\\\\/g,'/') at the production source BEFORE substitution; never push normalization to tests; cross-platform content is POSIX-only; applies to: computePathPrefix output, install-path rewrites, generated shim paths emitted into .md bodies; idempotent on POSIX so unconditional; mechanically enforced by local/normalize-path-in-content (eslint, src/**/*.cts; #1733)", - "line": 808 + "line": 785 }, { "id": "RULESET.CONTRIB.CLASSIFY.enhancement", "klass": "RULESET", "value": "requires approved-enhancement before implementation", - "line": 594 + "line": 573 }, { "id": "RULESET.CONTRIB.CLASSIFY.feature", "klass": "RULESET", "value": "requires approved-feature before implementation", - "line": 595 + "line": 574 }, { "id": "RULESET.CONTRIB.CLASSIFY.fix", "klass": "RULESET", "value": "requires confirmed-bug before implementation (legacy 'confirmed' label is back-compat only for duplicate-sweep exemption, not a valid implementation gate)", - "line": 593 - }, - { - "id": "RULESET.CONTRIB.GAP-TRACKING", - "klass": "RULESET", - "value": "a known gap is tracked only by an issue — a code comment, TODO, or changeset note is attached to code and dies with the refactor that deletes it, while the gap stays real (worked example: archived changeset clever-yaks-cheer.md recorded the shortFormToId backfill as 'tracked as a follow-up parity gap' + a // KNOWN GAP: comment marked the call site; the SDK retirement deleted both, no issue was opened, and it surfaced a year later as #3427). If you write 'tracked as a follow-up' anywhere, open the issue first and cite its number. Prose twin of #3473 criterion B6 (a guard needs a retirement condition)", - "line": 597 - }, - { - "id": "RULESET.CONTRIB.GATE.DELETED-MODULE-DISPOSITION", - "klass": "RULESET", - "value": "a PR deleting a module/package/lineage with a surviving counterpart carries a symbol-level disposition table for the deleted side — every declared name (exported or local) marked migrated (location named) | renamed → X (conventions normalized: verifyX → cmdVerifyX) | dropped because Y; SCREAMING_CASE constants ranked first (policy lives there; both confirmed ADR-0174 losses were bounds); the surviving tests passing is NOT evidence — the surviving tests belong to the surviving side (#3484, ADR-0174 behavior-carry-forward amendment; losses: #3427 shortFormToId tier + unresolved-deps warning, #3477 regexForKeyLinkPattern 512-char cap + nested-quantifier screen, epic #3473 B5 MAX_JSON_SEARCH_DEPTH=48). Method with mandatory positive control: docs/how-to/audit-a-retiring-lineage.md", - "line": 596 + "line": 572 }, { "id": "RULESET.CONTRIB.GATE.ORDER", "klass": "RULESET", "value": "issue-first -> approval-label -> code -> PR-link -> changeset/no-changelog", - "line": 592 + "line": 571 }, { "id": "RULESET.CR-THREAD-RESOLVE", "klass": "RULESET", "value": "after adding // allow-test-rule: to silence lint, resolve existing inline CR threads via graphql resolveReviewThread mutation before merge — open threads mislead future reviewers; pattern: gh api graphql -f query='mutation { resolveReviewThread(input:{threadId:\"PRRT_...\"}) { thread { isResolved } } }'", - "line": 586 + "line": 565 }, { "id": "RULESET.EMITTED_ATTRIBUTION", "klass": "RULESET", "value": "the emitted-artifact family (ADR-2719, epic #2719) — POST-CUTOVER (#2724, Phase 4). Historically tests/fixtures/golden-install-parity/*.json (19 path→hash manifests) + tests/workflow-size-baseline.json + tests/agent-size-baseline.json were all committed, PURE FUNCTIONS of the source tree whose correct merge was ALWAYS \"recompute\" — 140 of 143 conflicted-file instances across the open PR queue were these files. #2724 DELETES all three, the golden test (tests/golden-install-parity.test.cjs), the generator (scripts/gen-golden-install-parity-zcode.cjs), `npm run gen:golden`, `UPDATE_GOLDEN`, the merge-driver bridge (scripts/git-merge-regen-driver.cjs, `npm run setup:merge-driver`, the .gitattributes merge=gsd-regen block), and scripts/update-size-baseline.cjs (`npm run size:baseline`). The differential attribution check (tests/emitted-attribution.test.cjs + tests/emitted-provenance.test.cjs) is now the SOLE gate for emitted-artifact propagation AND size growth — no committed artifact, nothing to hand-merge, nothing to regenerate. `npm run regen:derived` still exists for what remains committed and derived: build, registry, ADR index, capability matrix, inventory manifest, manifest versions, and `tests/fixtures/install-tree/*.json` (now `npm run gen:install-tree`, folded into `regen:derived`). tests/fixtures/install-tree/*.json is DELIBERATELY EXCLUDED from the cutover (ADR-2719 §7): it conflicts on 0 of 7, its diffs are readable, and it preserves \"the installer stopped shipping X\" as a hard absolute failure — capturing it would convert that absolute into an attribution-free auto-resolve. The baseline the differential compares against is now published by `scripts/gen-emitted-baseline.cjs` on every push to `next` (cached, keyed on sha) and restored in PR lanes via `GSD_EMITTED_BASELINE`/`resolveBaseline()` (tests/helpers/emitted-baseline.cjs); a cache miss falls back to an in-job build via a throwaway `git worktree` (tests/helpers/emitted-runtime.cjs's `buildBaselineAtRef`). REMEDIATION IS PART OF THE GATE (#2778): the failure output names its own remedy, because a gate that states a requirement and withholds the means of satisfying it is a maintainer round-trip, not a gate — ADR-2719 §3's \"conspicuous declaration\" only works if the contributor can discover how to make it. Both failing branches name a NEW fragment to create under `tests/emitted-drift-acks/` (#2914; pick a name nobody else is using), say it may not exist yet (absence is the healthy steady state), print a minimal valid document, and repeat \"do NOT regenerate anything\" — post-#2724 there is nothing left to regenerate, and hunting for a deleted baseline is the predictable wrong guess. The two branches key on DIFFERENT spaces and each says which: the hash pass keys on the EMITTED PATH (always contains a `/`), the size ratchet keys on the BARE FILENAME (`currentSizes` writes `sizes[entry.name]` from readdirSync over `gsd-core/workflows/` + `agents/`). A stale-ack failure additionally says to delete the FILE when removing its last entry, since an empty-but-present ack parses fine yet signals nothing; post-#2789 it also offers CORRECTING the entry to name the ripple actually made, which is the other honest resolution and the one a contributor usually wants. NOT ack-able and deliberately given no ack text: the `NEW_FILE_CAP` branch, whose remedy is extraction. Text is sourced from one frozen `REMEDIATION` export in tests/helpers/emitted-diff.cjs whose example document is rendered from `ACK_VERSION` via `JSON.stringify`, so the taught schema cannot drift from the accepted one (a round-trip test feeds the printed document back through `parseAck`); the message teaches ONE canonical shape even though `parseAck` also accepts a bare-string reason and a missing `version` — liberal in what it accepts, conservative in what it sends. Note the ADR's Consequences originally called the #2724 migration \"terminal\"; #2778 corrected that — it is terminal only for a PR that grows no shipped file. #2914 replaced the single shared ack file with per-PR fragments under `tests/emitted-drift-acks/` — exactly the shape `.changeset/` already uses for the identical \"every PR rewrites one shared document\" conflict problem — so two PRs needing an ack can no longer collide with each other, and a fragment left on `next` after merge is inert rather than a shared cell; the legacy file is still read and unioned in for branches that predate the split, and a duplicate path key across two sources is a hard, loudly-reported error, never silent last-wins. `tests/emitted-drift-ack.json` (the LEGACY file specifically, NOT the fragment directory) must NEVER persist on `next` (#2914): every entry is scoped to the diff that introduced it, so once merged it is by definition already at the base — spent and inert regardless of shape — and a persistent copy makes that ONE file a shared merge-conflict cell across every open PR that also carries an ack, exactly the \"140 of 143\" cost this whole cutover exists to remove; a persisting FRAGMENT is harmless by construction and is deliberately not what this guard checks. This is enforced on `next` itself only, never as a PR-lane check: the `guard-no-ack-on-next` job in `.github/workflows/test.yml` (push-to-`next` trigger) runs `scripts/lint-emitted-drift-ack.cjs --guard-next` (`assertAbsentOnNext`), which fails on the LEGACY file's PRESENCE alone, valid or not — a PR-lane \"base ack must be absent\" check would red every open PR the instant a spent ack merged, which is the #2768 shape #2789 already ended. cf `RULESET.WORKFLOW_SIZE_BUDGET`, `RULESET.AGENT_SIZE_BUDGET`; see `### Emitted Artifact Provenance`", - "line": 569 + "line": 548 }, { "id": "RULESET.GENERATIVE-FIX", "klass": "RULESET", "value": "parallel implementations diverge silently when no parity test enforces equality at the test layer; for any new constant/array/parser shared between two parallel surfaces (two workflow surfaces, or a generated artifact and its hand-authored source), the same commit MUST add a parity assertion that fails when the two diverge; exemplar: tests/runtime-launcher-parity.test.cjs (asserts every workflow bash block uses the canonical gsd_run launcher)", - "line": 806 + "line": 783 }, { "id": "RULESET.GH.AUTH.DEFAULT", "klass": "RULESET", "value": "source .envrc GITHUB_TOKEN before gh; exception=ambient allowed only when user explicitly says machine-only fallback", - "line": 601 + "line": 578 }, { "id": "RULESET.HARNESS.test-memory-guard", "klass": "RULESET", "value": "~/.claude/hooks/test-memory-guard.sh fires on every Bash PreToolUse; if argv[0]∈{node|vitest|jest|mocha|tsx|ts-node|tap|ava|playwright|cypress} OR matches (npm|pnpm|yarn|bun) (run )?(t|test|tests|vitest|jest); blocks via hookSpecificOutput.permissionDecision=deny when sum(RSS of running matching procs, excluding tsserver|*-mcp|claude|Electron|...) ≥ 4 GiB OR when argv[0] basename matches a running process's argv[0]. Exception: node --version|-v|--help|-h|-p|-e are trivial probes and skip the check. Designed for a 24 GB Mac where prior accidental fan-out exhausted RAM", - "line": 845 + "line": 822 }, { "id": "RULESET.MANIFEST-CANONICAL-KEY", "klass": "RULESET", "value": "docs/INVENTORY-MANIFEST.json has a single top-level key: families; ALL EIGHT families.* arrays (agents/commands/workflows/references/cli_modules/hooks flat, plus workflow_modes/workflow_steps nested — #2996, epic #1671 Phase 6.5) are canonical, consumed by test suites — tests/inventory-manifest-sync.test.cjs reads all eight, edit-phase/enh-2380/enh-2430 tests read commands+workflows; the six flat families are keyed by BARE BASENAME while the two nested families are keyed by // path, deliberately, because two workflows may each own a same-named step file and a basename key would silently drop one under a JSON-equality comparison; recursion is bounded at exactly one named subdirectory, never a general walk; the family tables live ONCE in scripts/gen-inventory-manifest.cjs and are IMPORTED by the test (the test formerly redeclared them, a DEFECT.GENERATIVE-FIX divergence that let a new family be verified by nobody while still reporting green); the old generated date field and the stale top-level workflows key are both gone; regen via node scripts/gen-inventory-manifest.cjs --write, AFTER build:lib", - "line": 580 + "line": 559 }, { "id": "RULESET.PR-FLOW.docker-before-push", "klass": "RULESET", "value": "before ANY git push of any fix to any PR, run gsd-test (docker on the remote, mirrors ubuntu CI) and confirm exit 0. macOS-local node --test is NOT a substitute — many failures are platform-specific (path separators, case sensitivity, locale, fs semantics). Watchdog with Monitor on the output log; never set a sleep/timer and walk away. Source: user feedback 2026-05-16 — \"we don't set a timer we actively watch and record results in real time as possible\". SUPERSEDED 2026-07-17: 'confirm exit 0' is a false-green trap — piping/backgrounding can report exit 0 on a failed suite; gate on the verdict-line outcome:\"passed\" for the exact HEAD sha instead. See CLAUDE.md's gsd-test rule and the gsd-test-is-ref-based-commit-first predicate for the current, correct gating contract.", - "line": 847 + "line": 824 }, { "id": "RULESET.PR-FLOW.templates-mandatory", "klass": "RULESET", "value": "every gh pr create|edit|gh issue create|edit MUST first invoke the gh-templates-first skill and Read (Read tool, not Bash cat — k321 read-tracking) the matching template in .github/. Apply ALL required sections; never write freeform bodies. Repo enforces this via gsd-pr-template-policy GitHub Action which flags any non-templated body — the bot allows the PR to stay open only because authors are contributors-or-higher, but the warning is a real complaint that must be cured. Source: user feedback 2026-05-16 (multi-message escalation) — \"the whole reason i have that github action is because you fucking blow through and ignore using the templates\"", - "line": 849 + "line": 826 }, { "id": "RULESET.PR-SCOPE.one-concern-per-pr", "klass": "RULESET", "value": "split unrelated changes into separate PRs; cherry-pick doc changes to dedicated docs/ branch immediately, then force-push original to remove the commit", - "line": 582 + "line": 561 }, { "id": "RULESET.SHARED-HELPERS-LINT-VS-TEST", "klass": "RULESET", "value": "when a lint script and test suite both implement same constant (CANONICAL_TOOLS) or parser (parseFrontmatter, executionContextRefs), extract to scripts/*-helpers.cjs required by both — silent divergence otherwise", - "line": 577 + "line": 556 }, { "id": "RULESET.TESTS.CODERABBIT_FIX", "klass": "RULESET", "value": "prefer exported-function behavioral tests over source-grep; lint-no-source-grep rejects readFileSync source assertions without allow-test-rule", - "line": 608 + "line": 585 }, { "id": "RULESET.TESTS.boundary-coverage", "klass": "RULESET", "value": "tests MUST exercise inputs at and near the threshold/limit, not only trivial-fit and trivial-overflow; pick inputs where N ∈ {limit-1, limit, limit+1} and where pre-trim/pre-check accumulators ≈ effective limit; \"very small\" and \"very large\" inputs alone do not constitute edge-case coverage and routinely miss off-by-one + reservation-accounting bugs", - "line": 551 + "line": 530 }, { "id": "RULESET.TESTS.boundary-coverage.anti-pattern", "klass": "RULESET", "value": "test suites that pair budget:1_000_000 (trivially fits) with budget:1 (trivially overflows) and skip the boundary region; failure mode that shipped PR #3708 UNNEEDED_TRIM + FALSE_HARDFAIL regressions (commit 2df566ed, fixed bde1ae8f)", - "line": 554 + "line": 533 }, { "id": "RULESET.TESTS.boundary-coverage.fixtures", "klass": "RULESET", "value": "for any code with budget/limit/quota/threshold parameter, test suite MUST include: (a) input where SUT estimate == limit exactly, (b) input where estimate == limit - 1, (c) input where estimate == limit + 1, (d) input where any internal reserve/safety constant pushes baseline within reserve-distance of limit (catches early-pressure firing)", - "line": 553 + "line": 532 }, { "id": "RULESET.TESTS.clock-seam", "klass": "RULESET", "value": "concurrency logic must accept an optional {clock=Date} parameter; tests control time via t.mock.timers.enable(['Date']) + t.mock.timers.setTime(0) + t.mock.timers.tick(N); real OS scheduler races are not a permitted test pattern after ADR 456 (2026-05-28); real-race tests are deleted once deterministic seam tests cover the same logical path; clock.cjs realClock adds nowIso() (→ new Date(this.now()).toISOString()) and today() (→ nowIso().split('T')[0]) so all date-stamping in state.cjs routes through the seam; subprocess time-pin adapter: set GSD_TEST_MODE=1 + GSD_NOW_MS= in runGsdTools env to pin the date written by the SUT without touching real wall-clock (issue #474)", - "line": 558 + "line": 537 }, { "id": "RULESET.TESTS.coderabbit-fix-prefer", "klass": "RULESET", "value": "behavioral tests (call exported fn, capture JSON, assert typed fields) over source-grep", - "line": 549 + "line": 528 }, { "id": "RULESET.TESTS.delete-bad-tests", "klass": "RULESET", "value": "pass-always / vacuous-truth / source-grep / elapsed-time / real-race / permanent-allow-test-rule tests are DELETED and replaced with compliant tests in the same PR; not skipped, not commented out, not permanently exempted; replacement must cover the same logical path via typed-surface assertion or clock-seam pattern", - "line": 561 + "line": 540 }, { "id": "RULESET.TESTS.diagnostics", "klass": "RULESET", "value": "after JSON.parse, assert output shape (Array.isArray(output.phases)) with raw-output-prefix diagnostics before .map() — prevents opaque TypeErrors when CLI output shape changes", - "line": 550 + "line": 529 }, { "id": "RULESET.TESTS.escape-regex", "klass": "RULESET", "value": "new RegExp(\"prefix${var}\") must escapeRegex(var); phase-id.cjs exports escapeRegex (core.cjs re-export spine retired in epic #1267); phase IDs like 5.1 contain . which is metacharacter", - "line": 546 + "line": 525 }, { "id": "RULESET.TESTS.eslint-harness", "klass": "RULESET", "value": "ADR 452 (2026-05-28): ESLint flat config + typescript-eslint + eslint-plugin-n + eslint-plugin-no-only-tests + local plugin at eslint-rules/ (repo root, NOT scripts/eslint-rules/); replaces scripts/lint-*.cjs regex scanners (fully removed in #632); all three test-rigor rules now ship at error in tests/**/*.test.cjs scope: local/no-source-grep and local/no-magic-sleep-in-tests promoted by #3313, local/no-elapsed-assertion promoted by #3331 once #3314 delivered its ADR-456 §(a) precondition (epic #1885 was subsumed into epic #3053 and closed stale before this promotion landed)", - "line": 562 + "line": 541 }, { "id": "RULESET.TESTS.feedback-loop-convergence", "klass": "RULESET", "value": "when a feature's OUTPUT feeds back into its own INPUT (calibration, retry backoff, adaptive budgets, ratchets, any self-correcting signal), step-wise tests are NOT sufficient evidence of correctness: they assert `given X return Y` while the defect lives in the TRAJECTORY across iterations. Required: a closed-loop test that (a) drives the REAL end-to-end surface — not the pure core alone, since composition bugs live between surfaces — for N >= 2x the loop's window, (b) asserts convergence on the known-true value, (c) asserts the fixed point (an already-correct history must produce NO correction), and (d) asserts boundedness under an adversarial/oscillating history. Two defects shipped past a green ~26,800-test suite in epic #1952 for want of exactly this: calibration applied twice across two surfaces (factor^2, #2631) and calibration measured against its own corrected output so it oscillated to ~1.41 instead of converging on 2.0 (#2632). Every unit, boundary, property and round-trip test passed for both. HOW TO SPOT ONE (the detection tell, not a judgment call): the feature's own acceptance criterion carries a TEMPORAL QUANTIFIER — \"after N phases\", \"subsequent\", \"over time\", \"improves\", \"learns\", \"adapts\". That phrasing means the claim is about a TRAJECTORY, so a step-wise `given X return Y` test does not test the claim that was made. #1952's AC4 read \"After N phases, the error is computed and applied as a correction to SUBSEQUENT estimates\" — the tell was in plain sight and was still tested as a point. Survey of this repo (2026-07): estimation calibration is the ONLY true instance; size/mutation ratchets are exempt because they fail on both growth AND shrinkage (cannot self-satisfy), and retry ladders (node_repair_budget, plan_bounce_passes, provider_escalation) terminate rather than feed back. Test anchor: tests/estimate-loop-convergence.test.cjs", - "line": 552 + "line": 531 }, { "id": "RULESET.TESTS.guard-toplevel-readFileSync", "klass": "RULESET", "value": "module-level const src = readFileSync(...) throws before any test() registers — wrap in try/catch in test() or use lazy load", - "line": 548 + "line": 527 }, { "id": "RULESET.TESTS.mutation-score", "klass": "RULESET", "value": "Stryker runs incremental (--since origin/next) on ubuntu-latest/Node24 CI leg; default threshold 80% killed/total; surviving mutants in scope block merge unless path is listed in stryker.config.mjs with documented reason; treat surviving mutant as a failing test specification", - "line": 560 + "line": 539 }, { "id": "RULESET.TESTS.no-dead-regex-in-includes", "klass": "RULESET", "value": "src.includes(\"foo.*bar\") is always false — .* is regex metacharacter not wildcard; use new RegExp(...).test(src) or delete", - "line": 547 + "line": 526 }, { "id": "RULESET.TESTS.no-duplicate-fold-marker", "klass": "RULESET", "value": "local/no-duplicate-fold-marker ESLint AST rule (eslint-rules/no-duplicate-fold-marker.cjs, #3271) reports the 2nd and every later __foldDescribe(\"folded: ...\") call carrying a marker already seen in the SAME file, naming the first occurrence's line; error in tests/**/*.cjs. The key is the WHITESPACE-delimited token after folded:, NOT a [a-z0-9-]* slice — a slice truncates at \".\" and collides feat-443-effort-fast-mode.integration with feat-443-effort-fast-mode (two distinct suites coexisting in tests/model-resolver.test.cjs), and NOT the whole title, so a re-fold under a different batch label (\"B1 #1970\" vs \"B5 #1975\") is still caught. Deliberately silent on: a __foldDescribe title with no folded: prefix (the alias is reused for one ordinary describe in tests/review-default-reviewers-workflow.test.cjs), a plain describe(), a non-literal title, and the same marker in two DIFFERENT files (the defect class is intra-file).", - "line": 543 + "line": 522 }, { "id": "RULESET.TESTS.no-duplicate-fold-marker.why", "klass": "RULESET", "value": "consolidation epic #1969 folds are self-contained blocks, so a second verbatim copy parses, registers and PASSES twice — nothing reports it; #3271 found 25 such copies (~5,800 lines) in tests/install.test.cjs (18), tests/install-minimal-hooks.test.cjs (5) and tests/install-write-confinement.test.cjs (2), all from one stale-base re-application in 6d072435d (#1975 re-applying #1970's hunks, 2026-07-03). Ref DEFECT.GENERATIVE-FIX: the two copies drift apart silently when a contributor fixes one and leaves the other asserting the old behavior, with the suite still green.", - "line": 544 + "line": 523 }, { "id": "RULESET.TESTS.no-source-grep", "klass": "RULESET", "value": "local/no-source-grep ESLint AST rule (eslint-rules/no-source-grep.cjs) rejects readFileSync of a source .cjs/.js/.ts path bound to a var later hit with .includes()/.match()/.startsWith()/.endsWith()/.indexOf()/.search(); error in tests/**/*.test.cjs, warn in gsd-core/bin/**/*.cjs + scripts/**/*.cjs (ADR 452 retired the old regex script, removed for good in #632)", - "line": 540 + "line": 519 }, { "id": "RULESET.TESTS.no-source-grep.exemption", "klass": "RULESET", "value": "// allow-test-rule: with one-line justification; reserved for tests where the file content IS the product surface (STATE.md, config.toml, hooks.json, agent .md). Migration to typed-IR parser tracked in #2974.", - "line": 541 + "line": 520 }, { "id": "RULESET.TESTS.no-source-grep.tmp-file-traps", "klass": "RULESET", "value": "reading tmp files written by the SUT in tests still trips lint; round-trip through CLI (e.g. frontmatter get) instead of readFileSync+.includes()", - "line": 542 + "line": 521 }, { "id": "RULESET.TESTS.no-timing-assertion", "klass": "RULESET", "value": "do not assert on wall-clock elapsed time (Date.now() delta, performance.now(), process.hrtime() comparison); such assertions test the host machine not the SUT and flake on loaded CI runners; enforcement: local/no-elapsed-assertion ESLint rule, error (promoted by #3331 once #3314 delivered the ADR-456 §(a) reachability rule + deterministic backfill precondition); canonical replacement: clock-seam pattern with node:test mock.timers", - "line": 557 + "line": 536 }, { "id": "RULESET.TESTS.property-based-testing", "klass": "RULESET", "value": "modules implementing parsing / transformation / budget-limit / bijective contracts must include at least one fast-check (fc) property test asserting a domain invariant; invariant categories: round-trip, monotonicity, boundary-containment, idempotency; property tests live in *.test.cjs alongside unit tests; CI signal: Stryker mutation score below 80% blocks merge", - "line": 559 + "line": 538 }, { "id": "RULESET.TRIAGE-EXISTING-WORK", "klass": "RULESET", "value": "before writing agent brief for confirmed bug, check (1) local branches git branch -a | grep , (2) untracked/modified files on that branch, (3) stash, (4) open PRs with matching head branch — recover existing work rather than re-implement", - "line": 584 + "line": 563 }, { "id": "RULESET.WORKFLOW.COVERAGE-METADATA", "klass": "RULESET", "value": "#1602 SUMMARY frontmatter `coverage:` block (list of {id,description,requirement?,verification:[{kind∈unit|integration|e2e|automated_ui|manual_procedural|other, ref, status∈pass|fail|unknown}],human_judgment:bool,rationale?}) is the per-deliverable RTM consumed DETERMINISTICALLY by verify-work extract_tests via `gsd-tools uat classify-coverage --summary ` (src/coverage.cts → bin/lib/coverage.cjs). AUTHORING: execute-plan create_summary populates it from task results; every deliverable MUST be classified; fail-safe default = human_judgment:true + rationale. CLASSIFY CONTRACT: auto-pass (skip human) ONLY when human_judgment===false (strict boolean) AND verification non-empty AND every status==='pass' AND zero validation errors — else PRESENT to human. mode:legacy (no block) ⇒ byte-identical prose `## Accomplishments` fall-through; `coverage: []` ⇒ mode:coverage, zero entries (single-confirmation). Frozen IR: MODE/PRESENT_REASON/ERROR_CODE enums locked by tests/coverage-metadata-parser.test.cjs. extractFrontmatter CANNOT parse it (scalars-only `-` items) → dedicated parser, sibling of parseMustHavesBlock. Asymmetry by design: false-negative=redundant prompt (status quo); false-positive=shipped bug UAT existed to catch", - "line": 573 + "line": 552 }, { "id": "RULESET.WORKFLOW_EXECUTE_END_TO_END", "klass": "RULESET", "value": "standard for single-workflow commands is \"Execute end-to-end.\" (no bolded **Follow the X workflow** fragments); flag-dispatch routing uses \"execute the X workflow end-to-end.\" in routing bullets — convention verified live across ~20 commands/gsd/*.md files; no ADR currently documents this specific phrasing rule (ADR-0002 covers the adjacent but distinct command-contract/@-ref-resolution seam, not this convention)", - "line": 572 + "line": 551 }, { "id": "RULESET.WORKFLOW_EXECUTION_CONTEXT", "klass": "RULESET", "value": "@-ref in commands/gsd/*.md must resolve to an existing file on disk; regression test in tests/docs-update.test.cjs (folds former \\`bug-3135-capture-backlog-workflow\\`, consolidation epic #1969); INVENTORY.md row + INVENTORY-MANIFEST.json families.workflows must stay in sync; \"Invoked by\" attribution must move when a flag absorbs a micro-skill", - "line": 571 + "line": 550 }, { "id": "RULESET.WORKFLOW_FILE_NAMES", "klass": "RULESET", "value": "workflow files use hyphens; XML attributes must match (extract-learnings not extract_learnings); tests should pin exact hyphenated name", - "line": 570 + "line": 549 }, { "id": "RULESET.WORKFLOW_MARKDOWN.FENCES", "klass": "RULESET", "value": "preserve opening language fence when editing shell snippets in workflow markdown; malformed fence creates fresh CR threads (MD040)", - "line": 566 + "line": 545 }, { "id": "RULESET.WORKFLOW_SIZE_BUDGET", "klass": "RULESET", "value": "workflow size enforcement (#1074; BYTES not lines per #717; LF-normalized per #683) = differential attribution size ratchet (PRIMARY anti-creep since #2724/ADR-2719 §4: tests/emitted-attribution.test.cjs's real-tree test reports growth in any gsd-core/workflows/*.md with its exact byte delta vs `next`, no committed snapshot, requires an ack entry — a fragment under tests/emitted-drift-acks/, #2914; the legacy tests/emitted-drift-ack.json is still honored and unioned in) + loose tier hard caps (outer red lines, NEVER raised on approach: XL<=98304 / LARGE<=61440 / DEFAULT<=40960) + discuss-phase<32000; a file that grew fails the differential guard — add an ack entry naming the file and reason, justify the growth in the PR (or extract LAZILY-loaded content; eager @-imports don't reduce loaded context); crossing a hard cap means EXTRACT, not bump. The prior per-file baseline (tests/workflow-size-baseline.json, `npm run size:baseline`) is REMOVED by #2724. Its new-file cap (ADR-1610 Decision point 3, un-baselined files <=32768, the Codex anchor) is REVIVED inside the differential's size ratchet itself (`NEW_FILE_CAP` in tests/helpers/emitted-diff.cjs) rather than lost: \"not yet baselined\" is exactly \"present in sizeCurrent, absent from sizeBaseline\", a signal the ratchet already computes for its own reasons. NOT ack-able — same as the tier hard caps, the fix is extraction. Narrower than the original: this check cannot see XL/LARGE tiering (tests/workflow-size-budget.test.cjs's classification, invisible to the pure differential module), so a legitimately large NEW file must extract rather than tier in, one release earlier than an existing file would need to — a disclosed, deliberate simplification", - "line": 567 + "line": 546 }, { "id": "SESSION.2026-05-05", "klass": "SESSION", "value": "[PRED.k320..k331 introduced; DEFECT.SOURCE-GREP-IN-NEW-TESTS, DEFECT.CHANGESET-PR-FIELD-DRIFT, DEFECT.PHASE-DIR-PREFIX-DRIFT, DEFECT.PROMPT-INJECTION-SCAN-COLLISION; ADR-0002 thin-wrapper pattern findings folded into RULESET.WORKFLOW_*]", - "line": 835 + "line": 812 }, { "id": "SESSION.2026-05-05.sdk-bridge", "klass": "SESSION", "value": "PR #3158 SDK Runtime Bridge — observability isolation rule; strict-mode dispatchMode reporting invariant; transport decision ordering (guard before event emission); folded into Dispatch Policy Module glossary", - "line": 836 + "line": 813 }, { "id": "SESSION.2026-05-09", "klass": "SESSION", "value": "[8-PR triage wave, 7 merged + 1 subsumed; META.RULE.* introduced; WAVE.LESSON.* captured; k320/k322/k323/k326/k331 evidence; AI Ops Memory predicate format established]", - "line": 837 + "line": 814 }, { "id": "SESSION.2026-05-10", "klass": "SESSION", "value": "[ai-ops memory consolidation; release-notes standard taxonomy + templates; RELEASE-NOTES.* predicates introduced]", - "line": 838 + "line": 815 }, { "id": "SESSION.2026-05-13", "klass": "SESSION", "value": "[Shell Command Projection Module expansion (#3465-#3468); ADR-0009 superseded; new exports for subprocess dispatch and platform file I/O; phase-gated migration plan; PR #3464 three-gate invariant CI+CR+unresolved=0; PR #3470 stash-include-untracked rebase pattern]", - "line": 839 + "line": 816 }, { "id": "SESSION.2026-05-14", "klass": "SESSION", "value": "[#3095/PR #3490 EXEC.CLASSIFY.* introduced (Anthropic/Copilot/Codex/Gemini [runtime removed #1928] cross-runtime rate-limit sentinel coverage); #3489/PR #3499 DEFECT.STATE-TRAMPLE.idempotency-oracle (STATE.md current_phase field is oracle for state.complete-phase); #3488/PR #3501 DAG resolver same-phase short-form depends_on (shortFormToId index added to sdk/src/query/phase.ts); #3491/PR #3502 DEFECT.NESTED-GIT-INIT (gitWorktreeInfoInternal helper); #3493/PR #3500 extractCurrentMilestone generic Phase Details continuation past planned-milestone siblings; #3503/PR #3504 DEFECT.PATH-SUBSTRING-CHECK (trailing-slash anchor for homedir checks); #3346/PR #3505 codex AoT TOML leaf-key via extractFlatHookEventName; #3506/PR #3507 label-scoped stale-bot sub-job pattern; multi-PR triage operational lessons folded into PROC.TRIAGE.*; #3508 DEFECT.AGENT-ISOLATION-SILENT-FAIL; gsd-test image-missing auto-build (locally-built image via embedded heredoc Dockerfile); refined PRED.k322 threshold to 3 PRs/<10min]", - "line": 840 + "line": 817 }, { "id": "SESSION.2026-05-15", "klass": "SESSION", "value": "[#3537/PR #3538 DEFECT.PHASE-REGEX-FANOUT — phaseMarkdownRegexSource promoted to core.cjs and wired to 7 sites; parity-style regression test established as DEFECT.GENERATIVE-FIX exemplar; trek-e/gsd-test-runner#1 filed for DEFECT.GSD-TEST-MIRROR-POISONED — chown-back-before-exec legacy gap (poisoned holodeck mirror unstuck via authorized docker chown to remote 1000:1000); RULESET.PR-FLOW.* codified from project CLAUDE.md load-bearing rule; first dispatch under run-tests-before-create held cleanly (PR #3520 worker stopped on Docker exit 12 infra failure, orchestrator opened PR after unblock); CONTEXT.md refactored from 882 lines of mixed prose+predicates into ~500 lines of pure-predicate format with chronological session log]", - "line": 841 + "line": 818 }, { "id": "SESSION.2026-05-15.parallel-fix-dispatch", "klass": "SESSION", "value": "[#3542/PR #3546 prohibit git stash family in executor agents (shared refs/stash across worktrees); #3541/PR #3547 non-TTY resolution for installer prompt-user actions (default remove for SDK build artifacts, keep for skills/gsd-*/SKILL.md); #3545 filed for gsd-test-summary concurrent /tmp output collision; new predicates DEFECT.HOOK-OVER-ENFORCEMENT.read-tool-tracking, DEFECT.GSD-TEST-CONCURRENT-OUTPUT-COLLISION, DEFECT.SUBAGENT-LONG-RUNNING-BG-STALL, DEFECT.AGENT-RETIRED-SLASH-SYNTAX-DRIFT, PROC.PARALLEL-FIX-DISPATCH; agent-trust-but-verify caught /gsd-update retired-syntax comment slip in #3541 implementation before PR open]", - "line": 842 + "line": 819 }, { "id": "SESSION.2026-05-16", "klass": "SESSION", "value": "[multi-PR triage wave (#3577/3581/3640/3641/3642/3648/3649/3637/3639). Established global PreToolUse hook ~/.claude/hooks/test-memory-guard.sh denying new node/test spawns when sum(RSS of node|vitest|jest|...) >= 4 GiB on the 24 GB Mac OR when a same-runner process is already in argv[0] — hard deny via hookSpecificOutput.permissionDecision=deny. PR #3577 fix: revert config-ensure-section dispatch to CJS cmdConfigEnsureSection (SDK author wrote single-section semantics under a name whose legacy callers expect full-default config init); plus 3 SDK parity carve-outs (configNewProject defaults align with sdk/shared/config-defaults.manifest.json, return relative .planning/config.json path, drop quotes from Unknown config key, lead malformed-JSON error with \"Failed to read config.json:\"). PR #3649 fix: chunk node --test spawn at 28K argv ceiling (Windows CreateProcess lpCommandLine cap 32,767 was instantly aborting unchunked spawn of 546 paths). Chunking fix surfaced 14 pre-existing Windows-only test bugs (4010 pass / 14 fail; vs 0/0 before — entire suite was un-runnable on Windows). PRs #3639 + #3637 confirmed unable to stand alone (legitimately depend on Phase 6 scaffolding only present on feat/3575-enforcement-hardening) — user decision: cherry-pick into #3577 and close. Five other PRs each had ≤1 unresolved CR thread of the changeset-pr-number / null-vs-throw / implicit-Claude-runtime / docs-stale-guidance / hardcoded-tests-path family — all quick wins. New predicates: DEFECT.SDK-PORT-NAME-COLLISION, DEFECT.WINDOWS-ARGV-OVERFLOW, DEFECT.STACKED-PR-CANNOT-STAND-ALONE, DEFECT.CANARY-VERSION-LEAK, DEFECT.GSD-TEST-HOST-MID-RUN-DEATH, RULESET.HARNESS.test-memory-guard, RULESET.PR-FLOW.docker-before-push, RULESET.PR-FLOW.templates-mandatory]", - "line": 843 + "line": 820 }, { "id": "WAVE.LESSON.agent-narrative-unreliable", "klass": "WAVE", "value": "k095/k324 confirmed at scale: 5 of 8 agents terminated mid-monitor with stale claims requiring direct verification", - "line": 799 + "line": 776 }, { "id": "WAVE.LESSON.changelog-policy-violation-multiplier", "klass": "WAVE", "value": "brief contradicting CONTRIBUTING.md's changelog-fragment policy (\"CHANGELOG Entries — Drop a Fragment\" section) produced violations on 5 of 8 PRs (#3300, #3302, #3304, #3305, #3308); k326 + k320 capture", - "line": 796 + "line": 773 }, { "id": "WAVE.LESSON.cr-throttle-burst-correlation", "klass": "WAVE", "value": "8 PRs in <15min triggered k322 sustained-throttle on multiple PRs (#3306 worst case)", - "line": 797 + "line": 774 }, { "id": "WAVE.LESSON.k101-still-trips", "klass": "WAVE", "value": "even after CONTEXT.md k101 reinforcement, agent of record posted self-PR comment on close; k331 adds explicit close-time literal-instruction guard", - "line": 800 + "line": 777 }, { "id": "WAVE.LESSON.sibling-audit-overlap", "klass": "WAVE", "value": "k015-family parallel dispatch on #3297 + #3298 produced k323 add-backlog.md cross-PR overlap", - "line": 798 + "line": 775 }, { "id": "WORKSTREAM.INVARIANT.migrate-name", "klass": "WORKSTREAM", "value": "must normalize through canonical slug policy", - "line": 622 + "line": 599 }, { "id": "WORKSTREAM.INVARIANT.slug-contract", "klass": "WORKSTREAM", "value": "all .planning/workstreams/ must be addressable by set/get/status/complete", - "line": 623 + "line": 600 }, { "id": "WORKSTREAM.NAME.POLICY.cjs-module", "klass": "WORKSTREAM", "value": "gsd-core/bin/lib/workstream-name-policy.cjs owns toWorkstreamSlug + active-name/path-segment validation", - "line": 638 + "line": 615 }, { "id": "WORKSTREAM.POINTER.SEAM.cjs-module", "klass": "WORKSTREAM", "value": "gsd-core/bin/lib/active-workstream-store.cjs owns read/write self-heal for .planning/active-workstream", - "line": 639 + "line": 616 }, { "id": "WORKSTREAM.REGRESSION.test-anchor", "klass": "WORKSTREAM", "value": "tests/workstream.test.cjs::normalizes --migrate-name to a valid workstream slug", - "line": 624 + "line": 601 }, { "id": "WORKTREE.SEAM.caller-rule", "klass": "WORKTREE", "value": "verify.cjs must consume inspectWorktreeHealth for W017 classification; no ad-hoc porcelain parsing in callers", - "line": 632 + "line": 609 }, { "id": "WORKTREE.SEAM.current", "klass": "WORKTREE", "value": "Worktree Safety Policy Module", - "line": 616 + "line": 593 }, { "id": "WORKTREE.SEAM.decision-1", "klass": "WORKTREE", "value": "retain non-destructive default; destructive path only as explicit future opt-in scaffold", - "line": 620 + "line": 597 }, { "id": "WORKTREE.SEAM.default-prune-policy", "klass": "WORKTREE", "value": "metadata_prune_only (non-destructive)", - "line": 619 + "line": 596 }, { "id": "WORKTREE.SEAM.files", "klass": "WORKTREE", "value": "[gsd-core/bin/lib/worktree-safety.cjs]", - "line": 617 + "line": 594 }, { "id": "WORKTREE.SEAM.interface", "klass": "WORKTREE", "value": "[resolveWorktreeContext, parseWorktreePorcelain, planWorktreePrune, executeWorktreePrunePlan, planWorktreeRecordAgent, cmdWorktreeRecordAgent]", - "line": 618 + "line": 595 }, { "id": "WORKTREE.SEAM.invariant", "klass": "WORKTREE", "value": "parser failure must degrade to metadata_prune_only and never escalate to destructive removal", - "line": 630 + "line": 607 }, { "id": "WORKTREE.SEAM.inventory-interface", "klass": "WORKTREE", "value": "[listLinkedWorktreePaths, inspectWorktreeHealth]", - "line": 631 + "line": 608 }, { "id": "WORKTREE.SEAM.inventory-snapshot", "klass": "WORKTREE", "value": "snapshotWorktreeInventory(repoRoot,{staleAfterMs,nowMs}) is canonical linked-worktree health snapshot for callers", - "line": 634 + "line": 611 }, { "id": "WORKTREE.SEAM.test-anchor-w017", "klass": "WORKTREE", "value": "tests/orphan-worktree-detection.test.cjs + tests/worktree-safety-policy.test.cjs", - "line": 633 + "line": 610 }, { "id": "WORKTREE.SEAM.test-anchors", "klass": "WORKTREE", "value": "[resolveWorktreeContext:has_local_planning|linked_worktree|not_git_repo|main_worktree, planWorktreePrune:git_list_failed|worktrees_present|no_worktrees|parser_throw_fallback, executeWorktreePrunePlan:missing_plan|skip_passthrough|unsupported_action|metadata_prune_only]", - "line": 629 + "line": 606 }, { "id": "WORKTREE.SEAM.test-policy", "klass": "WORKTREE", "value": "cover all decision branches in policy module before changing prune behavior", - "line": 628 + "line": 605 } ], "duplicates": [] diff --git a/src/capability-lifecycle.cts b/src/capability-lifecycle.cts index ac3898664..bddc1c9f8 100644 --- a/src/capability-lifecycle.cts +++ b/src/capability-lifecycle.cts @@ -202,6 +202,31 @@ interface LifecycleOptions { /** Stamp written onto every capability-owned shared-config entry, for surgical removal. */ const CAP_MARKER = '_gsdCapability'; +/** + * #3514 (epic #1900 F21c): what KIND of pin (if any) the source content was pinned with — shared + * by the install and upgrade verdicts so the two cannot drift. Reaching the verdict with a + * supplied `--integrity` pin means the pin VERIFIED (the resolver throws on mismatch). A git + * `#sha:<40-hex-commit>` ref is the git analog of a hash pin — a commit checkout, verified by + * spelling: `/^sha:[0-9a-f]{7,40}$/i` accepts only a hex commit id, so a mutable ref + * (`#sha:main`, `#sha:v1`) is NOT counted as a pin (isolated review finding — a moving ref must + * never render as pinned). Everything else stages with no pin and must say so in the consent + * prompt. + */ +type IntegrityPin = 'sha512' | 'git-commit' | 'none'; + +const GIT_SHA_PIN_RE = /^sha:[0-9a-f]{7,40}$/i; + +function resolveIntegrityPin( + integrity: unknown, + parsed: { kind?: string; ref?: unknown }, +): IntegrityPin { + if (typeof integrity === 'string' && integrity.length > 0) return 'sha512'; + if (parsed.kind === 'git' && typeof parsed.ref === 'string' && GIT_SHA_PIN_RE.test(parsed.ref)) { + return 'git-commit'; + } + return 'none'; +} + /** Keys that must never be used as object indices (prototype-pollution guard). */ function isUnsafeKey(k: string): boolean { return k === '__proto__' || k === 'constructor' || k === 'prototype'; @@ -1023,6 +1048,7 @@ async function installCapability(spec: string, opts: LifecycleOptions): Promise< stagedDir, strictKnownRegistries, hostVersion, + integrityPin: resolveIntegrityPin(opts.integrity, parsedPre), }); if (!verdict.allowed) { @@ -1225,6 +1251,7 @@ async function upgradeCapability(spec: string, opts: LifecycleOptions): Promise< stagedDir, strictKnownRegistries, hostVersion, + integrityPin: resolveIntegrityPin(opts.integrity, parsedPre), }); if (!verdict.allowed) { return { status: 'blocked', disclosure: verdict.disclosure, blockReasons: verdict.blockReasons }; diff --git a/src/capability-source.cts b/src/capability-source.cts index 5c65d8ea2..55166fc34 100644 --- a/src/capability-source.cts +++ b/src/capability-source.cts @@ -212,8 +212,102 @@ function _setHttpsGetImpl(fn: HttpsGetImpl | null): void { // Injectable HTTP transport (test seam) // --------------------------------------------------------------------------- +/** + * #3514 (epic #1900 F21): pre-transport fetch gate. Runs BEFORE any bytes leave and before the + * (possibly injected) transport is invoked: + * + * - Non-`https` URLs are refused with a NAMED, actionable reason. parseSpec still CLASSIFIES an + * `http://` tarball URL as tarball (no parse-time hard reject — internal-mirror classification + * flows stay intact, the adopted split decision on the epic); the https-only transport then + * fails clearly instead of with a raw ERR_INVALID_PROTOCOL. + * - Loopback / link-local / metadata / unspecified hosts and `localhost`-form names are denied + * unconditionally — no legitimate capability install fetches them. RFC1918 private ranges are + * deliberately ALLOWED (internal mirrors are the legitimate use this leaves room for); the + * denylist is not an allowlist. + * + * Known limit (disclosed): the check is on the URL's HOST LITERAL. A public hostname that + * RESOLVES via DNS to a denied range (rebinding) is out of scope — `https.get` resolves + * internally and this gate does not double-resolve. + */ +function assertFetchableUrl(rawUrl: string): void { + let u: URL; + try { + u = new URL(rawUrl); + } catch { + throw new Error(`invalid capability source URL: "${rawUrl}"`); + } + if (u.protocol !== 'https:') { + throw new Error( + `capability source fetch requires https — refusing "${u.protocol}" URL (host "${u.hostname}"). ` + + 'The transport never fetches plaintext http; for an internal mirror use an https URL or a local path source.' + ); + } + // Node's URL.hostname KEEPS the brackets on IPv6 literals ('[::1]') — strip them so the + // v6 checks below see the bare address. A single trailing dot is FQDN syntax ('localhost.' + // is the same name as 'localhost'; some resolvers synthesize loopback for it) — strip it. + const host = u.hostname.toLowerCase().replace(/^\[|\]$/g, '').replace(/\.$/, ''); + const deniedHost = () => + new Error( + `refusing to fetch capability source from internal host "${host}" (loopback/link-local/metadata denylist)` + ); + if (host === 'localhost' || host.endsWith('.localhost')) { + throw deniedHost(); + } + // IPv4 literal: 127/8 loopback, 0/8 unspecified, 169.254/16 link-local (contains the cloud + // metadata IPs). Malformed literals that match no range simply fall through to a failed + // resolution downstream — denying is not needed for safety there. + const v4 = host.match(/^(\d{1,3})\.(\d{1,3})\.(\d{1,3})\.(\d{1,3})$/); + if (v4) { + const a = Number(v4[1]); + const b = Number(v4[2]); + if (a === 127 || a === 0 || (a === 169 && b === 254)) throw deniedHost(); + return; + } + // IPv6-literal checks. A registered domain NEVER contains ':' after bracket-strip, while every + // IPv6 literal does — that colon is the guard. Without it, the fe80::/10 prefix test below + // would deny legitimate domains like 'feather.internal' or 'february.example.com' (isolated + // review finding — a domain-shaped host must never reach the v6 branch). + if (!host.includes(':')) { + return; + } + // An IPv4-mapped address re-checks as IPv4 — in EITHER spelling, because the URL parser may + // preserve the dotted form ('::ffff:127.0.0.1') or normalize to hex hextets ('::ffff:7f00:1'). + // Otherwise deny ::1 (loopback), :: (unspecified), and fe80::/10 (link-local, hex prefix range + // fe80–febf). + if (host.startsWith('::ffff:')) { + const suffix = host.slice('::ffff:'.length); + let v4str: string | null = null; + if (/^\d{1,3}(?:\.\d{1,3}){3}$/.test(suffix)) { + v4str = suffix; + } else { + const hextets = suffix.split(':'); + if (hextets.length === 2 && /^[\da-f]{1,4}$/.test(hextets[0]) && /^[\da-f]{1,4}$/.test(hextets[1])) { + const hi = parseInt(hextets[0], 16); + const lo = parseInt(hextets[1], 16); + v4str = `${(hi >> 8) & 0xff}.${hi & 0xff}.${(lo >> 8) & 0xff}.${lo & 0xff}`; + } + } + if (v4str) { + assertFetchableUrl(`https://${v4str}/`); + return; + } + } + // fe80::/10 = hex prefix fe80..febf, i.e. /^fe[89ab]/ (third nibble 8-b has top two bits set). + if (host === '::1' || host === '::' || /^fe[89ab]/.test(host)) { + throw deniedHost(); + } +} + function realHttpsGet(url: string): Promise { return new Promise((resolve, reject) => { + // #3514: gate BEFORE the transport — scheme + host denylist refuse with named errors and + // never invoke the (possibly injected) transport. A throw inside the executor rejects. + try { + assertFetchableUrl(url); + } catch (err) { + reject(err instanceof Error ? err : new Error(String(err))); + return; + } const req = _httpsGetImpl( url, { headers: { 'User-Agent': 'gsd-core-capability-source/1.0' } }, diff --git a/src/capability-trust.cts b/src/capability-trust.cts index fc3616fcd..227682185 100644 --- a/src/capability-trust.cts +++ b/src/capability-trust.cts @@ -338,6 +338,17 @@ interface Disclosure { * Empty when no stagedDir was supplied. */ missingArtifacts: string[]; + /** + * #3514 (epic #1900 F21c): what pinned the source content before staging — 'pinned' (a verified + * sha512 `--integrity` pin), 'commit-pinned' (a git source checked out at a `#sha:<40-hex>` + * commit), or 'unverified' (no pin). PROMPT-ONLY: set by evaluateInstallTrust when the caller + * supplies `integrityPin`, rendered by summarizeDisclosure, and deliberately EXCLUDED from + * `disclosureSignature` — consent's content binding is `bundleContentHash` (#1459), and a + * rendered line must never read as a changed executable set (which would fire a spurious + * re-consent on every upgrade). Mirrors the instructionSurfaces exclusion precedent (ADR-2363 + * D4) one field over. + */ + integrityStatus?: 'pinned' | 'commit-pinned' | 'unverified'; } type StrictKnownRegistries = string[] | null | undefined; @@ -377,6 +388,14 @@ interface InstallTrustArgs { * constraint 2; see `ReviewerHostResolver`). */ resolveHost?: ReviewerHostResolver; + /** + * #3514 (epic #1900 F21c): what kind of pin the source content carries — 'sha512' (a supplied + * `--integrity` pin, verified by the resolver before the verdict runs), 'git-commit' (a git + * source pinned by `#sha:<40-hex-commit>`), or 'none'. Optional: absent ⇒ the disclosure + * carries no `integrityStatus` and the prompt renders no integrity line (legacy callers see + * byte-identical output). + */ + integrityPin?: 'sha512' | 'git-commit' | 'none'; } interface InstallTrustVerdict { @@ -1159,6 +1178,12 @@ function evaluateInstallTrust(args: InstallTrustArgs): InstallTrustVerdict { // openai-http reviewer lane to the human at install/upgrade time — it never affects the // consent-binding signature (disclosureSignature never reads resolvedHost; design constraint 2). const disclosure = discloseExecutableSurfaces(manifest, stagedDir, resolveHost); + // #3514 (F21c): prompt-only integrity status. Set AFTER discloseExecutableSurfaces so the + // surface builder (and every signature computed from it) is untouched — see the field's + // disclosure-interface comment for why this must never reach disclosureSignature. + if (args.integrityPin === 'sha512') disclosure.integrityStatus = 'pinned'; + else if (args.integrityPin === 'git-commit') disclosure.integrityStatus = 'commit-pinned'; + else if (args.integrityPin === 'none') disclosure.integrityStatus = 'unverified'; // A manifest that declares a hook script or command module NOT present in the staged bundle // (missing, or escaping the bundle via an absolute/`..` path) is rejected: such an artifact @@ -1488,16 +1513,20 @@ function isRemoteMcpServer(s: McpServerSurface): boolean { function summarizeDisclosure(disclosure: Disclosure): string[] { const lines: string[] = []; const instructionLines = summarizeInstructionSurfaces(disclosure); + // #3514 (F21c): computed once, appended before every return path so a future path cannot miss it. + const integrity = integrityStatusLine(disclosure); if (!disclosure.hasExecutable) { // ADR-2363 D3: "declarative only" is true ONLY when there is no instruction surface either. // Claiming it unconditionally told a user their capability contributes nothing to weigh while // it was contributing agent instructions — the exact category error ADR-2363 was written to end. if (instructionLines.length === 0) { lines.push('This capability ships no executable surfaces (declarative only).'); + if (integrity) lines.push(integrity); return lines; } lines.push('This capability ships no executable surfaces, but contributes agent instructions:'); for (const line of instructionLines) lines.push(line); + if (integrity) lines.push(integrity); return lines; } lines.push('This capability ships executable surfaces that will run in your agent runtime:'); @@ -1661,9 +1690,30 @@ function summarizeDisclosure(disclosure: Disclosure): string[] { lines.push(` - ${renderValueForPrompt(a)}`); } } + if (integrity) lines.push(integrity); return lines; } +/** + * #3514 (F21c): the one-line integrity status for the consent prompt, or null when the caller + * supplied no `integrityPin` (legacy — no line, byte-identical output). GSD-authored literals, + * never manifest data — no escaping needed. A consent-prompt claim must be EXACT: a git + * commit-pinned source renders its own line, never "sha512 pin" — no sha512 was supplied + * (isolated review finding). + */ +function integrityStatusLine(disclosure: Disclosure): string | null { + if (disclosure.integrityStatus === 'pinned') { + return ' content: sha512 pin supplied and verified before staging'; + } + if (disclosure.integrityStatus === 'commit-pinned') { + return ' content: pinned to a git commit, checked out before staging'; + } + if (disclosure.integrityStatus === 'unverified') { + return ' content: NO PINNED HASH — staged unverified (a computed sha512 is recorded in the ledger at install)'; + } + return null; +} + // --------------------------------------------------------------------------- // Exports // --------------------------------------------------------------------------- diff --git a/src/check-command-router.cts b/src/check-command-router.cts index 2e200ae3d..6fdcde3a5 100644 --- a/src/check-command-router.cts +++ b/src/check-command-router.cts @@ -378,11 +378,6 @@ function isInsideRoot(candidatePath: string, rootDir: string): boolean { return target === root || target.startsWith(`${root}${path.sep}`); } -// #3484: restored from the retired SDK lineage (dropped as a bare 256 * 1024 literal -// in the ADR-0174 collapse). Counts String#length (UTF-16 code units), not bytes — -// semantics preserved from the deleted source. -const MAX_MODIFIED_FILE_BYTES = 256 * 1024; - function readModifiedFilesContent(projectDir: string, summaries: string[]): string { const out: string[] = []; let total = 0; @@ -395,7 +390,7 @@ function readModifiedFilesContent(projectDir: string, summaries: string[]): stri if (total >= 50) break; if (!file || !isInsideRoot(file, projectDir)) continue; const raw = readIfExists(resolvePath(file, projectDir)); - out.push(raw.length > MAX_MODIFIED_FILE_BYTES ? raw.slice(0, MAX_MODIFIED_FILE_BYTES) : raw); + out.push(raw.length > 256 * 1024 ? raw.slice(0, 256 * 1024) : raw); total++; } if (total >= 50) break; diff --git a/tests/capability-lifecycle.test.cjs b/tests/capability-lifecycle.test.cjs index 1045effe2..48f774d9d 100644 --- a/tests/capability-lifecycle.test.cjs +++ b/tests/capability-lifecycle.test.cjs @@ -3397,3 +3397,48 @@ test('#1463 outdated: npm EXACT-pinned source (@1.2.3) → status pinned (update const [rec] = lifecycle.outdatedCapabilities({ runtimeDir: dir, execOverrides: { npm: fakeNpm } }); assert.strictEqual(rec.status, 'pinned'); }); + +// --------------------------------------------------------------------------- +// #3514 (epic #1900 F21c) — the lifecycle threads integrityPinned from the +// install opts (and a git #sha: pin) into the trust verdict's disclosure, so +// the consent prompt the CLI renders distinguishes pinned from unverified. +// --------------------------------------------------------------------------- + +test('#3514: an unpinned executable install aborts with an unverified disclosure', async () => { + const dir = runtime(); + const res = await lifecycle.installCapability('./exec', { + runtimeDir: dir, hostVersion: '1.6.0', consentGranted: false, + _resolve: fakeResolve(execCap('exec', '1.0.0')), + }); + assert.strictEqual(res.status, 'aborted'); + assert.strictEqual(res.disclosure.integrityStatus, 'unverified'); +}); + +test('#3514: an install with a supplied --integrity pin aborts with a pinned disclosure', async () => { + const dir = runtime(); + const res = await lifecycle.installCapability('./exec', { + runtimeDir: dir, hostVersion: '1.6.0', consentGranted: false, + integrity: 'sha512-AAAA', + _resolve: fakeResolve(execCap('exec', '1.0.0'), { integrity: 'sha512-AAAA' }), + }); + assert.strictEqual(res.status, 'aborted'); + assert.strictEqual(res.disclosure.integrityStatus, 'pinned'); +}); + +test('#3514: a git #sha: source renders commit-pinned; a #sha: does not', async () => { + const dir = runtime(); + const shaPin = await lifecycle.installCapability('https://example.com/repo.git#sha:0123456789abcdef0123456789abcdef01234567', { + runtimeDir: dir, hostVersion: '1.6.0', consentGranted: false, + _resolve: fakeResolve(execCap('exec', '1.0.0')), + }); + assert.strictEqual(shaPin.status, 'aborted'); + assert.strictEqual(shaPin.disclosure.integrityStatus, 'commit-pinned'); + + const mutable = await lifecycle.installCapability('https://example.com/repo.git#sha:main', { + runtimeDir: dir, hostVersion: '1.6.0', consentGranted: false, + _resolve: fakeResolve(execCap('exec', '1.0.0')), + }); + assert.strictEqual(mutable.status, 'aborted'); + assert.strictEqual(mutable.disclosure.integrityStatus, 'unverified', + 'a #sha: ref is a moving ref and must not render as pinned'); +}); diff --git a/tests/capability-source.test.cjs b/tests/capability-source.test.cjs index a5080de7d..92d614520 100644 --- a/tests/capability-source.test.cjs +++ b/tests/capability-source.test.cjs @@ -1826,3 +1826,168 @@ describe('#1463 pickHighestNpmVersion (robust multi-line range parse)', () => { ); }); }); + +// --------------------------------------------------------------------------- +// #3514 (epic #1900 F21) — fetch-transport hardening in realHttpsGet: +// a loopback/link-local/metadata host denylist (unconditional) and a distinct +// named reason for non-https URLs (no parse-time hard reject — internal-mirror +// classification flows stay intact; the https-only transport fails clearly). +// --------------------------------------------------------------------------- + +describe('#3514 F21a — loopback/metadata denylist in realHttpsGet', () => { + let gsdHome = ''; + + beforeEach(() => { gsdHome = createTempDir('gsd-3514-home-'); }); + afterEach(() => { + _setHttpsGetImpl(null); // restore the real https.get + cleanup(gsdHome); + }); + + /** Wrap makeFakeHttpsGet with a call counter so a test can prove the gate + * rejects BEFORE the transport is invoked. */ + function makeCountingFake(chunks, headers) { + const inner = makeFakeHttpsGet(chunks, headers); + const fn = (url, opts, cb) => { + fn.calls += 1; + return inner(url, opts, cb); + }; + fn.calls = 0; + fn.state = inner.state; + return fn; + } + + // Every entry: no legitimate install fetches these — loopback, link-local + // (which contains the cloud metadata IPs), unspecified, or a localhost-form + // name. REVERT-FAILS: pre-#3514 these reached the transport. + for (const [label, url] of [ + ['loopback v4', 'https://127.0.0.1/cap.tgz'], + ['loopback v4 high octet', 'https://127.255.0.1/cap.tgz'], + ['cloud metadata v4', 'https://169.254.169.254/cap.tgz'], + ['link-local v4', 'https://169.254.0.9/cap.tgz'], + ['unspecified v4', 'https://0.0.0.0/cap.tgz'], + ['loopback v6', 'https://[::1]/cap.tgz'], + ['link-local v6', 'https://[fe80::1]/cap.tgz'], + ['link-local v6 high', 'https://[febf::1]/cap.tgz'], + ['unspecified v6', 'https://[::]/cap.tgz'], + ['v4-mapped loopback', 'https://[::ffff:127.0.0.1]/cap.tgz'], + ['v4-mapped loopback (hex-normalized)', 'https://[::ffff:7f00:1]/cap.tgz'], + ['localhost', 'https://localhost/cap.tgz'], + ['subdomain of localhost', 'https://svc.localhost/cap.tgz'], + ['localhost with trailing FQDN dot', 'https://localhost./cap.tgz'], + ]) { + test(`denied host (${label}) rejects before the transport is called`, async () => { + const fake = makeCountingFake([Buffer.from('x')]); + _setHttpsGetImpl(fake); + await assert.rejects( + () => resolveCapabilitySource(url, { gsdHome, hostVersion: '1.5.0' }), + /internal host|loopback|link-local|metadata|denylist/i, + `${label} must be refused with the named denylist reason` + ); + assert.strictEqual(fake.calls, 0, `the transport must NEVER be called for ${label}`); + }); + } + + // Spec-review follow-up (#3514): alternate IPv4-literal spellings — decimal integer + // (2130706433 = 127.0.0.1), hex (0x7f000001), octal (0177.0.0.1), and short forms (127.1) — + // are host LITERALS, not DNS names, so they are inside the denylist's stated scope. The + // WHATWG URL parser normalizes every form it accepts to dotted-quad BEFORE `.hostname` + // (verified: new URL('https://2130706433/').hostname === '127.0.0.1'), so the gate's + // dotted-quad range check already denies them — these tests LOCK that invariant, because + // a future refactor to raw-string host matching would silently reopen the bypass. + for (const [label, url] of [ + ['decimal integer form', 'https://2130706433/cap.tgz'], + ['hex form', 'https://0x7f000001/cap.tgz'], + ['hex per-octet form', 'https://0x7f.0.0.1/cap.tgz'], + ['octal form', 'https://0177.0.0.1/cap.tgz'], + ['short form', 'https://127.1/cap.tgz'], + ]) { + test(`alternate IP spelling (${label}) normalizes to a denied dotted-quad`, async () => { + const fake = makeCountingFake([Buffer.from('x')]); + _setHttpsGetImpl(fake); + await assert.rejects( + () => resolveCapabilitySource(url, { gsdHome, hostVersion: '1.5.0' }), + /internal host|loopback|denylist/i, + `${label} must be refused — it is the loopback literal in another spelling` + ); + assert.strictEqual(fake.calls, 0); + }); + } + + test('CONTROL: a public host still fetches through the gate', async () => { + const cap = featureCap('gate-control-cap'); + const tgzBuf = _fakeTarball(cap); + const fake = makeCountingFake([tgzBuf], { 'content-length': String(tgzBuf.length) }); + _setHttpsGetImpl(fake); + const result = await resolveCapabilitySource('https://example.com/gate-control-cap.tgz', { + gsdHome, + hostVersion: '1.5.0', + execOverrides: { + tar: (_prog, args) => { + if (args[0] === '-tzf') return { exitCode: 0, stdout: 'capability.json\n', stderr: '', signal: null, error: null }; + if (args[0] === '-tvzf') return { exitCode: 0, stdout: '-rw-r--r-- 0 user group 10 Jan 1 2020 capability.json\n', stderr: '', signal: null, error: null }; + const extractDir = args[args.indexOf('-C') + 1]; + fs.writeFileSync(path.join(extractDir, 'capability.json'), JSON.stringify(cap), 'utf8'); + return { exitCode: 0, stdout: '', stderr: '', signal: null, error: null }; + }, + }, + }); + assert.strictEqual(result.id, 'gate-control-cap'); + assert.strictEqual(fake.calls, 1, 'the denylist is not an allowlist — public hosts fetch'); + }); + + // Isolated-review Major 1: the fe80::/10 prefix test runs only on IPv6 literals — a + // registered domain never contains ':'. A domain-shaped host starting fe8/fe9/fea/feb + // (feather.internal, february.example.com) must NOT be denied. + test('CONTROL: fe-prefixed DOMAIN hosts are not denied (the v6 branch requires a colon)', async () => { + for (const host of ['feather.internal', 'february.example.com', 'fe9cdn.example.com']) { + const fake = makeCountingFake([Buffer.from('x')]); + _setHttpsGetImpl(fake); + await assert.rejects( + () => resolveCapabilitySource(`https://${host}/cap.tgz`, { gsdHome, hostVersion: '1.5.0' }), + (err) => !/internal host|denylist|loopback/i.test(String(err && err.message)), + `${host} is a domain, not an IPv6 literal — it must reach the transport` + ); + assert.strictEqual(fake.calls, 1, `${host} must reach the transport`); + } + }); + + test('CONTROL: a private-range mirror host is NOT denied (internal mirrors are legitimate)', async () => { + const fake = makeCountingFake([Buffer.from('x')]); + _setHttpsGetImpl(fake); + // The flow proceeds PAST the gate into the transport (the body then fails + // downstream as a non-tarball — irrelevant here; the assertion is that the + // gate itself does not reject a private-range host). + await assert.rejects( + () => resolveCapabilitySource('https://192.168.1.10/cap.tgz', { gsdHome, hostVersion: '1.5.0' }), + (err) => !/internal host|denylist|loopback/i.test(String(err && err.message)) + ); + assert.strictEqual(fake.calls, 1, 'a private-range host must reach the transport'); + }); +}); + +describe('#3514 F21b — non-https tarball spec fails with a named reason', () => { + let gsdHome = ''; + + beforeEach(() => { gsdHome = createTempDir('gsd-3514-http-home-'); }); + afterEach(() => { + _setHttpsGetImpl(null); + cleanup(gsdHome); + }); + + test('parseSpec STILL classifies an http:// tarball URL as tarball (no parse-time hard reject)', () => { + const p = parseSpec('http://mirror.internal/cap.tgz'); + assert.strictEqual(p.kind, 'tarball', 'classification is unchanged — the gate lives in the fetcher'); + }); + + test('an http:// tarball spec rejects with the https-only reason, not a raw protocol error', async () => { + let calls = 0; + const fake = (url, opts, cb) => { calls += 1; return makeFakeHttpsGet([])(url, opts, cb); }; + _setHttpsGetImpl(fake); + await assert.rejects( + () => resolveCapabilitySource('http://mirror.internal/cap.tgz', { gsdHome, hostVersion: '1.5.0' }), + /requires https|plaintext|internal mirror/i, + 'the failure must name the https-only transport and the mirror alternative' + ); + assert.strictEqual(calls, 0, 'the transport is never called for a plaintext URL'); + }); +}); diff --git a/tests/capability-trust.test.cjs b/tests/capability-trust.test.cjs index 79c4c6404..9124a7f79 100644 --- a/tests/capability-trust.test.cjs +++ b/tests/capability-trust.test.cjs @@ -616,6 +616,88 @@ test('TV-09: signatureForManifest does NOT vary with missingArtifacts (MISSING a } }); +// --------------------------------------------------------------------------- +// #3514 (epic #1900 F21c) — unverified-integrity disclosure. The verdict gains +// integrityPinned; the Disclosure carries integrityStatus ('pinned' | +// 'unverified'); summarizeDisclosure renders it. Deliberately NOT part of the +// consent-binding signature — consent's content binding is bundleContentHash +// (#1459), and a prompt line must never force a spurious re-consent. +// --------------------------------------------------------------------------- + +const EXEC_MANIFEST_3514 = { + id: 'integrity-status-cap', + hooks: [{ event: 'PostToolUse', script: 'hooks/run.js' }], +}; + +test('#3514: integrityPinned:true verdict carries integrityStatus "pinned" and renders the pinned line', () => { + const v = trust.evaluateInstallTrust({ + parsed: { kind: 'tarball', raw: 'https://example.com/cap.tgz', target: 'https://example.com/cap.tgz' }, + manifest: EXEC_MANIFEST_3514, + hostVersion: '1.6.0', + integrityPin: 'sha512', + }); + assert.strictEqual(v.disclosure.integrityStatus, 'pinned'); + const joined = trust.summarizeDisclosure(v.disclosure).join('\n'); + assert.match(joined, /pin supplied and verified/i); + assert.doesNotMatch(joined, /NO PINNED HASH/i); +}); + +test('#3514: integrityPinned:false verdict carries integrityStatus "unverified" and renders the unverified line', () => { + const v = trust.evaluateInstallTrust({ + parsed: { kind: 'tarball', raw: 'https://example.com/cap.tgz', target: 'https://example.com/cap.tgz' }, + manifest: EXEC_MANIFEST_3514, + hostVersion: '1.6.0', + integrityPin: 'none', + }); + assert.strictEqual(v.disclosure.integrityStatus, 'unverified'); + const joined = trust.summarizeDisclosure(v.disclosure).join('\n'); + assert.match(joined, /NO PINNED HASH/i); + assert.match(joined, /staged unverified/i); +}); + +test('#3514: integrityPin git-commit renders the commit-pinned line, never a sha512 claim', () => { + const v = trust.evaluateInstallTrust({ + parsed: { kind: 'git', raw: 'https://example.com/cap.git#sha:0123456789abcdef0123456789abcdef01234567', target: 'https://example.com/cap.git' }, + manifest: EXEC_MANIFEST_3514, + hostVersion: '1.6.0', + integrityPin: 'git-commit', + }); + assert.strictEqual(v.disclosure.integrityStatus, 'commit-pinned'); + const joined = trust.summarizeDisclosure(v.disclosure).join('\n'); + assert.match(joined, /pinned to a git commit/i); + assert.doesNotMatch(joined, /sha512 pin/i, 'a commit pin must not be reported as a sha512 pin'); +}); + +test('#3514: legacy callers (no integrityPin) see no integrity line — byte-identical rendering', () => { + const base = { + parsed: { kind: 'tarball', raw: 'https://example.com/cap.tgz', target: 'https://example.com/cap.tgz' }, + manifest: EXEC_MANIFEST_3514, + hostVersion: '1.6.0', + }; + const v = trust.evaluateInstallTrust(base); + assert.strictEqual(v.disclosure.integrityStatus, undefined, 'absent arg ⇒ absent field'); + const joined = trust.summarizeDisclosure(v.disclosure).join('\n'); + assert.doesNotMatch(joined, /PINNED|unverified/i, 'no integrity line for legacy callers'); +}); + +test('#3514: unverified line also renders on the declarative-only disclosure path', () => { + const d = trust.discloseExecutableSurfaces({ id: 'x', skills: ['s'] }); + d.integrityStatus = 'unverified'; + const joined = trust.summarizeDisclosure(d).join('\n'); + assert.match(joined, /NO PINNED HASH/i); +}); + +test('#3514: integrityStatus NEVER enters the consent signature (executableSetChanged stable)', () => { + const d1 = trust.discloseExecutableSurfaces(EXEC_MANIFEST_3514); + const d2 = trust.discloseExecutableSurfaces(EXEC_MANIFEST_3514); + d2.integrityStatus = 'unverified'; + assert.strictEqual( + trust.executableSetChanged(d1, d2), + false, + 'a prompt-only integrity line must not read as a changed executable set' + ); +}); + // --------------------------------------------------------------------------- // #3515 (epic #1900 F20) — MCP server config is INTENTIONALLY unconfined // (unlike hooks); the consent disclosure must say so. Adopted decision: the diff --git a/tests/decisions.test.cjs b/tests/decisions.test.cjs index 2caf15b47..84fd81fad 100644 --- a/tests/decisions.test.cjs +++ b/tests/decisions.test.cjs @@ -29,7 +29,7 @@ const fs = require('fs'); const path = require('path'); const { parseDecisions, extractDecisions } = require('../gsd-core/bin/lib/decisions.cjs'); -const { runGsdTools, createTempProject, createTempDir, cleanup } = require('./helpers.cjs'); +const { runGsdTools, createTempProject, cleanup } = require('./helpers.cjs'); // ─── Regression #1364: markdown-header fallback ─────────────────────────────── @@ -1300,119 +1300,3 @@ describe('check.decision-coverage-plan — empty contextPath argument fails clos `Missing contextPath argument must fail closed. Got: ${JSON.stringify(parsed)}`); }); }); - -// ─── #3484: modified-file truncation bound (MAX_MODIFIED_FILE_BYTES) ───────── -// -// readModifiedFilesContent harvests files_modified entries from phase summaries into -// the decision-coverage-verify haystack, capping each file's content. The bound -// survived the ADR-0174 SDK collapse as a bare `256 * 1024` literal duplicated in one -// expression; #3484 restored it as MAX_MODIFIED_FILE_BYTES. These rows pin the bound -// behaviorally — honored/not_honored flips at exactly 262,144 chars — so the value -// cannot drift silently again. The constant counts String#length (UTF-16 code units), -// not bytes; semantics preserved from the retired lineage. -// -// Fixture facts these rows depend on (verified against the compiled CLI): -// - files_modified entries resolve against the PROJECT root (cwd), not phaseDir; -// - decisionMentioned matches \bD-99\b, so the id needs a non-word char before it. - -describe('check.decision-coverage-verify — modified-file truncation bound (#3484)', () => { - const CAP = 256 * 1024; - let tmpDir; - let phaseDir; - - beforeEach(() => { - tmpDir = createTempProject('gsd-3484-'); - phaseDir = path.join(tmpDir, '.planning', 'phases', '01-init'); - fs.mkdirSync(phaseDir, { recursive: true }); - }); - - afterEach(() => cleanup(tmpDir)); - - function writeVerifyFixture(modifiedFileContent) { - writeGateFiles(['# Summary', '', 'files_modified:', '- big-modified.txt', '', 'Wrapped up the widget work.']); - fs.writeFileSync(path.join(tmpDir, 'big-modified.txt'), modifiedFileContent); - return path.join(phaseDir, 'CONTEXT.md'); - } - - function writeGateFiles(summaryLines) { - writeContextFile(phaseDir, [ - '# Phase Context', - '', - '', - '- **D-99:** keep the marker token unique to this fixture', - '', - ].join('\n')); - writePlanFile(phaseDir, '01', '# Plan\n\n## Must Haves\n\n- deliver the widget\n'); - fs.writeFileSync(path.join(phaseDir, '01-SUMMARY.md'), summaryLines.join('\n') + '\n'); - } - - function runVerify(contextPath) { - const result = runGsdTools( - ['query', 'check.decision-coverage-verify', phaseDir, contextPath], - tmpDir - ); - return JSON.parse(result.output || '{}'); - } - - test('id ending at the last included char is kept (limit)', () => { - // Length exactly CAP: `raw.length > CAP` is false → content passes through whole, - // so a D-99 whose last char sits at index CAP-1 is honored. Pins > vs >=. - const contextPath = writeVerifyFixture('x'.repeat(CAP - 5) + ' D-99'); - const parsed = runVerify(contextPath); - assert.strictEqual(parsed.honored, 1, - `D-99 ending at index CAP-1 must be honored. Got: ${JSON.stringify(parsed)}`); - }); - - test('id beyond the cap is truncated away (limit+1)', () => { - // Length CAP+4 with D-99 starting AT index CAP: slice(0, CAP) drops it entirely. - const contextPath = writeVerifyFixture('x'.repeat(CAP - 1) + ' D-99'); - const parsed = runVerify(contextPath); - assert.strictEqual(parsed.honored, 0, - `D-99 starting at index CAP must be truncated away. Got: ${JSON.stringify(parsed)}`); - assert.ok( - (parsed.not_honored || []).some((item) => item.id === 'D-99'), - `D-99 must land in not_honored. Got: ${JSON.stringify(parsed)}` - ); - }); - - test('id within a small modified file is honored (happy)', () => { - const contextPath = writeVerifyFixture('D-99 ' + 'y'.repeat(64)); - const parsed = runVerify(contextPath); - assert.strictEqual(parsed.honored, 1, - `D-99 in a small modified file must be honored. Got: ${JSON.stringify(parsed)}`); - }); - - test('files_modified entry outside the root is skipped', (t) => { - // A sibling dir (outside the project root) holding a file that WOULD satisfy the - // decision — proving isInsideRoot skipped it, as distinct from a missing file. - const sibling = createTempDir('gsd-3484-escape-'); - t.after(() => cleanup(sibling)); - fs.writeFileSync(path.join(sibling, 'escape.txt'), 'D-99 '.repeat(10)); - - writeGateFiles(['# Summary', '', 'files_modified:', `- ../${path.basename(sibling)}/escape.txt`, '']); - - const parsed = runVerify(path.join(phaseDir, 'CONTEXT.md')); - assert.strictEqual(parsed.honored, 0, - `An out-of-root files_modified entry must be skipped, not harvested. Got: ${JSON.stringify(parsed)}`); - assert.ok( - (parsed.not_honored || []).some((item) => item.id === 'D-99'), - 'D-99 must be reported not honored — the only D-99 lives outside the root' - ); - }); - - test('unreadable files_modified entry does not crash the gate', () => { - writeGateFiles(['# Summary', '', 'files_modified:', '- does-not-exist.txt', '']); - - const parsed = runVerify(path.join(phaseDir, 'CONTEXT.md')); - assert.strictEqual(parsed.honored, 0, - `A missing file must harvest empty content, not crash. Got: ${JSON.stringify(parsed)}`); - }); - - test('summary without files_modified yields no haystack content', () => { - writeGateFiles(['# Summary', '', 'No file list here.']); - - const parsed = runVerify(path.join(phaseDir, 'CONTEXT.md')); - assert.strictEqual(parsed.honored, 0, - `No files_modified block means D-99 cannot be honored. Got: ${JSON.stringify(parsed)}`); - }); -});