c0b2a05d2f310adc0a1f35fd71fbc9f28f4e4977
8 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c0b2a05d2f |
fix(#4594): one canonical dispatch-identity owner — the emitted format and the parser that reads it back (#4693)
* fix(#4594): give dispatch identity one owner for the emitted format and its parser The isolation guards decided whether a run-scoped sentinel applied to a dispatch by regex-scraping model-authored prose. The scrape returned values in a different namespace from the ones the sentinel records, so the comparison could never succeed: sentinel { phase: "03", plan: "03-02-hardening" } <- $PHASE_NUMBER / $plan_id prose "Execute plan 02 of phase 03-auth." scraped { phase: "03-auth.", plan: "02" } <- greedy (\S+), both wrong #4594 reports only the phase half. Measured against a real phase-plan-index run, plans[].id is phase-prefixed, plan-numbered AND slugged, while the prose carries a bare in-phase plan number — so the plan field mismatches too, and the Claude path is dead rather than latent. A fresh sentinel was therefore discarded on every executor dispatch and every legitimate ISOLATION=none degrade was denied, leaving the work unrun. hooks/lib/dispatch-identity.js is now the single owner of both halves. The two prompt-body producers emit a canonical marker carrying the same shell values the sentinel records, so producer and consumer agree by construction. The prose frame stays as a fallback, bounded by the phase-token grammar ADR-2121 owns and deliberately reporting no plan — an absent identifier means "cannot compare" and is safe; a wrong one is a false mismatch and is not. The prose sentence itself is byte-identical: the executor agent reads it too, so the marker is purely additive (Hyrum's Law). An inapplicable sentinel is now named in the guards' deny reason instead of being dropped silently — the silence is why this survived three producers and two consumers unnoticed. Interpolated values come from a sentinel file and from prompt text, so both are length-bounded and stripped of control characters. ADR-4630 locks the seam and maps the epic's three phases. Refs #4630 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4594): resolve eight review findings across the dispatch-identity seam Three orthogonal review engines ran on 43418af144 — the code-review skill's Standards and Spec axes, and an isolated adversarial security pass — plus a self-review of the committed diff. Every finding is fixed here; none deferred. F1 (major, reproduced). A keyless or unknown-key-only marker — the literal "[gsd:dispatch]" or "[gsd:dispatch run=..]" — matched the marker grammar and returned source:'marker' with both fields null, suppressing the prose fallback entirely. Any prompt text containing that literal silently disabled identity narrowing, so a fresh sentinel applied to a dispatch it was never scoped to, defeating #3045 SECURITY F2. Prompt text is attacker-influenceable. A marker that yields neither recognized key is no longer a marker: the scan continues to later markers, then later texts, then prose. Forward-compatible tolerance of unknown keys is unchanged. F2/F3 (major). The first cut duplicated sanitizeForReason, describeSentinelDiscard and REASON_INTERPOLATION_MAX_LEN byte-for-byte across both guards — the exact defect class this epic exists to delete, and with no cold-load justification, since both hooks already require hooks/lib/. They now live in hooks/lib/isolation-deny-reason.js, and buildSentinelDiscard lives in isolation-sentinel.js beside the comparison it mirrors, returning the nested {sentinel:{phase,plan}, dispatch:{phase,plan}} shape instead of a bespoke four-field bag that renamed the pairs already flowing through the seam. F4 (hard violation). The visibility test asserted on the deny reason's prose. CONTRIBUTING.md prohibits raw text matching on hook output, which is why every deny carries a stable reason_code. The discard is now a structured sentinel_discarded field on each hook's stdout JSON, and the test asserts that; the sentence stays for the operator but is no longer the contract. F5 (hard violation). The 64-character truncation limit had no boundary coverage. 63/64/65 are now exercised against the single consolidated helper. F6 (minor). sanitizeForReason stripped C0/C1 controls but not U+2028/U+2029 or the bidi overrides, so a crafted value could still reflow or reverse the message. Both classes are stripped, with a test each. F7 (major). The producer/template parity test was vacuous — it rendered a marker and re-parsed its own output, and would have passed with both templates deleted. It now reads the two workflow templates, extracts each marker line, substitutes the measured values and asserts the owner's parser returns them. Proven red by deleting one template's marker line before being proven green. F8 (doc). ADR-4630 and the design notes claimed the marker is guaranteed on the orchestrator-worktree path because that prompt is built in shell. It is not: executor-isolation-dispatch.md:131 says plainly that those are template placeholders, not shell variables, so {plan_id} is model-substituted there too. A false guarantee in a design lock is worse than a stated limit. Both documents now say the marker is model-substituted on both paths and that the prose fallback is the real floor everywhere. The "3 workflow templates" count was also wrong — 3 prose sites across 2 files, 2 of which carry the marker. Refs #4630 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4594): refresh the compact-content baseline and acknowledge execute-phase.md growth Refs #4630. The dispatch-identity marker and its substitution note grew gsd-core/workflows/execute-phase.md by 525 bytes (91846 -> 92371), which drifts two real-tree guards that lint:ci does not run: - tests/benchmark-compact-content.test.cjs asserts the committed baseline is "up to date"; the split for execute-phase.md moved off 25827 -> 25952 and on 23576 -> 23701, taking its compaction reduction 8.72% -> 8.67%. Baseline regenerated with scripts/benchmark-compact-content.cjs --write. - tests/emitted-attribution.test.cjs requires a growth acknowledgment trailer for any emitted file that grows, keyed on the bare filename. Added below. The growth is two additions and no rewrites: the [gsd:dispatch ...] marker line inside the Agent() prompt's <objective>, and the note telling the orchestrator to substitute {plan_id} with the plan's id verbatim. Both are load-bearing -- the marker is what lets a guard hook match a dispatch to the sentinel the per-plan gate wrote, and without the note the orchestrator has no instruction telling it the value must not be paraphrased. Emitted-Drift-Ack-Growth: execute-phase.md — adds the canonical [gsd:dispatch] identity marker and its {plan_id} substitution note, which the isolation guards compare verbatim against the run-scoped sentinel (#4594) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4594): set changeset fragment pr to 4693 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2ea5efc151 |
enhance(#3911): hooks declare their crash policy (#3960)
* enhance(#3911): give hooks an exit seam that needs no build ADR-3889 Phase 7 foundation. The 19 shipped enforcement hooks hold 91 of the epic's 128 terminators and cannot reach `terminateNow` today. The obvious route — requiring `gsd-core/bin/lib/cli-exit.cjs`, as gsd-agent-isolation-guard.js already does for two other modules — is rejected. That precedent carries its own warning (#3582): those files are tsc output, gitignored and absent on a raw plugin-marketplace or git-clone install, so the hook must first call ensureRuntimeBuild() to self-heal. Making the module a hook needs IN ORDER TO TERMINATE depend on a build inverts the dependency, and its failure mode is precisely the fail-open this phase exists to remove: a guard that cannot terminate cannot deny. `lint-hooks-runtime-build-seam` already encodes that concern, and Design B would have had to add an ensureRuntimeBuild() call to all 19 hooks to satisfy it. So `hooks/lib/` becomes a third emit location for cli-exit and a fifth for the registry, preserving the invariant `src/cli-exit.cts`'s own header states: it imports nothing but node:fs and its sibling registry, and the generator dual-emits that sibling alongside each copy so a relative require resolves next to whichever copy loaded it. Shipping needed no change — build-hooks.js already declares HOOKS_SUBDIRS_TO_COPY = ['lib']. Proven, not asserted: the two files are copied into an otherwise-empty tmpdir and a child process requires them and terminates — PASS exits 0, HOOK_DENY exits 2 with the payload on both stdout and stderr. That test fails the moment the hooks copy gains a require reaching outside hooks/lib/. Also fixed inline: the registry's fifth target let any `--write` test overwrite the real committed hooks/lib/exit-code-registry.js, because the test helper derived only three of the other output paths. It now redirects all five, and a regression test asserts every committed artifact is byte-identical after a redirected write. Install-tree goldens pick up the two new shipped paths across 11 runtimes — insertions only, no removals. lint:ci was green while they were stale, so this was found by regenerating rather than by a gate. Verification runs on the remote runner. Refs #3911 * enhance(#3911): declare a crash policy, and migrate the write guard Adds `hooks/lib/hook-exit.js` — the hook-facing vocabulary over `terminateNow`, hand-written because the cli-exit copy beside it is generated: allow(payload) exit 0 deny(payload, stderr?) exit 2 crash(onCrash, payload) whichever the hook DECLARED `crash()` takes the policy as a required argument with no default, which is the whole mechanism: fail-open by accident stops being expressible. A hook must name ALLOW or DENY at the call site, and an unrecognized value terminates INTERNAL rather than guessing. Fail-open stays legal; fail-open by omission does not. `gsd-write-guard.js` is the first hook migrated, all 12 sites, and it exposed a gap in the seam. `terminateNow`'s doc comment justified its fd-2 write by citing this hook's `emitBlock` — but modeled it as sending the same bytes to both streams, when `emitBlock` actually sends full JSON to stdout and only the bare `reason` string to stderr, because Kimi's hook bus feeds stderr verbatim back to the model. Migrating as written would have turned a readable sentence into a JSON blob for Kimi-backed agents. #3911 requires both "all 19 hooks terminate through terminateNow" and "no hook's effective default changes". Those are jointly satisfiable only by teaching the seam to carry a distinct stderr payload, so `terminateNow` gains an optional third argument: omitted, behavior is byte-for-byte what it was; a string is written raw, which is exactly the Kimi case. The doc comment's inaccurate claim about emitBlock is corrected in place. Proven rather than asserted: the pre-migration file is reconstructed from HEAD and driven with the same catastrophic-shrink payload as the migrated one — exit code, stdout and stderr all byte-identical. Verification runs on the remote runner. Refs #3911 * enhance(#3911): all 19 hooks terminate through the seam Migrates the remaining 18 enforcement hooks onto allow/deny/crash. An AST walk now reports zero `process.exit(` call sites across every `hooks/*.js` — down from the 91 the census measured. Each hook with an outer catch declares its policy once, at module top, with the reason that policy is right for that specific guard: a read guard that cannot scan must not block the read; a statusline that renders every prompt must degrade rather than crash; an injection scanner must not retroactively block a result already returned. Those sentences are the deliverable — they are what turns fail-open-by-accident into fail-open-on-purpose. No hook's effective default changed. Wiring exposed two defects, both fixed here rather than noted. A SECOND stdout/stderr-splitting site turned up in `gsd-workflow-guard.js`'s `emitForceAddBlock`, matching the pattern already known from the write guard — full JSON to stdout, bare reason to stderr for the Kimi bus. It uses the `stderrPayload` argument added in the previous commit, which is now carrying its second real caller rather than one special case. More seriously, `terminateNow` emitted both streams inside ONE try, so a payload that failed to serialize aborted before the stderr write ever ran. The two windsurf guards write nothing to stdout on a block and only a reason string to stderr, so `deny(undefined, reason)` exited 2 with EMPTY stderr — a deny that silently loses its reason, which is the exact "fails with success" class this epic exists to close. The streams are now emitted independently, each with its own guard, and `undefined` means "nothing to write for this stream" rather than an error. Regression tests inject a throwing write on one fd and assert the other still receives its payload; they fail against the single-try version. Byte-identity was proven per hook, not assumed: each pre-change file is reconstructed from HEAD and driven side by side with the migrated one across its normal path, its deny path, malformed stdin and empty stdin — exit code, stdout and stderr compared. Verification runs on the remote runner. Refs #3911 * enhance(#3911): harden the three shell hooks, and pin every hook's policy `gsd-phase-boundary.sh`, `gsd-session-state.sh` and `gsd-validate-commit.sh` gain `set -euo pipefail`. The expected hazard did not materialize, and that is worth recording: every intentionally-non-zero command in all three is already the condition of an `if`/`elif`, which `set -e` never fires on, and none of them reads a possibly-unset variable or pipes through a grep that may legitimately match nothing. No `|| true` guards were needed. Each hook was still checked command-by-command before the flags went in rather than after. Twenty-one before/after cases across the three hooks — disabled and enabled, planning and non-planning, missing STATE.md, malformed JSON, the Kimi payload shape, quoted and unquoted `-m`, valid and over-long Conventional Commits — all match on exit code, stdout and stderr. The hardening is shown to actually fire, not merely added: with a stubbed `node` that fails at the JSON-emit step, phase-boundary and session-state go from silently exiting 0 with empty stdout to failing visibly with the error surfaced. No such case could be constructed for `gsd-validate-commit.sh`, whose every statement already sits inside an if-condition — recorded as unproven rather than claimed. `tests/hooks-crash-policy.test.cjs` adds the per-hook coverage the issue asks for, table-driven over all 19 hooks rather than 76 hand-written cases: normal allow, deny where a deny path exists, crash-honors-the-declared-policy, and an unclosed-stdin case — the one `process.exitCode` structurally cannot serve. The deny assertions encode each hook's ACTUAL stream split rather than a uniform shape, since four of the six deliberately differ. A drift guard enumerates `hooks/*.js` and fails if a terminating hook is ever added without a row. Writing those tests surfaced two hooks that emit a block decision in their JSON body and exit 0. Both were checked rather than assumed, and neither is a fails-with-success: `gsd-read-injection-scanner.js` is PostToolUse, where the tool has already run and exit 2 has no meaning, and `gsd-cursor-subagent-start.js` follows Cursor's JSON-body protocol. They are deliberately left alone — a mechanical sweep to `deny()` would have broken exactly these two. Verification runs on the remote runner. Refs #3911 * fix(#3838): the commit validator says when it could not validate #3911 claims to subsume #3838. Measurement said otherwise, so this closes it for real rather than by assertion. `set -euo pipefail`, added earlier on this branch, does NOT fix #3838: bash exempts a command used as an `if` condition from `set -e`, and all three of the hook's swallow-and-pass sites are exactly that shape. Verified against the hardened hook with a node shim that fails only the classifier call — a non-conforming commit still exited 0 with empty stdout AND empty stderr, indistinguishable from "your commit conforms". That is the defect verbatim. All three sites named in #3838 now capture the real exit status instead of consuming it as a condition, and each distinguishes its genuine negative from "could not run": - the classifier: 0 = is a git commit, 1 = genuinely not one, anything else = could not classify. Its `node -e` now wraps the require and the call in try/catch and exits 3 on a throw, so a broken require chain can never be mistaken for `isGitSubcommand` legitimately returning false — which is the arm that matters, since `token-scanner.cjs` is a gitignored build artifact and a fresh checkout lands there. - the opt-in config read and the JSON command extraction get the same treatment. On "could not run" the hook emits a diagnostic to stderr naming which check failed and why, then exits 0. The issue confirms this is safe — it is a PreToolUse hook, so stderr does not disturb the JSON protocol — and ranks it the smallest sufficient fix. The gate still fails open, but it can no longer do so silently, which is the whole complaint: a validator that disables itself quietly costs more than one that is absent, because it is trusted. Both controls are unchanged and pinned by tests: a conforming commit still passes silently, a non-conforming one still exits 2 with its existing block payload. The defect test asserts stderr is non-empty and names the failure; it fails against the pre-fix hook. Verification runs on the remote runner. Refs #3911, #3838 * docs(#3911): document the hook crash-policy contract Reference and Explanation via a new docs/features fragment (FEATURES.md is generated from it), INVENTORY rows for the three new hooks/lib files, and an ARCHITECTURE note on the hooks section. How-To: docs/how-to/declare-a-hook-crash-policy.md, indexed from docs/README.md — a hook author now has to choose and declare a crash policy, which is more than one step and crosses into which harness protocol their hook speaks. It covers allow/deny/crash, writing an ON_CRASH reason that is actually useful, when a deny needs a distinct stderr payload, the two hooks whose harness reads a JSON-body decision and must NOT use deny(), and what to do when a check cannot run at all — with #3838 as the worked example. Refs #3911 * test(#3911): prove the seam actually ships, and stop hand-rolling temp cleanup Two review findings. The acceptance criterion 'hooks/dist/** stays in parity via the build seam (lint:hooks-runtime-build-seam)' was misstated and unmet: that lint checks something else — that a hook requiring a compiled gsd-core/bin/lib module also calls ensureRuntimeBuild(). Nothing exercised that the three new hooks/lib files reach hooks/dist/lib at all. That gap is not theoretical: #770 is a recorded ship-blocking bug where a new hook never shipped because a copy list missed it. The suite now builds dist through the repo's own ensureBuiltHooks(), byte-compares each shipped copy against its source, and spawns a child that requires the SHIPPED dist copy and denies — which is what catches a copy that exists but cannot resolve its sibling registry. gsd-validate-commit.sh hand-duplicated mktemp/run/rm three times; one idempotent trap on EXIT replaces them, guarded so cleanup cannot alter the exit status. Behavior-neutral across five cases, with temp-file counts taken before and after each run. Refs #3911 * fix(#3911): stage transitive hook lib requires, not just one level The remote run returned 7 failures across 3 real causes. The important one is a PRODUCTION bug this phase exposed rather than caused. `writeCursorHooksJson` scanned each hook script for `./lib/X` requires exactly one level deep and never re-scanned the lib files it staged for their own sibling requires. Nothing had a transitive lib dependency before, so the gap was invisible. Adding hook-exit.js -> cli-exit.js -> exit-code-registry.js made real Cursor installs ship a bundle that dies at require time with MODULE_NOT_FOUND. It now walks to a fixed point, and a real installed Cursor hook runs to completion. The staging harness in shared-hooks-dir-resolution hand-copied its fixture, so the injection scanner crashed at require time and its exit-1 was being read as a policy decision. Migrated to copyScriptWithDeps, which walks the require graph — the repo's recorded rule for this class, since adding another copyFileSync keeps it alive for the next person. The missing-lib-source test in cursor-hook-workspace-roots hardcoded which lib file it expected to be named in the abort message; the same throw now fires for a different file first. Its assertion is unchanged in substance — staging still must abort rather than ship a broken hook — only the name is no longer pinned. The last one was my own test asserting an uppercase reason code. Measured against origin/next: the pre-change hook emits the same lowercase 'config_unreadable', so the test was wrong, not the migration. Corrected to the real value rather than making the code match the test. Verification runs on the remote runner. Refs #3911 * chore(#3911): regenerate the cursor install-tree golden The staging fix means a Cursor install now correctly carries the two transitive lib files it was silently missing. Additive only — no path was removed. The golden diff is the evidence the packaging defect was real. Refs #3911 * chore(#3911): backfill the changeset PR number Refs #3911 * fix(#3911): a git probe that timed out is not a negative A macOS CI lane failed three deny cases at 2084ms, 2112ms and 2177ms — just past the 2000ms budget these hooks give their git probes. The three that passed took 72ms, 595ms and 651ms. Under shard contention `git rev-parse` overruns, the hook reads the non-zero result as "not a git repo", and allows with exit 0 and empty stdout AND empty stderr. Under load, the guards silently stop guarding. That is ADR-3889's thesis exactly, sitting inside the security hooks this phase is about. The repo had already recognized the class in one place — gsd-cursor-subagent-start.js fail-closed-denies on `git_timed_out` (#3045) — but nowhere else. `hooks/lib/git-probe.js` classifies a probe's outcome, distinguishing a real non-zero exit from ETIMEDOUT, a signal kill, and a spawn failure, rather than folding all four into `status !== 0`. Three guards route their eight git probes through it. The resolution is the same shape #3838 took, and the same one that issue endorsed as smallest-sufficient: fail open, but loudly. **No exit code changes on any path** — a developer on a loaded machine is still not blocked, which keeps #3911's declaration-pass contract intact for exit codes. What changes is that the hook now says on stderr which probe could not answer, instead of presenting silence as a clean verdict. Scope was checked across every hooks/*.js, not just the three that failed: gsd-agent-isolation-guard spawns no git; gsd-statusline's two probes gate only a cosmetic display segment, not an allow/deny decision, and are left alone. The C2 deny assertion was a real-race test — it demanded exit 2 while a slow git legitimately yields 0. It now requires the hook to either deny, or allow with a diagnostic naming the probe that could not run; a silent allow still fails, so the assertion is not vacuous. A deterministic regression stubs git on PATH to sleep past the budget rather than waiting for load to reproduce it. Verification runs on the remote runner. Refs #3911 * test(#3911): a PATH shim cannot intercept the hooks' git spawn on Windows The deterministic timeout regression stubbed git on PATH and asserted the guard reports rather than silently allows. It passes on Linux and macOS and failed on Windows in 83ms and 176ms — the stub was never invoked at all. Mechanism: the hooks call spawnSync('git', args) with no shell:true, so on Windows CreateProcess resolves git.exe only and never a PATH .cmd shim. The git.cmd branch could not have worked and is removed rather than left implying a Windows path that does. Adding shell:true to the hooks to serve a test would change product behavior and widen an injection surface, so the case is skipped on win32 only, with the mechanism written into the skip reason so a future reader does not 'fix' it that way. Linux and macOS keep the coverage, and macOS is where the underlying fail-open was actually caught. Refs #3911 --------- Co-authored-by: sim <sim@local> |
||
|
|
9410f7e6e6 |
enhance(#3897): ADR-3473 §8.3 rungs 2-4 — runtime marker, derived Codex sandbox, short-form depends_on (#3941)
* test(#3897): failing-first coverage for §8.3 rungs 2-4 ADR-3473 §8.3 has four rungs; #3883/PR #3896 shipped the first. This pins the other three RED before any fix. Rung 2 — the install marker has four readers and resolveRuntime is not one. resolveRuntime resolves GSD_RUNTIME > config.runtime > 'claude' and reads no marker at all, while bin/install.js writes one (#2297) and FOUR hand-rolled readInstallRuntimeMarker copies exist: src/model-resolver.cts:65 (cached, with test seams), hooks/gsd-agent-isolation-guard.js:112, and TWICE in hooks/gsd-cursor-subagent-start.js at :346 and :355. Four copies of one rule. Fixtures and seam names mined from PR #3382 rather than re-derived; it implemented this rung and was closed "not on the merits". Rung 3 — the sandbox map, and the fallback that was the real defect. Measured across all 35 files in agents/, deriving workspace-write iff tools: declares Write or Edit: - all 11 CODEX_AGENT_SANDBOX entries derive to their mapped value exactly, zero disagreements — the map carries nothing the contract does not - 24 roles fall through `|| 'read-only'`, of which 16 declare Write or Edit So the map is redundant and the silent fallback is the defect. The maintainer chose to derive but hold those 16 at read-only pending the question of whether Codex enforces sandbox_mode or merely advises; HALT.md records it. T20 asserts the emitted sandbox_mode PER ROLE against a captured baseline, not in aggregate — an aggregate passes while one role silently widens, which is the proxy-instead-of-identity shape this repo names. T24 and T25 fail on a stale hold, so the hold list cannot rot into the subset map being deleted. Rung 4 — shortFormToId, recovered rather than invented. I nearly reported this as another wrong §8.3 claim: `git log -S shortFormToId` returns only documentation commits. That was the wrong instrument. Direct inspection of sdk/src/query/phase.ts at 11918dcc3^ shows five occurrences, and the tests match that code rather than a guess at its semantics — including first-write-wins on a duplicate short form. T43 asserts at the consumer's output: the emitted `waves` map from the real CLI, which pre-fix collapses to {"1":[...]} because every short-form edge is dropped. A unit assertion on resolveDependencyId would have passed throughout this defect's life. Observed RED, this tree: rung 2 11/11 fail — no marker rung, no seams rung 3 T23,T24,T25,T26,T30 fail; T28 fails (validate agents passes a TOML whose sandbox_mode disagrees — it checks presence only) rung 4 T42,T44 fail; T43,T49 fail with waves collapsed to a single wave 1 Green and staying green: T20/T21/T22/T27 as captured baselines, #3885's unresolvable-token warning and wave-verdict suppression, and #3785's display-mapping passthrough. If the third tier over-reaches, those go red — that is their job. Disclosed weakness: T45 (a canonical id with no dash is not short-form indexed) cannot be isolated behaviorally, because planMap always masks it. It is a non-crash boundary pin, weaker than the other rows, and is recorded as such rather than presented as equivalent. Design: .gsd/phase/feat-3897-adr3473-83-rungs/40-design.md Test matrix: .gsd/phase/feat-3897-adr3473-83-rungs/50-test-matrix.md Decision: .gsd/phase/feat-3897-adr3473-83-rungs/HALT.md Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enhance(#3897): §8.3 rungs 2-4 — one marker reader, a derived sandbox, the third depends_on tier ADR-3473 §8.3 has four rungs. #3883/PR #3896 shipped the first. These are the other three. Rung 2 — the install marker had four readers, and resolveRuntime was not one. resolveRuntime resolved GSD_RUNTIME > config.runtime > 'claude' and read no marker, while bin/install.js writes one (#2297) and four hand-rolled readInstallRuntimeMarker copies existed: src/model-resolver.cts (cached, with seams), hooks/gsd-agent-isolation-guard.js, and twice in hooks/gsd-cursor-subagent-start.js. model-resolver's was already the house idiom, so it was promoted rather than replaced: src/runtime-slash.cts now owns it, and model-resolver plus both hooks delegate. The hooks reach it through ensureRuntimeBuild(), the seam lint-hooks-runtime-build-seam enforces. No import cycle existed - checked both directions before moving anything. The marker is the THIRD rung: env > project config > marker > 'claude'. N1 was checked rather than assumed, and my first reading of it was wrong. A marker holding an unknown name comes back essentially verbatim, which looked like a validation gap. Measured against the env rung with the same inputs - including "../../etc/passwd" and "claude;rm -rf /" - the two are identical, because they share resolveRuntimeNameFromCandidates. N1 asks for exactly that, and it is met. The residual (the shared normalizer normalizes shape, it does not validate against the known-runtime set) is pre-existing on the env rung and plausibly deliberate, since a new runtime should not need a code change. The marker also does not widen the trust boundary in any real sense: it lives inside the install tree beside the code, so anyone who can write it can write runtime-slash.cjs itself. Rung 3 — the map was redundant; the silent fallback was the defect. Measured across all 35 files in agents/, deriving workspace-write iff tools: declares Write or Edit: all 11 CODEX_AGENT_SANDBOX entries derive to their mapped value exactly, zero disagreements. The map carried nothing the contract did not already have, so it is DELETED rather than reconciled. What was actually broken is `|| 'read-only'`, which silently under-granted 24 of 35 roles. 16 of those 24 declare Write or Edit and would widen under derivation. Per the maintainer's decision (HALT.md), they are held at read-only pending the question of whether Codex enforces sandbox_mode or merely advises. Emitted TOML is therefore byte-identical for all 35 roles - asserted per role, not in aggregate, because an aggregate passes while one role silently widens. The hold list self-invalidates. A hold whose role no longer derives broader fails, and so does a hold naming a role with no agents/<name>.md. Without that it would rot into exactly the hand-maintained subset map being deleted, and this commit's own ledger claim would become false over time. Both cases were proved by injecting them and watching them throw. Two committed tests asserted the deleted map's existence and contents. They were pinning the thing being removed, so the tests moved rather than the production code: the 11 role-value pairs survive as a test-local PRE_3897_CODEX_AGENT_SANDBOX baseline, and the assertions now drive the real derivation against real agents/*.md. The coverage is preserved; only its source moved out of production code. validate agents gains checkCodexSandboxPosture, mirroring the existing checkCodexModelPosture: each installed TOML's sandbox_mode must equal the role's expected value, failing with role, expected and found. It previously checked file presence and manifest completeness only, so a TOML whose sandbox_mode disagreed passed. Rung 4 — shortFormToId, recovered rather than invented. I nearly reported this as another wrong §8.3 claim: git log -S returns only documentation commits. Wrong instrument. sdk/src/query/phase.ts at 11918dcc3^ carries five occurrences, and the implementation here matches that code rather than a guess at its semantics - including first-write-wins on a duplicate short form, deterministic from the sorted plan order. It resolves the bare plan number: depends_on: ["01"] now reaches 26-01-auth-hardening. That is a control-flow change, not a diagnostic one - plans that silently collapsed into a single wave 1 now execute in their declared waves, and execute-phase.md consumes those wave values. In-phase only, by construction: the map is built from this phase's rawPlans, so a same-named short form in another phase does not resolve. #3785's display-mapping passthrough and #3885's unresolvable-token warning and wave-verdict suppression are untouched and stay green. If the third tier had over-reached, those are what would have caught it. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): close a fail-open I introduced, and wire the posture check to its command Two blockers from review. Both are mine, and one is a security regression my own change created. 1. A held role could escape its hold by editing its own frontmatter. The Codex install loop set the sandbox identity from the agent's frontmatter `name:` field rather than from its filename, so the hold lookup keyed off a value the file itself declares: deriveCodexSandboxMode('gsd-doc-writer', <real file>) -> read-only deriveCodexSandboxMode('gsd-doc-writer-x', <same file, name: edited>) -> workspace-write deriveCodexSandboxMode('GSD-Doc-Writer', <same file, name: recased>) -> workspace-write What makes this a blocker rather than a nit is the DIRECTION. The deleted CODEX_AGENT_SANDBOX map had the identical lookup-key quirk, but it was an allowlist: an unmatched key fell back to read-only, which is safe. The new scheme derives workspace-write from the tool contract and uses the hold as a subtraction, so the same mismatch fails OPEN. I converted a fail-closed quirk into a fail-open one and did not notice; the isolated reviewer proved it by execution. Neither safety net caught it. validateCodexSandboxHolds only checks that <key>.md exists, never that a file's derived identity matches its key. checkCodexSandboxPosture looks the canonical source up by the installed TOML's filename, finds nothing for a renamed agent, and treats it as a custom non-roster agent — silently no violation. The identity is now the FILENAME STEM, which is what validateCodexSandboxHolds already validates and what an attacker editing frontmatter cannot change without renaming the file — at which point the existing validator catches it. The lookup is case-insensitive so a recase does not slip past either. The frontmatter name still drives the TOML body and filename, unchanged; only the sandbox identity moved. All 35 roster files were checked: name matches filename stem everywhere, so a stricter "they must agree or throw" invariant would have been safe against real content. It is deliberately NOT added — it would abort an install on a tampered file where emitting a correctly-derived read-only TOML is the safer outcome. Recorded as a fork rather than decided silently. 2. checkCodexSandboxPosture was exported and never called. cmdValidateAgents (src/verify.cts) called checkAgentsInstalled and checkCodexModelPosture only; grep for the sandbox check in that file returned nothing. So criterion 3 — "validate agents fails on semantic drift, not only on missing files" — was unmet, and `validate agents` behaved exactly as before. That is ADR-3473 Decision 2's named shape: a declared policy with no executor. It also meant the T28 test asserted at the helper's return value while the COMMAND stayed broken — the ADR-3180 Decision 4(b) failure this epic exists to close, committed by me while enforcing it elsewhere in the same epic. Now wired as an additive `sandbox_posture` field beside `codex_posture`, following the sibling precedent exactly. Drift is report-only, not a non-zero exit, because that is what checkCodexModelPosture does — two sibling posture checks disagreeing about whether a violation is fatal would be its own defect. The choice is recorded in a comment rather than left implicit. A consumer-output test now drives the real CLI and asserts on the emitted JSON, and was shown failing before the wiring and passing after. Also corrected a stale artifact: the design's Known limit L1 still claimed rung 3 was not in this deliverable, written while it was halted and false once the maintainer unblocked it. Verified after both fixes: the three bypass probes all return read-only, the per-role table is 35/35 byte-identical, and both hold self-invalidation cases still throw. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3897): the marker rung, the derived sandbox, and the bare plan-number depends_on Reference: the runtime precedence ladder in docs/CLI-TOOLS.md gains the install marker rung; docs/COMMANDS.md documents validate agents' new sandbox_posture field; docs/reference/plan-md.md documents that depends_on accepts the bare plan number. Explanation: a docs/features fragment keyed id 3897, so it cannot collide with a concurrent PR hand-allocating a section number, regenerated into FEATURES.md. ADR-3473 §8.3 gains an ANSWER blockquote in the document's own correction style, recording what was measured and built against the section's 2026-08-26 correction - including the qualification that checkAgentsInstalled itself still checks presence only, and the semantic assertion lives in a sibling wired into validate agents rather than folded into it. No how-to. Both user-visible changes are zero-step: a non-Claude install resolving its own runtime, and plans executing in their declared waves, both happen without the user doing anything. docs/how-to/control-the-reported-host-runtime.md covers a DIFFERENT ladder (resolveReportedRuntime / agent_runtime) that this change does not touch, and was deliberately left alone rather than edited by association. No tutorial - nothing multi-step to walk through. docs/AGENTS.md unchanged: it documents Claude-side tools frontmatter, never Codex sandbox_mode, and the emitted tools contract did not change. The prompt layer documents depends_on only by example, not by schema, so nothing there needed editing - and few-shot-examples/plan-checker.md already showed depends_on: ['01'], which now actually resolves. Translated copies of plan-md.md are untouched; the project treats translations as community-maintained. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): move the sandbox derivation out of the installer, off the install path, and off a third parser The full suite came back with 26 failures across four files. Three distinct causes, mapped individually rather than assuming the first explained the rest. A. Requiring bin/install.js printed the GSD banner to stdout and corrupted `validate agents` JSON. Unexpected token '', "[36m ██"... is not valid JSON checkCodexSandboxPosture reached deriveCodexSandboxMode by lazily requiring bin/install.js, whose module load prints the ASCII banner. So the command emitted banner bytes before its JSON and every JSON consumer broke, including ten tests that predate this branch. src/ reaching into bin/ was backwards layering that happened to also be loud. The derivation now lives in src/codex-agent-toml.cts - the existing Codex TOML domain module, no new module and no six-gate ripple - and both bin/install.js and src/agent-install-check.cts import it. One owner, which is §8.3's rule applied to the fix for §8.3. B. The stale-hold throw fired on a legitimate partial source dir, and masked a security assertion. validateCodexSandboxHolds treated "this hold's .md is absent from the install SOURCE dir" as a stale hold and threw. A test fixture, or any partial install source, legitimately contains a couple of agents. Worse, it threw BEFORE the path-escape check, so a test asserting that a `../../evil` frontmatter name is rejected got my unrelated error instead of the traversal rejection it was written for. A fail-closed check of mine was hiding a real security check. The "no stale holds, shrink-only" invariant is a property of the repo's canonical agents/ roster, not of whatever directory an install happens to read. It is off the runtime path and enforced where it belongs, in the tests that already existed for it. A partial source dir now installs cleanly, and the evil-name case throws with its own escapes-configHome message again. C. T8 depended on ambient process.env state. The marker/env parity assertion round-tripped through live process.env. It now compares against resolveExplicitRuntime's already-exported dependency-injection parameter - deterministic and hermetic, same claim. Proven still falsifiable rather than assumed: with the marker rung's normalization temporarily bypassed the two rungs diverge ("codex\n../../etc/passwd" vs "codex-../../etc/passwd") and the assertion fails, then passes again once reverted. One correction folded in along the way. The first version of the move added private _extractFrontmatterAndBody/_extractFrontmatterField helpers to codex-agent-toml.cts - a THIRD copy of frontmatter extraction, where the graph already shows two (bin/install.js:2348, runtime-artifact-conversion.cts:893). Adding a third inside the epic whose thesis is one implementation per rule is not defensible. deriveCodexSandboxMode no longer parses anything: it takes (identity, toolsValue) and each caller supplies the tools value using the extractor it already has. Both helpers are deleted. The identity argument is still the filename stem, so the fail-open fix is untouched. Verified after all three: `validate agents --raw` emits parseable JSON with no banner and both posture fields; the four hold-bypass probes still return read-only; the per-role table is 35/35 byte-identical at 26 read-only / 9 workspace-write; the hold list is still 16. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): drop a dev-only transitive dep, make the derivation total, retire a stale fallback test Suite down to 7 failures from 26. Three more causes, mapped individually. A. My extractor import dragged in a script that does not exist in an installed tree. Cannot find module '../../../scripts/fix-slash-commands.cjs' Chain: src/agent-install-check.cts imported runtime-artifact-conversion.cjs, which requires command-roster.cjs, whose line 36 requires ../../../scripts/fix-slash-commands.cjs. That path exists in the repo and not in an install, so every test exercising a synthetic install dir died at module load. I picked that extractor for convenience without checking what it pulls in - the same mistake that produced the banner bug, one layer further out. agent-install-check now uses a single-purpose extractToolsLine on codex-agent-toml.cts. That is deliberately NOT a general frontmatter parser: we deleted those helpers a commit ago for good reason, and this reads one line. Verified from outside the repo root that requiring either module prints nothing and does not throw. B. A test pinned the deleted name-based fallback. 'defaults unknown agents to read-only' called generateCodexAgentToml with a fixture declaring tools: Read, Write, Edit. Under derivation an unknown agent with a writing contract correctly derives workspace-write - design row S6, a new writing role gets the contract, not the pin. The behavior it asserted was the silent fallback this rung deleted; identity no longer decides the sandbox. Replaced with two rows rather than a flipped string: no tools declared -> read-only (absence is not a grant), and Write/Edit declared -> workspace-write. Strictly more coverage than the row it replaces. C. The stale-hold check still threw per derivation call. Last commit took the roster-existence check off the install path, but deriveCodexSandboxMode itself still threw when a hold's role did not derive broader FOR THE CONTENT IT WAS HANDED - so it fired on any synthetic fixture for a held role. The throw is gone, and it cost nothing: if a held role's content does not derive broader, the hold pins read-only and derivation returns read-only anyway, so the hold is a no-op and there is nothing to fail about. The staleness invariant is a property of the real agents/ roster, and validateCodexSandboxHolds still enforces it there - confirmed against the real roster after the change, not assumed. deriveCodexSandboxMode is now total: every (identity, toolsValue) including undefined and null returns read-only or workspace-write, never throws. Verified: validate agents emits parseable JSON; the four hold-bypass probes return read-only; the per-role table is 35/35 at 26 read-only / 9 workspace-write. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3897): put the rung-3 decision in the shipped docs instead of pointing at an ignored path The ADR entry and the feature fragment both ended their rung-3 explanation with "see .gsd/phase/feat-3897-adr3473-83-rungs/45-decision-rung3-sandbox.md". That directory is gitignored (.gitignore:55), so the rationale for holding 16 roles at read-only was reachable only from the machine that produced it. A reader of the ADR got a pointer to nothing. Both now carry the reasoning inline: the criterion asks both that the sandbox derive from the declared tool contract and that no role gain a broader sandbox, and those cannot both hold, because a faithful derivation widens 16 roles the deleted map never listed and that fell through its silent read-only default. The resolution is derive-and-hold - the derivation owns the rule now, each hold is released as its enforcement question is answered, and a hold is reversible where a widened sandbox that turns out to be enforced is not. Checked before assuming this was a defect class: CONTEXT.md cites .gsd/phase/<slug>/40-design.md as its standard Design: provenance line in eight module entries, and four other shipped docs do the same. Citing a phase artifact is an established convention here, so those are left alone. What was wrong was specific to these two: they put load-bearing rationale behind the pointer instead of provenance. docs/FEATURES.md regenerated from the fragment via scripts/gen-features.cjs rather than hand-edited. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): close a fail-open, stop a silent mis-resolution, and read a declaration as a declaration Two orthogonal reviews on the shipped sha. Three of the findings are the same failure class this epic exists to close, committed inside it. 1. BLOCKER - the sandbox was decided for one identity and applied to another. bin/install.js derived sandbox_mode for the filename stem and then wrote the result to `${name}.toml`, where name comes from the file's own frontmatter. Make the two disagree and a HELD role's artifact goes wide: rename gsd-doc-writer.md -> gsd-doc-writer-v2.md, keep name: gsd-doc-writer -> stem is unheld, derives workspace-write, lands on gsd-doc-writer.toml add any gsd-*.md whose frontmatter name: is a held role -> clobbers that role's toml with workspace-write Both emit read-only on origin/next, because the deleted map was an allowlist and a miss fell back safe. This is a regression my change introduced. The previous review round moved the HOLD KEY off frontmatter to the filename stem and left the OUTPUT PATH on frontmatter; my own comment at install.js:6985 calls that value attacker-editable, four lines above the line that uses it as the filename. The decision is now made over BOTH candidate identities, most-restrictive wins: if either the stem or the emitted name is held, the mode is read-only. 2. MAJOR - hold matching was toLowerCase() only, so confusables escaped. Turkish dotted/dotless i, fullwidth, NFD, trailing space/NBSP/dot/newline, ./ and ../agents/ all slipped the hold and emitted workspace-write. Identities are now basenamed, trimmed of NBSP/zero-width/control characters, NFKC-normalized and lowercased - and anything still carrying a character outside [a-z0-9._-] is treated as suspicious and derives read-only. We do not enumerate confusables; every shipped roster file is ASCII, so refusing to widen on an identity we cannot recognize is fail-closed with no false positives on real content. 3. MAJOR - the short-form depends_on tier mis-resolved SILENTLY. shortFormToId keyed on the last dash-segment of any canonical id with no constraint that it is a plan number, so a phase holding 09-FIX-auth-PLAN.md made depends_on: ["auth"] bind at wave 2 with zero warnings. This is the worst shape in the epic: the unresolvable-token warning fires on a DROPPED token, so a MIS-RESOLVED one is invisible and the tool reports a confident wave assignment built from a wrong edge. A wrong edge is worse than a missing one. The segment must now match /^\d+$/, which is exactly the contract docs/reference/plan-md.md already documents. This tier was recovered verbatim from the retired SDK lineage, which carried the same defect; we are deliberately NOT preserving it bug-for-bug, and the comment says so, so the next reader does not "restore" it. 4. MAJOR - the derivation was reading a declaration as an absence. extractToolsLine read one line, so a YAML list-form tools: block returned only its first item. Two roster files use list form, and gsd-nyquist-auditor declares Write and Edit there - parsed as "- Read", found no write tool, and emitted read-only. Rung 3's headline claim is that sandbox_mode derives from the declared tool contract; that claim was false for 2 of 35 roles and materially wrong for 1. Reading a declaration as an absence is the silent-drop class this epic exists to close. Renamed extractToolsValue and taught it both shapes. gsd-nyquist-auditor now derives workspace-write and joins CODEX_SANDBOX_HOLDS as its 17th entry, per the standing derive-and-hold decision - so emitted TOML stays byte-identical at 26 read-only / 9 workspace-write while the hold list finally records every role that would widen. A previous pass declined this fix because it moved the count; that inverts the priority. Byte-identity is preserved THROUGH the hold, not by leaving a parser broken. Divergence check, because this is where that bug hides: both paths feeding sandbox derivation - install.js's emitter and checkCodexSandboxPosture - now route through the one extractor. The tools readers in runtime-artifact-conversion and install.js's other frontmatter call sites serve Claude-side emission and do not feed sandbox derivation. Also fixed, each real: the posture check's `found` used a naive whole-file regex where its own sibling uses the block-aware scanner, so prose inside developer_instructions produced a false violation; `found` skipped truncatePostureValue and leaked a 300-char value into validate agents output; deriveCodexSandboxMode's absolute never-throws claim was false for an object with a throwing toString; T49 could not falsify cross-phase leakage (its target phase had its own 01, so a globally-scoped map passed too); T20/N6 iterated a hardcoded table and pinned the FIXTURE size, so a 36th agent would be silently unchecked; three tests reimplemented the code they were testing instead of importing it; and T2-T4 deleted GSD_RUNTIME without restoring it. Verified: hold list 17, gsd-nyquist-auditor derives workspace-write unheld and emits read-only held, roster 35/35 at 26/9, depends_on ["auth"] no longer resolves while ["01"] still does, both identity-bypass cases and every confusable vector emit read-only. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3897): the hold list is 17, and the reason the 17th was missing The count read 16 because the derivation could not read the declaration it claimed to derive from: the tools reader was single-line, so a YAML list-form tools: block returned only its first item and gsd-nyquist-auditor's declared Write and Edit were read as an absence. Both the ADR entry and the feature fragment now carry the corrected count and the reason for it, rather than a silently updated number. Deriving from a declaration you cannot parse is not deriving, and a flattering count is worse than a wrong one because it looks settled. docs/FEATURES.md regenerated from the fragment. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3897): backfill changeset pr number Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bf2332e67c |
fix(#3582): route every hook's compiled-module require through the self-heal build seam (#3629)
* test(3582): failing-first cold-tree coverage and the seam drift lint On a plugin-channel install the compiled gsd-core/bin/lib/*.cjs are legitimately absent (ADR-457 build-at-publish; the npm package builds before publishing, a raw tree materialization never does). gsd-tools.cjs calls ensureRuntimeBuild() before requiring ./lib; no hook does, so the isolation guard's Cannot-find-module lands in its fail-closed catch and is misreported as an unreadable dispatch-isolation configuration, blocking every executor dispatch. These tests fail on that: cold-tree runs of the isolation guard, statusline, cursor guard and update worker, plus the seam's actionable build error surfacing instead of the generic misreport. Also adds the drift lint the acceptance criteria require, with a fixture proving it CAN fail — a guard never shown to fail is worthless. It is red here by design: it flags today's unfixed hooks, which is exactly the defect. * fix(3582): route every hook's compiled-module require through the self-heal seam RED proven at 5b174b0d: 11 failures — the cold-tree runs for the isolation guard, cursor guard and update worker, the fail-closed-with-actionable-message assertion, and the lint's own real-tree check. The compiled runtime library is produced by build:lib and gitignored (ADR-457, build-at-publish). The npm package builds before publishing; a plugin-marketplace or git-clone install materializes the raw tree and never does, so on that channel those modules are legitimately absent. The self-heal seam added by #2002 exists to heal exactly this, and the CLI entrypoint already calls it — no hook did. The isolation guard's Cannot-find-module therefore landed in its fail-closed catch and was reported as 'could not read or resolve dispatch-isolation configuration', so an ARTIFACT ABSENCE was misdiagnosed as an unreadable project config and every executor dispatch was blocked. All SEVEN affected files now call the seam before their first compiled require. The issue named four; a scan found six; implementing it surfaced a seventh — the shared isolation sentinel helper, used by BOTH guards, which requires two compiled modules itself and would have defeated the guards' own fix on a genuinely cold tree. Same defect class, so fixed here rather than left as a known-broken remainder. Failure posture is deliberately split by hook kind: - Gates (agent isolation guard, cursor subagent start) surface the seam's actionable build error distinctly instead of swallowing it into the generic text, and stay fail-closed — a genuinely unreadable project config still DENIES exactly as before. - Cosmetic and detached hooks (statusline, update worker, update check, update banner) DEGRADE rather than crash: the statusline draws on every render and the worker is a detached process, so a build failure there must not take down the prompt. The npm path is untouched: the seam's already-built fast path returns immediately, so prebuilt installs pay nothing and behave bit-for-bit as before. Adds a drift lint, wired into the CI lint chain, so the invariant is enforced rather than remembered — without it the next hook to add a compiled require reintroduces the class silently. It is proven able to fail: a fixture hook requiring a compiled module without the seam is flagged, and one that uses the seam is not. Verified directly — on the unfixed tree it named all seven offenders; with the fix it passes. While writing the lint's comment stripper, a naive whole-text block-comment regex ate its own fixture, because this repo's comments legitimately spell the compiled-lib glob whose star-slash reads as a comment opener. Rewritten as a line-based scanner with a regression test pinning that case. * fix(3582): test the three untested seam call sites and assert typed reason codes Two independent reviews converged on the same major gap: the fix wired the seam into seven files but only four had cold-tree tests. The adversarial pass put it plainly — deleting the shared isolation-sentinel helper's seam call would not have failed any test in the diff. That file was my own addition beyond the issue's four, so it shipped untested; that is now closed. - Shared isolation-sentinel helper: its seam call is only reached when .planning is NOT directly under cwd, and every existing cold-tree fixture puts it there, so the early return always fired first. Now covered, and proven load-bearing by mutation: with the call removed the spy records zero seam invocations and the test fails. - update-check hook and update-banner hook: cold-tree tests added asserting the DEGRADED VERDICT — the fallback cache filename, and silent suppression when the package name degrades to null — rather than merely 'did not throw'. The banner hook previously had no test file at all. Standards violation fixed: two tests asserted on free-form prose via assert.match against a JSON reason string, which CONTRIBUTING bans by name — its own BAD example is exactly that. The ESLint rule only covers readFileSync/spawnSync text, so tooling did not catch it. Both isolation guards now emit a machine-readable reason_code from a frozen enum, following the repo's existing REASON convention, and the tests assert that instead. The human-readable message is unchanged for operators; only the assertion target moved. The duplicated degrade boilerplate across the three cosmetic hooks was deliberately NOT extracted, and the reason is recorded at each site: both viable shapes — a path-parameterized helper, or a ceremony-only wrapper — defeat the drift lint's per-file literal co-occurrence check, so extracting would require the lint to special-case its own helper. Triplication is the lesser evil while the lint stays a co-occurrence scan. The lint's header now states what it does and does not catch (literal quoted requires only; hooks/ scan root), so a future reader does not over-trust a guard that a concatenated path or a require inside a non-hooks helper would evade. * chore(3582): regenerate the committed install-tree fixtures Adding a new shipped hook helper changed the install tree, and those fixtures are committed-and-derived (regen:derived / gen:install-tree), so 12 'install tree — <runtime>' tests failed on 541a1913. Regenerated rather than hand-edited. The delta across all 15 runtime fixtures is exactly two lines — the new helper under both its hooks/ and gsd-hooks/ install paths — and nothing else, so the regeneration pulled in no unrelated drift. This is the bookkeeping ripple a new file under hooks/ carries; it was not visible from lint:ci, which passed both before and after. * chore(3582): backfill changeset PR number (#3629) --------- Co-authored-by: sim <sim@local> |
||
|
|
58e5a5b581 |
fix(#3566): read the per-install .gsd-runtime marker above host-wide defaults in the isolation guards (#3589)
* test(#3566): pin per-install .gsd-runtime marker precedence in the isolation guard Failing-first regression for #3566: resolveRuntimeIdentity must consult the per-install marker (<install>/gsd-core/.gsd-runtime, written by every install since #2297) above the host-wide ~/.gsd/defaults.json whose leakage #2840 exists to prevent. In-process block drives the marker through the same _setInstallRuntimeMarkerForTests seam model-resolver.cts established. * fix(#3566): read the per-install .gsd-runtime marker above host-wide defaults in the isolation guard resolveRuntimeIdentity consulted ~/.gsd/defaults.json — the exact host-wide file whose runtime leakage #2840 exists to prevent — and never the per-install marker the installer has written for every runtime since #2297. On a 2-runtime machine the guard confidently resolved the WRONG runtime and silently went inert when that runtime declares no harnessIsolationFlag. Precedence is now GSD_RUNTIME > config.json runtime > .gsd-runtime marker > defaults.json, restoring #2840's design; the defaults rung stays last so single-runtime and pre-#2297 installs keep #3045 BLOCKER 2 behavior. * fix(#3566): apply the marker rung to the cursor subagent-start fallback; review fixes Review finding (spec pass): hooks/gsd-cursor-subagent-start.js's resolveFallbackIsolation mirrored the Claude hook's exact three-rung chain and shared the bug — same rung inserted between config.json and the host-wide defaults, same #2297-pattern seam, in-process regression + negative controls. Review finding (standards): dropped the one new raw-text assert.match on the block reason (CONTRIBUTING test-output rule); the reason-naming property stays pinned by the pre-existing #3045 row. * chore(#3566): add changeset fragment * chore(#3566): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
8f75e27554 |
fix(#3045): fail closed when an executor dispatch drops its resolved isolation (#3069)
* feat(#3045): deny an executor dispatch that drops its isolation flag Every isolation gate already resolved correctly. The resolved value then reached the executor through a prose instruction telling the model to substitute it into a call the model composes itself, and nothing verified the substitution. When it was dropped, the executor edited and committed in the user's primary checkout with no consent and no warning. A prose backstop would be the same class of artifact as the defect, so this is a shipped PreToolUse hook on the Agent tool. It fires at the instant of the call rather than being read once at the top of a workflow, which is the only placement the model cannot skip. The guard is inert unless it can positively establish that this is a GSD project, that the project resolves to harness isolation, and that the dispatch targets an executor. A non-GSD repo has no invariant to enforce. Where it cannot read the configuration at all, it denies rather than assuming, with its own reason -- a guard that cannot verify must not answer safe. A malformed payload allows rather than throwing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3045): extend the isolation guard to Cursor Cursor is the second of only two runtimes that resolve harness isolation, so shipping the guard for Claude alone left half the exposed surface unguarded while the changeset implied it was covered. The two runtimes fail differently. On Claude the harness flag is a per-dispatch kwarg the model must copy into a call it composes, and the defect is that it can be dropped. On Cursor the flag is --worktree, which applies to the whole session, and the subagent-start payload carries no isolation field at all. There is no flag to check, so the guard verifies the effective state instead: whether the workspace is genuinely running outside the user's primary checkout. That is a stronger check than the Claude one because it tests reality rather than intent, and it is commented so nobody later rewrites it into a flag check. Isolation is established two ways, either sufficient: the workspace resolves to a linked git worktree, or it sits under the worktree root Cursor manages. The second matters because a directory Cursor placed there is a legitimate isolated session even before it becomes a distinct git worktree, where linkage alone would report no repository. Detecting linkage required a new primitive rather than the existing context resolver. That resolver short-circuits on finding a local .planning directory before it ever compares the git directory to the common one -- and an isolation worktree normally has its own checked-out .planning. Reusing it would have read a correctly isolated session as unisolated and denied it, which is the failure direction that gets a guard switched off. The comparison is now its own shortcut-free function that the resolver delegates to after its own shortcut, so existing behavior is unchanged, and the case that would have broken is pinned. The subagent type is checked before any configuration is read, so an unreadable config cannot deny a dispatch this guard would never have enforced against. The input-schema comment on the Cursor hook documented only the fields common to every event and omitted the ones specific to this one. That omission cost a halt during this work; it now documents both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): enforce the resolved dispatch decision, not the host capability The guard keyed on the registry's dispatch.isolation, which says only that a runtime is CAPABLE of harness worktrees. The decision that actually governs a dispatch is the one the workflow resolves after gating, and that legitimately comes out as sequential in three documented cases: a project setting use_worktrees false, a per-plan submodule intersection, and the base-check auto-degrade. The workflow tells the model to omit the flag in exactly those cases, and the guard was denying every one of them. The third case matters most. The preceding fix made the base-check degrade on git timeouts and a missing git binary, where it had previously answered "safe". That correction is right, and it means a transient hang now degrades to sequential far more often than before -- so the two changes composed into a trap where the workflow behaved exactly as designed and the guard blocked it. The workflow already resolves isolation in shell, deterministically, which is what makes it a trustworthy source in a way the model-authored call is not. It now records that resolved value through a dedicated verb, and both guards read it first. A fresh record is authoritative, so sequential dispatches pass untouched. Absent or stale, the guards fall back to the capability check combined with the project's use_worktrees setting, which still covers the case that never reaches the workflow. Also widened the matcher to accept Task alongside Agent, since a host that names the tool Task would otherwise leave the guard silently inert while implying coverage; stopped assuming Claude when no runtime is declared, which is the shipped default and would have demanded a Claude-only argument elsewhere; and made a non-git project inert rather than denied, since advising a worktree session is not actionable without a repository. The original diagnosis never modeled sequential mode as legitimate. That omission is what let this through, and it is now recorded there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): record at resolution and bind the record to its dispatch Two independent reviews converged on the same failure: the guard was fail-open in a default install, so it did not catch the defect it exists to catch. A shipped project carries no runtime key, which made "runtime not confidently known" the common case rather than a corner one. A record asserting that isolation was required but carrying no flag then fell through to a capability lookup that answered "none", and the dispatch was allowed. The flag itself only arrived from a second shell block -- the same block a model dropping the argument would also skip. A test had pinned that behavior as intended. The record is now written by the resolver, as an unavoidable consequence of asking for the value, rather than by a step the model is told in prose to go and run. A guard against a prose-carried value cannot itself depend on prose. Mode, flag and identifiers are written together and atomically, so the flagless window is gone, and a record asserting isolation with no resolvable flag now denies instead of degrading. Runtime is also resolved from the installer's own recorded default, which makes confident resolution the normal case. The per-plan submodule gate degrades after the phase-level decision and never re-recorded, so a plan that legitimately ran sequentially was denied against a still-fresh phase record. It now records its own, scoped to the plan. A record also authorized any dispatch for four hours. One phase degrading to sequential could silently license an unisolated dispatch in the next. Records now carry phase and plan, the guards require them to match, and the window is minutes rather than hours -- the resolver rewrites it before every dispatch, so a long window bought nothing and only widened the hole. The flag validator rejected any value beginning with two dashes, which is exactly the form Cursor and Windsurf declare, so their real value could never have been stored. Writer and reader also derived the record path differently and diverged inside a linked worktree without local planning state. The predictable path remains a way to silence the control without leaving a trace in the diff. It grants no access an agent with shell does not already have, so it is documented as accepted rather than redesigned around. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): correct the staleness boundary and unmask a vacuous parity test The remote runner returned twenty failures. One was a real production defect the boundary case existed to catch: a record whose age exactly equalled the staleness window was treated as fresh, so it stayed authoritative for one tick past its own expiry. Freshness is now strictly inside the window. The parity test meant to stop the two guards' executor lists from drifting could never have failed. Its project fixture was a bare directory rather than a repository, so the non-git inert branch answered before the executor list was ever consulted. It asserted agreement it never actually measured. The fixture is now a real repository, like every sibling in the file. A test also asserted that Windsurf declares the worktree flag. It does not -- Windsurf resolves to no isolation by design, having no named concurrent dispatch to isolate. The test claimed a registry fact that was never true, and a comment in the resolver repeated it. Both corrected, and the test now proves what it should have all along: that the parser accepts any bare flag value, rather than one runtime's supposed value. The new guard was missing from the bundled-hook whitelist, which is the surface that decides what actually ships, and the per-plan gate had gained calls to the launcher without the preamble those calls require. The changeset carried parenthetical product descriptions the purity rule forbids. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3045): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3045): make the guard tests hold on Windows Two tests redirect HOME to control where the installer-persisted runtime default is read from. Node resolves the home directory from USERPROFILE on Windows and never consults HOME, so both silently read the real runner profile, found no recorded runtime, and asserted against a project the hook had not recognised. The production code was already correct in asking the platform rather than the variable; only the tests were wrong to assume one variable answers everywhere. The helpers now mirror the override onto both. The symlink spoofing test also created a directory symlink unconditionally, which needs elevated privileges on Windows. It survived on this runner, but it would fail on any host without them, so the creation is now attempted and the test skips explicitly when it cannot be done -- a bare return would have counted as a pass and hidden the gap. Skipping alone would have left the platform uncovered, so the behaviour it proves is now also driven in-process through an injected realpath, following the seam already used for the clock. That case no longer depends on privileges at all, and the end-to-end test keeps its original assertions wherever symlinks work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c7c2fe3c2b |
fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd (#2680)
* fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd gsd-cursor-session-start.js and gsd-cursor-stop.js both resolved the project as path.join(process.cwd(), '.planning', 'STATE.md'). Under the cursor-agent CLI, hooks are invoked with cwd set to the Cursor config dir (~/.cursor), not the workspace — so the lookup always missed. sessionStart could only ever emit the "no .planning/ workflow found" nudge and stop's verify-work reminder could never fire, even with .planning/STATE.md sitting in the workspace. Slash commands were unaffected, which is why only the hook layer looked blind. Both hooks already buffered stdin into `raw` and never parsed it; the payload's workspace_roots carries the real path. Multi-root was left open in the report ("first root vs any root"). Resolved forward: prefer the first root that actually carries .planning/STATE.md, so a workspace whose GSD project is not the first root still resolves — strictly better than first-root-only and identical to it in the single-root CLI case. Falls back to roots[0], then to cwd, keeping IDE behavior unchanged if the IDE ever invokes hooks from the workspace. The resolver is duplicated verbatim across the two scripts rather than shared via hooks/lib/: these hooks ship standalone, and a new hooks/lib/ file must be registered in the GENERATED installer's GSD_HOOK_LIB_FILES allowlist — the installer-omits-shipped-file class that yields MODULE_NOT_FOUND at runtime. Per CLAUDE.md "Generative Fix Divergence", the duplication carries a parity assertion so the copies cannot drift. Failing-first, demonstrated by direct invocation with cwd != workspace: pre-fix sessionStart -> "no .planning/ workflow found" stop -> {} post-fix sessionStart -> ".planning/STATE.md is present" stop -> reminder tests/fix-2587-cursor-hook-workspace-roots.test.cjs spawns the real scripts as child processes with a cwd lacking .planning/ and workspace_roots pointing at it. Boundary coverage on the roots array (0 / 1 / 2 entries), plus malformed-JSON fail-open, junk-entry filtering, the parity assertion, and a guard that neither script resolves .planning from cwd again. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2587): extend workspace_roots fix to subagentStart; keep cwd a candidate Three findings from the isolated review, all fixed. 1. MISSED SITE (high). gsd-cursor-subagent-start.js carried the identical defect at line 43 — its own header documents workspace_roots in the input schema, but it resolved .planning/ from process.cwd() anyway. Under the cursor-agent CLI that meant every Cursor subagent (planner, executor, verifier) started with "no .planning/ workflow found" and no phase context. The report named only sessionStart and stop; the defect class was wider. Verified pre-fix vs post-fix by direct invocation with cwd != workspace. 2. SEMANTIC NARROWING (medium). The first cut searched only workspace_roots and fell back to cwd solely when the array was EMPTY. So when roots were supplied but none carried .planning/ while cwd did, the hook reported absent — where the pre-fix code, which always used cwd, reported present. That contradicted the fallback's own stated intent of preserving IDE behavior. cwd is now a CANDIDATE in the search (`[...roots, process.cwd()]`), so the fix is a strict superset of both the old behavior and the CLI fix, never a narrowing. 3. STALE GOLDEN FIXTURES (high, would have failed CI). The golden-install-parity fixtures store a content hash per installed file; these three hooks appear in 13 of the 19 runtime fixtures. Regenerated via `npm run gen:golden` — the diff is exactly the three hook hashes in exactly those 13 runtimes. Tests extended: subagentStart resolution via workspace_roots; the stop hook's absent branch (previously only session-start's was covered); an explicit regression guard that a project at cwd is still found when roots miss; parity now asserts all THREE copies byte-identical; and the cwd guard sweeps the whole RESOLVING_HOOKS list so a future hook in this family cannot be left on cwd. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * refactor(#2587): extract cursor workspace resolution to a shared hooks/lib module The duplicate-plus-parity-test approach was the wrong call. The reported issue named two hooks; a third (subagentStart) had the identical defect. That is the signature of a systemic problem, and three copies of a resolver guarded by a parity assertion is a divergence risk maintained by hand rather than a fix. hooks/lib/cursor-workspace.js is now the single implementation. All three Cursor hooks require it; none defines a local copy. Divergence is prevented structurally instead of by asserting three copies stay byte-identical. The reason duplication looked necessary was real, and is fixed properly here rather than worked around: Cursor sets hostBehaviors.skipSharedHooksInstall (#2089), so it never reaches the installer's bulk hooks/lib copy — it was the ONE runtime shipping these hooks WITHOUT hooks/lib (verified against all 19 golden fixtures: cursor had the hook scripts, no lib). A naive require would have thrown MODULE_NOT_FOUND at load, BEFORE each hook's own try/catch, wedging every session on precisely the runtime this bug is about. writeCursorHooksJson (src/runtime-hooks-surface.cts) now stages the hooks/lib helpers the staged scripts actually require, discovered by scanning their require('./lib/…') calls rather than a hardcoded name — so a future helper cannot be silently omitted. This is narrower than flipping skipSharedHooksInstall, which would wrongly pull in every shared hook. cursor-workspace.js is also added to GSD_HOOK_LIB_FILES so uninstall and the manifest manage it for the runtimes that do receive hooks/lib. Verified against a REAL install (runMinimalInstall, cursor/global): the helper is staged, and all three INSTALLED hooks resolve the workspace end-to-end from a cwd that is not the project. Also closes the review gap that the stop hook was excluded from the cwd-candidate regression loop — it now sweeps RESOLVING_HOOKS. The byte-parity test is replaced by a structural guard (every hook requires the shared module, none redefines it) plus a new install test asserting the helper is staged and the installed hook actually loads against it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2587): fail loud on a missing hook lib source; drop unsubstituted version marker Two findings from the installer-focused review. H1 — the staging step's `if (!fs.existsSync(libSrc)) continue;` silently defeated the very guarantee it was added for. Reproduced: delete hooks/lib/cursor-workspace.js from source, run the cursor install — it exits 0, prints "Done!", and ships the three hook scripts with an EMPTY hooks/lib/. The installed hook then throws `Cannot find module './lib/cursor-workspace.js'` at load, before its own try/catch, wedging every session — and nothing surfaces until a user hits it. The scan protected against a required-but-UNLISTED helper while leaving required-but-MISSING wide open (typo, bad rebase, an accidental delete). It now throws: a missing helper source is a packaging bug and aborts the install. M1 — hooks/lib/cursor-workspace.js carried a `gsd-hook-version: <placeholder>` marker that NOTHING substitutes: copyLibDir stamps .sh files only, and writeCursorHooksJson's staging applies just the colon-to-dash rewrite. Verified the literal was reaching disk on both the bulk (--claude) and Cursor (--cursor) paths. hooks/lib/git-cmd.js — the only pre-existing hooks/lib/*.js — carries no such marker, so this was newly introduced, not inherited. Marker removed, matching that precedent, with a note on why. (The explanatory comment deliberately does not spell the token out, or it would reintroduce the literal.) M2 — the require-scan regex demanded the exact compact form, so `require( "./lib/x.js" )` would silently fail to stage its helper and compound H1. Now tolerant of interior whitespace and either quote style. Regression test added for H1 — the reviewer confirmed the invariant had zero coverage repo-wide: a source tree carrying the hooks but no hooks/lib/ must make writeCursorHooksJson throw rather than produce a broken install. Re-verified end to end: the missing-source case throws, no unsubstituted literal ships, and the installed hook still resolves the workspace from a foreign cwd. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2587): backfill changeset pr number (#2680) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b0d985ccb3 | feat(#2089): migrate cursor onto imperative adapter + hook-bus/dispatch upgrades |