* feat(#3045): deny an executor dispatch that drops its isolation flag Every isolation gate already resolved correctly. The resolved value then reached the executor through a prose instruction telling the model to substitute it into a call the model composes itself, and nothing verified the substitution. When it was dropped, the executor edited and committed in the user's primary checkout with no consent and no warning. A prose backstop would be the same class of artifact as the defect, so this is a shipped PreToolUse hook on the Agent tool. It fires at the instant of the call rather than being read once at the top of a workflow, which is the only placement the model cannot skip. The guard is inert unless it can positively establish that this is a GSD project, that the project resolves to harness isolation, and that the dispatch targets an executor. A non-GSD repo has no invariant to enforce. Where it cannot read the configuration at all, it denies rather than assuming, with its own reason -- a guard that cannot verify must not answer safe. A malformed payload allows rather than throwing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3045): extend the isolation guard to Cursor Cursor is the second of only two runtimes that resolve harness isolation, so shipping the guard for Claude alone left half the exposed surface unguarded while the changeset implied it was covered. The two runtimes fail differently. On Claude the harness flag is a per-dispatch kwarg the model must copy into a call it composes, and the defect is that it can be dropped. On Cursor the flag is --worktree, which applies to the whole session, and the subagent-start payload carries no isolation field at all. There is no flag to check, so the guard verifies the effective state instead: whether the workspace is genuinely running outside the user's primary checkout. That is a stronger check than the Claude one because it tests reality rather than intent, and it is commented so nobody later rewrites it into a flag check. Isolation is established two ways, either sufficient: the workspace resolves to a linked git worktree, or it sits under the worktree root Cursor manages. The second matters because a directory Cursor placed there is a legitimate isolated session even before it becomes a distinct git worktree, where linkage alone would report no repository. Detecting linkage required a new primitive rather than the existing context resolver. That resolver short-circuits on finding a local .planning directory before it ever compares the git directory to the common one -- and an isolation worktree normally has its own checked-out .planning. Reusing it would have read a correctly isolated session as unisolated and denied it, which is the failure direction that gets a guard switched off. The comparison is now its own shortcut-free function that the resolver delegates to after its own shortcut, so existing behavior is unchanged, and the case that would have broken is pinned. The subagent type is checked before any configuration is read, so an unreadable config cannot deny a dispatch this guard would never have enforced against. The input-schema comment on the Cursor hook documented only the fields common to every event and omitted the ones specific to this one. That omission cost a halt during this work; it now documents both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): enforce the resolved dispatch decision, not the host capability The guard keyed on the registry's dispatch.isolation, which says only that a runtime is CAPABLE of harness worktrees. The decision that actually governs a dispatch is the one the workflow resolves after gating, and that legitimately comes out as sequential in three documented cases: a project setting use_worktrees false, a per-plan submodule intersection, and the base-check auto-degrade. The workflow tells the model to omit the flag in exactly those cases, and the guard was denying every one of them. The third case matters most. The preceding fix made the base-check degrade on git timeouts and a missing git binary, where it had previously answered "safe". That correction is right, and it means a transient hang now degrades to sequential far more often than before -- so the two changes composed into a trap where the workflow behaved exactly as designed and the guard blocked it. The workflow already resolves isolation in shell, deterministically, which is what makes it a trustworthy source in a way the model-authored call is not. It now records that resolved value through a dedicated verb, and both guards read it first. A fresh record is authoritative, so sequential dispatches pass untouched. Absent or stale, the guards fall back to the capability check combined with the project's use_worktrees setting, which still covers the case that never reaches the workflow. Also widened the matcher to accept Task alongside Agent, since a host that names the tool Task would otherwise leave the guard silently inert while implying coverage; stopped assuming Claude when no runtime is declared, which is the shipped default and would have demanded a Claude-only argument elsewhere; and made a non-git project inert rather than denied, since advising a worktree session is not actionable without a repository. The original diagnosis never modeled sequential mode as legitimate. That omission is what let this through, and it is now recorded there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): record at resolution and bind the record to its dispatch Two independent reviews converged on the same failure: the guard was fail-open in a default install, so it did not catch the defect it exists to catch. A shipped project carries no runtime key, which made "runtime not confidently known" the common case rather than a corner one. A record asserting that isolation was required but carrying no flag then fell through to a capability lookup that answered "none", and the dispatch was allowed. The flag itself only arrived from a second shell block -- the same block a model dropping the argument would also skip. A test had pinned that behavior as intended. The record is now written by the resolver, as an unavoidable consequence of asking for the value, rather than by a step the model is told in prose to go and run. A guard against a prose-carried value cannot itself depend on prose. Mode, flag and identifiers are written together and atomically, so the flagless window is gone, and a record asserting isolation with no resolvable flag now denies instead of degrading. Runtime is also resolved from the installer's own recorded default, which makes confident resolution the normal case. The per-plan submodule gate degrades after the phase-level decision and never re-recorded, so a plan that legitimately ran sequentially was denied against a still-fresh phase record. It now records its own, scoped to the plan. A record also authorized any dispatch for four hours. One phase degrading to sequential could silently license an unisolated dispatch in the next. Records now carry phase and plan, the guards require them to match, and the window is minutes rather than hours -- the resolver rewrites it before every dispatch, so a long window bought nothing and only widened the hole. The flag validator rejected any value beginning with two dashes, which is exactly the form Cursor and Windsurf declare, so their real value could never have been stored. Writer and reader also derived the record path differently and diverged inside a linked worktree without local planning state. The predictable path remains a way to silence the control without leaving a trace in the diff. It grants no access an agent with shell does not already have, so it is documented as accepted rather than redesigned around. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): correct the staleness boundary and unmask a vacuous parity test The remote runner returned twenty failures. One was a real production defect the boundary case existed to catch: a record whose age exactly equalled the staleness window was treated as fresh, so it stayed authoritative for one tick past its own expiry. Freshness is now strictly inside the window. The parity test meant to stop the two guards' executor lists from drifting could never have failed. Its project fixture was a bare directory rather than a repository, so the non-git inert branch answered before the executor list was ever consulted. It asserted agreement it never actually measured. The fixture is now a real repository, like every sibling in the file. A test also asserted that Windsurf declares the worktree flag. It does not -- Windsurf resolves to no isolation by design, having no named concurrent dispatch to isolate. The test claimed a registry fact that was never true, and a comment in the resolver repeated it. Both corrected, and the test now proves what it should have all along: that the parser accepts any bare flag value, rather than one runtime's supposed value. The new guard was missing from the bundled-hook whitelist, which is the surface that decides what actually ships, and the per-plan gate had gained calls to the launcher without the preamble those calls require. The changeset carried parenthetical product descriptions the purity rule forbids. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3045): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3045): make the guard tests hold on Windows Two tests redirect HOME to control where the installer-persisted runtime default is read from. Node resolves the home directory from USERPROFILE on Windows and never consults HOME, so both silently read the real runner profile, found no recorded runtime, and asserted against a project the hook had not recognised. The production code was already correct in asking the platform rather than the variable; only the tests were wrong to assume one variable answers everywhere. The helpers now mirror the override onto both. The symlink spoofing test also created a directory symlink unconditionally, which needs elevated privileges on Windows. It survived on this runner, but it would fail on any host without them, so the creation is now attempted and the test skips explicitly when it cannot be done -- a bare return would have counted as a pass and hidden the gap. Skipping alone would have left the platform uncovered, so the behaviour it proves is now also driven in-process through an injected realpath, following the seam already used for the clock. That case no longer depends on privileges at all, and the end-to-end test keeps its original assertions wherever symlinks work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
561 lines
28 KiB
JavaScript
561 lines
28 KiB
JavaScript
#!/usr/bin/env node
|
|
// gsd-hook-version: {{GSD_VERSION}}
|
|
// gsd-cursor-subagent-start.js — Cursor subagentStart hook (ADR-1239 / #2089,
|
|
// isolation guard #3045)
|
|
//
|
|
// Cursor invokes this script when a subagent session starts.
|
|
// Protocol: JSON from Cursor on stdin; JSON response on stdout.
|
|
//
|
|
// Input schema (cursor subagentStart) — Cursor's hooks contract is a COMMON
|
|
// envelope shared by every hook, PLUS event-specific fields layered on top
|
|
// (cursor.com/docs/hooks, "Reference > Common schema"). A prior version of
|
|
// this comment documented only the envelope and omitted the event-specific
|
|
// fields entirely — that omission is exactly what caused #3045's isolation
|
|
// guard work to stall on a false schema conflict, so every field below is
|
|
// still read defensively (assume any of them may be absent/malformed):
|
|
// Common envelope (all hooks): conversation_id, generation_id, model,
|
|
// model_id, model_params, hook_event_name, cursor_version,
|
|
// workspace_roots (array of paths), user_email, transcript_path.
|
|
// (Some fields are omitted for app-lifecycle hooks; this script's own
|
|
// prior comment listed session_id/is_background_agent instead of
|
|
// model_id/model_params — the exact set observed is not guaranteed.)
|
|
// subagentStart-specific additions: subagent_id, subagent_type, task,
|
|
// parent_conversation_id, tool_call_id, subagent_model,
|
|
// is_parallel_worker, git_branch (optional).
|
|
//
|
|
// Output schema (cursor subagentStart):
|
|
// { additional_context?: string, permission?: "allow"|"deny", user_message?: string }
|
|
// "ask" is NOT a supported permission value for subagentStart — Cursor
|
|
// treats it as "deny". This script only ever emits "allow" (by omitting
|
|
// `permission`, preserving the pre-#3045 output shape) or an explicit
|
|
// "deny" with `user_message`.
|
|
//
|
|
// Behaviour:
|
|
// - Injects a brief GSD state reminder so subagents (planner, executor,
|
|
// verifier) have the current phase context (unchanged since #2587).
|
|
// - NEW (#3045): denies spawning a GSD executor subagent when this
|
|
// project's dispatch isolation resolves to "harness-worktree" but the
|
|
// session is NOT actually running isolated from the user's primary
|
|
// checkout. Cursor's `--worktree` is a SESSION-level flag (no per-call
|
|
// isolation parameter exists on `subagentStart`, unlike Claude's
|
|
// `Agent(isolation=...)` kwarg), so this guard verifies EFFECTIVE STATE
|
|
// instead of looking for a flag — see resolveIsolationDecision() below.
|
|
// - Fails open on a payload it cannot parse or that carries fields it does
|
|
// not need: never throws, never blocks a call it cannot evaluate.
|
|
// Isolation resolution itself fails CLOSED (denies) for the two cases
|
|
// that are load-bearing and are NOT the same as "cannot parse": (a) a
|
|
// GSD project resolved to harness-worktree whose isolation state cannot
|
|
// be verified, and (b) a harness-worktree GSD project dispatch with no
|
|
// usable subagent_type — a guard that cannot verify must not answer
|
|
// "safe" (#3050).
|
|
//
|
|
// Cursor docs: https://cursor.com/docs/hooks
|
|
|
|
'use strict';
|
|
|
|
const fs = require('fs');
|
|
const path = require('path');
|
|
const os = require('os');
|
|
|
|
// Workspace resolution is shared across the Cursor hooks (#2587) — see
|
|
// hooks/lib/cursor-workspace.js. Staged next to these scripts by
|
|
// writeCursorHooksJson so the require always resolves post-install.
|
|
const { resolveStatePath } = require('./lib/cursor-workspace.js');
|
|
const { readSentinel, VALID_ISOLATION, extractDispatchIdentifiers, sentinelAppliesToDispatch } = require('./lib/isolation-sentinel.js');
|
|
|
|
const MSG_PRESENT =
|
|
'GSD: Subagent session started — review .planning/STATE.md for the current phase and any blockers before acting.';
|
|
const MSG_ABSENT =
|
|
'GSD: Subagent session started — no .planning/ workflow found.';
|
|
|
|
// GSD's Cursor agent artifacts install with `destSubpath: "agents"`,
|
|
// `prefix: "gsd-"`, flat nesting, via the `convertClaudeAgentToCursorAgent`
|
|
// converter, and `hostIntegration.dispatch.namedDispatch === true`
|
|
// (gsd-core/bin/lib/capability-registry.cjs, runtimes.cursor) — i.e. Cursor
|
|
// dispatches named subagents by their real agent name, identically to
|
|
// Claude. So GSD's executor surfaces as subagent_type === "gsd-executor" on
|
|
// Cursor too, the same identifier hooks/gsd-agent-isolation-guard.js checks
|
|
// for on Claude. A Set, not a bare string compare, so a future sibling
|
|
// executor role can be added here without touching the matching logic below.
|
|
const EXECUTOR_SUBAGENT_TYPES = new Set(['gsd-executor']);
|
|
|
|
/**
|
|
* Runs `realpathFn`, never throwing. A path that cannot be resolved (does not
|
|
* exist, dangling symlink, ELOOP, ...) yields `null` rather than an
|
|
* exception — the caller decides what "cannot resolve" means for its own
|
|
* verdict (#3045 finding 2).
|
|
*
|
|
* `realpathFn` is injectable (defaults to `fs.realpathSync`), per the repo's
|
|
* dependency-injection seam convention (mirrors the `clock` seam elsewhere in
|
|
* these hooks) — this lets tests exercise the realpath-based spoof-resistance
|
|
* logic below with a fabricated symlink-resolution mapping, without ever
|
|
* creating a real filesystem symlink (directory symlinks require elevated
|
|
* privileges on unprivileged Windows CI).
|
|
*/
|
|
function realpathOrNull(p, realpathFn) {
|
|
try {
|
|
return realpathFn(p);
|
|
} catch {
|
|
return null;
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Resolve whether `root` is running in a session Cursor ISOLATED FOR THIS
|
|
* DISPATCH — i.e. a worktree the harness itself created and manages, not
|
|
* merely "some linked git worktree".
|
|
*
|
|
* #3045 security review (finding 3): "is a linked git worktree" is NOT "is
|
|
* isolated from the tree the human is using". A developer who opens Cursor
|
|
* directly in a hand-made `git worktree add` checkout — routine, see
|
|
* `.claude/worktrees/` in this very repo — is not protected by anything;
|
|
* nothing stops them from also editing that same checkout by hand. The ONLY
|
|
* signal that actually proves harness isolation is that `root` resolves
|
|
* under Cursor's OWN managed worktree root (`<cursor config dir>/worktrees`,
|
|
* i.e. `~/.cursor/worktrees` by default — `getGlobalConfigDir('cursor')`
|
|
* honors the `CURSOR_CONFIG_DIR` env override and `~` expansion for free).
|
|
* That is made NECESSARY AND SUFFICIENT below. Do NOT reinstate
|
|
* `resolveWorktreeLinkage`'s `linked_worktree_root` mode as an alternative
|
|
* OR'd proof of isolation — that is precisely the bypass finding 3 closed;
|
|
* a future "simplification" that merges it back in re-opens unconsented
|
|
* writes to the human's active checkout.
|
|
*
|
|
* `resolveWorktreeLinkage` is still called, but ONLY as a diagnostic: it
|
|
* distinguishes "confidently not isolated" from "git could not answer
|
|
* (timeout) — cannot determine" so the eventual deny reason stays
|
|
* actionable. Its result never flips `isolated`.
|
|
*
|
|
* Both the workspace root and the managed root are realpath'd before
|
|
* comparison (#3045 finding 2) — lexical `path.relative` alone is spoofable
|
|
* by a symlink or bind mount at either location, plantable by any process
|
|
* running with the user's permissions (including an agent already inside a
|
|
* legitimately isolated worktree, which has shell access by design).
|
|
* `fs.realpathSync` throwing (nonexistent path) never propagates — it
|
|
* degrades to "cannot resolve", never to "isolated". realpath also resolves
|
|
* the `CURSOR_CONFIG_DIR`-derived managed root itself (not just `root`), so a
|
|
* symlinked or case-differing `CURSOR_CONFIG_DIR` (case-insensitive
|
|
* filesystems normalize to on-disk casing via realpath's dirent walk, not
|
|
* string comparison) is covered on BOTH sides of the comparison, not only
|
|
* `root`'s.
|
|
*
|
|
* Returns `{ isolated: true|false, cannotDetermine: bool, notApplicable: bool }`.
|
|
* `notApplicable` (#3045 MAJOR 3) is true only for a confidently-not-a-git-repo
|
|
* root — see the `not_git_repo` branch below.
|
|
*
|
|
* `realpath` is injectable (`(p: string) => string`, throws like
|
|
* `fs.realpathSync` on an unresolvable path; defaults to the real
|
|
* `fs.realpathSync`) per the repo's clock-seam-style dependency-injection
|
|
* convention. This lets tests drive the exact spoof-resistance logic this
|
|
* function exists for (a symlink at the managed root pointing OUTSIDE it)
|
|
* with an injected resolution mapping, in-process, on every platform —
|
|
* without creating a real directory symlink, which requires elevated
|
|
* privileges on unprivileged Windows CI.
|
|
*/
|
|
function resolveIsolationEvidence(root, { realpath = fs.realpathSync } = {}) {
|
|
let managedRoot = null;
|
|
try {
|
|
// Sibling data/policy module, staged alongside this hook at install time
|
|
// (same pattern as hooks/gsd-statusline.js's requires of gsd-core/bin/lib/*).
|
|
const { getGlobalConfigDir } = require('../gsd-core/bin/lib/runtime-homes.cjs');
|
|
managedRoot = path.join(getGlobalConfigDir('cursor'), 'worktrees');
|
|
} catch {
|
|
managedRoot = null;
|
|
}
|
|
|
|
const realRoot = realpathOrNull(root, realpath);
|
|
const realManagedRoot = managedRoot === null ? null : realpathOrNull(managedRoot, realpath);
|
|
|
|
if (realRoot !== null && realManagedRoot !== null) {
|
|
const rel = path.relative(realManagedRoot, realRoot);
|
|
const underManagedRoot = rel === '' || (!rel.startsWith('..') && !path.isAbsolute(rel));
|
|
if (underManagedRoot) return { isolated: true, cannotDetermine: false, notApplicable: false };
|
|
}
|
|
|
|
// Not proven isolated by the only signal that counts. Resolve the
|
|
// diagnostic-only linkage check purely to make the deny reason legible —
|
|
// see the doc comment above; this NEVER flips `isolated`.
|
|
let linkageReason = null;
|
|
try {
|
|
// Sibling data/policy module, staged alongside this hook at install time.
|
|
const { resolveWorktreeLinkage } = require('../gsd-core/bin/lib/worktree-safety.cjs');
|
|
linkageReason = resolveWorktreeLinkage(root).reason;
|
|
} catch {
|
|
linkageReason = null;
|
|
}
|
|
|
|
if (realRoot === null) {
|
|
// `root` itself could not be resolved on disk. In the live hook this is
|
|
// defense in depth rather than a reachable path today: the GSD-project
|
|
// existence gate in resolveIsolationDecision already requires `root` to
|
|
// resolve (it must contain a readable `.planning/config.json`) before
|
|
// evidence is ever consulted, so a workspace root that plainly does not
|
|
// exist allows earlier as "not a GSD project" — never here. Kept anyway
|
|
// per finding 2's explicit directive: an unresolvable path must never
|
|
// silently read as "isolated".
|
|
return { isolated: false, cannotDetermine: true, notApplicable: false };
|
|
}
|
|
if (linkageReason === 'git_timed_out') {
|
|
return { isolated: false, cannotDetermine: true, notApplicable: false };
|
|
}
|
|
if (linkageReason === 'not_git_repo') {
|
|
// #3045 MAJOR 3: a confidently-non-git `root` has no primary git
|
|
// checkout to protect from an isolated-worktree bypass — Cursor's
|
|
// `--worktree` / `/worktree` (the deny message's own remediation) create
|
|
// a GIT worktree, so telling the user to start one is unactionable
|
|
// advice for a directory that isn't a git repo at all. Treat as INERT
|
|
// (allow) rather than a confident negative; this is distinct from
|
|
// `cannotDetermine` (git responded definitively here, it just said "not
|
|
// a repo") and from `isolated` (nothing was proven isolated) — it is its
|
|
// own "this guard's threat model does not apply" outcome.
|
|
return { isolated: false, cannotDetermine: false, notApplicable: true };
|
|
}
|
|
return { isolated: false, cannotDetermine: false, notApplicable: false };
|
|
}
|
|
|
|
/**
|
|
* Resolve every non-empty string entry of `workspace_roots` — ALL checkout
|
|
* paths Cursor is operating on for this hook invocation, not just the first.
|
|
*
|
|
* #3045 security review (finding 1): a multi-root Cursor workspace whose
|
|
* FIRST root is a non-GSD directory (or an isolated worktree) and whose
|
|
* SECOND root is the GSD project in the primary checkout must still be
|
|
* caught — every root is a directory the dispatched subagent can reach and
|
|
* write to, regardless of position. `hooks/lib/cursor-workspace.js` already
|
|
* established the "scan every root" precedent for its own (different)
|
|
* purpose; this is a parallel scan for isolation applicability, not a
|
|
* duplicate of that module's single-root-resolution job (it resolves ONE
|
|
* root to report state-file presence; this resolves the full set to decide
|
|
* whether ANY of them is an unconsented write target).
|
|
*
|
|
* Cursor runs hooks with cwd set to its own config dir (~/.cursor), NOT the
|
|
* workspace (hooks/lib/cursor-workspace.js), so `workspace_roots` is the
|
|
* only reliable source for "what directories is this dispatch actually in".
|
|
*
|
|
* #3045 MINOR: a RELATIVE entry is rejected (`path.isAbsolute`), not merely
|
|
* accepted-and-hoped — every downstream consumer (`.planning/config.json`
|
|
* existence check, `realpathOrNull`, `resolveWorktreeLinkage`) joins/resolves
|
|
* it against whatever the CURRENT PROCESS cwd happens to be, which for this
|
|
* hook is Cursor's own config dir (~/.cursor per the comment above), NOT the
|
|
* workspace. A relative root would therefore resolve against the wrong
|
|
* directory and — because a wrong/nonexistent `.planning/config.json` path
|
|
* reads as "not a GSD project" — silently ALLOW a dispatch this guard should
|
|
* have evaluated (fail OPEN). Filtering it out here instead makes it "not a
|
|
* resolvable workspace root", which degrades the SAME way (allow, step 2 of
|
|
* resolveIsolationDecision's applicability list) but for the honest reason.
|
|
*/
|
|
function getWorkspaceRoots(data) {
|
|
const roots = Array.isArray(data.workspace_roots) ? data.workspace_roots : [];
|
|
return roots.filter((r) => typeof r === 'string' && r.length > 0 && path.isAbsolute(r));
|
|
}
|
|
|
|
/**
|
|
* Decide whether to deny this subagentStart. Returns
|
|
* `{ action: 'allow' } | { action: 'deny', reason: string }`.
|
|
*
|
|
* Applicability (must positively determine all of the following to deny —
|
|
* otherwise allow):
|
|
* 1. `subagent_type` is not confidently a NON-executor (a present,
|
|
* non-empty string that isn't in EXECUTOR_SUBAGENT_TYPES short-circuits
|
|
* to allow immediately, before any project/isolation resolution runs —
|
|
* mirrors hooks/gsd-agent-isolation-guard.js checking subagent_type
|
|
* first, and matters here specifically: an unreadable config must never
|
|
* deny a dispatch this guard was never going to enforce against),
|
|
* 2. a workspace root is resolvable from `workspace_roots`,
|
|
* 3. that root is a GSD project (`.planning/config.json` exists there),
|
|
* 4. the resolved dispatch isolation is `harness-worktree`,
|
|
* 5. `subagent_type` identifies a GSD executor (or is missing/malformed —
|
|
* see the cannot-determine case below),
|
|
* 6. the session is NOT actually isolated (resolveIsolationEvidence).
|
|
*
|
|
* No workspace root at all degrades to allow (step 2), mirroring
|
|
* hooks/gsd-agent-isolation-guard.js's own "not a GSD project → allow"
|
|
* branch: project-existence is the gate that makes fail-closed apply in the
|
|
* first place, so being unable to even locate a candidate project is not
|
|
* itself a fail-closed trigger — it is the same "not a GSD project" shape
|
|
* that guard already treats as inert.
|
|
*
|
|
* Two DISTINCT fail-closed ("cannot determine") reasons per #3050's lesson
|
|
* that a guard which cannot verify must not answer "safe" — both scoped to
|
|
* "GSD project resolved to harness-worktree", never to a dispatch already
|
|
* confirmed to be a non-executor:
|
|
* - this project's dispatch-isolation configuration cannot be read/resolved
|
|
* (registry require/parse failure, or config.json unreadable),
|
|
* - `subagent_type` is missing or not a usable non-empty string on a
|
|
* dispatch this guard could not rule out as an executor.
|
|
*
|
|
* Isolation resolution (#3045 BLOCKER fix, see hooks/lib/isolation-sentinel.js):
|
|
* prefers the workflow's own PERSISTED per-dispatch decision (the sentinel
|
|
* `record-dispatch-isolation` writes) over re-deriving a host CAPABILITY from
|
|
* the registry. `none`/`orchestrator-worktree` from a fresh sentinel ALLOW
|
|
* immediately — sequential/orchestrator-managed dispatch is legitimate. An
|
|
* absent/stale sentinel falls back to `resolveFallbackIsolation` (registry +
|
|
* `workflow.use_worktrees`, runtime resolved GSD_RUNTIME env >
|
|
* .planning/config.json `runtime` key > 'cursor'). The default is
|
|
* confidently "cursor" here — UNLIKE hooks/gsd-agent-isolation-guard.js's own
|
|
* fallback, which must treat "no explicit signal" as cannot-determine
|
|
* because that hook installs across every `hostIntegration.hooksSurface ===
|
|
* 'settings-json'` runtime — because THIS script only ever runs as Cursor's
|
|
* own subagentStart hook; there is no other host it could be executing
|
|
* under, so defaulting to 'cursor' is a confirmed fact of the execution
|
|
* context, not a guess (#3045 MINOR — this note replaces a prior comment
|
|
* that inaccurately claimed to "mirror" runtime-slash.cjs's resolveRuntime,
|
|
* which defaults to 'claude'; the two intentionally diverge).
|
|
*
|
|
* #3045 security review (finding 1): applicability step 2 above now means
|
|
* "a workspace root is resolvable", plural — resolveIsolationDecision
|
|
* evaluates EVERY entry of `workspace_roots` via evaluateRootIsolation() and
|
|
* denies on the first one that fails. `subagent_type` is still resolved
|
|
* exactly once, up front, before any root is touched (applicability step 1
|
|
* stays a single check, not per-root — an unreadable config on one root must
|
|
* never even be attempted for a confirmed non-executor dispatch).
|
|
*/
|
|
function resolveIsolationDecision(data, { clock = Date, realpath = fs.realpathSync } = {}) {
|
|
const subagentType = data.subagent_type;
|
|
const isConfirmedNonExecutor = typeof subagentType === 'string'
|
|
&& subagentType.length > 0
|
|
&& !EXECUTOR_SUBAGENT_TYPES.has(subagentType);
|
|
if (isConfirmedNonExecutor) return { action: 'allow' };
|
|
|
|
const roots = getWorkspaceRoots(data);
|
|
if (roots.length === 0) return { action: 'allow' };
|
|
|
|
// #3045 SECURITY F2: best-effort plan/phase extraction from this
|
|
// dispatch's own `task` text (Cursor carries the same prompt content the
|
|
// Claude Agent() dispatch does — see extractDispatchIdentifiers), so a
|
|
// fresh sentinel that disagrees with THIS dispatch is treated as
|
|
// inapplicable rather than trusted.
|
|
const dispatchIds = extractDispatchIdentifiers(data.task);
|
|
|
|
for (const root of roots) {
|
|
const verdict = evaluateRootIsolation(root, subagentType, { clock, dispatchIds, realpath });
|
|
if (verdict.action === 'deny') return verdict;
|
|
}
|
|
return { action: 'allow' };
|
|
}
|
|
|
|
/**
|
|
* Conservative fallback resolution used when the #3045 sentinel is absent or
|
|
* stale for `root`: re-derive isolation from the registry CAPABILITY, gated
|
|
* by `workflow.use_worktrees` (config-schema key confirmed present in
|
|
* gsd-core/bin/shared/config-schema.manifest.json's validKeys, so it survives
|
|
* loadConfig's whitelist; read directly from the raw config.json here, same
|
|
* side-effect-free approach cmdConfigGet itself uses).
|
|
*
|
|
* #3045 MAJOR fix ("Cursor residual false-deny"): previously defaulted
|
|
* confidently to 'cursor' whenever no `GSD_RUNTIME`/config.json `runtime`
|
|
* signal existed, purely because this script only ever executes as Cursor's
|
|
* OWN `subagentStart` hook — true of the PROCESS, but not evidence the
|
|
* PROJECT itself declared an isolation requirement this guard can verify.
|
|
* Combined with a stale/absent sentinel (outside `execute-phase`, after
|
|
* `.gsd` cleanup, a phase running past the sentinel's staleness window, or a
|
|
* base-check-degraded run whose sentinel went stale before a fresh one was
|
|
* recorded), that default made every such `gsd-executor` dispatch resolve to
|
|
* "harness-worktree" and then hard-DENY unless the session happened to be
|
|
* running under Cursor's own managed worktree root — a false-deny of
|
|
* otherwise legitimate dispatches, unlike `hooks/gsd-agent-isolation-guard.js`,
|
|
* which degrades an undeterminable runtime to inert (#3045 MAJOR 2). Aligned
|
|
* here: an explicit signal is now required — `GSD_RUNTIME` > config.json
|
|
* `runtime` key > `~/.gsd/defaults.json` `runtime` (mirrors the Claude hook's
|
|
* `resolveRuntimeIdentity`; `bin/install.js`'s `writeNonClaudeDefaults`
|
|
* persists the installed runtime there for every non-Claude install,
|
|
* including Cursor, so a REAL Cursor+GSD install still resolves confidently
|
|
* — this only stops GUESSING 'cursor' for a project that never declared any
|
|
* runtime signal at all).
|
|
*/
|
|
function resolveFallbackIsolation(root, configPath) {
|
|
const { resolveRuntimeNameFromCandidates } = require('../gsd-core/bin/lib/runtime-name-policy.cjs');
|
|
const { runtimes } = require('../gsd-core/bin/lib/capability-registry.cjs');
|
|
|
|
let runtimeId = resolveRuntimeNameFromCandidates(process.env.GSD_RUNTIME);
|
|
const rawConfig = fs.readFileSync(configPath, 'utf-8');
|
|
const parsedConfig = JSON.parse(rawConfig);
|
|
if (!runtimeId && parsedConfig && typeof parsedConfig === 'object' && 'runtime' in parsedConfig) {
|
|
runtimeId = resolveRuntimeNameFromCandidates(parsedConfig.runtime) || null;
|
|
}
|
|
if (!runtimeId) {
|
|
try {
|
|
const defaultsPath = path.join(os.homedir(), '.gsd', 'defaults.json');
|
|
const defaultsParsed = JSON.parse(fs.readFileSync(defaultsPath, 'utf-8'));
|
|
if (defaultsParsed && typeof defaultsParsed === 'object' && 'runtime' in defaultsParsed) {
|
|
runtimeId = resolveRuntimeNameFromCandidates(defaultsParsed.runtime) || null;
|
|
}
|
|
} catch {
|
|
// Absent/unreadable ~/.gsd/defaults.json — no signal, fall through.
|
|
}
|
|
}
|
|
if (!runtimeId) {
|
|
// No explicit signal anywhere confirms this project resolved
|
|
// harness-worktree — degrade to inert rather than guess 'cursor'.
|
|
return 'none';
|
|
}
|
|
|
|
const runtimeEntry = runtimes != null ? runtimes[runtimeId] : null;
|
|
const declared = runtimeEntry?.runtime?.hostIntegration?.dispatch?.isolation ?? null;
|
|
let declaredIsolation = (typeof declared === 'string' && VALID_ISOLATION.has(declared)) ? declared : 'none';
|
|
|
|
if (declaredIsolation === 'harness-worktree' &&
|
|
parsedConfig && typeof parsedConfig === 'object' && parsedConfig.workflow &&
|
|
typeof parsedConfig.workflow === 'object' && parsedConfig.workflow.use_worktrees === false) {
|
|
declaredIsolation = 'none';
|
|
}
|
|
|
|
return declaredIsolation;
|
|
}
|
|
|
|
/**
|
|
* Applicability + isolation verdict for a SINGLE workspace root. Returns
|
|
* `{ action: 'allow' } | { action: 'deny', reason: string }`. Extracted from
|
|
* resolveIsolationDecision (#3045 finding 1) so every root in a multi-root
|
|
* workspace runs the identical check.
|
|
*/
|
|
function evaluateRootIsolation(root, subagentType, { clock = Date, dispatchIds = null, realpath = fs.realpathSync } = {}) {
|
|
const configPath = path.join(root, '.planning', 'config.json');
|
|
let isGsdProject;
|
|
try {
|
|
fs.accessSync(configPath, fs.constants.F_OK);
|
|
isGsdProject = true;
|
|
} catch {
|
|
isGsdProject = false;
|
|
}
|
|
if (!isGsdProject) return { action: 'allow' };
|
|
|
|
let declaredIsolation;
|
|
try {
|
|
// #3045 BLOCKER fix: a fresh sentinel is authoritative for THIS
|
|
// dispatch's actual resolved isolation — see the doc comment above.
|
|
// #3045 SECURITY F2: a fresh sentinel that names a DIFFERENT
|
|
// plan/phase than this dispatch is not applicable to it — fall through
|
|
// to the conservative fallback exactly as a stale sentinel would.
|
|
const sentinel = readSentinel(root, { clock });
|
|
declaredIsolation = (sentinel.present && !sentinel.stale && sentinelAppliesToDispatch(sentinel, dispatchIds))
|
|
? sentinel.isolation
|
|
: resolveFallbackIsolation(root, configPath);
|
|
} catch {
|
|
return {
|
|
action: 'deny',
|
|
reason:
|
|
`GSD subagent isolation guard: could not read or resolve this project's ` +
|
|
`dispatch-isolation configuration ('.planning/config.json' exists under "${root}"). ` +
|
|
`Refusing to allow this subagent to spawn without being able to verify whether ` +
|
|
`isolation is required — a guard that cannot verify must not answer "safe" (#3050). ` +
|
|
`Retry once the project configuration is readable.`,
|
|
};
|
|
}
|
|
|
|
if (declaredIsolation !== 'harness-worktree') return { action: 'allow' };
|
|
|
|
// isConfirmedNonExecutor already excluded "present, non-empty, unrecognized
|
|
// string" above — reaching here means subagentType is either the confirmed
|
|
// executor or missing/malformed (cannot rule it out).
|
|
if (typeof subagentType !== 'string' || subagentType.length === 0) {
|
|
return {
|
|
action: 'deny',
|
|
reason:
|
|
`GSD subagent isolation guard: this project's dispatch isolation resolves to ` +
|
|
`"harness-worktree", but the subagentStart payload for this dispatch carries no usable ` +
|
|
`subagent_type. Refusing to allow it to spawn without being able to confirm whether it ` +
|
|
`is a GSD executor — a guard that cannot verify must not answer "safe" (#3050).`,
|
|
};
|
|
}
|
|
|
|
const evidence = resolveIsolationEvidence(root, { realpath });
|
|
if (evidence.isolated) return { action: 'allow' };
|
|
if (evidence.notApplicable) return { action: 'allow' };
|
|
|
|
if (evidence.cannotDetermine) {
|
|
return {
|
|
action: 'deny',
|
|
reason:
|
|
`GSD subagent isolation guard: this project's dispatch isolation resolves to ` +
|
|
`"harness-worktree", but whether "${root}" is running in an isolated Cursor worktree ` +
|
|
`could not be determined (git did not respond). Refusing to allow subagent_type=` +
|
|
`"${subagentType}" to spawn without being able to verify isolation — a guard that ` +
|
|
`cannot verify must not answer "safe" (#3050). Retry once git is responsive.`,
|
|
};
|
|
}
|
|
|
|
return {
|
|
action: 'deny',
|
|
reason:
|
|
`GSD subagent isolation guard: this project's dispatch isolation resolves to ` +
|
|
`"harness-worktree", but subagent_type="${subagentType}" is about to spawn in "${root}", ` +
|
|
`which is not an isolated Cursor worktree — it would edit the user's primary checkout ` +
|
|
`directly, with no consent and no warning. Start an isolated session first (the ` +
|
|
`"--worktree" CLI flag or the "/worktree" chat command; Cursor manages these worktrees ` +
|
|
`under "~/.cursor/worktrees/") and retry.`,
|
|
};
|
|
}
|
|
|
|
/* istanbul ignore next -- stdin adapter, exercised via spawnSync in tests */
|
|
function main() {
|
|
let raw = '';
|
|
const stdinTimeout = setTimeout(() => {
|
|
process.exit(0);
|
|
}, 10000);
|
|
|
|
process.stdin.setEncoding('utf8');
|
|
process.stdin.on('data', (chunk) => { raw += chunk; });
|
|
process.stdin.on('end', () => {
|
|
clearTimeout(stdinTimeout);
|
|
|
|
let data = null;
|
|
try {
|
|
data = JSON.parse(raw);
|
|
} catch {
|
|
data = null;
|
|
}
|
|
|
|
// Resolve the state-reminder context ONCE, up front, so it can ride along
|
|
// with EITHER outcome below (#3045 MINOR: a deny previously dropped this
|
|
// reminder entirely — process.stdout.write for the deny branch returned
|
|
// before the additional_context block ever ran — instead of preserving it
|
|
// alongside the deny; the subagent still benefits from phase/blocker
|
|
// context even when its dispatch is refused).
|
|
let additionalContext = null;
|
|
try {
|
|
const statePath = resolveStatePath(raw);
|
|
const statePresent = fs.existsSync(statePath);
|
|
additionalContext = statePresent ? MSG_PRESENT : MSG_ABSENT;
|
|
} catch {
|
|
additionalContext = null;
|
|
}
|
|
|
|
if (data && typeof data === 'object') {
|
|
let decision = { action: 'allow' };
|
|
try {
|
|
decision = resolveIsolationDecision(data);
|
|
} catch {
|
|
// Defense in depth only: every verify-and-deny path above has its own
|
|
// explicit try/catch that resolves to a deny with a distinct reason.
|
|
// Anything reaching here is an unexpected failure outside those paths
|
|
// (e.g. malformed workspace_roots entries) — never crash the hook.
|
|
decision = { action: 'allow' };
|
|
}
|
|
if (decision.action === 'deny') {
|
|
const out = { permission: 'deny', user_message: decision.reason };
|
|
if (additionalContext !== null) out.additional_context = additionalContext;
|
|
process.stdout.write(JSON.stringify(out));
|
|
return;
|
|
}
|
|
}
|
|
|
|
process.stdout.write(JSON.stringify(additionalContext !== null ? { additional_context: additionalContext } : {}));
|
|
});
|
|
}
|
|
|
|
if (require.main === module) {
|
|
main();
|
|
}
|
|
|
|
// #3045 MAJOR ("clock seam is dead code" fix): exported so tests can
|
|
// `require()` this module and inject a `clock` (`{now(): number}`) directly
|
|
// per the repo's clock-seam convention, instead of racing real `Date.now()`
|
|
// across a spawned subprocess boundary.
|
|
module.exports = {
|
|
resolveIsolationDecision,
|
|
evaluateRootIsolation,
|
|
resolveFallbackIsolation,
|
|
resolveIsolationEvidence,
|
|
getWorkspaceRoots,
|
|
};
|