* feat(#3045): deny an executor dispatch that drops its isolation flag Every isolation gate already resolved correctly. The resolved value then reached the executor through a prose instruction telling the model to substitute it into a call the model composes itself, and nothing verified the substitution. When it was dropped, the executor edited and committed in the user's primary checkout with no consent and no warning. A prose backstop would be the same class of artifact as the defect, so this is a shipped PreToolUse hook on the Agent tool. It fires at the instant of the call rather than being read once at the top of a workflow, which is the only placement the model cannot skip. The guard is inert unless it can positively establish that this is a GSD project, that the project resolves to harness isolation, and that the dispatch targets an executor. A non-GSD repo has no invariant to enforce. Where it cannot read the configuration at all, it denies rather than assuming, with its own reason -- a guard that cannot verify must not answer safe. A malformed payload allows rather than throwing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3045): extend the isolation guard to Cursor Cursor is the second of only two runtimes that resolve harness isolation, so shipping the guard for Claude alone left half the exposed surface unguarded while the changeset implied it was covered. The two runtimes fail differently. On Claude the harness flag is a per-dispatch kwarg the model must copy into a call it composes, and the defect is that it can be dropped. On Cursor the flag is --worktree, which applies to the whole session, and the subagent-start payload carries no isolation field at all. There is no flag to check, so the guard verifies the effective state instead: whether the workspace is genuinely running outside the user's primary checkout. That is a stronger check than the Claude one because it tests reality rather than intent, and it is commented so nobody later rewrites it into a flag check. Isolation is established two ways, either sufficient: the workspace resolves to a linked git worktree, or it sits under the worktree root Cursor manages. The second matters because a directory Cursor placed there is a legitimate isolated session even before it becomes a distinct git worktree, where linkage alone would report no repository. Detecting linkage required a new primitive rather than the existing context resolver. That resolver short-circuits on finding a local .planning directory before it ever compares the git directory to the common one -- and an isolation worktree normally has its own checked-out .planning. Reusing it would have read a correctly isolated session as unisolated and denied it, which is the failure direction that gets a guard switched off. The comparison is now its own shortcut-free function that the resolver delegates to after its own shortcut, so existing behavior is unchanged, and the case that would have broken is pinned. The subagent type is checked before any configuration is read, so an unreadable config cannot deny a dispatch this guard would never have enforced against. The input-schema comment on the Cursor hook documented only the fields common to every event and omitted the ones specific to this one. That omission cost a halt during this work; it now documents both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): enforce the resolved dispatch decision, not the host capability The guard keyed on the registry's dispatch.isolation, which says only that a runtime is CAPABLE of harness worktrees. The decision that actually governs a dispatch is the one the workflow resolves after gating, and that legitimately comes out as sequential in three documented cases: a project setting use_worktrees false, a per-plan submodule intersection, and the base-check auto-degrade. The workflow tells the model to omit the flag in exactly those cases, and the guard was denying every one of them. The third case matters most. The preceding fix made the base-check degrade on git timeouts and a missing git binary, where it had previously answered "safe". That correction is right, and it means a transient hang now degrades to sequential far more often than before -- so the two changes composed into a trap where the workflow behaved exactly as designed and the guard blocked it. The workflow already resolves isolation in shell, deterministically, which is what makes it a trustworthy source in a way the model-authored call is not. It now records that resolved value through a dedicated verb, and both guards read it first. A fresh record is authoritative, so sequential dispatches pass untouched. Absent or stale, the guards fall back to the capability check combined with the project's use_worktrees setting, which still covers the case that never reaches the workflow. Also widened the matcher to accept Task alongside Agent, since a host that names the tool Task would otherwise leave the guard silently inert while implying coverage; stopped assuming Claude when no runtime is declared, which is the shipped default and would have demanded a Claude-only argument elsewhere; and made a non-git project inert rather than denied, since advising a worktree session is not actionable without a repository. The original diagnosis never modeled sequential mode as legitimate. That omission is what let this through, and it is now recorded there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): record at resolution and bind the record to its dispatch Two independent reviews converged on the same failure: the guard was fail-open in a default install, so it did not catch the defect it exists to catch. A shipped project carries no runtime key, which made "runtime not confidently known" the common case rather than a corner one. A record asserting that isolation was required but carrying no flag then fell through to a capability lookup that answered "none", and the dispatch was allowed. The flag itself only arrived from a second shell block -- the same block a model dropping the argument would also skip. A test had pinned that behavior as intended. The record is now written by the resolver, as an unavoidable consequence of asking for the value, rather than by a step the model is told in prose to go and run. A guard against a prose-carried value cannot itself depend on prose. Mode, flag and identifiers are written together and atomically, so the flagless window is gone, and a record asserting isolation with no resolvable flag now denies instead of degrading. Runtime is also resolved from the installer's own recorded default, which makes confident resolution the normal case. The per-plan submodule gate degrades after the phase-level decision and never re-recorded, so a plan that legitimately ran sequentially was denied against a still-fresh phase record. It now records its own, scoped to the plan. A record also authorized any dispatch for four hours. One phase degrading to sequential could silently license an unisolated dispatch in the next. Records now carry phase and plan, the guards require them to match, and the window is minutes rather than hours -- the resolver rewrites it before every dispatch, so a long window bought nothing and only widened the hole. The flag validator rejected any value beginning with two dashes, which is exactly the form Cursor and Windsurf declare, so their real value could never have been stored. Writer and reader also derived the record path differently and diverged inside a linked worktree without local planning state. The predictable path remains a way to silence the control without leaving a trace in the diff. It grants no access an agent with shell does not already have, so it is documented as accepted rather than redesigned around. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): correct the staleness boundary and unmask a vacuous parity test The remote runner returned twenty failures. One was a real production defect the boundary case existed to catch: a record whose age exactly equalled the staleness window was treated as fresh, so it stayed authoritative for one tick past its own expiry. Freshness is now strictly inside the window. The parity test meant to stop the two guards' executor lists from drifting could never have failed. Its project fixture was a bare directory rather than a repository, so the non-git inert branch answered before the executor list was ever consulted. It asserted agreement it never actually measured. The fixture is now a real repository, like every sibling in the file. A test also asserted that Windsurf declares the worktree flag. It does not -- Windsurf resolves to no isolation by design, having no named concurrent dispatch to isolate. The test claimed a registry fact that was never true, and a comment in the resolver repeated it. Both corrected, and the test now proves what it should have all along: that the parser accepts any bare flag value, rather than one runtime's supposed value. The new guard was missing from the bundled-hook whitelist, which is the surface that decides what actually ships, and the per-plan gate had gained calls to the launcher without the preamble those calls require. The changeset carried parenthetical product descriptions the purity rule forbids. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3045): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3045): make the guard tests hold on Windows Two tests redirect HOME to control where the installer-persisted runtime default is read from. Node resolves the home directory from USERPROFILE on Windows and never consults HOME, so both silently read the real runner profile, found no recorded runtime, and asserted against a project the hook had not recognised. The production code was already correct in asking the platform rather than the variable; only the tests were wrong to assume one variable answers everywhere. The helpers now mirror the override onto both. The symlink spoofing test also created a directory symlink unconditionally, which needs elevated privileges on Windows. It survived on this runner, but it would fail on any host without them, so the creation is now attempted and the test skips explicitly when it cannot be done -- a bare return would have counted as a pass and hidden the gap. Skipping alone would have left the platform uncovered, so the behaviour it proves is now also driven in-process through an injected realpath, following the seam already used for the clock. That case no longer depends on privileges at all, and the end-to-end test keeps its original assertions wherever symlinks work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
269 lines
13 KiB
JavaScript
269 lines
13 KiB
JavaScript
'use strict';
|
|
// hooks/lib/isolation-sentinel.js — shared sentinel reader for the #3045
|
|
// agent-dispatch isolation guards (hooks/gsd-agent-isolation-guard.js,
|
|
// hooks/gsd-cursor-subagent-start.js).
|
|
//
|
|
// #3045 BLOCKER: the guards previously keyed enforcement on the capability
|
|
// REGISTRY's `dispatch.isolation` ("this host CAN isolate"), not the
|
|
// workflow's resolved per-dispatch ISOLATION ("this dispatch SHOULD be
|
|
// isolated"). Sequential ISOLATION=none legitimately happens on a
|
|
// harness-worktree-capable host — project-level `workflow.use_worktrees:
|
|
// false`, the #2474 per-plan submodule degrade, and the #683/#3060
|
|
// base-check auto-degrade all resolve to `none` and are NOT bugs
|
|
// (gsd-core/workflows/execute-phase/steps/executor-isolation-dispatch.md:
|
|
// "Sequential mode … Omit isolation=\"worktree\" from the Agent call").
|
|
//
|
|
// The workflow already computes ISOLATION deterministically in shell before
|
|
// any executor dispatch. CORE REDESIGN (two-review follow-up): the PRIMARY
|
|
// write path is now `dispatch-isolation` itself (gsd-tools.cjs
|
|
// routeDispatchIsolation) — it persists mode + harnessFlag + phase/plan
|
|
// identifiers to the sentinel as an unconditional side effect of resolving
|
|
// them, since the workflow must call it to learn ISOLATION at all.
|
|
// `record-dispatch-isolation` (routeRecordDispatchIsolation) remains as an
|
|
// explicit fallback for the per-plan submodule-degrade override and shares
|
|
// the exact same atomic-write implementation.
|
|
//
|
|
// Sentinel path: `<cwd>/.gsd/dispatch-isolation-sentinel.json`. `.gsd` is
|
|
// gitignored (root `.gitignore`'s bare `.gsd` entry matches at any depth),
|
|
// and lives inside the checkout that ran the workflow — concurrent worktrees
|
|
// of the same repo are separate directories on disk, so each gets its own
|
|
// sentinel with no cross-worktree collision, no lock file needed.
|
|
//
|
|
// #3045 SECURITY F3 (LOW, accepted risk — documented, not redesigned for):
|
|
// the sentinel lives at a fixed, predictable, gitignored path inside the
|
|
// PRIMARY checkout. An agent already running inside a legitimately isolated
|
|
// worktree (i.e. one that has shell/filesystem access to the primary
|
|
// checkout's absolute path, which every worktree-isolated agent does by
|
|
// construction) could write a fabricated fresh `{isolation:"none"}` sentinel
|
|
// into the primary checkout to silence this guard for that checkout's next
|
|
// dispatch.
|
|
// - This grants NO NEW write privilege: that same agent could already edit
|
|
// any file in the primary checkout directly (worktree isolation protects
|
|
// the primary checkout from an UNTRUSTED executor's own dispatch, not
|
|
// from a trusted orchestrator process choosing to reach across).
|
|
// - It DOES add an evasion path: silencing the guard via the sentinel
|
|
// leaves no trace in `git status` (`.gsd/` is gitignored), where a direct
|
|
// edit to a tracked file would.
|
|
// Accepted because the threat model this guard defends against is an
|
|
// UNCONSENTED, UNVERIFIED dispatch — not a deliberately adversarial
|
|
// orchestrator process, which already has strictly more direct means to
|
|
// cause harm than forging this one file. If that threat model changes (e.g.
|
|
// executors become mutually distrusting / sandboxed from the orchestrator's
|
|
// own filesystem), the hardening path is a SESSION-KEYED sentinel written
|
|
// outside any worktree the executor can reach (e.g. under the harness's own
|
|
// config dir, keyed by a session/run id neither the executor nor a forged
|
|
// file can predict) rather than a path derivable from `cwd`.
|
|
|
|
const fs = require('fs');
|
|
const path = require('path');
|
|
|
|
// Isolation modes ADR-1239 declares (mirrors gsd-tools.cjs
|
|
// routeDispatchIsolation / routeRecordDispatchIsolation).
|
|
const VALID_ISOLATION = new Set(['harness-worktree', 'orchestrator-worktree', 'none']);
|
|
|
|
const SENTINEL_RELATIVE_PATH = path.join('.gsd', 'dispatch-isolation-sentinel.json');
|
|
|
|
// #3045 SECURITY F2 fix: how long a written sentinel is trusted as "this
|
|
// dispatch's decision" before a reader falls back to the conservative
|
|
// registry+config check.
|
|
//
|
|
// Previously 4h, on the theory that a slow multi-wave phase execution could
|
|
// span well over an hour. That reasoning no longer holds: the #3045 CORE
|
|
// REDESIGN makes `dispatch-isolation` (gsd-tools.cjs routeDispatchIsolation)
|
|
// the sole write path, called as a side effect of resolving ISOLATION — and
|
|
// the workflow now re-resolves (and therefore re-records) immediately before
|
|
// EVERY plan's dispatch, at the per-plan worktree gate
|
|
// (execute-phase/steps/per-plan-worktree-gate.md), not once per phase. A long
|
|
// trust window no longer buys the workflow anything and only widens the
|
|
// window in which a stale sentinel from an EARLIER, DIFFERENT phase/plan
|
|
// (e.g. one that legitimately degraded to `none`) could be misread as
|
|
// authorizing a LATER dispatch that never got its own fresh record (a model
|
|
// skipping the kwarg on a harness-worktree phase while a same-session
|
|
// same-project stale `none` from a prior phase is still "fresh" by the old
|
|
// 4h window).
|
|
//
|
|
// 10 minutes generously covers the real latency between a per-plan gate's
|
|
// resolve call and that same plan's `Agent()`/`Task()` dispatch (worktree
|
|
// creation, orphan-worktree sweep, base-check, prompt composition) — all
|
|
// bounded, sub-minute operations per their own repo-mandated subprocess
|
|
// timeouts — while being far too short for a sentinel to survive into a
|
|
// later, unrelated phase.
|
|
const SENTINEL_STALE_MS = 10 * 60 * 1000; // 10 minutes
|
|
|
|
function sentinelPath(cwd) {
|
|
return path.join(cwd, SENTINEL_RELATIVE_PATH);
|
|
}
|
|
|
|
/**
|
|
* Resolve the project root a sentinel should be read from/written to, using
|
|
* the SAME derivation gsd-tools.cjs's dispatcher applies to every `--cwd`
|
|
* before invoking a route handler: `findProjectRoot(resolveMainWorktreeCwd(cwd))`
|
|
* (gsd-core/bin/gsd-tools.cjs main(), :3506/:3603 — `record-dispatch-isolation`
|
|
* and `dispatch-isolation` are not in SKIP_ROOT_RESOLUTION, so every write
|
|
* goes through both steps).
|
|
*
|
|
* #3045 MINOR fix: the guard hooks previously read the sentinel from the raw
|
|
* `data.cwd` / `workspace_roots[i]` the harness reports, with NO equivalent
|
|
* resolution. For a linked worktree that does not itself own a `.planning/`
|
|
* (the common shape — `.planning/` lives in the main worktree only), the
|
|
* writer resolves up to the MAIN worktree and writes there, while the reader
|
|
* checked `.planning/config.json` at the raw (unresolved) linked-worktree
|
|
* path, found nothing, and silently treated the dispatch as "not a GSD
|
|
* project" (inert allow) — the guard was reading a sentinel that was never
|
|
* written where it looked. Deriving both sides through this one function
|
|
* closes that divergence.
|
|
*
|
|
* `findProjectRoot`/`resolveWorktreeRoot` are read from the sibling
|
|
* `gsd-core/bin/lib/*.cjs` modules staged alongside these hooks at install
|
|
* time (same pattern the guard hooks already use for
|
|
* capability-registry.cjs/runtime-name-policy.cjs) — two directories up from
|
|
* `hooks/lib/` (`hooks/lib/isolation-sentinel.js` -> `hooks/` -> repo/install
|
|
* root -> `gsd-core/bin/lib/`), mirroring the one-directory-up requires the
|
|
* top-level `hooks/*.js` guard scripts already use successfully.
|
|
*
|
|
* Never throws; any resolution failure (module missing, git unavailable,
|
|
* git timeout) degrades to the raw `cwd` unchanged — the caller's existing
|
|
* "sentinel absent -> conservative fallback" path already covers that safely.
|
|
*/
|
|
function resolveSentinelRoot(cwd) {
|
|
try {
|
|
if (fs.existsSync(path.join(cwd, '.planning'))) {
|
|
return cwd;
|
|
}
|
|
const { resolveWorktreeRoot } = require('../../gsd-core/bin/lib/worktree-safety.cjs');
|
|
const { root } = resolveWorktreeRoot(cwd);
|
|
const { findProjectRoot } = require('../../gsd-core/bin/lib/project-root.cjs');
|
|
return findProjectRoot(root);
|
|
} catch {
|
|
return cwd;
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Read and validate the dispatch-isolation sentinel for `cwd`. Never throws.
|
|
* `cwd` is resolved through `resolveSentinelRoot` first (#3045 MINOR — see
|
|
* its doc comment), so callers may pass the raw, unresolved dispatch cwd
|
|
* directly.
|
|
*
|
|
* Returns one of:
|
|
* { present: false }
|
|
* { present: true, stale: true, malformed: true }
|
|
* { present: true, stale: true, malformed: false, isolation, harnessFlag, phase, plan, writtenAt }
|
|
* { present: true, stale: false, malformed: false, isolation, harnessFlag, phase, plan, writtenAt }
|
|
*
|
|
* A malformed/unparseable sentinel is treated as STALE, never fatal — the
|
|
* caller's conservative fallback path covers both "absent" and "stale"
|
|
* identically.
|
|
*
|
|
* `clock` is injectable (`{ now(): number }`, defaults to the real `Date`)
|
|
* per the repo's clock-seam convention, so staleness is testable without
|
|
* asserting on wall-clock time.
|
|
*/
|
|
function readSentinel(cwd, { clock = Date } = {}) {
|
|
const root = resolveSentinelRoot(cwd);
|
|
let raw;
|
|
try {
|
|
raw = fs.readFileSync(sentinelPath(root), 'utf-8');
|
|
} catch {
|
|
return { present: false };
|
|
}
|
|
|
|
let parsed;
|
|
try {
|
|
parsed = JSON.parse(raw);
|
|
} catch {
|
|
return { present: true, stale: true, malformed: true };
|
|
}
|
|
|
|
if (
|
|
!parsed || typeof parsed !== 'object' ||
|
|
!VALID_ISOLATION.has(parsed.isolation) ||
|
|
typeof parsed.written_at !== 'number' || !Number.isFinite(parsed.written_at)
|
|
) {
|
|
return { present: true, stale: true, malformed: true };
|
|
}
|
|
|
|
const harnessFlag = typeof parsed.harness_flag === 'string' && parsed.harness_flag.length > 0
|
|
? parsed.harness_flag
|
|
: null;
|
|
const phase = typeof parsed.phase === 'string' && parsed.phase.length > 0 ? parsed.phase : null;
|
|
// #3045 SECURITY F2: `plan` was not previously part of the sentinel shape.
|
|
// Recorded so a phase-level-only sentinel (plan: null) is distinguishable
|
|
// from a plan-scoped one — see the guards' dispatch-matching logic, which
|
|
// treats a plan/phase MISMATCH (both sides present and disagreeing) as "no
|
|
// applicable sentinel", not an allow.
|
|
const plan = typeof parsed.plan === 'string' && parsed.plan.length > 0 ? parsed.plan : null;
|
|
|
|
const now = clock.now();
|
|
const age = now - parsed.written_at;
|
|
// Negative age beyond a small tolerance means the sentinel claims to be
|
|
// written in the future — never trust it, but still surface the parsed
|
|
// fields so callers can log an actionable reason.
|
|
const stale = age >= SENTINEL_STALE_MS || age < -5000;
|
|
|
|
return {
|
|
present: true,
|
|
stale,
|
|
malformed: false,
|
|
isolation: parsed.isolation,
|
|
harnessFlag,
|
|
phase,
|
|
plan,
|
|
writtenAt: parsed.written_at,
|
|
};
|
|
}
|
|
|
|
/**
|
|
* #3045 SECURITY F2: extract the `{plan, phase}` a specific Agent()/Task()
|
|
* dispatch is FOR, from the one place that data is reliably embedded today —
|
|
* the dispatch prompt/description text (`execute-phase.md`'s Agent() block
|
|
* uses the literal shape `description="Execute plan {plan_number} of phase
|
|
* {phase_number}"`, and the prompt body's `<objective>` repeats "Execute plan
|
|
* {plan_number} of phase {phase_number}-{phase_name}." verbatim — the SAME
|
|
* text the orchestrator-worktree EXECUTOR_PROMPT template and Cursor's `task`
|
|
* field carry, since Cursor dispatches the same prompt content). There is no
|
|
* structured per-dispatch kwarg carrying plan/phase identifiers today (#3045
|
|
* would need a larger dispatch-protocol change to add one) — this is
|
|
* therefore a best-effort, NOT a guaranteed, extraction: a dispatch whose
|
|
* text doesn't match the expected shape returns `{ plan: null, phase: null }`
|
|
* and the caller must NOT treat that as a mismatch (see
|
|
* `sentinelAppliesToDispatch`).
|
|
*/
|
|
function extractDispatchIdentifiers(text) {
|
|
if (typeof text !== 'string' || text.length === 0) return { plan: null, phase: null };
|
|
const m = /execute\s+plan\s+(\S+)\s+of\s+phase\s+(\S+)/i.exec(text);
|
|
if (!m) return { plan: null, phase: null };
|
|
return { plan: m[1], phase: m[2] };
|
|
}
|
|
|
|
/**
|
|
* #3045 SECURITY F2: does a fresh, non-malformed sentinel apply to THIS
|
|
* dispatch? `dispatchIds` is the `{plan, phase}` extracted from the
|
|
* dispatch's own text via `extractDispatchIdentifiers` (or manually supplied
|
|
* by a caller with a more reliable source).
|
|
*
|
|
* Returns false (mismatch — "no applicable sentinel") ONLY when both sides
|
|
* carry a value for the SAME identifier and they disagree. Any side missing
|
|
* a value (sentinel predates this fix, or the dispatch text didn't match the
|
|
* expected shape) is treated as "cannot compare" and does NOT itself produce
|
|
* a mismatch — this stays a defense-in-depth narrowing of an otherwise-fresh
|
|
* sentinel's applicability, not a new fail-open/fail-closed axis on its own.
|
|
*/
|
|
function sentinelAppliesToDispatch(sentinel, dispatchIds) {
|
|
if (!sentinel || !dispatchIds) return true;
|
|
if (sentinel.phase && dispatchIds.phase && sentinel.phase !== dispatchIds.phase) return false;
|
|
if (sentinel.plan && dispatchIds.plan && sentinel.plan !== dispatchIds.plan) return false;
|
|
return true;
|
|
}
|
|
|
|
module.exports = {
|
|
VALID_ISOLATION,
|
|
SENTINEL_RELATIVE_PATH,
|
|
SENTINEL_STALE_MS,
|
|
sentinelPath,
|
|
resolveSentinelRoot,
|
|
readSentinel,
|
|
extractDispatchIdentifiers,
|
|
sentinelAppliesToDispatch,
|
|
};
|