9faacc0c153a88f939ef07ced74d56dc468f0708
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8f75e27554 |
fix(#3045): fail closed when an executor dispatch drops its resolved isolation (#3069)
* feat(#3045): deny an executor dispatch that drops its isolation flag Every isolation gate already resolved correctly. The resolved value then reached the executor through a prose instruction telling the model to substitute it into a call the model composes itself, and nothing verified the substitution. When it was dropped, the executor edited and committed in the user's primary checkout with no consent and no warning. A prose backstop would be the same class of artifact as the defect, so this is a shipped PreToolUse hook on the Agent tool. It fires at the instant of the call rather than being read once at the top of a workflow, which is the only placement the model cannot skip. The guard is inert unless it can positively establish that this is a GSD project, that the project resolves to harness isolation, and that the dispatch targets an executor. A non-GSD repo has no invariant to enforce. Where it cannot read the configuration at all, it denies rather than assuming, with its own reason -- a guard that cannot verify must not answer safe. A malformed payload allows rather than throwing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3045): extend the isolation guard to Cursor Cursor is the second of only two runtimes that resolve harness isolation, so shipping the guard for Claude alone left half the exposed surface unguarded while the changeset implied it was covered. The two runtimes fail differently. On Claude the harness flag is a per-dispatch kwarg the model must copy into a call it composes, and the defect is that it can be dropped. On Cursor the flag is --worktree, which applies to the whole session, and the subagent-start payload carries no isolation field at all. There is no flag to check, so the guard verifies the effective state instead: whether the workspace is genuinely running outside the user's primary checkout. That is a stronger check than the Claude one because it tests reality rather than intent, and it is commented so nobody later rewrites it into a flag check. Isolation is established two ways, either sufficient: the workspace resolves to a linked git worktree, or it sits under the worktree root Cursor manages. The second matters because a directory Cursor placed there is a legitimate isolated session even before it becomes a distinct git worktree, where linkage alone would report no repository. Detecting linkage required a new primitive rather than the existing context resolver. That resolver short-circuits on finding a local .planning directory before it ever compares the git directory to the common one -- and an isolation worktree normally has its own checked-out .planning. Reusing it would have read a correctly isolated session as unisolated and denied it, which is the failure direction that gets a guard switched off. The comparison is now its own shortcut-free function that the resolver delegates to after its own shortcut, so existing behavior is unchanged, and the case that would have broken is pinned. The subagent type is checked before any configuration is read, so an unreadable config cannot deny a dispatch this guard would never have enforced against. The input-schema comment on the Cursor hook documented only the fields common to every event and omitted the ones specific to this one. That omission cost a halt during this work; it now documents both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): enforce the resolved dispatch decision, not the host capability The guard keyed on the registry's dispatch.isolation, which says only that a runtime is CAPABLE of harness worktrees. The decision that actually governs a dispatch is the one the workflow resolves after gating, and that legitimately comes out as sequential in three documented cases: a project setting use_worktrees false, a per-plan submodule intersection, and the base-check auto-degrade. The workflow tells the model to omit the flag in exactly those cases, and the guard was denying every one of them. The third case matters most. The preceding fix made the base-check degrade on git timeouts and a missing git binary, where it had previously answered "safe". That correction is right, and it means a transient hang now degrades to sequential far more often than before -- so the two changes composed into a trap where the workflow behaved exactly as designed and the guard blocked it. The workflow already resolves isolation in shell, deterministically, which is what makes it a trustworthy source in a way the model-authored call is not. It now records that resolved value through a dedicated verb, and both guards read it first. A fresh record is authoritative, so sequential dispatches pass untouched. Absent or stale, the guards fall back to the capability check combined with the project's use_worktrees setting, which still covers the case that never reaches the workflow. Also widened the matcher to accept Task alongside Agent, since a host that names the tool Task would otherwise leave the guard silently inert while implying coverage; stopped assuming Claude when no runtime is declared, which is the shipped default and would have demanded a Claude-only argument elsewhere; and made a non-git project inert rather than denied, since advising a worktree session is not actionable without a repository. The original diagnosis never modeled sequential mode as legitimate. That omission is what let this through, and it is now recorded there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): record at resolution and bind the record to its dispatch Two independent reviews converged on the same failure: the guard was fail-open in a default install, so it did not catch the defect it exists to catch. A shipped project carries no runtime key, which made "runtime not confidently known" the common case rather than a corner one. A record asserting that isolation was required but carrying no flag then fell through to a capability lookup that answered "none", and the dispatch was allowed. The flag itself only arrived from a second shell block -- the same block a model dropping the argument would also skip. A test had pinned that behavior as intended. The record is now written by the resolver, as an unavoidable consequence of asking for the value, rather than by a step the model is told in prose to go and run. A guard against a prose-carried value cannot itself depend on prose. Mode, flag and identifiers are written together and atomically, so the flagless window is gone, and a record asserting isolation with no resolvable flag now denies instead of degrading. Runtime is also resolved from the installer's own recorded default, which makes confident resolution the normal case. The per-plan submodule gate degrades after the phase-level decision and never re-recorded, so a plan that legitimately ran sequentially was denied against a still-fresh phase record. It now records its own, scoped to the plan. A record also authorized any dispatch for four hours. One phase degrading to sequential could silently license an unisolated dispatch in the next. Records now carry phase and plan, the guards require them to match, and the window is minutes rather than hours -- the resolver rewrites it before every dispatch, so a long window bought nothing and only widened the hole. The flag validator rejected any value beginning with two dashes, which is exactly the form Cursor and Windsurf declare, so their real value could never have been stored. Writer and reader also derived the record path differently and diverged inside a linked worktree without local planning state. The predictable path remains a way to silence the control without leaving a trace in the diff. It grants no access an agent with shell does not already have, so it is documented as accepted rather than redesigned around. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): correct the staleness boundary and unmask a vacuous parity test The remote runner returned twenty failures. One was a real production defect the boundary case existed to catch: a record whose age exactly equalled the staleness window was treated as fresh, so it stayed authoritative for one tick past its own expiry. Freshness is now strictly inside the window. The parity test meant to stop the two guards' executor lists from drifting could never have failed. Its project fixture was a bare directory rather than a repository, so the non-git inert branch answered before the executor list was ever consulted. It asserted agreement it never actually measured. The fixture is now a real repository, like every sibling in the file. A test also asserted that Windsurf declares the worktree flag. It does not -- Windsurf resolves to no isolation by design, having no named concurrent dispatch to isolate. The test claimed a registry fact that was never true, and a comment in the resolver repeated it. Both corrected, and the test now proves what it should have all along: that the parser accepts any bare flag value, rather than one runtime's supposed value. The new guard was missing from the bundled-hook whitelist, which is the surface that decides what actually ships, and the per-plan gate had gained calls to the launcher without the preamble those calls require. The changeset carried parenthetical product descriptions the purity rule forbids. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3045): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3045): make the guard tests hold on Windows Two tests redirect HOME to control where the installer-persisted runtime default is read from. Node resolves the home directory from USERPROFILE on Windows and never consults HOME, so both silently read the real runner profile, found no recorded runtime, and asserted against a project the hook had not recognised. The production code was already correct in asking the platform rather than the variable; only the tests were wrong to assume one variable answers everywhere. The helpers now mirror the override onto both. The symlink spoofing test also created a directory symlink unconditionally, which needs elevated privileges on Windows. It survived on this runner, but it would fail on any host without them, so the creation is now attempted and the test skips explicitly when it cannot be done -- a bare return would have counted as a pass and hidden the gap. Skipping alone would have left the platform uncovered, so the behaviour it proves is now also driven in-process through an injected realpath, following the seam already used for the clock. That case no longer depends on privileges at all, and the end-to-end test keeps its original assertions wherever symlinks work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c7c2fe3c2b |
fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd (#2680)
* fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd gsd-cursor-session-start.js and gsd-cursor-stop.js both resolved the project as path.join(process.cwd(), '.planning', 'STATE.md'). Under the cursor-agent CLI, hooks are invoked with cwd set to the Cursor config dir (~/.cursor), not the workspace — so the lookup always missed. sessionStart could only ever emit the "no .planning/ workflow found" nudge and stop's verify-work reminder could never fire, even with .planning/STATE.md sitting in the workspace. Slash commands were unaffected, which is why only the hook layer looked blind. Both hooks already buffered stdin into `raw` and never parsed it; the payload's workspace_roots carries the real path. Multi-root was left open in the report ("first root vs any root"). Resolved forward: prefer the first root that actually carries .planning/STATE.md, so a workspace whose GSD project is not the first root still resolves — strictly better than first-root-only and identical to it in the single-root CLI case. Falls back to roots[0], then to cwd, keeping IDE behavior unchanged if the IDE ever invokes hooks from the workspace. The resolver is duplicated verbatim across the two scripts rather than shared via hooks/lib/: these hooks ship standalone, and a new hooks/lib/ file must be registered in the GENERATED installer's GSD_HOOK_LIB_FILES allowlist — the installer-omits-shipped-file class that yields MODULE_NOT_FOUND at runtime. Per CLAUDE.md "Generative Fix Divergence", the duplication carries a parity assertion so the copies cannot drift. Failing-first, demonstrated by direct invocation with cwd != workspace: pre-fix sessionStart -> "no .planning/ workflow found" stop -> {} post-fix sessionStart -> ".planning/STATE.md is present" stop -> reminder tests/fix-2587-cursor-hook-workspace-roots.test.cjs spawns the real scripts as child processes with a cwd lacking .planning/ and workspace_roots pointing at it. Boundary coverage on the roots array (0 / 1 / 2 entries), plus malformed-JSON fail-open, junk-entry filtering, the parity assertion, and a guard that neither script resolves .planning from cwd again. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2587): extend workspace_roots fix to subagentStart; keep cwd a candidate Three findings from the isolated review, all fixed. 1. MISSED SITE (high). gsd-cursor-subagent-start.js carried the identical defect at line 43 — its own header documents workspace_roots in the input schema, but it resolved .planning/ from process.cwd() anyway. Under the cursor-agent CLI that meant every Cursor subagent (planner, executor, verifier) started with "no .planning/ workflow found" and no phase context. The report named only sessionStart and stop; the defect class was wider. Verified pre-fix vs post-fix by direct invocation with cwd != workspace. 2. SEMANTIC NARROWING (medium). The first cut searched only workspace_roots and fell back to cwd solely when the array was EMPTY. So when roots were supplied but none carried .planning/ while cwd did, the hook reported absent — where the pre-fix code, which always used cwd, reported present. That contradicted the fallback's own stated intent of preserving IDE behavior. cwd is now a CANDIDATE in the search (`[...roots, process.cwd()]`), so the fix is a strict superset of both the old behavior and the CLI fix, never a narrowing. 3. STALE GOLDEN FIXTURES (high, would have failed CI). The golden-install-parity fixtures store a content hash per installed file; these three hooks appear in 13 of the 19 runtime fixtures. Regenerated via `npm run gen:golden` — the diff is exactly the three hook hashes in exactly those 13 runtimes. Tests extended: subagentStart resolution via workspace_roots; the stop hook's absent branch (previously only session-start's was covered); an explicit regression guard that a project at cwd is still found when roots miss; parity now asserts all THREE copies byte-identical; and the cwd guard sweeps the whole RESOLVING_HOOKS list so a future hook in this family cannot be left on cwd. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * refactor(#2587): extract cursor workspace resolution to a shared hooks/lib module The duplicate-plus-parity-test approach was the wrong call. The reported issue named two hooks; a third (subagentStart) had the identical defect. That is the signature of a systemic problem, and three copies of a resolver guarded by a parity assertion is a divergence risk maintained by hand rather than a fix. hooks/lib/cursor-workspace.js is now the single implementation. All three Cursor hooks require it; none defines a local copy. Divergence is prevented structurally instead of by asserting three copies stay byte-identical. The reason duplication looked necessary was real, and is fixed properly here rather than worked around: Cursor sets hostBehaviors.skipSharedHooksInstall (#2089), so it never reaches the installer's bulk hooks/lib copy — it was the ONE runtime shipping these hooks WITHOUT hooks/lib (verified against all 19 golden fixtures: cursor had the hook scripts, no lib). A naive require would have thrown MODULE_NOT_FOUND at load, BEFORE each hook's own try/catch, wedging every session on precisely the runtime this bug is about. writeCursorHooksJson (src/runtime-hooks-surface.cts) now stages the hooks/lib helpers the staged scripts actually require, discovered by scanning their require('./lib/…') calls rather than a hardcoded name — so a future helper cannot be silently omitted. This is narrower than flipping skipSharedHooksInstall, which would wrongly pull in every shared hook. cursor-workspace.js is also added to GSD_HOOK_LIB_FILES so uninstall and the manifest manage it for the runtimes that do receive hooks/lib. Verified against a REAL install (runMinimalInstall, cursor/global): the helper is staged, and all three INSTALLED hooks resolve the workspace end-to-end from a cwd that is not the project. Also closes the review gap that the stop hook was excluded from the cwd-candidate regression loop — it now sweeps RESOLVING_HOOKS. The byte-parity test is replaced by a structural guard (every hook requires the shared module, none redefines it) plus a new install test asserting the helper is staged and the installed hook actually loads against it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2587): fail loud on a missing hook lib source; drop unsubstituted version marker Two findings from the installer-focused review. H1 — the staging step's `if (!fs.existsSync(libSrc)) continue;` silently defeated the very guarantee it was added for. Reproduced: delete hooks/lib/cursor-workspace.js from source, run the cursor install — it exits 0, prints "Done!", and ships the three hook scripts with an EMPTY hooks/lib/. The installed hook then throws `Cannot find module './lib/cursor-workspace.js'` at load, before its own try/catch, wedging every session — and nothing surfaces until a user hits it. The scan protected against a required-but-UNLISTED helper while leaving required-but-MISSING wide open (typo, bad rebase, an accidental delete). It now throws: a missing helper source is a packaging bug and aborts the install. M1 — hooks/lib/cursor-workspace.js carried a `gsd-hook-version: <placeholder>` marker that NOTHING substitutes: copyLibDir stamps .sh files only, and writeCursorHooksJson's staging applies just the colon-to-dash rewrite. Verified the literal was reaching disk on both the bulk (--claude) and Cursor (--cursor) paths. hooks/lib/git-cmd.js — the only pre-existing hooks/lib/*.js — carries no such marker, so this was newly introduced, not inherited. Marker removed, matching that precedent, with a note on why. (The explanatory comment deliberately does not spell the token out, or it would reintroduce the literal.) M2 — the require-scan regex demanded the exact compact form, so `require( "./lib/x.js" )` would silently fail to stage its helper and compound H1. Now tolerant of interior whitespace and either quote style. Regression test added for H1 — the reviewer confirmed the invariant had zero coverage repo-wide: a source tree carrying the hooks but no hooks/lib/ must make writeCursorHooksJson throw rather than produce a broken install. Re-verified end to end: the missing-source case throws, no unsubstituted literal ships, and the installed hook still resolves the workspace from a foreign cwd. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2587): backfill changeset pr number (#2680) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
dacc23137a |
feat(3347): opt-in auto-update of knowledge graph after main HEAD advances
Closes #3347 Config: - Add graphify.auto_update (default false) to manifests: sdk/shared/config-defaults.manifest.json, config-schema.manifest.json Hook: - hooks/gsd-graphify-update.sh — PostToolUse Bash matcher - Gates: tool_name=Bash, HEAD-advancing git op, CI=unset, in git repo, current branch == default branch (git.base_branch override or main/ master/trunk fallback), graphify.enabled && graphify.auto_update both true, graphify on PATH, no live PID lock - Writes .planning/graphs/.last-build-status.json with status=running synchronously, then detaches hooks/lib/gsd-graphify-rebuild.sh - hooks/lib/gsd-graphify-rebuild.sh — detached rebuild runner - PID-lock acquire + trap-on-exit cleanup - graphify update . then cp graphify-out/* → .planning/graphs/ - Status file rewritten to status=ok|failed with exit_code, duration_ms, head_at_build - Portable detach (subshell + disown, no setsid dependency) Installer: - bin/install.js: register hook as PostToolUse Bash matcher (5s timeout) - Add to gsdHooks uninstall list and expectedShHooks warning list Planner / researcher status surface (issue #3347 reviewer must-have AC): - agents/gsd-planner.md and agents/gsd-phase-researcher.md load_graph_context steps now read .last-build-status.json and surface: running → "rebuild in flight"; failed → "auto-rebuild FAILED at {ts}, context is from prior build"; ok with stale head_at_build → "HEAD has advanced since last build" Settings: - get-shit-done/workflows/settings.md adds "Graph auto-update" question with No-Recommended default; bullets and update_config block updated Inventory: - docs/INVENTORY.md hook count 12 → 13 with new row - docs/INVENTORY-MANIFEST.json regenerated Tests: - tests/feat-3347-graphify-auto-update-config.test.cjs (8 tests): isValidConfigKey accepts graphify.auto_update, CANONICAL_CONFIG_DEFAULTS default false, config-set round-trip, sibling key preservation - tests/feat-3347-graphify-auto-update-hook.test.cjs (18 tests): all bail paths (non-Bash, non-HEAD-advancing, enabled=false, auto_update=false, CI=true, non-default-branch, missing graphify bin, live-PID lock), dispatch path with mock graphify bin (sync running status + detached transition to ok/failed), stale-PID lock, all five HEAD-advancing command matchers, git.base_branch override Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
7827e1ddee |
fix(#3129): replace bypassed bash regex with token-walk git-cmd.js classifier (#3141)
* fix(#3129): replace bypassed bash regex with token-walk git-cmd.js classifier Root cause: gsd-validate-commit.sh used: if [[ "$CMD" =~ ^git[[:space:]]+commit ]] This regex silently bypasses Conventional Commits enforcement for: git -C /path commit -m ... (working-directory prefix) GIT_AUTHOR_NAME=x git commit (env-var prefix) /usr/bin/git commit -m ... (full-path executable) Fix: introduces hooks/lib/git-cmd.js with isGitSubcommand(cmd, sub) — a token-walk classifier that handles all four forms by: 1. Skipping leading VAR=VALUE env assignments 2. Validating the git executable (basename check for full-path support) 3. Consuming git global options (-C <path>, --git-dir=, -p, etc.) 4. Checking the subcommand token The hook delegates to this classifier via node shell-out. node is already called twice in this hook (config check + JSON parse), so no new runtime dependency. This becomes the single source of truth for all hooks that gate on git subcommands (pre-commit-review-gate, post-push-verify, etc.). Regression test: 27 assertions — tokenize correctness, 12 must-match cases (including all 3 bypass forms), 8 must-not-match cases, 3 source checks. All are real behavioral tests, not string comparisons. Suite: 7035/7035. Closes #3129. * fix(lint+hook+changeset): allow-test-rule, fix HOOK_DIR quote injection, fix changeset pr+typo |