Commit Graph

11 Commits

Author SHA1 Message Date
allcounter
723ea08dc2 fix(#4016): imperative-override injection patterns tolerate filler words (#4061)
* fix(#4016): imperative-override patterns tolerate filler words

The narrow imperative-override family tolerates no filler between the
verb and the noun, so a planted "Forget all of your instructions"
(measured in a real public transcript) matched none of the 14 patterns
and both consuming hooks stayed silent.

One combined filler-tolerant pattern is appended; the narrow four stay
untouched to keep the change merge-friendly. Known trade-offs, disclosed
in #4016: linter-doc prose like "ignore rules on a single line" now
trips a LOW advisory, and the overlap with the narrow patterns means one
sentence can count twice toward severity thresholds.

Regression tests assert the previously-missed phrasings fire in BOTH
consuming hooks (gsd-prompt-guard and gsd-read-injection-scanner), not
just in the raw pattern list, per the agent brief in #4016.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015aGr6fvmDMLznT7TvTXrsb

* chore(#4016): changeset fragment for PR #4061

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015aGr6fvmDMLznT7TvTXrsb

* test(#4016): pin the disclosed linter-doc FP as single-pattern LOW, never blocking

Review follow-up on PR #4061: the combined filler-tolerant pattern's
disclosed false-positive class (linter-doc prose such as "use
eslint-disable-next-line to ignore rules on a single line") was
documented in prose only. Two tests now pin it:

- the prose matches exactly ONE shared pattern (the #4016 combined
  pattern, not a narrow one), so it cannot silently start double-counting
  toward the 3+ HIGH threshold;
- through the real gsd-read-injection-scanner subprocess with
  security.injection_blocking=true, the prose yields a single-finding
  LOW advisory and no block decision — with an in-test positive control
  proving a 3+-pattern payload DOES block in the same directory, so the
  non-blocking assertion cannot pass vacuously.

Samples are fragment-built like the existing SAMPLES rows so this file's
own diff does not trip the CI injection scanner.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#4016): replace the five narrow imperative-override patterns with one superset

The first cut appended a filler-tolerant combined pattern next to the five
narrow verb patterns. Both consumers count one finding per matching pattern
toward the severity threshold, so the overlap made one sentence count twice:
"Ignore previous instructions. Forget your instructions." scored 2 (LOW) on
next and 3 (HIGH, blockable) on the branch. It also left `override` out of
the combined pattern.

Replace the narrow family (ignore x2, disregard, forget, override) with ONE
superset pattern over ignore|disregard|forget|discard|override. At least one
filler (all|of|the|your|my|system|previous|prior|above|earlier) must sit
between verb and noun, enforced by a lookahead with no repetition; the two
noun-less/bare forms the old list accepted (`disregard (all) previous`,
`forget instructions`) are kept as explicit tails so the new pattern is a
strict superset. Bare "override rules" / "ignore instructions" are ordinary
repo prose (6 measured hits across docs and source) and stay unmatched.

Corpus measurement over 3019 .md/.js/.cjs/.mjs files (injection-sample tests
excluded): the old family hit 2 lines, the new pattern hits 3, the only new
one being a documented injection example in planner-reversibility.md that
the old family missed (the issue's own class).

Tests: SAMPLES reshaped to the 10-entry list; superset proof table (17 legacy
phrasings, each matching exactly one pattern); five issue phrasings including
`override all of your previous instructions` counted exactly once through
both hook subprocesses; double-count regression (1 finding, LOW); design pin
that bare verb+noun matches nothing; linter-doc FP pin split into bare
(silent) and determined (single LOW, never blocks). All fragment-built; the
CI prompt-injection scanner reports 0 findings on every touched file.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012rvqL1wk6s8dQEXsFm4RB1

* chore(#4016): changeset body in the canonical bold-lead format

.changeset/README.md Format: a leading bold change sentence, then an em-dash
explanation. Also drops the verbatim planted phrase from the body so the
rendered CHANGELOG line does not trip the pattern it describes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012rvqL1wk6s8dQEXsFm4RB1

* fix(#4016): render a bounded pattern label in the prompt-guard advisory, pin plural prompts

Review round 4 of PR #4061 left two nits open.

1. gsd-prompt-guard.js pushed `pattern.source` verbatim into its typed
   finding and, through renderFinding, into the user-facing advisory. With
   the #4016 superset pattern that source is 300 characters, so a genuine hit
   surfaced an advisory dominated by a raw regex dump. The read scanner has
   trimmed its equivalent since #3523 (`\s+` -> `-`, strip `()\`, cut at 50).
   That transform is hoisted into hooks/lib/injection-patterns.js as
   `describePattern` and used by BOTH hooks, so one finding renders the same
   label everywhere. Byte-identical to the scanner's old inline output for
   all 10 patterns (measured). No new staging dependency: both hooks already
   require this module.

2. The noun alternation `prompts?` had no positive coverage for the plural
   branch. One filler-regression row now exercises `... previous prompts ...`
   and runs through the existing once-per-hook, exactly-one-pattern loops.

The parity test's prompt-guard count assertion moves off substring-matching
the advisory prose onto the typed `findings` surface added in #3546, per
CONTRIBUTING's raw-text-matching prohibition. New test: the superset source
exceeds the bound (positive control), the prompt guard never embeds it, and
both hooks carry the identical label in `findings[0].match`.

Tests: parity, read-scanner, kimi field-shadowing, prompt-injection-scan,
hooks-crash-policy, dead-exports: 206 run, 196 pass, 0 fail, 10 pre-existing
platform skips. eslint clean; changeset lint ok; hooks runtime-build-seam lint
ok; the CI prompt-injection scanner reports 0 findings on the PR diff.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016GGp8kEB5zCDmJ6TYHP1Nj

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 06:58:45 -04:00
Behruz Nassre Esfahani
4499933807 fix(#3802): resolve the heredoc body before validating the commit subject (#3816)
* fix(#3802): resolve the heredoc body before validating the commit subject

With hooks.community: true, gsd-validate-commit.sh blocked EVERY heredoc-form
commit with CONVENTIONAL_COMMITS_VIOLATION regardless of the message, including
Claude Code's own documented idiom:

    git commit -m "$(cat <<'EOF'
    feat(auth): add login flow
    EOF
    )"

Reproduced before changing anything: conforming heredoc -> exit 2; plain
-m "feat(auth): add login flow" -> exit 0.

Root cause is the extraction regex `-m[[:space:]]+"([^"]+)"`. Bash `[^"]`
matches newlines, so the capture ran from the quote after -m to the FINAL quote
at `)"`, swallowing the whole span. `head -1` then returned the literal
`$(cat <<'EOF'` as the subject, which can never satisfy Conventional Commits.

Fixed by not answering a regex bug with another regex. hooks/lib/git-cmd.js
already exists because "a naive regex misses all three" invocation forms, and
extractBranchArgument is the established precedent for pulling an argument off a
git command line. extractCommitSubject joins it on the same tokenizeShellLike
seam — which, checked first, already returns the entire heredoc span as ONE
token, leaving only "resolve the body to its first line" as new logic.

Because the walk starts at the subcommand, `git -C <path> commit` and
env-prefixed invocations now extract correctly too — forms the raw string scan
never handled.

Deliberately unchanged, and pinned as such: a glued `-mfeat: x` and
`--message=...` still yield no message, exactly as the regex left them. The fix
stays scoped to the reported defect rather than widening on a true observation.

Two things I got wrong and corrected by measuring rather than reasoning:

  - I expected `git commit -m ""` to be blocked. Checked against the ORIGINAL
    hook: allowed before, allowed now, identical. The scanner drops the empty
    token so it takes the null path. My expectation was wrong, not the code.
  - That exposed a false comment I had just written, claiming the exit-status
    split prevents silently allowing `-m ""`. It does not. The split IS
    load-bearing, but for a heredoc whose body's first line is blank, which
    resolves to an empty subject and is correctly blocked. The comment now names
    the real case and records that `-m ""` is not it.

Tests at both layers: 9 unit rows on extractCommitSubject beside its sibling in
tests/worktree-safety.test.cjs, and 5 behavioral rows piping real PreToolUse
payloads through the hook in tests/hooks-opt-in.test.cjs. Replacing
firstLineOfMessageArg with a plain first-line return reds 8 of them across both
files. (A first mutation attempt silently no-opped and reported green — the
mutated body is echoed in the transcript for the run that counted.)

Out of scope, per the issue: the hooks.commit_types config surface, split off by
the maintainer as #3811 and explicitly sequenced after this.

Verified: `npm run lint:ci` exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3802): confine the fix to heredoc resolution, closing four regressions

Codex review of the first attempt. It was right, and the finding is one my own
rules already name: a true observation is not a licence to widen the diff.

The first attempt replaced the shell's `-m` extraction with a token walk. That
looked like the better abstraction — this module exists precisely because a
naive regex misses invocation forms — but selecting WHICH argument is the
message was never the defect, and changing it regressed four forms that
upstream allowed, plus opened a bypass:

  - `git commit -- -m WIP`             -- introduces pathspecs; `-m` is a path
  - `git commit --amend && echo -m WIP` a later command's flag became the message
  - `git commit -m "" --allow-empty-message`  the shared scanner drops empty
                                        tokens, so the next flag became the
                                        message
  - `git commit -m WIP`                unquoted argument
  - `-m "WIP notes <<EOF\nfix: smuggled subject"` was ALLOWED — the opener was
    recognised unanchored, so validation skipped past the real, non-conforming
    subject. An enforcement bypass, not a misclassification.

Now confined to the actual defect. The shell's `-m` capture is restored byte for
byte, and only the subject-from-message step is delegated, to a PURE STRING
helper `resolveCommitSubject()` that never tokenizes. Verified as a differential
against the upstream hook run inside the real tree: the only behaviours that
change are the two intended heredoc rows (2 -> 0); all four forms above read
identical, and the bypass case blocks.

That differential also corrected my own control. An earlier comparison ran the
upstream hook from a scratch directory, where its `lib/` could not resolve
`../../gsd-core/bin/lib/token-scanner.cjs`, so the classifier failed open and
reported exit 0 for everything. That made a real regression look pre-existing.
Re-run inside the tree, `<<-"TAG"` (a double-quoted tag nested in the
double-quoted argument) is genuinely pre-existing — the capture truncates — and
is now recorded as a known limitation rather than silently "fixed".

Also fixed from the review:
  - `<<-` strips leading TABS from body lines; returning the raw line blocked a
    conforming message.
  - a non-identifier tag such as `END-MSG` is a valid bash word and was rejected.
  - an immediately-following terminator is an EMPTY message, not a subject.
  - a node/library failure now falls back to the previous `head -1` instead of
    skipping validation, so a broken extractor degrades to old behaviour rather
    than becoming a new silent-allow path.

Tests strengthened per the review: the opener-spelling rows now assert BOTH
directions per spelling, since "conforming passes" alone would also pass if the
resolver returned an empty subject for a spelling it failed to parse. Added
differential rows pinning the five previously-allowed forms, and a row for the
bypass. Dropped two rows whose comments claimed the raw scan could not handle
`-C`/env-prefix invocations — it could; the claim was wrong.

Replacing resolveCommitSubject with a plain first-line return reds 9 rows across
both files. (Mutant body echoed in the transcript; an earlier mutation attempt
on this branch silently no-opped and reported green.)

Verified: `npm run lint:ci` exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3802): keep the installed hook runtime-neutral

`hooks/lib/git-cmd.js` ships into every runtime, including hermes and qwen,
where tests/install.test.cjs enforces that no Claude reference leaks into the
installed tree. My JSDoc named the idiom after the runtime that documents it.

Reworded to describe the SHAPE rather than the vendor; the runtime is still
named in the changeset, which feeds CHANGELOG.md where such references are
allowed, and in the tests, which are not installed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#3802): backfill changeset pr number

The fragment shipped with the documented `pr: 0` placeholder, which the
changeset lint treats as always-silent, because the number does not exist until
the PR is opened. Backfilled to 3816 now that it does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3802): close the truncated-capture hole, add the required test artifacts

Review round 1. Major 3 was the one that mattered, and it disproved a claim I
had stated in falsifiable form — the PR body said only two behaviours change;
the differential found five.

Major 3 — an embedded `"` truncates the `-m` capture, so the resolver received a
PREFIX of the real subject and the length gate measured the wrong string. Before
this fix the whole form was blocked outright, so the gate was unreachable; the
fix opened the path and then mismeasured it. A new enforcement hole, so it is
CLOSED here rather than declared.

Closed precisely rather than bluntly. A first attempt refused to resolve any body
with no terminator, which also blocked commits whose SUBJECT was intact and whose
quote sat further down the body — a false positive of its own. Truncation is only
fatal to the line it lands IN, and a captured line is complete exactly when
another line follows it, because the capture kept its newline. So an unterminated
body whose subject line is followed by more text stays measurable; only a subject
line running to the end of a truncated capture falls back to the opener, which
fails the format gate exactly as this form did before the fix.

Major 1 — fast-check property rows for the new parser, via the shared seeded
setup helper rather than requiring fast-check directly, per repo convention:
totality (a security property here, since an exception on this path fails OPEN),
idempotency, and that the result is always a single line drawn from the input —
the third catches a resolver that concatenated or trimmed while satisfying the
first two.

Major 2 — the 72-char gate is now exercised at {71, 72, 73} on the RESOLVED
heredoc subject, with the fixture length asserted so a mis-built fixture cannot
silently pass. 92 chars did not show which side of `> 72` the code sits on.

Minor 1 — leading blank body lines are skipped, as git's cleanup=whitespace does.
A conforming commit written that way was still blocked, which is the same defect
class #3802 reports.

Nit 1 — a backslash-escaped delimiter (`<<\EOF`) is now the same delimiter rather
than failing closed on a delimiter that includes the backslash.

Nit 5 — changeset trimmed from 2,208 chars of design note to the user-visible
change.

Mutation discipline, including a correction to my own: dropping the truncation
guard reds the unit rows, and the pre-review naive shape reds the hook-level row
too. My first mutant did NOT distinguish the hook row — removing the guard made
an empty slice and blocked for an unrelated reason, so the row passed and looked
proven. Only mutating to the actual pre-review shape showed it discriminates.

Verified: `npm run lint:ci` exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3802): measure the subject as git does — strip trailing whitespace, split CRLF

git's cleanup=whitespace strips whitespace at BOTH ends of a line; the
resolver handled only the leading direction, so a 72-char subject with
trailing spaces measured 75 and stayed blocked — the defect class #3802
reports, surviving one round further (review of #3816, Major 2). The
resolved subject now drops trailing spaces and tabs; the plain non-
heredoc path is untouched, keeping the fix confined to heredoc
resolution. The length-gate boundary rows gain dirty fixtures: 72+3
trailing spaces passes, 73+1 stays blocked on LENGTH.

split('\n') left \r on every body line, so on CRLF input the delimiter
never matched: the truncation guard was inert, an empty message resolved
to 'EOF\r', and a real 72-char subject measured 73. Split on /\r?\n/
(Minor 3).

The three property tests never reached the parser — the pinned-seed
fc.string corpus contained no newline and no opener, so every property
reduced to f(s) === s (Major 1). The generator now constructs heredoc-
shaped input (all opener spellings, <<- tabs, optional terminator, CRLF)
and each property asserts a floor on inputs its corpus actually resolved.
All new rows proved failing-first against the pre-fix resolver.

Also records the unquoted-delimiter expansion limit as one JSDoc
sentence (Informational 5).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3802): close two recognition bypasses, pin the dquoted-delimiter limit

Codex whole-PR review found two enforcement bypasses in the resolver:

- The opener's path prefix was \S*, which accepted `id;/bin/cat` — the
  resolver then validated the heredoc BODY while bash runs `id` first
  and git's real subject is id's OUTPUT. The prefix is now a
  path-character class; any shell metacharacter fails recognition and
  the form falls back to the opener line and the format gate.

- The blank-line skip used JavaScript trim(), whose Unicode whitespace
  class skips lines git KEEPS: a NBSP first body line resolved to the
  SECOND line while git's real subject is the NBSP line (verified
  against git stripspace — the c2a0 bytes survive). Blank is now git's
  ASCII space/tab only; a Unicode-blank line is returned and fails the
  format gate, the same fail-closed direction git takes.

Both proven failing-first at resolver AND hook level. Also: the
<<"TAG" spelling is recorded as a documented limit — the -m capture
stops at the delimiter's own quote so the caller can never deliver it
(fail closed; widening the capture would change every embedded-quote
case) — with a hook-level row pinning the limit; and the derivation
property no longer accepts '' unconditionally, only for heredoc-shaped
input, so a conditional constant-'' regression can't satisfy the corpus
floor unnoticed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3802): recognition whitespace is ASCII, and '' answers to the generator

Codex round 2: the opener's \s accepted Unicode whitespace bash does
not split on — $(<NBSP>/bin/cat was recognized here while bash reads
<NBSP>/bin/cat as the executable NAME, so recognition claimed a
substitution that does not run cat. Every whitespace position in the
recognition is now [ \t], the same ASCII rule as the blank-line skip,
proven failing-first.

The derivation property's ''-acceptance now consults GENERATION-TIME
metadata: the heredoc generator records whether it built an empty
message (terminator reachable, all scanned lines ASCII-blank, <<- tab
stripping accounted for), and '' is accepted exactly then — a resolver
conditionally degrading to '' on non-empty heredocs now fails, closing
the residual round-1 permissiveness without re-deriving resolver logic.

The changeset no longer overstates the opener spellings: it names the
capture-deliverable set and the documented <<"EOF" limit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3802): nothing after the terminator escapes measurement

Round-3 BLOCKER: everything after the heredoc terminator was silently
discarded, so `-m "$(cat <<'EOF'\nfeat: ok\nEOF\n) <200 a's>"` —
one 200+ char real subject once bash substitutes — measured 8 chars and
dodged COMMIT_SUBJECT_TOO_LONG, a hole the base did not have. The
canonical idiom's tail is exactly one closing-paren line; any other tail
now falls back to the opener line and the format gate, the pre-fix
behaviour for the whole form. Proven failing-first at resolver and hook
level, including the glued-text and second-substitution variants.

Also from round 3: `cat<<'EOF'` (no space) is legal bash and now
resolves — the token before << is still literally cat; the env-prefixed
and option-terminated spellings join the JSDoc KNOWN LIMIT list instead
(fail closed, modelling bash prefix words is cost with no reported
user); the changeset states the embedded-quote truncation limit for the
message body, not just the <<"EOF" spelling; the dquoted unit and hook
rows now cross-reference each other; and the fast-check setup helper's
docstring no longer claims property-file exclusivity.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3802): glued text outside the closing quote must not shrink the measurement

Codex on the round-3 guard: bash concatenates -m "$(…)"suffix into ONE
argument, but the capture holds only the quoted part — so the resolver
measured the heredoc body (8 chars) for a 200+ char real subject, a
net-new length-gate bypass the base did not have (base measured the
opener and blocked). When the closing quote is followed by anything but
whitespace or end-of-command, the hook now skips the resolver and keeps
the pre-fix first-line subject: the heredoc form fails the format gate
exactly as on base, and the plain single-line form keeps base behavior
unchanged — both pinned as differential rows, the glued-suffix row
proven failing-first against the unguarded script.

The property generator's ''-oracle now models the post-terminator guard
it previously predated: expectEmpty requires the FIRST reachable
terminator to be followed by the one canonical closing-paren line, so a
resolver regressing to '' on a non-canonical tail (e.g. a body line that
doubles as an early terminator) fails the derivation property instead of
being blessed by stale metadata.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: retrigger CI — the previous wave never started (Actions queue stall)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3802): only resolve a heredoc whose body bash does not rewrite

Round-4 review found two net-new enforcement bypasses: commands the base
hook blocked (exit 2) that this branch allowed (exit 0). Both reproduced as
a base-vs-head differential against the real hook, not inferred.

The predicate "may I resolve this?" was computed from the resolver's input
string alone, while two of its determinants live outside that string:

  1. WHICH -m quote arm produced the input. Inside -m '...' bash performs no
     command substitution, so $(cat <<'EOF' is literal text and git's real
     subject is the opener line. The resolver ran on both arms, so all four
     delimiter spellings went 2 -> 0 on the sq arm — reachable by the
     ordinary slip of typing ' for ". The hook now records MSG_QUOTE and
     gates the resolver on dq; sq keeps head -1, exact base parity.

  2. WHETHER the delimiter suppresses expansion. Only <<'D', <<"D" and <<\D
     do; a bare <<D is expanded by bash before git sees it. Resolving the
     literal dodged the format gate (feat: $UNSET_VAR reaches git as feat:)
     and the length gate (feat: ${LONG} reaches it at any length). The
     opener regex now separates the backslash-quoted and bare alternatives
     and refuses the bare one — the same fail-closed rule the metacharacter,
     truncation and post-terminator guards already follow.

A test row asserted exit 0 for a bare-delimiter body, so the suite defended
the second bypass and the fix could not land without editing a test that
read as intentional. That row and its two unit counterparts now assert the
block, per RULESET.TESTS.delete-bad-tests. Two unrelated rows used <<-EOF
to exercise tab stripping; they move to <<-'EOF' so each tests what it names.

Scoping the adjacency guard to the matched arm — required by the fix above —
also removes a spurious block (round-4 Minor 1): a double-quoted heredoc
whose body mentioned a glued single-quoted token tripped the sq arm.

The JSDoc claimed <<"EOF" was unreachable through the caller and that the
bare-delimiter gap was pre-existing. Round 4 disproved both; both corrected
here, along with the matching changeset sentence.

Verified: 7 bypass commands now block at head (was allow), the #3802 fix and
plain-form parity are unchanged across 8 control commands, hooks-opt-in 44/44,
worktree-safety 401/401, property-test non-vacuity 73/200 against a floor of
20, lint:ci exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG

* fix(#3802): resolve only where the captured text is provably git's subject

Codex review of the full PR found two more inputs where the validated text
is not the subject git receives, both net-new bypasses (base 2 -> head 0),
plus one escalation of round-4 Minor 2. All reproduced here against the real
hook and confirmed against real commits before fixing.

BLOCKER — the matched -m need not be git's message. The capture is a search
over the whole command and the double-quoted arm runs first, so it could
select a -m that is not the subject at all. git concatenates multiple -m
values and takes the FIRST as the subject, so

    git commit -m 'WIP first' -m "$(cat <<'EOF' … )"

commits the subject `WIP first` while the hook validated the heredoc. Same
for an unquoted earlier -m, for a heredoc after `--` (a pathspec, not a
message), and for one belonging to a later `&& echo`. The mis-selection is
pre-existing; resolving it is what made it a bypass. The hook now resolves
only when nothing before the matched -m could have been an earlier message,
an end-of-options marker, or another command.

BLOCKER — cleanup mode is part of the predicate. The resolver skips leading
blank lines and strips trailing whitespace because git's DEFAULT
cleanup=whitespace does. Under --cleanup=verbatim git does neither, so a
72-char subject plus three trailing spaces is committed at 75 bytes while
the hook measured 72 — COMMIT_SUBJECT_TOO_LONG dodged. This one hides from
`git log --pretty=%s`, which strips trailing whitespace in its own output;
the raw commit object shows 75 vs 72. Any named mode other than whitespace,
in either the --cleanup= or -c commit.cleanup= form, now refuses to resolve.

MAJOR — recognition trusted any path ending in /cat, so a planted
`../evil/cat` printing `WIP injected` had its heredoc body validated while
git's real subject was `WIP injected`. Only a bare `cat` or an absolute path
is recognised now. A bare `cat` shadowed on PATH is a documented residual and
is not fixable from a string — nor a meaningful boundary, since planting an
executable already allows running git directly.

The changeset and the JSDoc both asserted that a `"` anywhere in the message
blocks. Measured false: a `"` on a later body line resolves fine, because the
subject completes before the truncation point; only a `"` in the subject line
blocks. The changeset also listed <<"EOF" as covered when it measures 2/2.
Both rewritten to claim only what is measured, and the residual false
positives are now named.

Verified: 4 + 2 + 3 new bypass commands now block, with non-vacuity controls
proving the default path still resolves; all round-4 maintainer blockers stay
closed; the #3802 fix and plain-form parity unchanged across 7 controls;
hooks-opt-in 47/47, worktree-safety 402/402, property non-vacuity 73/200
against a floor of 20, lint:ci exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG

* fix(#3802): scope the cleanup-mode guard to the command outside the message

The guard scanned the whole $CMD for `--cleanup=` / `commit.cleanup=`,
and the heredoc BODY sits verbatim inside $CMD, so any conforming
message that merely MENTIONED the token was refused, fell back to the
opener line, and was blocked with CONVENTIONAL_COMMITS_VIOLATION. These
are ordinary English in this repository, whose own hooks and docs
discuss cleanup modes constantly. Reproduced against the real hook:
`fix: document commit.cleanup=strip behavior` blocked, the same message
without the token allowed (review of #3816, round 5 — BLOCKER).

Scoping to $MSG_PREFIX alone, as prescribed, would have reopened the
round-4 length-gate bypass the guard exists for: git accepts the flag on
EITHER side of -m, and `git commit -m "<heredoc>" --cleanup=verbatim` is
caught today only because the scan is command-wide. Measured, not
assumed. The scan now covers MSG_PREFIX + MSG_SUFFIX — the whole command
minus the one span that is message text — joined with a space so a token
cannot be forged across the seam.

Swept the guard class rather than the reported instance. The adjacency
guard does not share the defect: an in-body `-m "foo"bar` is refused by
the already-documented embedded-quote capture limit (any `"` in the
subject line truncates the capture), and an in-body `-m ` without quotes
resolves and is allowed. Deliberately untouched.

Both directions pinned failing-first: the three false-positive rows red
against the unscoped guard, and the trailing-flag row reds against
prefix-only scoping. Each mutation was echoed back to prove it landed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XogDtuuuGQEfsWaLSaZCLB

* fix(#3802): read commit options the way bash hands them to git

Round 6 reported the adjacency guard scanning all of $CMD for a glued
`-m "..."`, so a glued -m belonging to a chained-after command refused a
heredoc that was never truncated. Glue is a property of the ONE character
following the matched span, so that character is now the whole window.
Separators and redirections are excluded because bash does not
concatenate across them: in `-m "msg"&& echo hi` the argument ends at the
quote, so there is no truncated capture to defend against.

An independent full-PR pass then found three accept-direction defects
this PR had introduced in earlier rounds, each measured against a real
commit by reading the raw commit object — `git log --pretty=%s` strips
the trailing whitespace that makes the length wrong and hides it:

  --cle=verbatim         git accepts any unambiguous prefix of a long
                         option, so the mode was set by a token that is
                         not the literal --cleanup. 75-byte subject
                         recorded, 72 measured.
  -am 'WIP first'        git reads this as -a -m, so the real subject is
                         `WIP first` and the heredoc is only the second
                         message. The scan looked for a standalone -m.
  --clean""up=verbatim   bash removes quotes before git sees the
  -""m                   argument, so a spliced spelling is the same
                         option and matched no literal.

The two option-name scans now read their window with quote characters
removed, which is what bash does to it, and the cleanup class covers
git's abbreviations. The adjacency test deliberately keeps the raw text:
it asks about a literal character position, not an option name.

Narrowing the cleanup window to git's own command segment was tried and
reverted. `;`, `&` and `|` end a command only outside quotes, and this is
a substring scan, not a parse: an unconditional trim cut the window short
on `--author "a&b"`, and a quote-aware trim still cut it on `--author
a\&b`. Each hid a real trailing --cleanup=verbatim and accepted a 75-byte
subject. The resulting false positive — a --cleanup carried by a chained
command refuses the commit — is documented and pinned instead. Refusing a
commit git would take is recoverable; accepting an over-long subject is
not.

Sixteen rows in tests/hooks-opt-in.test.cjs. Seven mutations, including
both reverted narrowings, so no dead end can be reintroduced silently.

* fix(#3802): close six accept-direction bypasses in the resolve guards

Round 7's FIRST-MESSAGE GUARD Major does not reproduce. Measured against the
real hook in a complete tree at the reviewed head: the classifier gate runs
before any guard, so `git add -A && git commit …` (git->add stops on a
non-commit subcommand) and `cd dir && git commit …` (the first executable is
not git) exit 0 without a guard being evaluated. The control is the proof — a
subject the bare form blocks with CONVENTIONAL_COMMITS_VIOLATION exits 0 in
both chained forms, so the hook never validated them and cannot be
over-blocking them. The guard is unchanged; scoping this scan to $MSG_PREFIX
alone is what reopened the round-4 trailing-flag bypass.

The class was real, though, one shape further out: `FOO=bar; git commit …` IS
classified and then refused, because assignment detection is prefix-anchored
and the tokenizer does not split operators. Pinned as a counterexample and
disclosed rather than generalised away; narrowing it means changing
isGitSubcommand, the shared git-commit detector every gating hook uses, and it
fails closed.

Six accept-direction bypasses are fixed. Each let the hook resolve and ALLOW a
commit whose real subject the rules refuse; the three that turn on git's
recorded subject were confirmed against the RAW COMMIT OBJECT, since
`git log --pretty=%s` strips trailing whitespace and hid two of them:

  --cleanup=whitespace -m <72+spaces> --cleanup=verbatim  git kept 75 bytes
  -mWIP -m <heredoc>                                      git recorded `WIP`
  --mes=WIP -m <heredoc>                                  git recorded `WIP`
  -\m WIP -m <heredoc>                                    git recorded `WIP`
  git commit --amend --no-edit \n echo -m <heredoc>        echo's argument read
  --squash=HEAD -m <heredoc>                              `squash! …`

Causes: one BASH_REMATCH inspected only the FIRST cleanup directive while git
applies the last, so multiplicity now refuses rather than guesses at an
argument order a substring scan cannot recover; the option scan required a
trailing space or `=`, missing attached values and long-option abbreviations;
dequoting removed quotes but not the syntactic backslashes bash also removes;
the separator scan omitted newline; and --squash/--fixup have git compose the
subject, so the supplied message is not the subject at all. Every fix widens
refusal, the direction this file documents as recoverable.

The multiplicity count first broke the hook outright: the script runs under
`set -euo pipefail` and grep exits 1 when it matches nothing, which is the
common case, so every ordinary commit died at exit 1 with no verdict. Guarded,
and only caught because the probe runs the real hook rather than the scan.

Five new rows, all five proven red against the pre-fix hook, each carrying a
non-vacuity assertion that the canonical single-`-m` heredoc still resolves.
Changeset corrected on three counts: "all fail-closed" was wrong (persistent
commit.cleanup fails OPEN, as do the -C/-c/-F/-t message sources), "global
options are all walked through" was too broad, and the chained-before claim
now states what is measured.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018FUAVz49BghqxoJgwt7EW9

* fix(#3802): stop the separator and glue classes matching a literal backslash

Round 8's Major, with two corrections to its account.

`;`, `&` and `|` are metacharacters inside `[[ ]]`, so an inline bracket
class must escape each one. POSIX bracket expressions have no escape
mechanism of their own, so on bash 3.2 -- the system /bin/bash on macOS,
already a supported target here per the `declare -A` ban in
tests/install.test.cjs -- those backslashes reach the regex engine and add
a literal `\` to the class. bash 4+ consumes them, which is why this is
invisible on a modern bash. The hazard is specific to bracket
expressions: `\(` outside one is made literal correctly on every version,
and the subject validator and the `-m` capture classes were checked and
are unaffected.

The prescribed fix is not taken, because it does not parse. Inline
`[;&|]` is a bash SYNTAX ERROR on 3.2 and on 5.3 alike -- the backslashes
exist to get the metacharacters past the `[[ ]]` parser, so removing them
leaves an unparseable script. Each class is held in a variable and
expanded unquoted on the right of `=~` instead, which is a plain regex on
both versions.

One root cause, consequences in BOTH directions. The reported half is the
separator scan over-blocking. The half not reported is the accept
direction, and it is the more serious: the glue class is NEGATED, so on
bash 3.2 a backslash-glued suffix fell inside the exclusion and the hook
RESOLVED a heredoc it should have declined -- measured exit 0 on 3.2
against the unfixed hook, exit 2 everywhere else, with a letter-glued
control refused in all four cells.

The reported repro is not actually fixed by this, and the changeset says
so. A `\`-newline line continuation carries a literal newline, which the
round-7 separator guard refuses on every bash, so that shape stays
blocked with or without this change. Narrowing the newline guard is not
attempted: telling a continuation from a separator by substring scan is
the class that was tried twice in earlier rounds and reverted both times,
and an escaped backslash sitting immediately before a real newline is
indistinguishable from a continuation. Disclosed as a known fail-closed
limit instead.

Every new row runs under each bash on the machine. Against the unfixed
hook both bash 3.2 rows go red while all four bash 5.3 rows stay green --
written the ordinary way these rows would run under PATH bash, pass
against the broken hook, and prove nothing. Two non-vacuity controls per
interpreter prove the validator is reached rather than passing
everything. All 8 rows of the established differential harness are
byte-identical before and after on both versions: no regression, no new
refusal.

* fix(#3802): remove the $ of a dollar-quote from the option-name scans

Independent round-8 review, accept direction.

The option-name windows are dequoted so they match "the command as bash
hands it to git" -- round 6 removed quote characters, round 7 removed
syntactic backslashes. Both passes missed that bash has two further
quoting forms whose introducer is a `$`: `$'...'` and `$"..."`. Removing
the quote characters alone left that `$` stranded INSIDE the option name,
so `-$"m"` dequoted to `-$m` and matched no literal, while bash passed a
real `-m` to git.

Measured on bash 3.2.57 and 5.3.15 against a real repository: the hook
allowed

    git commit --allow-empty -$"m" WIP -m "$(cat <<'EOF'
    fix: a perfectly ordinary conforming subject
    EOF
    )"

with exit 0, and `git cat-file -p HEAD` recorded the subject `WIP`.

The comparison that establishes this is HEAD-internal, not a differential:
the same command spelled `-m WIP` is refused (exit 2). The merge-base
refuses EVERY heredoc form, including a perfectly conforming one, so its
exit 2 on this input says nothing about whether any guard fired -- it is
the absence of the feature, not a working check. The same miss covered
`$'m'`, spliced `--message`, `--cleanup`, `--squash` and `--fixup`.

An option NAME finished by a command substitution -- `--clean$(printf
up)=verbatim` -- is a different problem and gets its own guard: bash runs
a program to complete the name, so the argv git receives is not derivable
from this string at all, and resolution is refused rather than guessed.
The guard is scoped to the NAME: the class is a `-`-leading token whose
characters up to the substitution contain no `=`. A substitution
supplying a VALUE -- the ordinary `--author="$(git config user.name)"`,
spaced or glued, in either window -- is untouched and still resolves,
pinned in both directions. It is a SHAPE, not a segmentation of the
command line; segmenting was tried twice in earlier rounds and reverted
both times, and that reasoning stands.

Both new rows fail against the unfixed tree with their own assertions,
proven in a complete worktree at the previous head rather than a hook
copied out of its tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy

* fix(#3802): recognise a canonical cat, not any absolute path ending in /cat

Independent round-8 review, accept direction.

Round 4 restricted heredoc-opener recognition to an absolute path, after a
relative `./cat` was measured being trusted to echo its stdin. It stopped
at "absolute", so any absolute path ENDING in `/cat` was still trusted --
the same claim the round-4 reasoning had rejected one spelling earlier.

Measured on bash 3.2.57 and 5.3.15 against a real commit: with an
executable at `/.../fake-cat/cat` printing `WIP injected`, the hook
validated the conforming heredoc body and allowed the commit (exit 0)
while `git cat-file -p HEAD` recorded the subject `WIP injected`. The
same command through `./cat` was already refused, which is the control
that shows this is the round-4 class one spelling out rather than a new
one.

Recognition is now the canonical system locations -- bare `cat`,
`/bin/cat`, `/usr/bin/cat` -- which is the only identity claim a string
can support. `/usr/local/bin` is deliberately excluded: it is
user-writable on ordinary machines, which is the plantable case this
guard exists for. Anything else falls back to the opener line and the
format gate: fail closed, exactly the pre-fix behaviour for the form.

The pre-existing residual is unchanged and still documented: a bare `cat`
shadowed earlier on PATH is indistinguishable here, and is not a
meaningful boundary -- anyone able to plant an executable on PATH can run
`git commit` directly. This hook stays an authoring guard, not a security
control.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy

* fix(#3802): an option name carrying a shell expansion is unresolvable

Independent review, round 9, accept direction. Four more spellings, and a
change of strategy that is the actual point of this commit.

Rounds 6, 7 and 8 each tried to EMULATE what bash does to an argument
before git sees it -- round 6 removed quote characters, round 7 syntactic
backslashes, round 8 the `$` that introduces a dollar-quote -- and each
round review found another transform that had been missed. Round 9 found
four more. All measured on bash 3.2.57 and 5.3.15 against a real
repository, each with the plain spelling of the same command as its
control (refused, exit 2) and `git cat-file -p HEAD` for the subject git
actually recorded:

    -$'\155' WIP        hook 0, real subject `WIP`   ANSI-C octal -> m
    -$'\x6d' WIP        hook 0, real subject `WIP`   ANSI-C hex   -> m
    -`printf m` WIP     hook 0, real subject `WIP`   backtick substitution
    x= … -${x}m WIP     hook 0, real subject `WIP`   parameter expansion
    -? WIP              hook 0, real subject `WIP`   pathname expansion

and the same class through the cleanup guard, where git recorded a
75-character subject the length gate had measured as 72:

    --cle$'\141'nup=verbatim, --clean`printf up`=verbatim, --cle?nup=verbatim

The last two settle it. An option name finished by a PARAMETER expansion
depends on a variable's value at run time; one finished by a PATHNAME
expansion depends on the contents of the working directory. Neither is
derivable from the command string at any level of effort, so emulation
cannot be completed -- not "has not been completed yet". A fifth patch in
that direction would have the same shape as the previous four.

The rule is therefore no longer "normalise it and match the literal". It
is: an option NAME carrying a shell expansion or quoting construct is
UNRESOLVABLE, and unresolvable refuses. One rule covers every spelling
above and every spelling nobody has thought of yet, in the fail-closed
direction. The dequoting passes are kept rather than replaced: they still
normalise the deterministic removals, so the guards RECOGNISE
`--clean""up=` and `-\m` as the options they are instead of merely
refusing them, which keeps the existing rows meaningful.

Scope is unchanged and still pinned in both directions: the class is a
`-`-leading token whose characters up to the construct contain no `=`, so
a construct supplying a VALUE -- `--author="$(git config user.name)"`,
the backtick spelling, `--date="${NOW}"`, a glob character inside an
author string, a pathspec after `--` -- still resolves. Nine such forms
are asserted to pass beside the seven that must refuse.

The class is bracket-only and holds no backslash, per round 8: a POSIX
bracket expression has no escape mechanism, and a backslash written
inside one becomes a literal member on bash 3.2.

The new rows fail against the previous head with their own assertion
message, in a complete worktree with the lib built, not a copied hook.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy

* docs(#3802): disclose and pin the two spellings the round-9 class over-blocks

A scoped review of the round-9 class asked one question -- does it refuse a
conforming heredoc commit that the previous head accepted -- and found two
spellings that it does. Both measured on bash 3.2.57 and 5.3.15, previous
head 518d97b64 exit 0, current head exit 2:

    git commit -S$SIGNING_KEY -m <conforming heredoc>
    git commit -m <conforming heredoc> -- -*.txt

Disclosed and pinned rather than narrowed, for two reasons.

Narrowing is not available cheaply. Dropping the bare `$` member reopens
`-$xm`: with `xm=m` bash hands git a real `-m`, which is the parameter
expansion bypass the round-9 commit exists to close. Skipping tokens after
`--` means deciding where git's options end from a substring scan, which
is the class this file has already reverted twice for opening
accept-direction holes -- a `--` inside a quoted value (`--author "a -- b"`)
would truncate the window and hide a real trailing directive.

And the limits are narrower than they look, because in both cases the
spelling a developer actually reaches for still resolves:

    -S "$KEY"  and  --gpg-sign="$KEY"        resolve
    '-*.txt', "-*.txt", ':(exclude)-*.txt'   resolve

The pathspec one is worth stating precisely: a glob only reaches git AS a
pathspec when it is quoted, because an unquoted one is expanded by the
shell before git is executed. So the refused spelling is not passing a
glob to git at all, and the spellings that do are unaffected.

Refusing a commit git would take is the recoverable direction; accepting a
non-conforming subject is not. That is the trade this file already makes
everywhere else, and it is made explicitly here.

Nine rows pin the working spellings beside the three that refuse, so a
later narrowing cannot silently drop the cases that must keep working.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy

* fix(#3802): join backslash-newline continuations before the resolve guards

Round 9's Major, with a correction to its diagnosis.

The cited bracket classes at :223 and :260 no longer exist -- round 8
moved both into SEP_CLASS and GLUE_CLASS, and a lone backslash before -m
resolves (exit 0) at the reviewed head on both bash 3.2.57 and 5.3.15.
What refuses the repro is the NEWLINE a `\`-continuation carries: round
7's separator guard reads any newline in a window as a command boundary,
and `git commit \` newline `  -m "$(cat <<'EOF' …` was refused for that
reason. Round 8 disclosed it as a fail-closed limit; round 9 calls the
idiom common and the limit a Major, and it is fixed here.

It was left as a limit because "is this newline a continuation" looked
like the segmentation question this file has reverted twice. It is not:
bash's rule is local and character-level. A newline preceded by an ODD
run of backslashes is a continuation and bash removes both; an EVEN run
(`\\` then newline) is a literal backslash followed by a real newline,
which IS a separator. Both scan windows are joined that way immediately
after they are cut from the command and before any dequote copy is
derived, in three bash-3.2-safe parameter expansions: every `\\` pair is
parked on \x01, any backslash-newline that remains is a lone one and is
removed, then the pairs are restored.

Measured on both bashes, both directions:

    git commit \<nl>  -m <heredoc>                 2 -> 0   the fix
    git commit \\<nl>  -m <heredoc>                2 -> 2   literal \ + real separator
    git commit<nl>  -m <heredoc>                   2 -> 2   bare newline
    -m <heredoc>\<nl>suffix                        2 -> 2   bash glues it; the glue guard sees it glued
    git commit … \<nl>  --allow-empty<nl>echo -m … 2 -> 2   the REAL newline still separates

The prescribed `[\;&|]` is not taken: a backslash written inside a
bracket expression becomes a literal member on bash 3.2, which is the
round-8 defect from the other side.

Rows run under each bash on the machine. The fix row fails against the
previous head in a complete worktree with the lib built; the four control
rows were measured against that same head and were already refused, so
they pin existing behaviour rather than the change.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 03:51:43 +00:00
Tom Boucher
d24e22b156 enhance(#3912): gsd-tools declares outcomes, pinned at v1 (#3983)
* enhance(#3912): gsd-tools declares outcomes, pinned at v1

ADR-3889 §4. Phase 6 already moved error()'s terminator onto the seam, so what
remained was the declaration — and the pin that makes it invisible today.

The census corrected two documented figures before any code changed.
ERROR_REASON has exactly 25 members (the ADR and epic were right; an earlier
note of mine claiming 23 was wrong and is corrected). And output({error}) is
**64 sites across 9 files, not the 60 ADR-2980 ratified** — the module shape
holds but the total drifted +4: frontmatter 7 not 6, phase 4 not 2, roadmap 3
not 2. That matters because this phase's criterion demands the pin be asserted
over the enumerated population rather than sampled; asserting over a stale 60
would leave four sites unpinned while claiming full coverage, which is the
shape of failure this epic exists to remove.

The issue does not state the fact that shapes the design: output() never
touches the exit code. Confirmed by reading it — it writes fd 1 and returns.
So a declared outcome for those 64 sites had nowhere to be READ. The mapping
was never the work; wiring somewhere for the declaration to land was.

The seam already existed twice over. cli-exit.cts holds two globalThis-Symbol
cells, each because the module is emitted to three locations and a module-level
`let` would let instances disagree, and runMain already maps a code returned by
main(). A third cell inherits that solution. output() records DEGRADED for any
{error} payload — key-order agnostic, which is exactly why the "42 sites"
figure undercounts — and runMain projects the cell only when main() returns
nothing, so an explicit return still wins.

error() maps its reason through a table over the closed 25-member enum, leaving
all 278 call sites untouched; 226 of them pass no reason at all. The version
gate lives in error(), NOT in projectOutcome: registered names are
version-invariant there, so mapping a reason straight through would make USAGE
project to 64 under v1 and break the pin on its first line. projectOutcome is
left exactly as Phase 2 shipped it, DEGRADED's 0/80 asymmetry included.

Proven rather than asserted. v1 is byte-identical across three real CLI paths —
config-get plain, config-get --json-errors, and an output({error}) path —
matching exit code and exact bytes against the pre-change build. Under
GSD_EXIT_CONTRACT=v2 the same commands now exit 66 (CONFIG_KEY_NOT_FOUND ->
NO_INPUT) and 80 (DEGRADED), both looked up through the registry. An
anti-vacuity test pins that v1 and v2 genuinely differ for at least one reason,
because without it a mapping where everything projects to 1 under both versions
would satisfy every other assertion and the declaration would be theatre.

A1 iterates all 25 enum members and A3 asserts over the measured 64-site
population, so a 26th reason or a 65th site fails until it is given a mapping —
the drift guard this phase needs, given ADR-2980's own count had drifted +4
unnoticed.

Verification runs on the remote runner.

Refs #3912

* fix(#3912): the outcome cell must never lower an exit code

The remote run caught a fail-open that this phase introduced, in the phase
whose entire purpose is removing fail-opens.

`state validate --strict` on a missing STATE.md exited **0** where it must exit
1. Mechanism: `runMain` projected the pending outcome whenever `main()` returned
void, and under v1 DEGRADED projects to 0 — so a `process.exitCode` already set
non-zero by the command was clobbered down to success. Confirmed live against a
fixture, before and after.

This refutes a review conclusion recorded earlier in this phase, that the cell
was "fail-closed and can never mask a failure as success". It could, and did.
Recording that plainly so the assumption is not repeated: the cell's danger was
never only that it might add a failure — it was that projecting it
unconditionally overwrites whatever decision came before.

Projection is now guarded: it may set a code only when none is set, and an
already-non-zero exit code always wins. The full precedence — explicit `main()`
return, then an existing non-zero exitCode, then the declared outcome — is
written at the projection site. A regression test drives a void return with a
pre-set non-zero code and a pending DEGRADED, and fails against the pre-fix
build.

The second failure was my test encoding the wrong contract, not a code defect.
It asserted `output({found:false, error: undefined})` records DEGRADED because
the KEY is present. `JSON.stringify` drops undefined, so the payload the user
receives is `{"found":false}` — carrying no error at all, and calling that
degraded would hand back exit 80 under v2 for output that reads as clean. The
discriminator is a serializable error VALUE, not key presence. The test now
pins `{error: undefined}` as explicitly NOT degraded, and the design doc's
wording is tightened to match.

Verification runs on the remote runner.

Refs #3912

* docs(#3912): the versioned exit contract, and a flag defect the docs found

Diataxis pass for Phase 8, plus a real fix that only surfaced because writing
the how-to meant running its own examples.

The docs. ADR-2980's "Revisit if" clause asked for exactly the versioned
projection this phase provides, so it gets an amendment naming #3912 /
ADR-3889 section 4 as that boundary: v1 stays 0 byte-for-byte, v2 projects
DEGRADED to 80. The amendment also records the count drift rather than
restating a stale figure — the ADR ratified 60 output({error}) sites in 9
modules; the AST-measured population is 64 across the same 9 (frontmatter 7
not 6, phase 4 not 2, roadmap 3 not 2). The pin is asserted over the
enumerated 64. json-errors.md gains the outcome-declaration reference,
including the precedence order a review pass got wrong and the suite refuted:
an explicit main() return, then an already-set non-zero process.exitCode, then
the declared outcome. Projection may only ever set a code, never lower one.

A how-to is owed here and is written, not skipped. Under v1 nothing changes,
so the audience is an operator opting into v2 and needing to know what the
codes mean for a CI gate — a migration, which is how-to shaped. It covers
turning v2 on, the code table, why 80 is "ran and reported a condition" rather
than a crash, and how to split a gate that treats any non-zero as fatal. No
tutorial: there is no new entry point to learn, and under the default contract
a reader would be walked through observing nothing.

The defect. Running the how-to's own Step 1 example returned

    $ gsd-tools --exit-contract=v2 state validate --strict
    Error: Unknown command: --exit-contract=v2          (exit 64)

while the same flag trailing the subcommand worked and exited 80. The flag
half-worked, by argv position. resolveContractVersion scans argv
non-destructively, so the token survived into the dispatcher, which treats
argv[2] as the command name. --json-errors had already solved precisely this
at gsd-tools.cjs:4455, under a comment naming the hazard verbatim: "The argv
splice must happen here too, otherwise the dispatcher below sees
--json-errors as an unknown command." The later flag never got the same
treatment.

Fixed rather than documented around: the version is resolved first — which
memoizes the cell and makes an invalid value throw early — and then every
--exit-contract= token is spliced out of the dispatcher's argv copy.
--exit-contract is now listed in TOP_LEVEL_USAGE, where it never was. The
regression test pins leading position, trailing position, agreement between
the two, and a loud failure on v3 rather than a silent fall back to v1.

Neither review engine would have caught this: the defect is invisible in the
diff, because the diff does not touch argv handling. It surfaced only from
running the documentation's own example. Writing a how-to is an execution pass.

Verification runs on the remote runner.

Refs #3912

* fix(#3912): the flag splice has to run before the run-with-timeout return

An isolated review of the previous commit found that the fix did not deliver
what it claimed, and that two of its own tests were weak. All three findings
reproduced by execution before any change was made.

The fix was placed below a return. main() intercepts `run-with-timeout` at
gsd-tools.cjs:4436 and returns from there — above both the --json-errors block
and the --exit-contract splice added in the previous commit. So the flag still
died in leading position for that one command:

    $ gsd-tools --exit-contract=v2 run-with-timeout 5 -- node -e "..."
    Error: Unknown command: run-with-timeout        (exit 64, child never ran)

The previous commit message and the test's describe-block both claimed
position-independence unconditionally. That was an overclaim, not a gap left
open, and it is the part worth naming: the fix was verified by hand on the
commands I happened to think of, and `run-with-timeout` returns before the
code I was verifying.

Both global-flag blocks now run above the interception, with a comment naming
it so a later edit cannot slide them back down. Moving --json-errors up fixes
the identical pre-existing bug for that flag, verified failing beforehand
(exit 1, sdk_unknown_command). Fixing the sibling is deliberate: same defect,
same block, and a known-broken twin next to a fixed one is not a resting state.

Two tests were not pulling their weight. The invalid-value test was vacuous —
it passed against the pre-fix build, because `--exit-contract=v3` already
exited 1 there and already printed the resolve error lazily through
error() -> getContractVersion. Both its assertions held before the fix, so it
pinned nothing. The real discriminator is that the pre-fix build emits BOTH
"Unknown command: --exit-contract=v3" and the resolve error, while the fixed
build emits only the latter; the test now asserts that absence.

The leading-position and leading==trailing tests asserted proxies — "not 64",
"no Unknown command", "the two agree" — none of which pin a value, and all of
which would survive both positions being identically broken. With a .planning
directory and no STATE.md, state-snapshot exits exactly 80 under v2 and 0
under v1 in both positions. Those numbers are pinned now. The multi-token case
the descending splice loop exists for is covered too, and run-with-timeout has
regression tests for both flags.

The lesson is narrower than "test more". Hand-verifying the production
behavior does not verify that the test would have caught its absence. The
pre-fix binary has to be run against the test's own assertions.

Investigated and deliberately not changed: splicing before --cwd parsing
degrades one diagnostic from "Missing value for --cwd" to "Invalid --cwd:
<path>", but that is pre-existing — verified on the pre-fix build via
--json-errors, which already did it. This change joins the pattern rather than
creating it, and both forms exit 64 on malformed input either way.

Verification runs on the remote runner.

Refs #3912

* chore(#3912): backfill changeset pr numbers to 3983

* test(#3912): pin the reason-table invariant as set equality, not a count

A graph-backed review flagged the unchecked lookup in
expectedErrorCode3912. Investigated by execution: the drift guard DOES
hold — for an unmapped reason under v2 the production error() yields 1
while the table yields undefined, so the assertion fails. Not a
correctness defect, and deliberately NOT made tolerant, since a tolerant
lookup would destroy the guard.

Two real problems remained. The guard asserted the wrong invariant: it
counted the TABLE's keys at 25 rather than checking they match the
ENUM's values, so a renamed member keeps the count at 25 and slips past,
and a 26th member leaves the table at 25 and slips past too. Both were
then caught only indirectly, by an undefined mismatch producing 'must
exit undefined'. It is now a sorted set equality, so the failure names
the specific missing or extra reason.

And the comment above it described a '?? FAIL' fallback that does not
exist anywhere in the function. It now states what the code actually
does, verified by running it rather than by reading it.

Refs #3912

---------

Co-authored-by: sim <sim@local>
2026-08-28 08:09:05 -04:00
Tom Boucher
2ea5efc151 enhance(#3911): hooks declare their crash policy (#3960)
* enhance(#3911): give hooks an exit seam that needs no build

ADR-3889 Phase 7 foundation. The 19 shipped enforcement hooks hold 91 of the
epic's 128 terminators and cannot reach `terminateNow` today.

The obvious route — requiring `gsd-core/bin/lib/cli-exit.cjs`, as
gsd-agent-isolation-guard.js already does for two other modules — is rejected.
That precedent carries its own warning (#3582): those files are tsc output,
gitignored and absent on a raw plugin-marketplace or git-clone install, so the
hook must first call ensureRuntimeBuild() to self-heal. Making the module a
hook needs IN ORDER TO TERMINATE depend on a build inverts the dependency, and
its failure mode is precisely the fail-open this phase exists to remove: a
guard that cannot terminate cannot deny. `lint-hooks-runtime-build-seam`
already encodes that concern, and Design B would have had to add an
ensureRuntimeBuild() call to all 19 hooks to satisfy it.

So `hooks/lib/` becomes a third emit location for cli-exit and a fifth for the
registry, preserving the invariant `src/cli-exit.cts`'s own header states: it
imports nothing but node:fs and its sibling registry, and the generator
dual-emits that sibling alongside each copy so a relative require resolves next
to whichever copy loaded it. Shipping needed no change — build-hooks.js already
declares HOOKS_SUBDIRS_TO_COPY = ['lib'].

Proven, not asserted: the two files are copied into an otherwise-empty tmpdir
and a child process requires them and terminates — PASS exits 0, HOOK_DENY
exits 2 with the payload on both stdout and stderr. That test fails the moment
the hooks copy gains a require reaching outside hooks/lib/.

Also fixed inline: the registry's fifth target let any `--write` test overwrite
the real committed hooks/lib/exit-code-registry.js, because the test helper
derived only three of the other output paths. It now redirects all five, and a
regression test asserts every committed artifact is byte-identical after a
redirected write.

Install-tree goldens pick up the two new shipped paths across 11 runtimes —
insertions only, no removals. lint:ci was green while they were stale, so this
was found by regenerating rather than by a gate.

Verification runs on the remote runner.

Refs #3911

* enhance(#3911): declare a crash policy, and migrate the write guard

Adds `hooks/lib/hook-exit.js` — the hook-facing vocabulary over `terminateNow`,
hand-written because the cli-exit copy beside it is generated:

  allow(payload)          exit 0
  deny(payload, stderr?)  exit 2
  crash(onCrash, payload) whichever the hook DECLARED

`crash()` takes the policy as a required argument with no default, which is the
whole mechanism: fail-open by accident stops being expressible. A hook must
name ALLOW or DENY at the call site, and an unrecognized value terminates
INTERNAL rather than guessing. Fail-open stays legal; fail-open by omission
does not.

`gsd-write-guard.js` is the first hook migrated, all 12 sites, and it exposed a
gap in the seam. `terminateNow`'s doc comment justified its fd-2 write by
citing this hook's `emitBlock` — but modeled it as sending the same bytes to
both streams, when `emitBlock` actually sends full JSON to stdout and only the
bare `reason` string to stderr, because Kimi's hook bus feeds stderr verbatim
back to the model. Migrating as written would have turned a readable sentence
into a JSON blob for Kimi-backed agents.

#3911 requires both "all 19 hooks terminate through terminateNow" and "no
hook's effective default changes". Those are jointly satisfiable only by
teaching the seam to carry a distinct stderr payload, so `terminateNow` gains
an optional third argument: omitted, behavior is byte-for-byte what it was; a
string is written raw, which is exactly the Kimi case. The doc comment's
inaccurate claim about emitBlock is corrected in place.

Proven rather than asserted: the pre-migration file is reconstructed from HEAD
and driven with the same catastrophic-shrink payload as the migrated one —
exit code, stdout and stderr all byte-identical.

Verification runs on the remote runner.

Refs #3911

* enhance(#3911): all 19 hooks terminate through the seam

Migrates the remaining 18 enforcement hooks onto allow/deny/crash. An AST walk
now reports zero `process.exit(` call sites across every `hooks/*.js` — down
from the 91 the census measured.

Each hook with an outer catch declares its policy once, at module top, with the
reason that policy is right for that specific guard: a read guard that cannot
scan must not block the read; a statusline that renders every prompt must
degrade rather than crash; an injection scanner must not retroactively block a
result already returned. Those sentences are the deliverable — they are what
turns fail-open-by-accident into fail-open-on-purpose. No hook's effective
default changed.

Wiring exposed two defects, both fixed here rather than noted.

A SECOND stdout/stderr-splitting site turned up in `gsd-workflow-guard.js`'s
`emitForceAddBlock`, matching the pattern already known from the write guard —
full JSON to stdout, bare reason to stderr for the Kimi bus. It uses the
`stderrPayload` argument added in the previous commit, which is now carrying
its second real caller rather than one special case.

More seriously, `terminateNow` emitted both streams inside ONE try, so a
payload that failed to serialize aborted before the stderr write ever ran. The
two windsurf guards write nothing to stdout on a block and only a reason string
to stderr, so `deny(undefined, reason)` exited 2 with EMPTY stderr — a deny
that silently loses its reason, which is the exact "fails with success" class
this epic exists to close. The streams are now emitted independently, each with
its own guard, and `undefined` means "nothing to write for this stream" rather
than an error. Regression tests inject a throwing write on one fd and assert
the other still receives its payload; they fail against the single-try version.

Byte-identity was proven per hook, not assumed: each pre-change file is
reconstructed from HEAD and driven side by side with the migrated one across
its normal path, its deny path, malformed stdin and empty stdin — exit code,
stdout and stderr compared.

Verification runs on the remote runner.

Refs #3911

* enhance(#3911): harden the three shell hooks, and pin every hook's policy

`gsd-phase-boundary.sh`, `gsd-session-state.sh` and `gsd-validate-commit.sh`
gain `set -euo pipefail`.

The expected hazard did not materialize, and that is worth recording: every
intentionally-non-zero command in all three is already the condition of an
`if`/`elif`, which `set -e` never fires on, and none of them reads a
possibly-unset variable or pipes through a grep that may legitimately match
nothing. No `|| true` guards were needed. Each hook was still checked
command-by-command before the flags went in rather than after.

Twenty-one before/after cases across the three hooks — disabled and enabled,
planning and non-planning, missing STATE.md, malformed JSON, the Kimi payload
shape, quoted and unquoted `-m`, valid and over-long Conventional Commits —
all match on exit code, stdout and stderr.

The hardening is shown to actually fire, not merely added: with a stubbed
`node` that fails at the JSON-emit step, phase-boundary and session-state go
from silently exiting 0 with empty stdout to failing visibly with the error
surfaced. No such case could be constructed for `gsd-validate-commit.sh`,
whose every statement already sits inside an if-condition — recorded as
unproven rather than claimed.

`tests/hooks-crash-policy.test.cjs` adds the per-hook coverage the issue asks
for, table-driven over all 19 hooks rather than 76 hand-written cases: normal
allow, deny where a deny path exists, crash-honors-the-declared-policy, and an
unclosed-stdin case — the one `process.exitCode` structurally cannot serve. The
deny assertions encode each hook's ACTUAL stream split rather than a uniform
shape, since four of the six deliberately differ. A drift guard enumerates
`hooks/*.js` and fails if a terminating hook is ever added without a row.

Writing those tests surfaced two hooks that emit a block decision in their JSON
body and exit 0. Both were checked rather than assumed, and neither is a
fails-with-success: `gsd-read-injection-scanner.js` is PostToolUse, where the
tool has already run and exit 2 has no meaning, and `gsd-cursor-subagent-start.js`
follows Cursor's JSON-body protocol. They are deliberately left alone — a
mechanical sweep to `deny()` would have broken exactly these two.

Verification runs on the remote runner.

Refs #3911

* fix(#3838): the commit validator says when it could not validate

#3911 claims to subsume #3838. Measurement said otherwise, so this closes it
for real rather than by assertion.

`set -euo pipefail`, added earlier on this branch, does NOT fix #3838: bash
exempts a command used as an `if` condition from `set -e`, and all three of the
hook's swallow-and-pass sites are exactly that shape. Verified against the
hardened hook with a node shim that fails only the classifier call — a
non-conforming commit still exited 0 with empty stdout AND empty stderr,
indistinguishable from "your commit conforms". That is the defect verbatim.

All three sites named in #3838 now capture the real exit status instead of
consuming it as a condition, and each distinguishes its genuine negative from
"could not run":

- the classifier: 0 = is a git commit, 1 = genuinely not one, anything else =
  could not classify. Its `node -e` now wraps the require and the call in
  try/catch and exits 3 on a throw, so a broken require chain can never be
  mistaken for `isGitSubcommand` legitimately returning false — which is the
  arm that matters, since `token-scanner.cjs` is a gitignored build artifact
  and a fresh checkout lands there.
- the opt-in config read and the JSON command extraction get the same
  treatment.

On "could not run" the hook emits a diagnostic to stderr naming which check
failed and why, then exits 0. The issue confirms this is safe — it is a
PreToolUse hook, so stderr does not disturb the JSON protocol — and ranks it
the smallest sufficient fix. The gate still fails open, but it can no longer do
so silently, which is the whole complaint: a validator that disables itself
quietly costs more than one that is absent, because it is trusted.

Both controls are unchanged and pinned by tests: a conforming commit still
passes silently, a non-conforming one still exits 2 with its existing block
payload. The defect test asserts stderr is non-empty and names the failure; it
fails against the pre-fix hook.

Verification runs on the remote runner.

Refs #3911, #3838

* docs(#3911): document the hook crash-policy contract

Reference and Explanation via a new docs/features fragment (FEATURES.md is
generated from it), INVENTORY rows for the three new hooks/lib files, and an
ARCHITECTURE note on the hooks section.

How-To: docs/how-to/declare-a-hook-crash-policy.md, indexed from docs/README.md
— a hook author now has to choose and declare a crash policy, which is more
than one step and crosses into which harness protocol their hook speaks. It
covers allow/deny/crash, writing an ON_CRASH reason that is actually useful,
when a deny needs a distinct stderr payload, the two hooks whose harness reads
a JSON-body decision and must NOT use deny(), and what to do when a check
cannot run at all — with #3838 as the worked example.

Refs #3911

* test(#3911): prove the seam actually ships, and stop hand-rolling temp cleanup

Two review findings.

The acceptance criterion 'hooks/dist/** stays in parity via the build seam
(lint:hooks-runtime-build-seam)' was misstated and unmet: that lint checks
something else — that a hook requiring a compiled gsd-core/bin/lib module also
calls ensureRuntimeBuild(). Nothing exercised that the three new hooks/lib
files reach hooks/dist/lib at all. That gap is not theoretical: #770 is a
recorded ship-blocking bug where a new hook never shipped because a copy list
missed it. The suite now builds dist through the repo's own ensureBuiltHooks(),
byte-compares each shipped copy against its source, and spawns a child that
requires the SHIPPED dist copy and denies — which is what catches a copy that
exists but cannot resolve its sibling registry.

gsd-validate-commit.sh hand-duplicated mktemp/run/rm three times; one idempotent
trap on EXIT replaces them, guarded so cleanup cannot alter the exit status.
Behavior-neutral across five cases, with temp-file counts taken before and
after each run.

Refs #3911

* fix(#3911): stage transitive hook lib requires, not just one level

The remote run returned 7 failures across 3 real causes.

The important one is a PRODUCTION bug this phase exposed rather than caused.
`writeCursorHooksJson` scanned each hook script for `./lib/X` requires exactly
one level deep and never re-scanned the lib files it staged for their own
sibling requires. Nothing had a transitive lib dependency before, so the gap
was invisible. Adding hook-exit.js -> cli-exit.js -> exit-code-registry.js
made real Cursor installs ship a bundle that dies at require time with
MODULE_NOT_FOUND. It now walks to a fixed point, and a real installed Cursor
hook runs to completion.

The staging harness in shared-hooks-dir-resolution hand-copied its fixture, so
the injection scanner crashed at require time and its exit-1 was being read as
a policy decision. Migrated to copyScriptWithDeps, which walks the require
graph — the repo's recorded rule for this class, since adding another
copyFileSync keeps it alive for the next person.

The missing-lib-source test in cursor-hook-workspace-roots hardcoded which lib
file it expected to be named in the abort message; the same throw now fires for
a different file first. Its assertion is unchanged in substance — staging still
must abort rather than ship a broken hook — only the name is no longer pinned.

The last one was my own test asserting an uppercase reason code. Measured
against origin/next: the pre-change hook emits the same lowercase
'config_unreadable', so the test was wrong, not the migration. Corrected to the
real value rather than making the code match the test.

Verification runs on the remote runner.

Refs #3911

* chore(#3911): regenerate the cursor install-tree golden

The staging fix means a Cursor install now correctly carries the two
transitive lib files it was silently missing. Additive only — no path was
removed. The golden diff is the evidence the packaging defect was real.

Refs #3911

* chore(#3911): backfill the changeset PR number

Refs #3911

* fix(#3911): a git probe that timed out is not a negative

A macOS CI lane failed three deny cases at 2084ms, 2112ms and 2177ms — just
past the 2000ms budget these hooks give their git probes. The three that passed
took 72ms, 595ms and 651ms. Under shard contention `git rev-parse` overruns,
the hook reads the non-zero result as "not a git repo", and allows with exit 0
and empty stdout AND empty stderr. Under load, the guards silently stop
guarding. That is ADR-3889's thesis exactly, sitting inside the security hooks
this phase is about.

The repo had already recognized the class in one place — gsd-cursor-subagent-start.js
fail-closed-denies on `git_timed_out` (#3045) — but nowhere else.

`hooks/lib/git-probe.js` classifies a probe's outcome, distinguishing a real
non-zero exit from ETIMEDOUT, a signal kill, and a spawn failure, rather than
folding all four into `status !== 0`. Three guards route their eight git probes
through it.

The resolution is the same shape #3838 took, and the same one that issue
endorsed as smallest-sufficient: fail open, but loudly. **No exit code changes
on any path** — a developer on a loaded machine is still not blocked, which
keeps #3911's declaration-pass contract intact for exit codes. What changes is
that the hook now says on stderr which probe could not answer, instead of
presenting silence as a clean verdict.

Scope was checked across every hooks/*.js, not just the three that failed:
gsd-agent-isolation-guard spawns no git; gsd-statusline's two probes gate only
a cosmetic display segment, not an allow/deny decision, and are left alone.

The C2 deny assertion was a real-race test — it demanded exit 2 while a slow
git legitimately yields 0. It now requires the hook to either deny, or allow
with a diagnostic naming the probe that could not run; a silent allow still
fails, so the assertion is not vacuous. A deterministic regression stubs git on
PATH to sleep past the budget rather than waiting for load to reproduce it.

Verification runs on the remote runner.

Refs #3911

* test(#3911): a PATH shim cannot intercept the hooks' git spawn on Windows

The deterministic timeout regression stubbed git on PATH and asserted the
guard reports rather than silently allows. It passes on Linux and macOS and
failed on Windows in 83ms and 176ms — the stub was never invoked at all.

Mechanism: the hooks call spawnSync('git', args) with no shell:true, so on
Windows CreateProcess resolves git.exe only and never a PATH .cmd shim. The
git.cmd branch could not have worked and is removed rather than left implying
a Windows path that does. Adding shell:true to the hooks to serve a test would
change product behavior and widen an injection surface, so the case is skipped
on win32 only, with the mechanism written into the skip reason so a future
reader does not 'fix' it that way.

Linux and macOS keep the coverage, and macOS is where the underlying fail-open
was actually caught.

Refs #3911

---------

Co-authored-by: sim <sim@local>
2026-08-27 22:21:10 -04:00
Tom Boucher
bf2332e67c fix(#3582): route every hook's compiled-module require through the self-heal build seam (#3629)
* test(3582): failing-first cold-tree coverage and the seam drift lint

On a plugin-channel install the compiled gsd-core/bin/lib/*.cjs are legitimately
absent (ADR-457 build-at-publish; the npm package builds before publishing, a raw
tree materialization never does). gsd-tools.cjs calls ensureRuntimeBuild() before
requiring ./lib; no hook does, so the isolation guard's Cannot-find-module lands in
its fail-closed catch and is misreported as an unreadable dispatch-isolation
configuration, blocking every executor dispatch.

These tests fail on that: cold-tree runs of the isolation guard, statusline, cursor
guard and update worker, plus the seam's actionable build error surfacing instead of
the generic misreport.

Also adds the drift lint the acceptance criteria require, with a fixture proving it
CAN fail — a guard never shown to fail is worthless. It is red here by design: it
flags today's unfixed hooks, which is exactly the defect.

* fix(3582): route every hook's compiled-module require through the self-heal seam

RED proven at 5b174b0d: 11 failures — the cold-tree runs for the isolation guard,
cursor guard and update worker, the fail-closed-with-actionable-message assertion, and
the lint's own real-tree check.

The compiled runtime library is produced by build:lib and gitignored (ADR-457,
build-at-publish). The npm package builds before publishing; a plugin-marketplace or
git-clone install materializes the raw tree and never does, so on that channel those
modules are legitimately absent. The self-heal seam added by #2002 exists to heal exactly
this, and the CLI entrypoint already calls it — no hook did. The isolation guard's
Cannot-find-module therefore landed in its fail-closed catch and was reported as
'could not read or resolve dispatch-isolation configuration', so an ARTIFACT ABSENCE was
misdiagnosed as an unreadable project config and every executor dispatch was blocked.

All SEVEN affected files now call the seam before their first compiled require. The issue
named four; a scan found six; implementing it surfaced a seventh — the shared isolation
sentinel helper, used by BOTH guards, which requires two compiled modules itself and
would have defeated the guards' own fix on a genuinely cold tree. Same defect class, so
fixed here rather than left as a known-broken remainder.

Failure posture is deliberately split by hook kind:
- Gates (agent isolation guard, cursor subagent start) surface the seam's actionable
  build error distinctly instead of swallowing it into the generic text, and stay
  fail-closed — a genuinely unreadable project config still DENIES exactly as before.
- Cosmetic and detached hooks (statusline, update worker, update check, update banner)
  DEGRADE rather than crash: the statusline draws on every render and the worker is a
  detached process, so a build failure there must not take down the prompt.

The npm path is untouched: the seam's already-built fast path returns immediately, so
prebuilt installs pay nothing and behave bit-for-bit as before.

Adds a drift lint, wired into the CI lint chain, so the invariant is enforced rather than
remembered — without it the next hook to add a compiled require reintroduces the class
silently. It is proven able to fail: a fixture hook requiring a compiled module without
the seam is flagged, and one that uses the seam is not. Verified directly — on the
unfixed tree it named all seven offenders; with the fix it passes.

While writing the lint's comment stripper, a naive whole-text block-comment regex ate its
own fixture, because this repo's comments legitimately spell the compiled-lib glob whose
star-slash reads as a comment opener. Rewritten as a line-based scanner with a regression
test pinning that case.

* fix(3582): test the three untested seam call sites and assert typed reason codes

Two independent reviews converged on the same major gap: the fix wired the seam into
seven files but only four had cold-tree tests. The adversarial pass put it plainly —
deleting the shared isolation-sentinel helper's seam call would not have failed any test
in the diff. That file was my own addition beyond the issue's four, so it shipped
untested; that is now closed.

- Shared isolation-sentinel helper: its seam call is only reached when .planning is NOT
  directly under cwd, and every existing cold-tree fixture puts it there, so the early
  return always fired first. Now covered, and proven load-bearing by mutation: with the
  call removed the spy records zero seam invocations and the test fails.
- update-check hook and update-banner hook: cold-tree tests added asserting the DEGRADED
  VERDICT — the fallback cache filename, and silent suppression when the package name
  degrades to null — rather than merely 'did not throw'. The banner hook previously had
  no test file at all.

Standards violation fixed: two tests asserted on free-form prose via assert.match against
a JSON reason string, which CONTRIBUTING bans by name — its own BAD example is exactly
that. The ESLint rule only covers readFileSync/spawnSync text, so tooling did not catch
it. Both isolation guards now emit a machine-readable reason_code from a frozen enum,
following the repo's existing REASON convention, and the tests assert that instead. The
human-readable message is unchanged for operators; only the assertion target moved.

The duplicated degrade boilerplate across the three cosmetic hooks was deliberately NOT
extracted, and the reason is recorded at each site: both viable shapes — a
path-parameterized helper, or a ceremony-only wrapper — defeat the drift lint's per-file
literal co-occurrence check, so extracting would require the lint to special-case its own
helper. Triplication is the lesser evil while the lint stays a co-occurrence scan.

The lint's header now states what it does and does not catch (literal quoted requires
only; hooks/ scan root), so a future reader does not over-trust a guard that a
concatenated path or a require inside a non-hooks helper would evade.

* chore(3582): regenerate the committed install-tree fixtures

Adding a new shipped hook helper changed the install tree, and those fixtures are
committed-and-derived (regen:derived / gen:install-tree), so 12 'install tree — <runtime>'
tests failed on 541a1913. Regenerated rather than hand-edited.

The delta across all 15 runtime fixtures is exactly two lines — the new helper under both
its hooks/ and gsd-hooks/ install paths — and nothing else, so the regeneration pulled in
no unrelated drift.

This is the bookkeeping ripple a new file under hooks/ carries; it was not visible from
lint:ci, which passed both before and after.

* chore(3582): backfill changeset PR number (#3629)

---------

Co-authored-by: sim <sim@local>
2026-08-18 14:11:23 -04:00
Tom Boucher
268ca7e32d fix(#3504): harden hook injection patterns and force-add guard (#3510)
* test(#3504): add failing-first parity, fail-closed, and bypass suites

* fix(#3504): harden hook injection patterns and force-add guard

* test(#3504): stage the scanner lib dependency in shared-hooks fixture

* chore(#3504): backfill changeset pr number

* test(#3504): build parity samples from fragments for the ci scan

---------

Co-authored-by: sim <sim@local>
2026-08-14 21:19:35 -04:00
Tom Boucher
470389f3a2 chore(#3212): tokenizer-first for stateful grammars — a shared scanner — Phase 3 (#3424)
* feat(#3414): promote git-cmd.js token-walk into a shared scanner, fix #3169

Phase 3 of epic #3212 (ADR-3212 §4). New src/token-scanner.cts generalizes
hooks/lib/git-cmd.js's proven token-walk (#3129 — "has not re-opened"):
tokenizeShellLike (quote-aware shell tokenizer, byte-identical port) and
indentWidth (bullet-nesting depth).

git-cmd.js migrates onto tokenizeShellLike with zero behavior change
(parity-asserted against every existing #3129 fixture in
tests/worktree-safety.test.cjs's folded block); isGitSubcommand's phases
1-3 (env-prefix skip, executable check, global-option consume) extracted
into skipToSubcommand, shared with the new extractBranchArgument (git
checkout -b / git branch <name>) — a new capability exercising the seam
on the domain the ADR names, not a migration of existing duplicated logic
(none existed).

Fixes #3169: src/decisions.cts's parseDecisionLines couldn't distinguish
a cross-reference bullet nested under an open decision from a fresh
malformed declaration attempt. An earlier bold-run-content-classification
design was tried and disproven against the repo's own existing FIX-B
fixtures (D-02, "no colon no dash") before being adopted — both have
identical shape under any content-only rule. Nesting depth (via
indentWidth) is the actual distinguishing signal: a bullet indented
deeper than the currently-open decision's own bullet is elaboration,
folded into its text like a continuation line, never tested against the
parse-miss guard. A bullet at the same-or-shallower indent is unchanged.

Scope-narrowing disclosed, not silent: of the ADR's four named bugs
(#3197, #3169, #2570, #2528), three no longer need this phase's work.
were independently fixed and closed since the ADR was authored — #2570's
fix is already a correctly-bounded regex per the ADR's own decidability
test (no scanner needed); #2528's fix is a deliberate, twice-reviewed
non-scanner design (its own code comment records a scanner-based attempt
that regressed a symmetric case and was reverted) that this phase does
not disturb. Only #3169 required new work.

get_impact: isGitSubcommand CRITICAL/196 affected symbols,
parseDecisionLines CRITICAL/164 affected symbols (ADR §6 due diligence).

Six-gate ripple: .gitignore, eslint.config.mjs, docs/INVENTORY.md,
docs/INVENTORY-MANIFEST.json (regenerated), CONTEXT.md glossary.

Design: .gsd/phase/chore-3414-tokenizer-first-seam/40-design.md
Test matrix: .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3414): add required fast-check property tests per code review

TESTING-STANDARDS.md:169 requires at least one fast-check property test
for any module that implements parsing — src/token-scanner.cts had none,
an orthogonal Standards-axis review finding. Adds two seeded property
tests (mirroring Phase 1/2's fast-check-setup.cjs convention):
indentWidth counts exactly a generated leading-space run; tokenizeShellLike
round-trips a generated array of whitespace/quote-free words joined with
single spaces.

The design doc's own "no property test needed" rationale was wrong — it
argued no algebraic law applied, but the standard is unconditional for
parsing modules regardless of whether one "feels" applicable. Corrected
in .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md.

Also fixes two Spec-axis wording drifts the same review found between
the design doc and the shipped code (doc-only, no behavior change):
extractBranchArgument's documented signature dropped an unused
subVariants parameter that was never implemented, and the #3169
fail-first fixture description corrected from "15-decision plan via
cmdDecisionCoverageVerify" to the actual compact 3-decision analog via
the real blocking gate, check.decision-coverage-plan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3414): add changeset for #3169 fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3414): backfill changeset pr number to 3424

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-13 23:08:34 -04:00
Tom Boucher
8f75e27554 fix(#3045): fail closed when an executor dispatch drops its resolved isolation (#3069)
* feat(#3045): deny an executor dispatch that drops its isolation flag

Every isolation gate already resolved correctly. The resolved value then reached
the executor through a prose instruction telling the model to substitute it into
a call the model composes itself, and nothing verified the substitution. When it
was dropped, the executor edited and committed in the user's primary checkout
with no consent and no warning.

A prose backstop would be the same class of artifact as the defect, so this is a
shipped PreToolUse hook on the Agent tool. It fires at the instant of the call
rather than being read once at the top of a workflow, which is the only placement
the model cannot skip.

The guard is inert unless it can positively establish that this is a GSD project,
that the project resolves to harness isolation, and that the dispatch targets an
executor. A non-GSD repo has no invariant to enforce. Where it cannot read the
configuration at all, it denies rather than assuming, with its own reason -- a
guard that cannot verify must not answer safe. A malformed payload allows rather
than throwing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#3045): extend the isolation guard to Cursor

Cursor is the second of only two runtimes that resolve harness isolation, so
shipping the guard for Claude alone left half the exposed surface unguarded while
the changeset implied it was covered.

The two runtimes fail differently. On Claude the harness flag is a per-dispatch
kwarg the model must copy into a call it composes, and the defect is that it can
be dropped. On Cursor the flag is --worktree, which applies to the whole session,
and the subagent-start payload carries no isolation field at all. There is no
flag to check, so the guard verifies the effective state instead: whether the
workspace is genuinely running outside the user's primary checkout. That is a
stronger check than the Claude one because it tests reality rather than intent,
and it is commented so nobody later rewrites it into a flag check.

Isolation is established two ways, either sufficient: the workspace resolves to a
linked git worktree, or it sits under the worktree root Cursor manages. The
second matters because a directory Cursor placed there is a legitimate isolated
session even before it becomes a distinct git worktree, where linkage alone would
report no repository.

Detecting linkage required a new primitive rather than the existing context
resolver. That resolver short-circuits on finding a local .planning directory
before it ever compares the git directory to the common one -- and an isolation
worktree normally has its own checked-out .planning. Reusing it would have read a
correctly isolated session as unisolated and denied it, which is the failure
direction that gets a guard switched off. The comparison is now its own
shortcut-free function that the resolver delegates to after its own shortcut, so
existing behavior is unchanged, and the case that would have broken is pinned.

The subagent type is checked before any configuration is read, so an unreadable
config cannot deny a dispatch this guard would never have enforced against.

The input-schema comment on the Cursor hook documented only the fields common to
every event and omitted the ones specific to this one. That omission cost a
halt during this work; it now documents both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3045): enforce the resolved dispatch decision, not the host capability

The guard keyed on the registry's dispatch.isolation, which says only that a
runtime is CAPABLE of harness worktrees. The decision that actually governs a
dispatch is the one the workflow resolves after gating, and that legitimately
comes out as sequential in three documented cases: a project setting
use_worktrees false, a per-plan submodule intersection, and the base-check
auto-degrade. The workflow tells the model to omit the flag in exactly those
cases, and the guard was denying every one of them.

The third case matters most. The preceding fix made the base-check degrade on
git timeouts and a missing git binary, where it had previously answered "safe".
That correction is right, and it means a transient hang now degrades to
sequential far more often than before -- so the two changes composed into a trap
where the workflow behaved exactly as designed and the guard blocked it.

The workflow already resolves isolation in shell, deterministically, which is
what makes it a trustworthy source in a way the model-authored call is not. It
now records that resolved value through a dedicated verb, and both guards read
it first. A fresh record is authoritative, so sequential dispatches pass
untouched. Absent or stale, the guards fall back to the capability check
combined with the project's use_worktrees setting, which still covers the case
that never reaches the workflow.

Also widened the matcher to accept Task alongside Agent, since a host that names
the tool Task would otherwise leave the guard silently inert while implying
coverage; stopped assuming Claude when no runtime is declared, which is the
shipped default and would have demanded a Claude-only argument elsewhere; and
made a non-git project inert rather than denied, since advising a worktree
session is not actionable without a repository.

The original diagnosis never modeled sequential mode as legitimate. That
omission is what let this through, and it is now recorded there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3045): record at resolution and bind the record to its dispatch

Two independent reviews converged on the same failure: the guard was fail-open in
a default install, so it did not catch the defect it exists to catch. A shipped
project carries no runtime key, which made "runtime not confidently known" the
common case rather than a corner one. A record asserting that isolation was
required but carrying no flag then fell through to a capability lookup that
answered "none", and the dispatch was allowed. The flag itself only arrived from
a second shell block -- the same block a model dropping the argument would also
skip. A test had pinned that behavior as intended.

The record is now written by the resolver, as an unavoidable consequence of
asking for the value, rather than by a step the model is told in prose to go and
run. A guard against a prose-carried value cannot itself depend on prose. Mode,
flag and identifiers are written together and atomically, so the flagless window
is gone, and a record asserting isolation with no resolvable flag now denies
instead of degrading. Runtime is also resolved from the installer's own recorded
default, which makes confident resolution the normal case.

The per-plan submodule gate degrades after the phase-level decision and never
re-recorded, so a plan that legitimately ran sequentially was denied against a
still-fresh phase record. It now records its own, scoped to the plan.

A record also authorized any dispatch for four hours. One phase degrading to
sequential could silently license an unisolated dispatch in the next. Records
now carry phase and plan, the guards require them to match, and the window is
minutes rather than hours -- the resolver rewrites it before every dispatch, so
a long window bought nothing and only widened the hole.

The flag validator rejected any value beginning with two dashes, which is exactly
the form Cursor and Windsurf declare, so their real value could never have been
stored. Writer and reader also derived the record path differently and diverged
inside a linked worktree without local planning state.

The predictable path remains a way to silence the control without leaving a trace
in the diff. It grants no access an agent with shell does not already have, so it
is documented as accepted rather than redesigned around.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3045): correct the staleness boundary and unmask a vacuous parity test

The remote runner returned twenty failures. One was a real production defect the
boundary case existed to catch: a record whose age exactly equalled the staleness
window was treated as fresh, so it stayed authoritative for one tick past its own
expiry. Freshness is now strictly inside the window.

The parity test meant to stop the two guards' executor lists from drifting could
never have failed. Its project fixture was a bare directory rather than a
repository, so the non-git inert branch answered before the executor list was
ever consulted. It asserted agreement it never actually measured. The fixture is
now a real repository, like every sibling in the file.

A test also asserted that Windsurf declares the worktree flag. It does not --
Windsurf resolves to no isolation by design, having no named concurrent dispatch
to isolate. The test claimed a registry fact that was never true, and a comment
in the resolver repeated it. Both corrected, and the test now proves what it
should have all along: that the parser accepts any bare flag value, rather than
one runtime's supposed value.

The new guard was missing from the bundled-hook whitelist, which is the surface
that decides what actually ships, and the per-plan gate had gained calls to the
launcher without the preamble those calls require. The changeset carried
parenthetical product descriptions the purity rule forbids.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3045): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3045): make the guard tests hold on Windows

Two tests redirect HOME to control where the installer-persisted runtime default
is read from. Node resolves the home directory from USERPROFILE on Windows and
never consults HOME, so both silently read the real runner profile, found no
recorded runtime, and asserted against a project the hook had not recognised. The
production code was already correct in asking the platform rather than the
variable; only the tests were wrong to assume one variable answers everywhere.
The helpers now mirror the override onto both.

The symlink spoofing test also created a directory symlink unconditionally, which
needs elevated privileges on Windows. It survived on this runner, but it would
fail on any host without them, so the creation is now attempted and the test
skips explicitly when it cannot be done -- a bare return would have counted as a
pass and hidden the gap.

Skipping alone would have left the platform uncovered, so the behaviour it proves
is now also driven in-process through an injected realpath, following the seam
already used for the clock. That case no longer depends on privileges at all, and
the end-to-end test keeps its original assertions wherever symlinks work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 23:42:16 -04:00
Tom Boucher
c7c2fe3c2b fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd (#2680)
* fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd

gsd-cursor-session-start.js and gsd-cursor-stop.js both resolved the project as
path.join(process.cwd(), '.planning', 'STATE.md'). Under the cursor-agent CLI,
hooks are invoked with cwd set to the Cursor config dir (~/.cursor), not the
workspace — so the lookup always missed. sessionStart could only ever emit the
"no .planning/ workflow found" nudge and stop's verify-work reminder could never
fire, even with .planning/STATE.md sitting in the workspace. Slash commands were
unaffected, which is why only the hook layer looked blind.

Both hooks already buffered stdin into `raw` and never parsed it; the payload's
workspace_roots carries the real path.

Multi-root was left open in the report ("first root vs any root"). Resolved
forward: prefer the first root that actually carries .planning/STATE.md, so a
workspace whose GSD project is not the first root still resolves — strictly
better than first-root-only and identical to it in the single-root CLI case.
Falls back to roots[0], then to cwd, keeping IDE behavior unchanged if the IDE
ever invokes hooks from the workspace.

The resolver is duplicated verbatim across the two scripts rather than shared via
hooks/lib/: these hooks ship standalone, and a new hooks/lib/ file must be
registered in the GENERATED installer's GSD_HOOK_LIB_FILES allowlist — the
installer-omits-shipped-file class that yields MODULE_NOT_FOUND at runtime. Per
CLAUDE.md "Generative Fix Divergence", the duplication carries a parity assertion
so the copies cannot drift.

Failing-first, demonstrated by direct invocation with cwd != workspace:
  pre-fix  sessionStart -> "no .planning/ workflow found"   stop -> {}
  post-fix sessionStart -> ".planning/STATE.md is present"  stop -> reminder

tests/fix-2587-cursor-hook-workspace-roots.test.cjs spawns the real scripts as
child processes with a cwd lacking .planning/ and workspace_roots pointing at it.
Boundary coverage on the roots array (0 / 1 / 2 entries), plus malformed-JSON
fail-open, junk-entry filtering, the parity assertion, and a guard that neither
script resolves .planning from cwd again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* fix(#2587): extend workspace_roots fix to subagentStart; keep cwd a candidate

Three findings from the isolated review, all fixed.

1. MISSED SITE (high). gsd-cursor-subagent-start.js carried the identical
   defect at line 43 — its own header documents workspace_roots in the input
   schema, but it resolved .planning/ from process.cwd() anyway. Under the
   cursor-agent CLI that meant every Cursor subagent (planner, executor,
   verifier) started with "no .planning/ workflow found" and no phase context.
   The report named only sessionStart and stop; the defect class was wider.
   Verified pre-fix vs post-fix by direct invocation with cwd != workspace.

2. SEMANTIC NARROWING (medium). The first cut searched only workspace_roots and
   fell back to cwd solely when the array was EMPTY. So when roots were supplied
   but none carried .planning/ while cwd did, the hook reported absent — where
   the pre-fix code, which always used cwd, reported present. That contradicted
   the fallback's own stated intent of preserving IDE behavior. cwd is now a
   CANDIDATE in the search (`[...roots, process.cwd()]`), so the fix is a strict
   superset of both the old behavior and the CLI fix, never a narrowing.

3. STALE GOLDEN FIXTURES (high, would have failed CI). The golden-install-parity
   fixtures store a content hash per installed file; these three hooks appear in
   13 of the 19 runtime fixtures. Regenerated via `npm run gen:golden` — the
   diff is exactly the three hook hashes in exactly those 13 runtimes.

Tests extended: subagentStart resolution via workspace_roots; the stop hook's
absent branch (previously only session-start's was covered); an explicit
regression guard that a project at cwd is still found when roots miss; parity now
asserts all THREE copies byte-identical; and the cwd guard sweeps the whole
RESOLVING_HOOKS list so a future hook in this family cannot be left on cwd.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* refactor(#2587): extract cursor workspace resolution to a shared hooks/lib module

The duplicate-plus-parity-test approach was the wrong call. The reported issue
named two hooks; a third (subagentStart) had the identical defect. That is the
signature of a systemic problem, and three copies of a resolver guarded by a
parity assertion is a divergence risk maintained by hand rather than a fix.

hooks/lib/cursor-workspace.js is now the single implementation. All three Cursor
hooks require it; none defines a local copy. Divergence is prevented
structurally instead of by asserting three copies stay byte-identical.

The reason duplication looked necessary was real, and is fixed properly here
rather than worked around: Cursor sets hostBehaviors.skipSharedHooksInstall
(#2089), so it never reaches the installer's bulk hooks/lib copy — it was the
ONE runtime shipping these hooks WITHOUT hooks/lib (verified against all 19
golden fixtures: cursor had the hook scripts, no lib). A naive require would
have thrown MODULE_NOT_FOUND at load, BEFORE each hook's own try/catch, wedging
every session on precisely the runtime this bug is about.

writeCursorHooksJson (src/runtime-hooks-surface.cts) now stages the hooks/lib
helpers the staged scripts actually require, discovered by scanning their
require('./lib/…') calls rather than a hardcoded name — so a future helper
cannot be silently omitted. This is narrower than flipping
skipSharedHooksInstall, which would wrongly pull in every shared hook.
cursor-workspace.js is also added to GSD_HOOK_LIB_FILES so uninstall and the
manifest manage it for the runtimes that do receive hooks/lib.

Verified against a REAL install (runMinimalInstall, cursor/global): the helper
is staged, and all three INSTALLED hooks resolve the workspace end-to-end from a
cwd that is not the project.

Also closes the review gap that the stop hook was excluded from the
cwd-candidate regression loop — it now sweeps RESOLVING_HOOKS. The byte-parity
test is replaced by a structural guard (every hook requires the shared module,
none redefines it) plus a new install test asserting the helper is staged and
the installed hook actually loads against it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* fix(#2587): fail loud on a missing hook lib source; drop unsubstituted version marker

Two findings from the installer-focused review.

H1 — the staging step's `if (!fs.existsSync(libSrc)) continue;` silently defeated
the very guarantee it was added for. Reproduced: delete hooks/lib/cursor-workspace.js
from source, run the cursor install — it exits 0, prints "Done!", and ships the
three hook scripts with an EMPTY hooks/lib/. The installed hook then throws
`Cannot find module './lib/cursor-workspace.js'` at load, before its own
try/catch, wedging every session — and nothing surfaces until a user hits it.
The scan protected against a required-but-UNLISTED helper while leaving
required-but-MISSING wide open (typo, bad rebase, an accidental delete).
It now throws: a missing helper source is a packaging bug and aborts the install.

M1 — hooks/lib/cursor-workspace.js carried a `gsd-hook-version: <placeholder>`
marker that NOTHING substitutes: copyLibDir stamps .sh files only, and
writeCursorHooksJson's staging applies just the colon-to-dash rewrite. Verified
the literal was reaching disk on both the bulk (--claude) and Cursor
(--cursor) paths. hooks/lib/git-cmd.js — the only pre-existing hooks/lib/*.js —
carries no such marker, so this was newly introduced, not inherited. Marker
removed, matching that precedent, with a note on why. (The explanatory comment
deliberately does not spell the token out, or it would reintroduce the literal.)

M2 — the require-scan regex demanded the exact compact form, so
`require( "./lib/x.js" )` would silently fail to stage its helper and compound
H1. Now tolerant of interior whitespace and either quote style.

Regression test added for H1 — the reviewer confirmed the invariant had zero
coverage repo-wide: a source tree carrying the hooks but no hooks/lib/ must make
writeCursorHooksJson throw rather than produce a broken install.

Re-verified end to end: the missing-source case throws, no unsubstituted literal
ships, and the installed hook still resolves the workspace from a foreign cwd.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* chore(#2587): backfill changeset pr number (#2680)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 19:57:58 -04:00
Tom Boucher
dacc23137a feat(3347): opt-in auto-update of knowledge graph after main HEAD advances
Closes #3347

Config:
- Add graphify.auto_update (default false) to manifests:
  sdk/shared/config-defaults.manifest.json, config-schema.manifest.json

Hook:
- hooks/gsd-graphify-update.sh — PostToolUse Bash matcher
  - Gates: tool_name=Bash, HEAD-advancing git op, CI=unset, in git repo,
    current branch == default branch (git.base_branch override or main/
    master/trunk fallback), graphify.enabled && graphify.auto_update both
    true, graphify on PATH, no live PID lock
  - Writes .planning/graphs/.last-build-status.json with status=running
    synchronously, then detaches hooks/lib/gsd-graphify-rebuild.sh
- hooks/lib/gsd-graphify-rebuild.sh — detached rebuild runner
  - PID-lock acquire + trap-on-exit cleanup
  - graphify update . then cp graphify-out/* → .planning/graphs/
  - Status file rewritten to status=ok|failed with exit_code, duration_ms,
    head_at_build
- Portable detach (subshell + disown, no setsid dependency)

Installer:
- bin/install.js: register hook as PostToolUse Bash matcher (5s timeout)
- Add to gsdHooks uninstall list and expectedShHooks warning list

Planner / researcher status surface (issue #3347 reviewer must-have AC):
- agents/gsd-planner.md and agents/gsd-phase-researcher.md
  load_graph_context steps now read .last-build-status.json and surface:
  running → "rebuild in flight"; failed → "auto-rebuild FAILED at {ts},
  context is from prior build"; ok with stale head_at_build → "HEAD has
  advanced since last build"

Settings:
- get-shit-done/workflows/settings.md adds "Graph auto-update" question
  with No-Recommended default; bullets and update_config block updated

Inventory:
- docs/INVENTORY.md hook count 12 → 13 with new row
- docs/INVENTORY-MANIFEST.json regenerated

Tests:
- tests/feat-3347-graphify-auto-update-config.test.cjs (8 tests):
  isValidConfigKey accepts graphify.auto_update, CANONICAL_CONFIG_DEFAULTS
  default false, config-set round-trip, sibling key preservation
- tests/feat-3347-graphify-auto-update-hook.test.cjs (18 tests):
  all bail paths (non-Bash, non-HEAD-advancing, enabled=false,
  auto_update=false, CI=true, non-default-branch, missing graphify bin,
  live-PID lock), dispatch path with mock graphify bin (sync running
  status + detached transition to ok/failed), stale-PID lock, all five
  HEAD-advancing command matchers, git.base_branch override

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-15 11:59:43 -04:00
Tom Boucher
7827e1ddee fix(#3129): replace bypassed bash regex with token-walk git-cmd.js classifier (#3141)
* fix(#3129): replace bypassed bash regex with token-walk git-cmd.js classifier

Root cause: gsd-validate-commit.sh used:
  if [[ "$CMD" =~ ^git[[:space:]]+commit ]]
This regex silently bypasses Conventional Commits enforcement for:
  git -C /path commit -m ...     (working-directory prefix)
  GIT_AUTHOR_NAME=x git commit   (env-var prefix)
  /usr/bin/git commit -m ...     (full-path executable)

Fix: introduces hooks/lib/git-cmd.js with isGitSubcommand(cmd, sub) —
a token-walk classifier that handles all four forms by:
  1. Skipping leading VAR=VALUE env assignments
  2. Validating the git executable (basename check for full-path support)
  3. Consuming git global options (-C <path>, --git-dir=, -p, etc.)
  4. Checking the subcommand token

The hook delegates to this classifier via node shell-out. node is
already called twice in this hook (config check + JSON parse), so no
new runtime dependency.

This becomes the single source of truth for all hooks that gate on
git subcommands (pre-commit-review-gate, post-push-verify, etc.).

Regression test: 27 assertions — tokenize correctness, 12 must-match
cases (including all 3 bypass forms), 8 must-not-match cases, 3 source
checks. All are real behavioral tests, not string comparisons.
Suite: 7035/7035. Closes #3129.

* fix(lint+hook+changeset): allow-test-rule, fix HOOK_DIR quote injection, fix changeset pr+typo
2026-05-05 15:02:15 -04:00