c4b6dbd486e9bb802132088098f24cd5d05db3ac
11 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
723ea08dc2 |
fix(#4016): imperative-override injection patterns tolerate filler words (#4061)
* fix(#4016): imperative-override patterns tolerate filler words The narrow imperative-override family tolerates no filler between the verb and the noun, so a planted "Forget all of your instructions" (measured in a real public transcript) matched none of the 14 patterns and both consuming hooks stayed silent. One combined filler-tolerant pattern is appended; the narrow four stay untouched to keep the change merge-friendly. Known trade-offs, disclosed in #4016: linter-doc prose like "ignore rules on a single line" now trips a LOW advisory, and the overlap with the narrow patterns means one sentence can count twice toward severity thresholds. Regression tests assert the previously-missed phrasings fire in BOTH consuming hooks (gsd-prompt-guard and gsd-read-injection-scanner), not just in the raw pattern list, per the agent brief in #4016. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015aGr6fvmDMLznT7TvTXrsb * chore(#4016): changeset fragment for PR #4061 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015aGr6fvmDMLznT7TvTXrsb * test(#4016): pin the disclosed linter-doc FP as single-pattern LOW, never blocking Review follow-up on PR #4061: the combined filler-tolerant pattern's disclosed false-positive class (linter-doc prose such as "use eslint-disable-next-line to ignore rules on a single line") was documented in prose only. Two tests now pin it: - the prose matches exactly ONE shared pattern (the #4016 combined pattern, not a narrow one), so it cannot silently start double-counting toward the 3+ HIGH threshold; - through the real gsd-read-injection-scanner subprocess with security.injection_blocking=true, the prose yields a single-finding LOW advisory and no block decision — with an in-test positive control proving a 3+-pattern payload DOES block in the same directory, so the non-blocking assertion cannot pass vacuously. Samples are fragment-built like the existing SAMPLES rows so this file's own diff does not trip the CI injection scanner. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#4016): replace the five narrow imperative-override patterns with one superset The first cut appended a filler-tolerant combined pattern next to the five narrow verb patterns. Both consumers count one finding per matching pattern toward the severity threshold, so the overlap made one sentence count twice: "Ignore previous instructions. Forget your instructions." scored 2 (LOW) on next and 3 (HIGH, blockable) on the branch. It also left `override` out of the combined pattern. Replace the narrow family (ignore x2, disregard, forget, override) with ONE superset pattern over ignore|disregard|forget|discard|override. At least one filler (all|of|the|your|my|system|previous|prior|above|earlier) must sit between verb and noun, enforced by a lookahead with no repetition; the two noun-less/bare forms the old list accepted (`disregard (all) previous`, `forget instructions`) are kept as explicit tails so the new pattern is a strict superset. Bare "override rules" / "ignore instructions" are ordinary repo prose (6 measured hits across docs and source) and stay unmatched. Corpus measurement over 3019 .md/.js/.cjs/.mjs files (injection-sample tests excluded): the old family hit 2 lines, the new pattern hits 3, the only new one being a documented injection example in planner-reversibility.md that the old family missed (the issue's own class). Tests: SAMPLES reshaped to the 10-entry list; superset proof table (17 legacy phrasings, each matching exactly one pattern); five issue phrasings including `override all of your previous instructions` counted exactly once through both hook subprocesses; double-count regression (1 finding, LOW); design pin that bare verb+noun matches nothing; linter-doc FP pin split into bare (silent) and determined (single LOW, never blocks). All fragment-built; the CI prompt-injection scanner reports 0 findings on every touched file. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012rvqL1wk6s8dQEXsFm4RB1 * chore(#4016): changeset body in the canonical bold-lead format .changeset/README.md Format: a leading bold change sentence, then an em-dash explanation. Also drops the verbatim planted phrase from the body so the rendered CHANGELOG line does not trip the pattern it describes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012rvqL1wk6s8dQEXsFm4RB1 * fix(#4016): render a bounded pattern label in the prompt-guard advisory, pin plural prompts Review round 4 of PR #4061 left two nits open. 1. gsd-prompt-guard.js pushed `pattern.source` verbatim into its typed finding and, through renderFinding, into the user-facing advisory. With the #4016 superset pattern that source is 300 characters, so a genuine hit surfaced an advisory dominated by a raw regex dump. The read scanner has trimmed its equivalent since #3523 (`\s+` -> `-`, strip `()\`, cut at 50). That transform is hoisted into hooks/lib/injection-patterns.js as `describePattern` and used by BOTH hooks, so one finding renders the same label everywhere. Byte-identical to the scanner's old inline output for all 10 patterns (measured). No new staging dependency: both hooks already require this module. 2. The noun alternation `prompts?` had no positive coverage for the plural branch. One filler-regression row now exercises `... previous prompts ...` and runs through the existing once-per-hook, exactly-one-pattern loops. The parity test's prompt-guard count assertion moves off substring-matching the advisory prose onto the typed `findings` surface added in #3546, per CONTRIBUTING's raw-text-matching prohibition. New test: the superset source exceeds the bound (positive control), the prompt guard never embeds it, and both hooks carry the identical label in `findings[0].match`. Tests: parity, read-scanner, kimi field-shadowing, prompt-injection-scan, hooks-crash-policy, dead-exports: 206 run, 196 pass, 0 fail, 10 pre-existing platform skips. eslint clean; changeset lint ok; hooks runtime-build-seam lint ok; the CI prompt-injection scanner reports 0 findings on the PR diff. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016GGp8kEB5zCDmJ6TYHP1Nj --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
4499933807 |
fix(#3802): resolve the heredoc body before validating the commit subject (#3816)
* fix(#3802): resolve the heredoc body before validating the commit subject With hooks.community: true, gsd-validate-commit.sh blocked EVERY heredoc-form commit with CONVENTIONAL_COMMITS_VIOLATION regardless of the message, including Claude Code's own documented idiom: git commit -m "$(cat <<'EOF' feat(auth): add login flow EOF )" Reproduced before changing anything: conforming heredoc -> exit 2; plain -m "feat(auth): add login flow" -> exit 0. Root cause is the extraction regex `-m[[:space:]]+"([^"]+)"`. Bash `[^"]` matches newlines, so the capture ran from the quote after -m to the FINAL quote at `)"`, swallowing the whole span. `head -1` then returned the literal `$(cat <<'EOF'` as the subject, which can never satisfy Conventional Commits. Fixed by not answering a regex bug with another regex. hooks/lib/git-cmd.js already exists because "a naive regex misses all three" invocation forms, and extractBranchArgument is the established precedent for pulling an argument off a git command line. extractCommitSubject joins it on the same tokenizeShellLike seam — which, checked first, already returns the entire heredoc span as ONE token, leaving only "resolve the body to its first line" as new logic. Because the walk starts at the subcommand, `git -C <path> commit` and env-prefixed invocations now extract correctly too — forms the raw string scan never handled. Deliberately unchanged, and pinned as such: a glued `-mfeat: x` and `--message=...` still yield no message, exactly as the regex left them. The fix stays scoped to the reported defect rather than widening on a true observation. Two things I got wrong and corrected by measuring rather than reasoning: - I expected `git commit -m ""` to be blocked. Checked against the ORIGINAL hook: allowed before, allowed now, identical. The scanner drops the empty token so it takes the null path. My expectation was wrong, not the code. - That exposed a false comment I had just written, claiming the exit-status split prevents silently allowing `-m ""`. It does not. The split IS load-bearing, but for a heredoc whose body's first line is blank, which resolves to an empty subject and is correctly blocked. The comment now names the real case and records that `-m ""` is not it. Tests at both layers: 9 unit rows on extractCommitSubject beside its sibling in tests/worktree-safety.test.cjs, and 5 behavioral rows piping real PreToolUse payloads through the hook in tests/hooks-opt-in.test.cjs. Replacing firstLineOfMessageArg with a plain first-line return reds 8 of them across both files. (A first mutation attempt silently no-opped and reported green — the mutated body is echoed in the transcript for the run that counted.) Out of scope, per the issue: the hooks.commit_types config surface, split off by the maintainer as #3811 and explicitly sequenced after this. Verified: `npm run lint:ci` exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3802): confine the fix to heredoc resolution, closing four regressions Codex review of the first attempt. It was right, and the finding is one my own rules already name: a true observation is not a licence to widen the diff. The first attempt replaced the shell's `-m` extraction with a token walk. That looked like the better abstraction — this module exists precisely because a naive regex misses invocation forms — but selecting WHICH argument is the message was never the defect, and changing it regressed four forms that upstream allowed, plus opened a bypass: - `git commit -- -m WIP` -- introduces pathspecs; `-m` is a path - `git commit --amend && echo -m WIP` a later command's flag became the message - `git commit -m "" --allow-empty-message` the shared scanner drops empty tokens, so the next flag became the message - `git commit -m WIP` unquoted argument - `-m "WIP notes <<EOF\nfix: smuggled subject"` was ALLOWED — the opener was recognised unanchored, so validation skipped past the real, non-conforming subject. An enforcement bypass, not a misclassification. Now confined to the actual defect. The shell's `-m` capture is restored byte for byte, and only the subject-from-message step is delegated, to a PURE STRING helper `resolveCommitSubject()` that never tokenizes. Verified as a differential against the upstream hook run inside the real tree: the only behaviours that change are the two intended heredoc rows (2 -> 0); all four forms above read identical, and the bypass case blocks. That differential also corrected my own control. An earlier comparison ran the upstream hook from a scratch directory, where its `lib/` could not resolve `../../gsd-core/bin/lib/token-scanner.cjs`, so the classifier failed open and reported exit 0 for everything. That made a real regression look pre-existing. Re-run inside the tree, `<<-"TAG"` (a double-quoted tag nested in the double-quoted argument) is genuinely pre-existing — the capture truncates — and is now recorded as a known limitation rather than silently "fixed". Also fixed from the review: - `<<-` strips leading TABS from body lines; returning the raw line blocked a conforming message. - a non-identifier tag such as `END-MSG` is a valid bash word and was rejected. - an immediately-following terminator is an EMPTY message, not a subject. - a node/library failure now falls back to the previous `head -1` instead of skipping validation, so a broken extractor degrades to old behaviour rather than becoming a new silent-allow path. Tests strengthened per the review: the opener-spelling rows now assert BOTH directions per spelling, since "conforming passes" alone would also pass if the resolver returned an empty subject for a spelling it failed to parse. Added differential rows pinning the five previously-allowed forms, and a row for the bypass. Dropped two rows whose comments claimed the raw scan could not handle `-C`/env-prefix invocations — it could; the claim was wrong. Replacing resolveCommitSubject with a plain first-line return reds 9 rows across both files. (Mutant body echoed in the transcript; an earlier mutation attempt on this branch silently no-opped and reported green.) Verified: `npm run lint:ci` exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3802): keep the installed hook runtime-neutral `hooks/lib/git-cmd.js` ships into every runtime, including hermes and qwen, where tests/install.test.cjs enforces that no Claude reference leaks into the installed tree. My JSDoc named the idiom after the runtime that documents it. Reworded to describe the SHAPE rather than the vendor; the runtime is still named in the changeset, which feeds CHANGELOG.md where such references are allowed, and in the tests, which are not installed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#3802): backfill changeset pr number The fragment shipped with the documented `pr: 0` placeholder, which the changeset lint treats as always-silent, because the number does not exist until the PR is opened. Backfilled to 3816 now that it does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3802): close the truncated-capture hole, add the required test artifacts Review round 1. Major 3 was the one that mattered, and it disproved a claim I had stated in falsifiable form — the PR body said only two behaviours change; the differential found five. Major 3 — an embedded `"` truncates the `-m` capture, so the resolver received a PREFIX of the real subject and the length gate measured the wrong string. Before this fix the whole form was blocked outright, so the gate was unreachable; the fix opened the path and then mismeasured it. A new enforcement hole, so it is CLOSED here rather than declared. Closed precisely rather than bluntly. A first attempt refused to resolve any body with no terminator, which also blocked commits whose SUBJECT was intact and whose quote sat further down the body — a false positive of its own. Truncation is only fatal to the line it lands IN, and a captured line is complete exactly when another line follows it, because the capture kept its newline. So an unterminated body whose subject line is followed by more text stays measurable; only a subject line running to the end of a truncated capture falls back to the opener, which fails the format gate exactly as this form did before the fix. Major 1 — fast-check property rows for the new parser, via the shared seeded setup helper rather than requiring fast-check directly, per repo convention: totality (a security property here, since an exception on this path fails OPEN), idempotency, and that the result is always a single line drawn from the input — the third catches a resolver that concatenated or trimmed while satisfying the first two. Major 2 — the 72-char gate is now exercised at {71, 72, 73} on the RESOLVED heredoc subject, with the fixture length asserted so a mis-built fixture cannot silently pass. 92 chars did not show which side of `> 72` the code sits on. Minor 1 — leading blank body lines are skipped, as git's cleanup=whitespace does. A conforming commit written that way was still blocked, which is the same defect class #3802 reports. Nit 1 — a backslash-escaped delimiter (`<<\EOF`) is now the same delimiter rather than failing closed on a delimiter that includes the backslash. Nit 5 — changeset trimmed from 2,208 chars of design note to the user-visible change. Mutation discipline, including a correction to my own: dropping the truncation guard reds the unit rows, and the pre-review naive shape reds the hook-level row too. My first mutant did NOT distinguish the hook row — removing the guard made an empty slice and blocked for an unrelated reason, so the row passed and looked proven. Only mutating to the actual pre-review shape showed it discriminates. Verified: `npm run lint:ci` exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#3802): measure the subject as git does — strip trailing whitespace, split CRLF git's cleanup=whitespace strips whitespace at BOTH ends of a line; the resolver handled only the leading direction, so a 72-char subject with trailing spaces measured 75 and stayed blocked — the defect class #3802 reports, surviving one round further (review of #3816, Major 2). The resolved subject now drops trailing spaces and tabs; the plain non- heredoc path is untouched, keeping the fix confined to heredoc resolution. The length-gate boundary rows gain dirty fixtures: 72+3 trailing spaces passes, 73+1 stays blocked on LENGTH. split('\n') left \r on every body line, so on CRLF input the delimiter never matched: the truncation guard was inert, an empty message resolved to 'EOF\r', and a real 72-char subject measured 73. Split on /\r?\n/ (Minor 3). The three property tests never reached the parser — the pinned-seed fc.string corpus contained no newline and no opener, so every property reduced to f(s) === s (Major 1). The generator now constructs heredoc- shaped input (all opener spellings, <<- tabs, optional terminator, CRLF) and each property asserts a floor on inputs its corpus actually resolved. All new rows proved failing-first against the pre-fix resolver. Also records the unquoted-delimiter expansion limit as one JSDoc sentence (Informational 5). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3802): close two recognition bypasses, pin the dquoted-delimiter limit Codex whole-PR review found two enforcement bypasses in the resolver: - The opener's path prefix was \S*, which accepted `id;/bin/cat` — the resolver then validated the heredoc BODY while bash runs `id` first and git's real subject is id's OUTPUT. The prefix is now a path-character class; any shell metacharacter fails recognition and the form falls back to the opener line and the format gate. - The blank-line skip used JavaScript trim(), whose Unicode whitespace class skips lines git KEEPS: a NBSP first body line resolved to the SECOND line while git's real subject is the NBSP line (verified against git stripspace — the c2a0 bytes survive). Blank is now git's ASCII space/tab only; a Unicode-blank line is returned and fails the format gate, the same fail-closed direction git takes. Both proven failing-first at resolver AND hook level. Also: the <<"TAG" spelling is recorded as a documented limit — the -m capture stops at the delimiter's own quote so the caller can never deliver it (fail closed; widening the capture would change every embedded-quote case) — with a hook-level row pinning the limit; and the derivation property no longer accepts '' unconditionally, only for heredoc-shaped input, so a conditional constant-'' regression can't satisfy the corpus floor unnoticed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3802): recognition whitespace is ASCII, and '' answers to the generator Codex round 2: the opener's \s accepted Unicode whitespace bash does not split on — $(<NBSP>/bin/cat was recognized here while bash reads <NBSP>/bin/cat as the executable NAME, so recognition claimed a substitution that does not run cat. Every whitespace position in the recognition is now [ \t], the same ASCII rule as the blank-line skip, proven failing-first. The derivation property's ''-acceptance now consults GENERATION-TIME metadata: the heredoc generator records whether it built an empty message (terminator reachable, all scanned lines ASCII-blank, <<- tab stripping accounted for), and '' is accepted exactly then — a resolver conditionally degrading to '' on non-empty heredocs now fails, closing the residual round-1 permissiveness without re-deriving resolver logic. The changeset no longer overstates the opener spellings: it names the capture-deliverable set and the documented <<"EOF" limit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3802): nothing after the terminator escapes measurement Round-3 BLOCKER: everything after the heredoc terminator was silently discarded, so `-m "$(cat <<'EOF'\nfeat: ok\nEOF\n) <200 a's>"` — one 200+ char real subject once bash substitutes — measured 8 chars and dodged COMMIT_SUBJECT_TOO_LONG, a hole the base did not have. The canonical idiom's tail is exactly one closing-paren line; any other tail now falls back to the opener line and the format gate, the pre-fix behaviour for the whole form. Proven failing-first at resolver and hook level, including the glued-text and second-substitution variants. Also from round 3: `cat<<'EOF'` (no space) is legal bash and now resolves — the token before << is still literally cat; the env-prefixed and option-terminated spellings join the JSDoc KNOWN LIMIT list instead (fail closed, modelling bash prefix words is cost with no reported user); the changeset states the embedded-quote truncation limit for the message body, not just the <<"EOF" spelling; the dquoted unit and hook rows now cross-reference each other; and the fast-check setup helper's docstring no longer claims property-file exclusivity. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3802): glued text outside the closing quote must not shrink the measurement Codex on the round-3 guard: bash concatenates -m "$(…)"suffix into ONE argument, but the capture holds only the quoted part — so the resolver measured the heredoc body (8 chars) for a 200+ char real subject, a net-new length-gate bypass the base did not have (base measured the opener and blocked). When the closing quote is followed by anything but whitespace or end-of-command, the hook now skips the resolver and keeps the pre-fix first-line subject: the heredoc form fails the format gate exactly as on base, and the plain single-line form keeps base behavior unchanged — both pinned as differential rows, the glued-suffix row proven failing-first against the unguarded script. The property generator's ''-oracle now models the post-terminator guard it previously predated: expectEmpty requires the FIRST reachable terminator to be followed by the one canonical closing-paren line, so a resolver regressing to '' on a non-canonical tail (e.g. a body line that doubles as an early terminator) fails the derivation property instead of being blessed by stale metadata. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: retrigger CI — the previous wave never started (Actions queue stall) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#3802): only resolve a heredoc whose body bash does not rewrite Round-4 review found two net-new enforcement bypasses: commands the base hook blocked (exit 2) that this branch allowed (exit 0). Both reproduced as a base-vs-head differential against the real hook, not inferred. The predicate "may I resolve this?" was computed from the resolver's input string alone, while two of its determinants live outside that string: 1. WHICH -m quote arm produced the input. Inside -m '...' bash performs no command substitution, so $(cat <<'EOF' is literal text and git's real subject is the opener line. The resolver ran on both arms, so all four delimiter spellings went 2 -> 0 on the sq arm — reachable by the ordinary slip of typing ' for ". The hook now records MSG_QUOTE and gates the resolver on dq; sq keeps head -1, exact base parity. 2. WHETHER the delimiter suppresses expansion. Only <<'D', <<"D" and <<\D do; a bare <<D is expanded by bash before git sees it. Resolving the literal dodged the format gate (feat: $UNSET_VAR reaches git as feat:) and the length gate (feat: ${LONG} reaches it at any length). The opener regex now separates the backslash-quoted and bare alternatives and refuses the bare one — the same fail-closed rule the metacharacter, truncation and post-terminator guards already follow. A test row asserted exit 0 for a bare-delimiter body, so the suite defended the second bypass and the fix could not land without editing a test that read as intentional. That row and its two unit counterparts now assert the block, per RULESET.TESTS.delete-bad-tests. Two unrelated rows used <<-EOF to exercise tab stripping; they move to <<-'EOF' so each tests what it names. Scoping the adjacency guard to the matched arm — required by the fix above — also removes a spurious block (round-4 Minor 1): a double-quoted heredoc whose body mentioned a glued single-quoted token tripped the sq arm. The JSDoc claimed <<"EOF" was unreachable through the caller and that the bare-delimiter gap was pre-existing. Round 4 disproved both; both corrected here, along with the matching changeset sentence. Verified: 7 bypass commands now block at head (was allow), the #3802 fix and plain-form parity are unchanged across 8 control commands, hooks-opt-in 44/44, worktree-safety 401/401, property-test non-vacuity 73/200 against a floor of 20, lint:ci exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG * fix(#3802): resolve only where the captured text is provably git's subject Codex review of the full PR found two more inputs where the validated text is not the subject git receives, both net-new bypasses (base 2 -> head 0), plus one escalation of round-4 Minor 2. All reproduced here against the real hook and confirmed against real commits before fixing. BLOCKER — the matched -m need not be git's message. The capture is a search over the whole command and the double-quoted arm runs first, so it could select a -m that is not the subject at all. git concatenates multiple -m values and takes the FIRST as the subject, so git commit -m 'WIP first' -m "$(cat <<'EOF' … )" commits the subject `WIP first` while the hook validated the heredoc. Same for an unquoted earlier -m, for a heredoc after `--` (a pathspec, not a message), and for one belonging to a later `&& echo`. The mis-selection is pre-existing; resolving it is what made it a bypass. The hook now resolves only when nothing before the matched -m could have been an earlier message, an end-of-options marker, or another command. BLOCKER — cleanup mode is part of the predicate. The resolver skips leading blank lines and strips trailing whitespace because git's DEFAULT cleanup=whitespace does. Under --cleanup=verbatim git does neither, so a 72-char subject plus three trailing spaces is committed at 75 bytes while the hook measured 72 — COMMIT_SUBJECT_TOO_LONG dodged. This one hides from `git log --pretty=%s`, which strips trailing whitespace in its own output; the raw commit object shows 75 vs 72. Any named mode other than whitespace, in either the --cleanup= or -c commit.cleanup= form, now refuses to resolve. MAJOR — recognition trusted any path ending in /cat, so a planted `../evil/cat` printing `WIP injected` had its heredoc body validated while git's real subject was `WIP injected`. Only a bare `cat` or an absolute path is recognised now. A bare `cat` shadowed on PATH is a documented residual and is not fixable from a string — nor a meaningful boundary, since planting an executable already allows running git directly. The changeset and the JSDoc both asserted that a `"` anywhere in the message blocks. Measured false: a `"` on a later body line resolves fine, because the subject completes before the truncation point; only a `"` in the subject line blocks. The changeset also listed <<"EOF" as covered when it measures 2/2. Both rewritten to claim only what is measured, and the residual false positives are now named. Verified: 4 + 2 + 3 new bypass commands now block, with non-vacuity controls proving the default path still resolves; all round-4 maintainer blockers stay closed; the #3802 fix and plain-form parity unchanged across 7 controls; hooks-opt-in 47/47, worktree-safety 402/402, property non-vacuity 73/200 against a floor of 20, lint:ci exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG * fix(#3802): scope the cleanup-mode guard to the command outside the message The guard scanned the whole $CMD for `--cleanup=` / `commit.cleanup=`, and the heredoc BODY sits verbatim inside $CMD, so any conforming message that merely MENTIONED the token was refused, fell back to the opener line, and was blocked with CONVENTIONAL_COMMITS_VIOLATION. These are ordinary English in this repository, whose own hooks and docs discuss cleanup modes constantly. Reproduced against the real hook: `fix: document commit.cleanup=strip behavior` blocked, the same message without the token allowed (review of #3816, round 5 — BLOCKER). Scoping to $MSG_PREFIX alone, as prescribed, would have reopened the round-4 length-gate bypass the guard exists for: git accepts the flag on EITHER side of -m, and `git commit -m "<heredoc>" --cleanup=verbatim` is caught today only because the scan is command-wide. Measured, not assumed. The scan now covers MSG_PREFIX + MSG_SUFFIX — the whole command minus the one span that is message text — joined with a space so a token cannot be forged across the seam. Swept the guard class rather than the reported instance. The adjacency guard does not share the defect: an in-body `-m "foo"bar` is refused by the already-documented embedded-quote capture limit (any `"` in the subject line truncates the capture), and an in-body `-m ` without quotes resolves and is allowed. Deliberately untouched. Both directions pinned failing-first: the three false-positive rows red against the unscoped guard, and the trailing-flag row reds against prefix-only scoping. Each mutation was echoed back to prove it landed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XogDtuuuGQEfsWaLSaZCLB * fix(#3802): read commit options the way bash hands them to git Round 6 reported the adjacency guard scanning all of $CMD for a glued `-m "..."`, so a glued -m belonging to a chained-after command refused a heredoc that was never truncated. Glue is a property of the ONE character following the matched span, so that character is now the whole window. Separators and redirections are excluded because bash does not concatenate across them: in `-m "msg"&& echo hi` the argument ends at the quote, so there is no truncated capture to defend against. An independent full-PR pass then found three accept-direction defects this PR had introduced in earlier rounds, each measured against a real commit by reading the raw commit object — `git log --pretty=%s` strips the trailing whitespace that makes the length wrong and hides it: --cle=verbatim git accepts any unambiguous prefix of a long option, so the mode was set by a token that is not the literal --cleanup. 75-byte subject recorded, 72 measured. -am 'WIP first' git reads this as -a -m, so the real subject is `WIP first` and the heredoc is only the second message. The scan looked for a standalone -m. --clean""up=verbatim bash removes quotes before git sees the -""m argument, so a spliced spelling is the same option and matched no literal. The two option-name scans now read their window with quote characters removed, which is what bash does to it, and the cleanup class covers git's abbreviations. The adjacency test deliberately keeps the raw text: it asks about a literal character position, not an option name. Narrowing the cleanup window to git's own command segment was tried and reverted. `;`, `&` and `|` end a command only outside quotes, and this is a substring scan, not a parse: an unconditional trim cut the window short on `--author "a&b"`, and a quote-aware trim still cut it on `--author a\&b`. Each hid a real trailing --cleanup=verbatim and accepted a 75-byte subject. The resulting false positive — a --cleanup carried by a chained command refuses the commit — is documented and pinned instead. Refusing a commit git would take is recoverable; accepting an over-long subject is not. Sixteen rows in tests/hooks-opt-in.test.cjs. Seven mutations, including both reverted narrowings, so no dead end can be reintroduced silently. * fix(#3802): close six accept-direction bypasses in the resolve guards Round 7's FIRST-MESSAGE GUARD Major does not reproduce. Measured against the real hook in a complete tree at the reviewed head: the classifier gate runs before any guard, so `git add -A && git commit …` (git->add stops on a non-commit subcommand) and `cd dir && git commit …` (the first executable is not git) exit 0 without a guard being evaluated. The control is the proof — a subject the bare form blocks with CONVENTIONAL_COMMITS_VIOLATION exits 0 in both chained forms, so the hook never validated them and cannot be over-blocking them. The guard is unchanged; scoping this scan to $MSG_PREFIX alone is what reopened the round-4 trailing-flag bypass. The class was real, though, one shape further out: `FOO=bar; git commit …` IS classified and then refused, because assignment detection is prefix-anchored and the tokenizer does not split operators. Pinned as a counterexample and disclosed rather than generalised away; narrowing it means changing isGitSubcommand, the shared git-commit detector every gating hook uses, and it fails closed. Six accept-direction bypasses are fixed. Each let the hook resolve and ALLOW a commit whose real subject the rules refuse; the three that turn on git's recorded subject were confirmed against the RAW COMMIT OBJECT, since `git log --pretty=%s` strips trailing whitespace and hid two of them: --cleanup=whitespace -m <72+spaces> --cleanup=verbatim git kept 75 bytes -mWIP -m <heredoc> git recorded `WIP` --mes=WIP -m <heredoc> git recorded `WIP` -\m WIP -m <heredoc> git recorded `WIP` git commit --amend --no-edit \n echo -m <heredoc> echo's argument read --squash=HEAD -m <heredoc> `squash! …` Causes: one BASH_REMATCH inspected only the FIRST cleanup directive while git applies the last, so multiplicity now refuses rather than guesses at an argument order a substring scan cannot recover; the option scan required a trailing space or `=`, missing attached values and long-option abbreviations; dequoting removed quotes but not the syntactic backslashes bash also removes; the separator scan omitted newline; and --squash/--fixup have git compose the subject, so the supplied message is not the subject at all. Every fix widens refusal, the direction this file documents as recoverable. The multiplicity count first broke the hook outright: the script runs under `set -euo pipefail` and grep exits 1 when it matches nothing, which is the common case, so every ordinary commit died at exit 1 with no verdict. Guarded, and only caught because the probe runs the real hook rather than the scan. Five new rows, all five proven red against the pre-fix hook, each carrying a non-vacuity assertion that the canonical single-`-m` heredoc still resolves. Changeset corrected on three counts: "all fail-closed" was wrong (persistent commit.cleanup fails OPEN, as do the -C/-c/-F/-t message sources), "global options are all walked through" was too broad, and the chained-before claim now states what is measured. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018FUAVz49BghqxoJgwt7EW9 * fix(#3802): stop the separator and glue classes matching a literal backslash Round 8's Major, with two corrections to its account. `;`, `&` and `|` are metacharacters inside `[[ ]]`, so an inline bracket class must escape each one. POSIX bracket expressions have no escape mechanism of their own, so on bash 3.2 -- the system /bin/bash on macOS, already a supported target here per the `declare -A` ban in tests/install.test.cjs -- those backslashes reach the regex engine and add a literal `\` to the class. bash 4+ consumes them, which is why this is invisible on a modern bash. The hazard is specific to bracket expressions: `\(` outside one is made literal correctly on every version, and the subject validator and the `-m` capture classes were checked and are unaffected. The prescribed fix is not taken, because it does not parse. Inline `[;&|]` is a bash SYNTAX ERROR on 3.2 and on 5.3 alike -- the backslashes exist to get the metacharacters past the `[[ ]]` parser, so removing them leaves an unparseable script. Each class is held in a variable and expanded unquoted on the right of `=~` instead, which is a plain regex on both versions. One root cause, consequences in BOTH directions. The reported half is the separator scan over-blocking. The half not reported is the accept direction, and it is the more serious: the glue class is NEGATED, so on bash 3.2 a backslash-glued suffix fell inside the exclusion and the hook RESOLVED a heredoc it should have declined -- measured exit 0 on 3.2 against the unfixed hook, exit 2 everywhere else, with a letter-glued control refused in all four cells. The reported repro is not actually fixed by this, and the changeset says so. A `\`-newline line continuation carries a literal newline, which the round-7 separator guard refuses on every bash, so that shape stays blocked with or without this change. Narrowing the newline guard is not attempted: telling a continuation from a separator by substring scan is the class that was tried twice in earlier rounds and reverted both times, and an escaped backslash sitting immediately before a real newline is indistinguishable from a continuation. Disclosed as a known fail-closed limit instead. Every new row runs under each bash on the machine. Against the unfixed hook both bash 3.2 rows go red while all four bash 5.3 rows stay green -- written the ordinary way these rows would run under PATH bash, pass against the broken hook, and prove nothing. Two non-vacuity controls per interpreter prove the validator is reached rather than passing everything. All 8 rows of the established differential harness are byte-identical before and after on both versions: no regression, no new refusal. * fix(#3802): remove the $ of a dollar-quote from the option-name scans Independent round-8 review, accept direction. The option-name windows are dequoted so they match "the command as bash hands it to git" -- round 6 removed quote characters, round 7 removed syntactic backslashes. Both passes missed that bash has two further quoting forms whose introducer is a `$`: `$'...'` and `$"..."`. Removing the quote characters alone left that `$` stranded INSIDE the option name, so `-$"m"` dequoted to `-$m` and matched no literal, while bash passed a real `-m` to git. Measured on bash 3.2.57 and 5.3.15 against a real repository: the hook allowed git commit --allow-empty -$"m" WIP -m "$(cat <<'EOF' fix: a perfectly ordinary conforming subject EOF )" with exit 0, and `git cat-file -p HEAD` recorded the subject `WIP`. The comparison that establishes this is HEAD-internal, not a differential: the same command spelled `-m WIP` is refused (exit 2). The merge-base refuses EVERY heredoc form, including a perfectly conforming one, so its exit 2 on this input says nothing about whether any guard fired -- it is the absence of the feature, not a working check. The same miss covered `$'m'`, spliced `--message`, `--cleanup`, `--squash` and `--fixup`. An option NAME finished by a command substitution -- `--clean$(printf up)=verbatim` -- is a different problem and gets its own guard: bash runs a program to complete the name, so the argv git receives is not derivable from this string at all, and resolution is refused rather than guessed. The guard is scoped to the NAME: the class is a `-`-leading token whose characters up to the substitution contain no `=`. A substitution supplying a VALUE -- the ordinary `--author="$(git config user.name)"`, spaced or glued, in either window -- is untouched and still resolves, pinned in both directions. It is a SHAPE, not a segmentation of the command line; segmenting was tried twice in earlier rounds and reverted both times, and that reasoning stands. Both new rows fail against the unfixed tree with their own assertions, proven in a complete worktree at the previous head rather than a hook copied out of its tree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy * fix(#3802): recognise a canonical cat, not any absolute path ending in /cat Independent round-8 review, accept direction. Round 4 restricted heredoc-opener recognition to an absolute path, after a relative `./cat` was measured being trusted to echo its stdin. It stopped at "absolute", so any absolute path ENDING in `/cat` was still trusted -- the same claim the round-4 reasoning had rejected one spelling earlier. Measured on bash 3.2.57 and 5.3.15 against a real commit: with an executable at `/.../fake-cat/cat` printing `WIP injected`, the hook validated the conforming heredoc body and allowed the commit (exit 0) while `git cat-file -p HEAD` recorded the subject `WIP injected`. The same command through `./cat` was already refused, which is the control that shows this is the round-4 class one spelling out rather than a new one. Recognition is now the canonical system locations -- bare `cat`, `/bin/cat`, `/usr/bin/cat` -- which is the only identity claim a string can support. `/usr/local/bin` is deliberately excluded: it is user-writable on ordinary machines, which is the plantable case this guard exists for. Anything else falls back to the opener line and the format gate: fail closed, exactly the pre-fix behaviour for the form. The pre-existing residual is unchanged and still documented: a bare `cat` shadowed earlier on PATH is indistinguishable here, and is not a meaningful boundary -- anyone able to plant an executable on PATH can run `git commit` directly. This hook stays an authoring guard, not a security control. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy * fix(#3802): an option name carrying a shell expansion is unresolvable Independent review, round 9, accept direction. Four more spellings, and a change of strategy that is the actual point of this commit. Rounds 6, 7 and 8 each tried to EMULATE what bash does to an argument before git sees it -- round 6 removed quote characters, round 7 syntactic backslashes, round 8 the `$` that introduces a dollar-quote -- and each round review found another transform that had been missed. Round 9 found four more. All measured on bash 3.2.57 and 5.3.15 against a real repository, each with the plain spelling of the same command as its control (refused, exit 2) and `git cat-file -p HEAD` for the subject git actually recorded: -$'\155' WIP hook 0, real subject `WIP` ANSI-C octal -> m -$'\x6d' WIP hook 0, real subject `WIP` ANSI-C hex -> m -`printf m` WIP hook 0, real subject `WIP` backtick substitution x= … -${x}m WIP hook 0, real subject `WIP` parameter expansion -? WIP hook 0, real subject `WIP` pathname expansion and the same class through the cleanup guard, where git recorded a 75-character subject the length gate had measured as 72: --cle$'\141'nup=verbatim, --clean`printf up`=verbatim, --cle?nup=verbatim The last two settle it. An option name finished by a PARAMETER expansion depends on a variable's value at run time; one finished by a PATHNAME expansion depends on the contents of the working directory. Neither is derivable from the command string at any level of effort, so emulation cannot be completed -- not "has not been completed yet". A fifth patch in that direction would have the same shape as the previous four. The rule is therefore no longer "normalise it and match the literal". It is: an option NAME carrying a shell expansion or quoting construct is UNRESOLVABLE, and unresolvable refuses. One rule covers every spelling above and every spelling nobody has thought of yet, in the fail-closed direction. The dequoting passes are kept rather than replaced: they still normalise the deterministic removals, so the guards RECOGNISE `--clean""up=` and `-\m` as the options they are instead of merely refusing them, which keeps the existing rows meaningful. Scope is unchanged and still pinned in both directions: the class is a `-`-leading token whose characters up to the construct contain no `=`, so a construct supplying a VALUE -- `--author="$(git config user.name)"`, the backtick spelling, `--date="${NOW}"`, a glob character inside an author string, a pathspec after `--` -- still resolves. Nine such forms are asserted to pass beside the seven that must refuse. The class is bracket-only and holds no backslash, per round 8: a POSIX bracket expression has no escape mechanism, and a backslash written inside one becomes a literal member on bash 3.2. The new rows fail against the previous head with their own assertion message, in a complete worktree with the lib built, not a copied hook. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy * docs(#3802): disclose and pin the two spellings the round-9 class over-blocks A scoped review of the round-9 class asked one question -- does it refuse a conforming heredoc commit that the previous head accepted -- and found two spellings that it does. Both measured on bash 3.2.57 and 5.3.15, previous head 518d97b64 exit 0, current head exit 2: git commit -S$SIGNING_KEY -m <conforming heredoc> git commit -m <conforming heredoc> -- -*.txt Disclosed and pinned rather than narrowed, for two reasons. Narrowing is not available cheaply. Dropping the bare `$` member reopens `-$xm`: with `xm=m` bash hands git a real `-m`, which is the parameter expansion bypass the round-9 commit exists to close. Skipping tokens after `--` means deciding where git's options end from a substring scan, which is the class this file has already reverted twice for opening accept-direction holes -- a `--` inside a quoted value (`--author "a -- b"`) would truncate the window and hide a real trailing directive. And the limits are narrower than they look, because in both cases the spelling a developer actually reaches for still resolves: -S "$KEY" and --gpg-sign="$KEY" resolve '-*.txt', "-*.txt", ':(exclude)-*.txt' resolve The pathspec one is worth stating precisely: a glob only reaches git AS a pathspec when it is quoted, because an unquoted one is expanded by the shell before git is executed. So the refused spelling is not passing a glob to git at all, and the spellings that do are unaffected. Refusing a commit git would take is the recoverable direction; accepting a non-conforming subject is not. That is the trade this file already makes everywhere else, and it is made explicitly here. Nine rows pin the working spellings beside the three that refuse, so a later narrowing cannot silently drop the cases that must keep working. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy * fix(#3802): join backslash-newline continuations before the resolve guards Round 9's Major, with a correction to its diagnosis. The cited bracket classes at :223 and :260 no longer exist -- round 8 moved both into SEP_CLASS and GLUE_CLASS, and a lone backslash before -m resolves (exit 0) at the reviewed head on both bash 3.2.57 and 5.3.15. What refuses the repro is the NEWLINE a `\`-continuation carries: round 7's separator guard reads any newline in a window as a command boundary, and `git commit \` newline ` -m "$(cat <<'EOF' …` was refused for that reason. Round 8 disclosed it as a fail-closed limit; round 9 calls the idiom common and the limit a Major, and it is fixed here. It was left as a limit because "is this newline a continuation" looked like the segmentation question this file has reverted twice. It is not: bash's rule is local and character-level. A newline preceded by an ODD run of backslashes is a continuation and bash removes both; an EVEN run (`\\` then newline) is a literal backslash followed by a real newline, which IS a separator. Both scan windows are joined that way immediately after they are cut from the command and before any dequote copy is derived, in three bash-3.2-safe parameter expansions: every `\\` pair is parked on \x01, any backslash-newline that remains is a lone one and is removed, then the pairs are restored. Measured on both bashes, both directions: git commit \<nl> -m <heredoc> 2 -> 0 the fix git commit \\<nl> -m <heredoc> 2 -> 2 literal \ + real separator git commit<nl> -m <heredoc> 2 -> 2 bare newline -m <heredoc>\<nl>suffix 2 -> 2 bash glues it; the glue guard sees it glued git commit … \<nl> --allow-empty<nl>echo -m … 2 -> 2 the REAL newline still separates The prescribed `[\;&|]` is not taken: a backslash written inside a bracket expression becomes a literal member on bash 3.2, which is the round-8 defect from the other side. Rows run under each bash on the machine. The fix row fails against the previous head in a complete worktree with the lib built; the four control rows were measured against that same head and were already refused, so they pin existing behaviour rather than the change. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
d24e22b156 |
enhance(#3912): gsd-tools declares outcomes, pinned at v1 (#3983)
* enhance(#3912): gsd-tools declares outcomes, pinned at v1 ADR-3889 §4. Phase 6 already moved error()'s terminator onto the seam, so what remained was the declaration — and the pin that makes it invisible today. The census corrected two documented figures before any code changed. ERROR_REASON has exactly 25 members (the ADR and epic were right; an earlier note of mine claiming 23 was wrong and is corrected). And output({error}) is **64 sites across 9 files, not the 60 ADR-2980 ratified** — the module shape holds but the total drifted +4: frontmatter 7 not 6, phase 4 not 2, roadmap 3 not 2. That matters because this phase's criterion demands the pin be asserted over the enumerated population rather than sampled; asserting over a stale 60 would leave four sites unpinned while claiming full coverage, which is the shape of failure this epic exists to remove. The issue does not state the fact that shapes the design: output() never touches the exit code. Confirmed by reading it — it writes fd 1 and returns. So a declared outcome for those 64 sites had nowhere to be READ. The mapping was never the work; wiring somewhere for the declaration to land was. The seam already existed twice over. cli-exit.cts holds two globalThis-Symbol cells, each because the module is emitted to three locations and a module-level `let` would let instances disagree, and runMain already maps a code returned by main(). A third cell inherits that solution. output() records DEGRADED for any {error} payload — key-order agnostic, which is exactly why the "42 sites" figure undercounts — and runMain projects the cell only when main() returns nothing, so an explicit return still wins. error() maps its reason through a table over the closed 25-member enum, leaving all 278 call sites untouched; 226 of them pass no reason at all. The version gate lives in error(), NOT in projectOutcome: registered names are version-invariant there, so mapping a reason straight through would make USAGE project to 64 under v1 and break the pin on its first line. projectOutcome is left exactly as Phase 2 shipped it, DEGRADED's 0/80 asymmetry included. Proven rather than asserted. v1 is byte-identical across three real CLI paths — config-get plain, config-get --json-errors, and an output({error}) path — matching exit code and exact bytes against the pre-change build. Under GSD_EXIT_CONTRACT=v2 the same commands now exit 66 (CONFIG_KEY_NOT_FOUND -> NO_INPUT) and 80 (DEGRADED), both looked up through the registry. An anti-vacuity test pins that v1 and v2 genuinely differ for at least one reason, because without it a mapping where everything projects to 1 under both versions would satisfy every other assertion and the declaration would be theatre. A1 iterates all 25 enum members and A3 asserts over the measured 64-site population, so a 26th reason or a 65th site fails until it is given a mapping — the drift guard this phase needs, given ADR-2980's own count had drifted +4 unnoticed. Verification runs on the remote runner. Refs #3912 * fix(#3912): the outcome cell must never lower an exit code The remote run caught a fail-open that this phase introduced, in the phase whose entire purpose is removing fail-opens. `state validate --strict` on a missing STATE.md exited **0** where it must exit 1. Mechanism: `runMain` projected the pending outcome whenever `main()` returned void, and under v1 DEGRADED projects to 0 — so a `process.exitCode` already set non-zero by the command was clobbered down to success. Confirmed live against a fixture, before and after. This refutes a review conclusion recorded earlier in this phase, that the cell was "fail-closed and can never mask a failure as success". It could, and did. Recording that plainly so the assumption is not repeated: the cell's danger was never only that it might add a failure — it was that projecting it unconditionally overwrites whatever decision came before. Projection is now guarded: it may set a code only when none is set, and an already-non-zero exit code always wins. The full precedence — explicit `main()` return, then an existing non-zero exitCode, then the declared outcome — is written at the projection site. A regression test drives a void return with a pre-set non-zero code and a pending DEGRADED, and fails against the pre-fix build. The second failure was my test encoding the wrong contract, not a code defect. It asserted `output({found:false, error: undefined})` records DEGRADED because the KEY is present. `JSON.stringify` drops undefined, so the payload the user receives is `{"found":false}` — carrying no error at all, and calling that degraded would hand back exit 80 under v2 for output that reads as clean. The discriminator is a serializable error VALUE, not key presence. The test now pins `{error: undefined}` as explicitly NOT degraded, and the design doc's wording is tightened to match. Verification runs on the remote runner. Refs #3912 * docs(#3912): the versioned exit contract, and a flag defect the docs found Diataxis pass for Phase 8, plus a real fix that only surfaced because writing the how-to meant running its own examples. The docs. ADR-2980's "Revisit if" clause asked for exactly the versioned projection this phase provides, so it gets an amendment naming #3912 / ADR-3889 section 4 as that boundary: v1 stays 0 byte-for-byte, v2 projects DEGRADED to 80. The amendment also records the count drift rather than restating a stale figure — the ADR ratified 60 output({error}) sites in 9 modules; the AST-measured population is 64 across the same 9 (frontmatter 7 not 6, phase 4 not 2, roadmap 3 not 2). The pin is asserted over the enumerated 64. json-errors.md gains the outcome-declaration reference, including the precedence order a review pass got wrong and the suite refuted: an explicit main() return, then an already-set non-zero process.exitCode, then the declared outcome. Projection may only ever set a code, never lower one. A how-to is owed here and is written, not skipped. Under v1 nothing changes, so the audience is an operator opting into v2 and needing to know what the codes mean for a CI gate — a migration, which is how-to shaped. It covers turning v2 on, the code table, why 80 is "ran and reported a condition" rather than a crash, and how to split a gate that treats any non-zero as fatal. No tutorial: there is no new entry point to learn, and under the default contract a reader would be walked through observing nothing. The defect. Running the how-to's own Step 1 example returned $ gsd-tools --exit-contract=v2 state validate --strict Error: Unknown command: --exit-contract=v2 (exit 64) while the same flag trailing the subcommand worked and exited 80. The flag half-worked, by argv position. resolveContractVersion scans argv non-destructively, so the token survived into the dispatcher, which treats argv[2] as the command name. --json-errors had already solved precisely this at gsd-tools.cjs:4455, under a comment naming the hazard verbatim: "The argv splice must happen here too, otherwise the dispatcher below sees --json-errors as an unknown command." The later flag never got the same treatment. Fixed rather than documented around: the version is resolved first — which memoizes the cell and makes an invalid value throw early — and then every --exit-contract= token is spliced out of the dispatcher's argv copy. --exit-contract is now listed in TOP_LEVEL_USAGE, where it never was. The regression test pins leading position, trailing position, agreement between the two, and a loud failure on v3 rather than a silent fall back to v1. Neither review engine would have caught this: the defect is invisible in the diff, because the diff does not touch argv handling. It surfaced only from running the documentation's own example. Writing a how-to is an execution pass. Verification runs on the remote runner. Refs #3912 * fix(#3912): the flag splice has to run before the run-with-timeout return An isolated review of the previous commit found that the fix did not deliver what it claimed, and that two of its own tests were weak. All three findings reproduced by execution before any change was made. The fix was placed below a return. main() intercepts `run-with-timeout` at gsd-tools.cjs:4436 and returns from there — above both the --json-errors block and the --exit-contract splice added in the previous commit. So the flag still died in leading position for that one command: $ gsd-tools --exit-contract=v2 run-with-timeout 5 -- node -e "..." Error: Unknown command: run-with-timeout (exit 64, child never ran) The previous commit message and the test's describe-block both claimed position-independence unconditionally. That was an overclaim, not a gap left open, and it is the part worth naming: the fix was verified by hand on the commands I happened to think of, and `run-with-timeout` returns before the code I was verifying. Both global-flag blocks now run above the interception, with a comment naming it so a later edit cannot slide them back down. Moving --json-errors up fixes the identical pre-existing bug for that flag, verified failing beforehand (exit 1, sdk_unknown_command). Fixing the sibling is deliberate: same defect, same block, and a known-broken twin next to a fixed one is not a resting state. Two tests were not pulling their weight. The invalid-value test was vacuous — it passed against the pre-fix build, because `--exit-contract=v3` already exited 1 there and already printed the resolve error lazily through error() -> getContractVersion. Both its assertions held before the fix, so it pinned nothing. The real discriminator is that the pre-fix build emits BOTH "Unknown command: --exit-contract=v3" and the resolve error, while the fixed build emits only the latter; the test now asserts that absence. The leading-position and leading==trailing tests asserted proxies — "not 64", "no Unknown command", "the two agree" — none of which pin a value, and all of which would survive both positions being identically broken. With a .planning directory and no STATE.md, state-snapshot exits exactly 80 under v2 and 0 under v1 in both positions. Those numbers are pinned now. The multi-token case the descending splice loop exists for is covered too, and run-with-timeout has regression tests for both flags. The lesson is narrower than "test more". Hand-verifying the production behavior does not verify that the test would have caught its absence. The pre-fix binary has to be run against the test's own assertions. Investigated and deliberately not changed: splicing before --cwd parsing degrades one diagnostic from "Missing value for --cwd" to "Invalid --cwd: <path>", but that is pre-existing — verified on the pre-fix build via --json-errors, which already did it. This change joins the pattern rather than creating it, and both forms exit 64 on malformed input either way. Verification runs on the remote runner. Refs #3912 * chore(#3912): backfill changeset pr numbers to 3983 * test(#3912): pin the reason-table invariant as set equality, not a count A graph-backed review flagged the unchecked lookup in expectedErrorCode3912. Investigated by execution: the drift guard DOES hold — for an unmapped reason under v2 the production error() yields 1 while the table yields undefined, so the assertion fails. Not a correctness defect, and deliberately NOT made tolerant, since a tolerant lookup would destroy the guard. Two real problems remained. The guard asserted the wrong invariant: it counted the TABLE's keys at 25 rather than checking they match the ENUM's values, so a renamed member keeps the count at 25 and slips past, and a 26th member leaves the table at 25 and slips past too. Both were then caught only indirectly, by an undefined mismatch producing 'must exit undefined'. It is now a sorted set equality, so the failure names the specific missing or extra reason. And the comment above it described a '?? FAIL' fallback that does not exist anywhere in the function. It now states what the code actually does, verified by running it rather than by reading it. Refs #3912 --------- Co-authored-by: sim <sim@local> |
||
|
|
2ea5efc151 |
enhance(#3911): hooks declare their crash policy (#3960)
* enhance(#3911): give hooks an exit seam that needs no build ADR-3889 Phase 7 foundation. The 19 shipped enforcement hooks hold 91 of the epic's 128 terminators and cannot reach `terminateNow` today. The obvious route — requiring `gsd-core/bin/lib/cli-exit.cjs`, as gsd-agent-isolation-guard.js already does for two other modules — is rejected. That precedent carries its own warning (#3582): those files are tsc output, gitignored and absent on a raw plugin-marketplace or git-clone install, so the hook must first call ensureRuntimeBuild() to self-heal. Making the module a hook needs IN ORDER TO TERMINATE depend on a build inverts the dependency, and its failure mode is precisely the fail-open this phase exists to remove: a guard that cannot terminate cannot deny. `lint-hooks-runtime-build-seam` already encodes that concern, and Design B would have had to add an ensureRuntimeBuild() call to all 19 hooks to satisfy it. So `hooks/lib/` becomes a third emit location for cli-exit and a fifth for the registry, preserving the invariant `src/cli-exit.cts`'s own header states: it imports nothing but node:fs and its sibling registry, and the generator dual-emits that sibling alongside each copy so a relative require resolves next to whichever copy loaded it. Shipping needed no change — build-hooks.js already declares HOOKS_SUBDIRS_TO_COPY = ['lib']. Proven, not asserted: the two files are copied into an otherwise-empty tmpdir and a child process requires them and terminates — PASS exits 0, HOOK_DENY exits 2 with the payload on both stdout and stderr. That test fails the moment the hooks copy gains a require reaching outside hooks/lib/. Also fixed inline: the registry's fifth target let any `--write` test overwrite the real committed hooks/lib/exit-code-registry.js, because the test helper derived only three of the other output paths. It now redirects all five, and a regression test asserts every committed artifact is byte-identical after a redirected write. Install-tree goldens pick up the two new shipped paths across 11 runtimes — insertions only, no removals. lint:ci was green while they were stale, so this was found by regenerating rather than by a gate. Verification runs on the remote runner. Refs #3911 * enhance(#3911): declare a crash policy, and migrate the write guard Adds `hooks/lib/hook-exit.js` — the hook-facing vocabulary over `terminateNow`, hand-written because the cli-exit copy beside it is generated: allow(payload) exit 0 deny(payload, stderr?) exit 2 crash(onCrash, payload) whichever the hook DECLARED `crash()` takes the policy as a required argument with no default, which is the whole mechanism: fail-open by accident stops being expressible. A hook must name ALLOW or DENY at the call site, and an unrecognized value terminates INTERNAL rather than guessing. Fail-open stays legal; fail-open by omission does not. `gsd-write-guard.js` is the first hook migrated, all 12 sites, and it exposed a gap in the seam. `terminateNow`'s doc comment justified its fd-2 write by citing this hook's `emitBlock` — but modeled it as sending the same bytes to both streams, when `emitBlock` actually sends full JSON to stdout and only the bare `reason` string to stderr, because Kimi's hook bus feeds stderr verbatim back to the model. Migrating as written would have turned a readable sentence into a JSON blob for Kimi-backed agents. #3911 requires both "all 19 hooks terminate through terminateNow" and "no hook's effective default changes". Those are jointly satisfiable only by teaching the seam to carry a distinct stderr payload, so `terminateNow` gains an optional third argument: omitted, behavior is byte-for-byte what it was; a string is written raw, which is exactly the Kimi case. The doc comment's inaccurate claim about emitBlock is corrected in place. Proven rather than asserted: the pre-migration file is reconstructed from HEAD and driven with the same catastrophic-shrink payload as the migrated one — exit code, stdout and stderr all byte-identical. Verification runs on the remote runner. Refs #3911 * enhance(#3911): all 19 hooks terminate through the seam Migrates the remaining 18 enforcement hooks onto allow/deny/crash. An AST walk now reports zero `process.exit(` call sites across every `hooks/*.js` — down from the 91 the census measured. Each hook with an outer catch declares its policy once, at module top, with the reason that policy is right for that specific guard: a read guard that cannot scan must not block the read; a statusline that renders every prompt must degrade rather than crash; an injection scanner must not retroactively block a result already returned. Those sentences are the deliverable — they are what turns fail-open-by-accident into fail-open-on-purpose. No hook's effective default changed. Wiring exposed two defects, both fixed here rather than noted. A SECOND stdout/stderr-splitting site turned up in `gsd-workflow-guard.js`'s `emitForceAddBlock`, matching the pattern already known from the write guard — full JSON to stdout, bare reason to stderr for the Kimi bus. It uses the `stderrPayload` argument added in the previous commit, which is now carrying its second real caller rather than one special case. More seriously, `terminateNow` emitted both streams inside ONE try, so a payload that failed to serialize aborted before the stderr write ever ran. The two windsurf guards write nothing to stdout on a block and only a reason string to stderr, so `deny(undefined, reason)` exited 2 with EMPTY stderr — a deny that silently loses its reason, which is the exact "fails with success" class this epic exists to close. The streams are now emitted independently, each with its own guard, and `undefined` means "nothing to write for this stream" rather than an error. Regression tests inject a throwing write on one fd and assert the other still receives its payload; they fail against the single-try version. Byte-identity was proven per hook, not assumed: each pre-change file is reconstructed from HEAD and driven side by side with the migrated one across its normal path, its deny path, malformed stdin and empty stdin — exit code, stdout and stderr compared. Verification runs on the remote runner. Refs #3911 * enhance(#3911): harden the three shell hooks, and pin every hook's policy `gsd-phase-boundary.sh`, `gsd-session-state.sh` and `gsd-validate-commit.sh` gain `set -euo pipefail`. The expected hazard did not materialize, and that is worth recording: every intentionally-non-zero command in all three is already the condition of an `if`/`elif`, which `set -e` never fires on, and none of them reads a possibly-unset variable or pipes through a grep that may legitimately match nothing. No `|| true` guards were needed. Each hook was still checked command-by-command before the flags went in rather than after. Twenty-one before/after cases across the three hooks — disabled and enabled, planning and non-planning, missing STATE.md, malformed JSON, the Kimi payload shape, quoted and unquoted `-m`, valid and over-long Conventional Commits — all match on exit code, stdout and stderr. The hardening is shown to actually fire, not merely added: with a stubbed `node` that fails at the JSON-emit step, phase-boundary and session-state go from silently exiting 0 with empty stdout to failing visibly with the error surfaced. No such case could be constructed for `gsd-validate-commit.sh`, whose every statement already sits inside an if-condition — recorded as unproven rather than claimed. `tests/hooks-crash-policy.test.cjs` adds the per-hook coverage the issue asks for, table-driven over all 19 hooks rather than 76 hand-written cases: normal allow, deny where a deny path exists, crash-honors-the-declared-policy, and an unclosed-stdin case — the one `process.exitCode` structurally cannot serve. The deny assertions encode each hook's ACTUAL stream split rather than a uniform shape, since four of the six deliberately differ. A drift guard enumerates `hooks/*.js` and fails if a terminating hook is ever added without a row. Writing those tests surfaced two hooks that emit a block decision in their JSON body and exit 0. Both were checked rather than assumed, and neither is a fails-with-success: `gsd-read-injection-scanner.js` is PostToolUse, where the tool has already run and exit 2 has no meaning, and `gsd-cursor-subagent-start.js` follows Cursor's JSON-body protocol. They are deliberately left alone — a mechanical sweep to `deny()` would have broken exactly these two. Verification runs on the remote runner. Refs #3911 * fix(#3838): the commit validator says when it could not validate #3911 claims to subsume #3838. Measurement said otherwise, so this closes it for real rather than by assertion. `set -euo pipefail`, added earlier on this branch, does NOT fix #3838: bash exempts a command used as an `if` condition from `set -e`, and all three of the hook's swallow-and-pass sites are exactly that shape. Verified against the hardened hook with a node shim that fails only the classifier call — a non-conforming commit still exited 0 with empty stdout AND empty stderr, indistinguishable from "your commit conforms". That is the defect verbatim. All three sites named in #3838 now capture the real exit status instead of consuming it as a condition, and each distinguishes its genuine negative from "could not run": - the classifier: 0 = is a git commit, 1 = genuinely not one, anything else = could not classify. Its `node -e` now wraps the require and the call in try/catch and exits 3 on a throw, so a broken require chain can never be mistaken for `isGitSubcommand` legitimately returning false — which is the arm that matters, since `token-scanner.cjs` is a gitignored build artifact and a fresh checkout lands there. - the opt-in config read and the JSON command extraction get the same treatment. On "could not run" the hook emits a diagnostic to stderr naming which check failed and why, then exits 0. The issue confirms this is safe — it is a PreToolUse hook, so stderr does not disturb the JSON protocol — and ranks it the smallest sufficient fix. The gate still fails open, but it can no longer do so silently, which is the whole complaint: a validator that disables itself quietly costs more than one that is absent, because it is trusted. Both controls are unchanged and pinned by tests: a conforming commit still passes silently, a non-conforming one still exits 2 with its existing block payload. The defect test asserts stderr is non-empty and names the failure; it fails against the pre-fix hook. Verification runs on the remote runner. Refs #3911, #3838 * docs(#3911): document the hook crash-policy contract Reference and Explanation via a new docs/features fragment (FEATURES.md is generated from it), INVENTORY rows for the three new hooks/lib files, and an ARCHITECTURE note on the hooks section. How-To: docs/how-to/declare-a-hook-crash-policy.md, indexed from docs/README.md — a hook author now has to choose and declare a crash policy, which is more than one step and crosses into which harness protocol their hook speaks. It covers allow/deny/crash, writing an ON_CRASH reason that is actually useful, when a deny needs a distinct stderr payload, the two hooks whose harness reads a JSON-body decision and must NOT use deny(), and what to do when a check cannot run at all — with #3838 as the worked example. Refs #3911 * test(#3911): prove the seam actually ships, and stop hand-rolling temp cleanup Two review findings. The acceptance criterion 'hooks/dist/** stays in parity via the build seam (lint:hooks-runtime-build-seam)' was misstated and unmet: that lint checks something else — that a hook requiring a compiled gsd-core/bin/lib module also calls ensureRuntimeBuild(). Nothing exercised that the three new hooks/lib files reach hooks/dist/lib at all. That gap is not theoretical: #770 is a recorded ship-blocking bug where a new hook never shipped because a copy list missed it. The suite now builds dist through the repo's own ensureBuiltHooks(), byte-compares each shipped copy against its source, and spawns a child that requires the SHIPPED dist copy and denies — which is what catches a copy that exists but cannot resolve its sibling registry. gsd-validate-commit.sh hand-duplicated mktemp/run/rm three times; one idempotent trap on EXIT replaces them, guarded so cleanup cannot alter the exit status. Behavior-neutral across five cases, with temp-file counts taken before and after each run. Refs #3911 * fix(#3911): stage transitive hook lib requires, not just one level The remote run returned 7 failures across 3 real causes. The important one is a PRODUCTION bug this phase exposed rather than caused. `writeCursorHooksJson` scanned each hook script for `./lib/X` requires exactly one level deep and never re-scanned the lib files it staged for their own sibling requires. Nothing had a transitive lib dependency before, so the gap was invisible. Adding hook-exit.js -> cli-exit.js -> exit-code-registry.js made real Cursor installs ship a bundle that dies at require time with MODULE_NOT_FOUND. It now walks to a fixed point, and a real installed Cursor hook runs to completion. The staging harness in shared-hooks-dir-resolution hand-copied its fixture, so the injection scanner crashed at require time and its exit-1 was being read as a policy decision. Migrated to copyScriptWithDeps, which walks the require graph — the repo's recorded rule for this class, since adding another copyFileSync keeps it alive for the next person. The missing-lib-source test in cursor-hook-workspace-roots hardcoded which lib file it expected to be named in the abort message; the same throw now fires for a different file first. Its assertion is unchanged in substance — staging still must abort rather than ship a broken hook — only the name is no longer pinned. The last one was my own test asserting an uppercase reason code. Measured against origin/next: the pre-change hook emits the same lowercase 'config_unreadable', so the test was wrong, not the migration. Corrected to the real value rather than making the code match the test. Verification runs on the remote runner. Refs #3911 * chore(#3911): regenerate the cursor install-tree golden The staging fix means a Cursor install now correctly carries the two transitive lib files it was silently missing. Additive only — no path was removed. The golden diff is the evidence the packaging defect was real. Refs #3911 * chore(#3911): backfill the changeset PR number Refs #3911 * fix(#3911): a git probe that timed out is not a negative A macOS CI lane failed three deny cases at 2084ms, 2112ms and 2177ms — just past the 2000ms budget these hooks give their git probes. The three that passed took 72ms, 595ms and 651ms. Under shard contention `git rev-parse` overruns, the hook reads the non-zero result as "not a git repo", and allows with exit 0 and empty stdout AND empty stderr. Under load, the guards silently stop guarding. That is ADR-3889's thesis exactly, sitting inside the security hooks this phase is about. The repo had already recognized the class in one place — gsd-cursor-subagent-start.js fail-closed-denies on `git_timed_out` (#3045) — but nowhere else. `hooks/lib/git-probe.js` classifies a probe's outcome, distinguishing a real non-zero exit from ETIMEDOUT, a signal kill, and a spawn failure, rather than folding all four into `status !== 0`. Three guards route their eight git probes through it. The resolution is the same shape #3838 took, and the same one that issue endorsed as smallest-sufficient: fail open, but loudly. **No exit code changes on any path** — a developer on a loaded machine is still not blocked, which keeps #3911's declaration-pass contract intact for exit codes. What changes is that the hook now says on stderr which probe could not answer, instead of presenting silence as a clean verdict. Scope was checked across every hooks/*.js, not just the three that failed: gsd-agent-isolation-guard spawns no git; gsd-statusline's two probes gate only a cosmetic display segment, not an allow/deny decision, and are left alone. The C2 deny assertion was a real-race test — it demanded exit 2 while a slow git legitimately yields 0. It now requires the hook to either deny, or allow with a diagnostic naming the probe that could not run; a silent allow still fails, so the assertion is not vacuous. A deterministic regression stubs git on PATH to sleep past the budget rather than waiting for load to reproduce it. Verification runs on the remote runner. Refs #3911 * test(#3911): a PATH shim cannot intercept the hooks' git spawn on Windows The deterministic timeout regression stubbed git on PATH and asserted the guard reports rather than silently allows. It passes on Linux and macOS and failed on Windows in 83ms and 176ms — the stub was never invoked at all. Mechanism: the hooks call spawnSync('git', args) with no shell:true, so on Windows CreateProcess resolves git.exe only and never a PATH .cmd shim. The git.cmd branch could not have worked and is removed rather than left implying a Windows path that does. Adding shell:true to the hooks to serve a test would change product behavior and widen an injection surface, so the case is skipped on win32 only, with the mechanism written into the skip reason so a future reader does not 'fix' it that way. Linux and macOS keep the coverage, and macOS is where the underlying fail-open was actually caught. Refs #3911 --------- Co-authored-by: sim <sim@local> |
||
|
|
bf2332e67c |
fix(#3582): route every hook's compiled-module require through the self-heal build seam (#3629)
* test(3582): failing-first cold-tree coverage and the seam drift lint On a plugin-channel install the compiled gsd-core/bin/lib/*.cjs are legitimately absent (ADR-457 build-at-publish; the npm package builds before publishing, a raw tree materialization never does). gsd-tools.cjs calls ensureRuntimeBuild() before requiring ./lib; no hook does, so the isolation guard's Cannot-find-module lands in its fail-closed catch and is misreported as an unreadable dispatch-isolation configuration, blocking every executor dispatch. These tests fail on that: cold-tree runs of the isolation guard, statusline, cursor guard and update worker, plus the seam's actionable build error surfacing instead of the generic misreport. Also adds the drift lint the acceptance criteria require, with a fixture proving it CAN fail — a guard never shown to fail is worthless. It is red here by design: it flags today's unfixed hooks, which is exactly the defect. * fix(3582): route every hook's compiled-module require through the self-heal seam RED proven at 5b174b0d: 11 failures — the cold-tree runs for the isolation guard, cursor guard and update worker, the fail-closed-with-actionable-message assertion, and the lint's own real-tree check. The compiled runtime library is produced by build:lib and gitignored (ADR-457, build-at-publish). The npm package builds before publishing; a plugin-marketplace or git-clone install materializes the raw tree and never does, so on that channel those modules are legitimately absent. The self-heal seam added by #2002 exists to heal exactly this, and the CLI entrypoint already calls it — no hook did. The isolation guard's Cannot-find-module therefore landed in its fail-closed catch and was reported as 'could not read or resolve dispatch-isolation configuration', so an ARTIFACT ABSENCE was misdiagnosed as an unreadable project config and every executor dispatch was blocked. All SEVEN affected files now call the seam before their first compiled require. The issue named four; a scan found six; implementing it surfaced a seventh — the shared isolation sentinel helper, used by BOTH guards, which requires two compiled modules itself and would have defeated the guards' own fix on a genuinely cold tree. Same defect class, so fixed here rather than left as a known-broken remainder. Failure posture is deliberately split by hook kind: - Gates (agent isolation guard, cursor subagent start) surface the seam's actionable build error distinctly instead of swallowing it into the generic text, and stay fail-closed — a genuinely unreadable project config still DENIES exactly as before. - Cosmetic and detached hooks (statusline, update worker, update check, update banner) DEGRADE rather than crash: the statusline draws on every render and the worker is a detached process, so a build failure there must not take down the prompt. The npm path is untouched: the seam's already-built fast path returns immediately, so prebuilt installs pay nothing and behave bit-for-bit as before. Adds a drift lint, wired into the CI lint chain, so the invariant is enforced rather than remembered — without it the next hook to add a compiled require reintroduces the class silently. It is proven able to fail: a fixture hook requiring a compiled module without the seam is flagged, and one that uses the seam is not. Verified directly — on the unfixed tree it named all seven offenders; with the fix it passes. While writing the lint's comment stripper, a naive whole-text block-comment regex ate its own fixture, because this repo's comments legitimately spell the compiled-lib glob whose star-slash reads as a comment opener. Rewritten as a line-based scanner with a regression test pinning that case. * fix(3582): test the three untested seam call sites and assert typed reason codes Two independent reviews converged on the same major gap: the fix wired the seam into seven files but only four had cold-tree tests. The adversarial pass put it plainly — deleting the shared isolation-sentinel helper's seam call would not have failed any test in the diff. That file was my own addition beyond the issue's four, so it shipped untested; that is now closed. - Shared isolation-sentinel helper: its seam call is only reached when .planning is NOT directly under cwd, and every existing cold-tree fixture puts it there, so the early return always fired first. Now covered, and proven load-bearing by mutation: with the call removed the spy records zero seam invocations and the test fails. - update-check hook and update-banner hook: cold-tree tests added asserting the DEGRADED VERDICT — the fallback cache filename, and silent suppression when the package name degrades to null — rather than merely 'did not throw'. The banner hook previously had no test file at all. Standards violation fixed: two tests asserted on free-form prose via assert.match against a JSON reason string, which CONTRIBUTING bans by name — its own BAD example is exactly that. The ESLint rule only covers readFileSync/spawnSync text, so tooling did not catch it. Both isolation guards now emit a machine-readable reason_code from a frozen enum, following the repo's existing REASON convention, and the tests assert that instead. The human-readable message is unchanged for operators; only the assertion target moved. The duplicated degrade boilerplate across the three cosmetic hooks was deliberately NOT extracted, and the reason is recorded at each site: both viable shapes — a path-parameterized helper, or a ceremony-only wrapper — defeat the drift lint's per-file literal co-occurrence check, so extracting would require the lint to special-case its own helper. Triplication is the lesser evil while the lint stays a co-occurrence scan. The lint's header now states what it does and does not catch (literal quoted requires only; hooks/ scan root), so a future reader does not over-trust a guard that a concatenated path or a require inside a non-hooks helper would evade. * chore(3582): regenerate the committed install-tree fixtures Adding a new shipped hook helper changed the install tree, and those fixtures are committed-and-derived (regen:derived / gen:install-tree), so 12 'install tree — <runtime>' tests failed on 541a1913. Regenerated rather than hand-edited. The delta across all 15 runtime fixtures is exactly two lines — the new helper under both its hooks/ and gsd-hooks/ install paths — and nothing else, so the regeneration pulled in no unrelated drift. This is the bookkeeping ripple a new file under hooks/ carries; it was not visible from lint:ci, which passed both before and after. * chore(3582): backfill changeset PR number (#3629) --------- Co-authored-by: sim <sim@local> |
||
|
|
268ca7e32d |
fix(#3504): harden hook injection patterns and force-add guard (#3510)
* test(#3504): add failing-first parity, fail-closed, and bypass suites * fix(#3504): harden hook injection patterns and force-add guard * test(#3504): stage the scanner lib dependency in shared-hooks fixture * chore(#3504): backfill changeset pr number * test(#3504): build parity samples from fragments for the ci scan --------- Co-authored-by: sim <sim@local> |
||
|
|
470389f3a2 |
chore(#3212): tokenizer-first for stateful grammars — a shared scanner — Phase 3 (#3424)
* feat(#3414): promote git-cmd.js token-walk into a shared scanner, fix #3169 Phase 3 of epic #3212 (ADR-3212 §4). New src/token-scanner.cts generalizes hooks/lib/git-cmd.js's proven token-walk (#3129 — "has not re-opened"): tokenizeShellLike (quote-aware shell tokenizer, byte-identical port) and indentWidth (bullet-nesting depth). git-cmd.js migrates onto tokenizeShellLike with zero behavior change (parity-asserted against every existing #3129 fixture in tests/worktree-safety.test.cjs's folded block); isGitSubcommand's phases 1-3 (env-prefix skip, executable check, global-option consume) extracted into skipToSubcommand, shared with the new extractBranchArgument (git checkout -b / git branch <name>) — a new capability exercising the seam on the domain the ADR names, not a migration of existing duplicated logic (none existed). Fixes #3169: src/decisions.cts's parseDecisionLines couldn't distinguish a cross-reference bullet nested under an open decision from a fresh malformed declaration attempt. An earlier bold-run-content-classification design was tried and disproven against the repo's own existing FIX-B fixtures (D-02, "no colon no dash") before being adopted — both have identical shape under any content-only rule. Nesting depth (via indentWidth) is the actual distinguishing signal: a bullet indented deeper than the currently-open decision's own bullet is elaboration, folded into its text like a continuation line, never tested against the parse-miss guard. A bullet at the same-or-shallower indent is unchanged. Scope-narrowing disclosed, not silent: of the ADR's four named bugs (#3197, #3169, #2570, #2528), three no longer need this phase's work. were independently fixed and closed since the ADR was authored — #2570's fix is already a correctly-bounded regex per the ADR's own decidability test (no scanner needed); #2528's fix is a deliberate, twice-reviewed non-scanner design (its own code comment records a scanner-based attempt that regressed a symmetric case and was reverted) that this phase does not disturb. Only #3169 required new work. get_impact: isGitSubcommand CRITICAL/196 affected symbols, parseDecisionLines CRITICAL/164 affected symbols (ADR §6 due diligence). Six-gate ripple: .gitignore, eslint.config.mjs, docs/INVENTORY.md, docs/INVENTORY-MANIFEST.json (regenerated), CONTEXT.md glossary. Design: .gsd/phase/chore-3414-tokenizer-first-seam/40-design.md Test matrix: .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3414): add required fast-check property tests per code review TESTING-STANDARDS.md:169 requires at least one fast-check property test for any module that implements parsing — src/token-scanner.cts had none, an orthogonal Standards-axis review finding. Adds two seeded property tests (mirroring Phase 1/2's fast-check-setup.cjs convention): indentWidth counts exactly a generated leading-space run; tokenizeShellLike round-trips a generated array of whitespace/quote-free words joined with single spaces. The design doc's own "no property test needed" rationale was wrong — it argued no algebraic law applied, but the standard is unconditional for parsing modules regardless of whether one "feels" applicable. Corrected in .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md. Also fixes two Spec-axis wording drifts the same review found between the design doc and the shipped code (doc-only, no behavior change): extractBranchArgument's documented signature dropped an unused subVariants parameter that was never implemented, and the #3169 fail-first fixture description corrected from "15-decision plan via cmdDecisionCoverageVerify" to the actual compact 3-decision analog via the real blocking gate, check.decision-coverage-plan. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3414): add changeset for #3169 fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3414): backfill changeset pr number to 3424 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
8f75e27554 |
fix(#3045): fail closed when an executor dispatch drops its resolved isolation (#3069)
* feat(#3045): deny an executor dispatch that drops its isolation flag Every isolation gate already resolved correctly. The resolved value then reached the executor through a prose instruction telling the model to substitute it into a call the model composes itself, and nothing verified the substitution. When it was dropped, the executor edited and committed in the user's primary checkout with no consent and no warning. A prose backstop would be the same class of artifact as the defect, so this is a shipped PreToolUse hook on the Agent tool. It fires at the instant of the call rather than being read once at the top of a workflow, which is the only placement the model cannot skip. The guard is inert unless it can positively establish that this is a GSD project, that the project resolves to harness isolation, and that the dispatch targets an executor. A non-GSD repo has no invariant to enforce. Where it cannot read the configuration at all, it denies rather than assuming, with its own reason -- a guard that cannot verify must not answer safe. A malformed payload allows rather than throwing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3045): extend the isolation guard to Cursor Cursor is the second of only two runtimes that resolve harness isolation, so shipping the guard for Claude alone left half the exposed surface unguarded while the changeset implied it was covered. The two runtimes fail differently. On Claude the harness flag is a per-dispatch kwarg the model must copy into a call it composes, and the defect is that it can be dropped. On Cursor the flag is --worktree, which applies to the whole session, and the subagent-start payload carries no isolation field at all. There is no flag to check, so the guard verifies the effective state instead: whether the workspace is genuinely running outside the user's primary checkout. That is a stronger check than the Claude one because it tests reality rather than intent, and it is commented so nobody later rewrites it into a flag check. Isolation is established two ways, either sufficient: the workspace resolves to a linked git worktree, or it sits under the worktree root Cursor manages. The second matters because a directory Cursor placed there is a legitimate isolated session even before it becomes a distinct git worktree, where linkage alone would report no repository. Detecting linkage required a new primitive rather than the existing context resolver. That resolver short-circuits on finding a local .planning directory before it ever compares the git directory to the common one -- and an isolation worktree normally has its own checked-out .planning. Reusing it would have read a correctly isolated session as unisolated and denied it, which is the failure direction that gets a guard switched off. The comparison is now its own shortcut-free function that the resolver delegates to after its own shortcut, so existing behavior is unchanged, and the case that would have broken is pinned. The subagent type is checked before any configuration is read, so an unreadable config cannot deny a dispatch this guard would never have enforced against. The input-schema comment on the Cursor hook documented only the fields common to every event and omitted the ones specific to this one. That omission cost a halt during this work; it now documents both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): enforce the resolved dispatch decision, not the host capability The guard keyed on the registry's dispatch.isolation, which says only that a runtime is CAPABLE of harness worktrees. The decision that actually governs a dispatch is the one the workflow resolves after gating, and that legitimately comes out as sequential in three documented cases: a project setting use_worktrees false, a per-plan submodule intersection, and the base-check auto-degrade. The workflow tells the model to omit the flag in exactly those cases, and the guard was denying every one of them. The third case matters most. The preceding fix made the base-check degrade on git timeouts and a missing git binary, where it had previously answered "safe". That correction is right, and it means a transient hang now degrades to sequential far more often than before -- so the two changes composed into a trap where the workflow behaved exactly as designed and the guard blocked it. The workflow already resolves isolation in shell, deterministically, which is what makes it a trustworthy source in a way the model-authored call is not. It now records that resolved value through a dedicated verb, and both guards read it first. A fresh record is authoritative, so sequential dispatches pass untouched. Absent or stale, the guards fall back to the capability check combined with the project's use_worktrees setting, which still covers the case that never reaches the workflow. Also widened the matcher to accept Task alongside Agent, since a host that names the tool Task would otherwise leave the guard silently inert while implying coverage; stopped assuming Claude when no runtime is declared, which is the shipped default and would have demanded a Claude-only argument elsewhere; and made a non-git project inert rather than denied, since advising a worktree session is not actionable without a repository. The original diagnosis never modeled sequential mode as legitimate. That omission is what let this through, and it is now recorded there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): record at resolution and bind the record to its dispatch Two independent reviews converged on the same failure: the guard was fail-open in a default install, so it did not catch the defect it exists to catch. A shipped project carries no runtime key, which made "runtime not confidently known" the common case rather than a corner one. A record asserting that isolation was required but carrying no flag then fell through to a capability lookup that answered "none", and the dispatch was allowed. The flag itself only arrived from a second shell block -- the same block a model dropping the argument would also skip. A test had pinned that behavior as intended. The record is now written by the resolver, as an unavoidable consequence of asking for the value, rather than by a step the model is told in prose to go and run. A guard against a prose-carried value cannot itself depend on prose. Mode, flag and identifiers are written together and atomically, so the flagless window is gone, and a record asserting isolation with no resolvable flag now denies instead of degrading. Runtime is also resolved from the installer's own recorded default, which makes confident resolution the normal case. The per-plan submodule gate degrades after the phase-level decision and never re-recorded, so a plan that legitimately ran sequentially was denied against a still-fresh phase record. It now records its own, scoped to the plan. A record also authorized any dispatch for four hours. One phase degrading to sequential could silently license an unisolated dispatch in the next. Records now carry phase and plan, the guards require them to match, and the window is minutes rather than hours -- the resolver rewrites it before every dispatch, so a long window bought nothing and only widened the hole. The flag validator rejected any value beginning with two dashes, which is exactly the form Cursor and Windsurf declare, so their real value could never have been stored. Writer and reader also derived the record path differently and diverged inside a linked worktree without local planning state. The predictable path remains a way to silence the control without leaving a trace in the diff. It grants no access an agent with shell does not already have, so it is documented as accepted rather than redesigned around. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): correct the staleness boundary and unmask a vacuous parity test The remote runner returned twenty failures. One was a real production defect the boundary case existed to catch: a record whose age exactly equalled the staleness window was treated as fresh, so it stayed authoritative for one tick past its own expiry. Freshness is now strictly inside the window. The parity test meant to stop the two guards' executor lists from drifting could never have failed. Its project fixture was a bare directory rather than a repository, so the non-git inert branch answered before the executor list was ever consulted. It asserted agreement it never actually measured. The fixture is now a real repository, like every sibling in the file. A test also asserted that Windsurf declares the worktree flag. It does not -- Windsurf resolves to no isolation by design, having no named concurrent dispatch to isolate. The test claimed a registry fact that was never true, and a comment in the resolver repeated it. Both corrected, and the test now proves what it should have all along: that the parser accepts any bare flag value, rather than one runtime's supposed value. The new guard was missing from the bundled-hook whitelist, which is the surface that decides what actually ships, and the per-plan gate had gained calls to the launcher without the preamble those calls require. The changeset carried parenthetical product descriptions the purity rule forbids. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3045): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3045): make the guard tests hold on Windows Two tests redirect HOME to control where the installer-persisted runtime default is read from. Node resolves the home directory from USERPROFILE on Windows and never consults HOME, so both silently read the real runner profile, found no recorded runtime, and asserted against a project the hook had not recognised. The production code was already correct in asking the platform rather than the variable; only the tests were wrong to assume one variable answers everywhere. The helpers now mirror the override onto both. The symlink spoofing test also created a directory symlink unconditionally, which needs elevated privileges on Windows. It survived on this runner, but it would fail on any host without them, so the creation is now attempted and the test skips explicitly when it cannot be done -- a bare return would have counted as a pass and hidden the gap. Skipping alone would have left the platform uncovered, so the behaviour it proves is now also driven in-process through an injected realpath, following the seam already used for the clock. That case no longer depends on privileges at all, and the end-to-end test keeps its original assertions wherever symlinks work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c7c2fe3c2b |
fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd (#2680)
* fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd gsd-cursor-session-start.js and gsd-cursor-stop.js both resolved the project as path.join(process.cwd(), '.planning', 'STATE.md'). Under the cursor-agent CLI, hooks are invoked with cwd set to the Cursor config dir (~/.cursor), not the workspace — so the lookup always missed. sessionStart could only ever emit the "no .planning/ workflow found" nudge and stop's verify-work reminder could never fire, even with .planning/STATE.md sitting in the workspace. Slash commands were unaffected, which is why only the hook layer looked blind. Both hooks already buffered stdin into `raw` and never parsed it; the payload's workspace_roots carries the real path. Multi-root was left open in the report ("first root vs any root"). Resolved forward: prefer the first root that actually carries .planning/STATE.md, so a workspace whose GSD project is not the first root still resolves — strictly better than first-root-only and identical to it in the single-root CLI case. Falls back to roots[0], then to cwd, keeping IDE behavior unchanged if the IDE ever invokes hooks from the workspace. The resolver is duplicated verbatim across the two scripts rather than shared via hooks/lib/: these hooks ship standalone, and a new hooks/lib/ file must be registered in the GENERATED installer's GSD_HOOK_LIB_FILES allowlist — the installer-omits-shipped-file class that yields MODULE_NOT_FOUND at runtime. Per CLAUDE.md "Generative Fix Divergence", the duplication carries a parity assertion so the copies cannot drift. Failing-first, demonstrated by direct invocation with cwd != workspace: pre-fix sessionStart -> "no .planning/ workflow found" stop -> {} post-fix sessionStart -> ".planning/STATE.md is present" stop -> reminder tests/fix-2587-cursor-hook-workspace-roots.test.cjs spawns the real scripts as child processes with a cwd lacking .planning/ and workspace_roots pointing at it. Boundary coverage on the roots array (0 / 1 / 2 entries), plus malformed-JSON fail-open, junk-entry filtering, the parity assertion, and a guard that neither script resolves .planning from cwd again. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2587): extend workspace_roots fix to subagentStart; keep cwd a candidate Three findings from the isolated review, all fixed. 1. MISSED SITE (high). gsd-cursor-subagent-start.js carried the identical defect at line 43 — its own header documents workspace_roots in the input schema, but it resolved .planning/ from process.cwd() anyway. Under the cursor-agent CLI that meant every Cursor subagent (planner, executor, verifier) started with "no .planning/ workflow found" and no phase context. The report named only sessionStart and stop; the defect class was wider. Verified pre-fix vs post-fix by direct invocation with cwd != workspace. 2. SEMANTIC NARROWING (medium). The first cut searched only workspace_roots and fell back to cwd solely when the array was EMPTY. So when roots were supplied but none carried .planning/ while cwd did, the hook reported absent — where the pre-fix code, which always used cwd, reported present. That contradicted the fallback's own stated intent of preserving IDE behavior. cwd is now a CANDIDATE in the search (`[...roots, process.cwd()]`), so the fix is a strict superset of both the old behavior and the CLI fix, never a narrowing. 3. STALE GOLDEN FIXTURES (high, would have failed CI). The golden-install-parity fixtures store a content hash per installed file; these three hooks appear in 13 of the 19 runtime fixtures. Regenerated via `npm run gen:golden` — the diff is exactly the three hook hashes in exactly those 13 runtimes. Tests extended: subagentStart resolution via workspace_roots; the stop hook's absent branch (previously only session-start's was covered); an explicit regression guard that a project at cwd is still found when roots miss; parity now asserts all THREE copies byte-identical; and the cwd guard sweeps the whole RESOLVING_HOOKS list so a future hook in this family cannot be left on cwd. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * refactor(#2587): extract cursor workspace resolution to a shared hooks/lib module The duplicate-plus-parity-test approach was the wrong call. The reported issue named two hooks; a third (subagentStart) had the identical defect. That is the signature of a systemic problem, and three copies of a resolver guarded by a parity assertion is a divergence risk maintained by hand rather than a fix. hooks/lib/cursor-workspace.js is now the single implementation. All three Cursor hooks require it; none defines a local copy. Divergence is prevented structurally instead of by asserting three copies stay byte-identical. The reason duplication looked necessary was real, and is fixed properly here rather than worked around: Cursor sets hostBehaviors.skipSharedHooksInstall (#2089), so it never reaches the installer's bulk hooks/lib copy — it was the ONE runtime shipping these hooks WITHOUT hooks/lib (verified against all 19 golden fixtures: cursor had the hook scripts, no lib). A naive require would have thrown MODULE_NOT_FOUND at load, BEFORE each hook's own try/catch, wedging every session on precisely the runtime this bug is about. writeCursorHooksJson (src/runtime-hooks-surface.cts) now stages the hooks/lib helpers the staged scripts actually require, discovered by scanning their require('./lib/…') calls rather than a hardcoded name — so a future helper cannot be silently omitted. This is narrower than flipping skipSharedHooksInstall, which would wrongly pull in every shared hook. cursor-workspace.js is also added to GSD_HOOK_LIB_FILES so uninstall and the manifest manage it for the runtimes that do receive hooks/lib. Verified against a REAL install (runMinimalInstall, cursor/global): the helper is staged, and all three INSTALLED hooks resolve the workspace end-to-end from a cwd that is not the project. Also closes the review gap that the stop hook was excluded from the cwd-candidate regression loop — it now sweeps RESOLVING_HOOKS. The byte-parity test is replaced by a structural guard (every hook requires the shared module, none redefines it) plus a new install test asserting the helper is staged and the installed hook actually loads against it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2587): fail loud on a missing hook lib source; drop unsubstituted version marker Two findings from the installer-focused review. H1 — the staging step's `if (!fs.existsSync(libSrc)) continue;` silently defeated the very guarantee it was added for. Reproduced: delete hooks/lib/cursor-workspace.js from source, run the cursor install — it exits 0, prints "Done!", and ships the three hook scripts with an EMPTY hooks/lib/. The installed hook then throws `Cannot find module './lib/cursor-workspace.js'` at load, before its own try/catch, wedging every session — and nothing surfaces until a user hits it. The scan protected against a required-but-UNLISTED helper while leaving required-but-MISSING wide open (typo, bad rebase, an accidental delete). It now throws: a missing helper source is a packaging bug and aborts the install. M1 — hooks/lib/cursor-workspace.js carried a `gsd-hook-version: <placeholder>` marker that NOTHING substitutes: copyLibDir stamps .sh files only, and writeCursorHooksJson's staging applies just the colon-to-dash rewrite. Verified the literal was reaching disk on both the bulk (--claude) and Cursor (--cursor) paths. hooks/lib/git-cmd.js — the only pre-existing hooks/lib/*.js — carries no such marker, so this was newly introduced, not inherited. Marker removed, matching that precedent, with a note on why. (The explanatory comment deliberately does not spell the token out, or it would reintroduce the literal.) M2 — the require-scan regex demanded the exact compact form, so `require( "./lib/x.js" )` would silently fail to stage its helper and compound H1. Now tolerant of interior whitespace and either quote style. Regression test added for H1 — the reviewer confirmed the invariant had zero coverage repo-wide: a source tree carrying the hooks but no hooks/lib/ must make writeCursorHooksJson throw rather than produce a broken install. Re-verified end to end: the missing-source case throws, no unsubstituted literal ships, and the installed hook still resolves the workspace from a foreign cwd. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2587): backfill changeset pr number (#2680) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
dacc23137a |
feat(3347): opt-in auto-update of knowledge graph after main HEAD advances
Closes #3347 Config: - Add graphify.auto_update (default false) to manifests: sdk/shared/config-defaults.manifest.json, config-schema.manifest.json Hook: - hooks/gsd-graphify-update.sh — PostToolUse Bash matcher - Gates: tool_name=Bash, HEAD-advancing git op, CI=unset, in git repo, current branch == default branch (git.base_branch override or main/ master/trunk fallback), graphify.enabled && graphify.auto_update both true, graphify on PATH, no live PID lock - Writes .planning/graphs/.last-build-status.json with status=running synchronously, then detaches hooks/lib/gsd-graphify-rebuild.sh - hooks/lib/gsd-graphify-rebuild.sh — detached rebuild runner - PID-lock acquire + trap-on-exit cleanup - graphify update . then cp graphify-out/* → .planning/graphs/ - Status file rewritten to status=ok|failed with exit_code, duration_ms, head_at_build - Portable detach (subshell + disown, no setsid dependency) Installer: - bin/install.js: register hook as PostToolUse Bash matcher (5s timeout) - Add to gsdHooks uninstall list and expectedShHooks warning list Planner / researcher status surface (issue #3347 reviewer must-have AC): - agents/gsd-planner.md and agents/gsd-phase-researcher.md load_graph_context steps now read .last-build-status.json and surface: running → "rebuild in flight"; failed → "auto-rebuild FAILED at {ts}, context is from prior build"; ok with stale head_at_build → "HEAD has advanced since last build" Settings: - get-shit-done/workflows/settings.md adds "Graph auto-update" question with No-Recommended default; bullets and update_config block updated Inventory: - docs/INVENTORY.md hook count 12 → 13 with new row - docs/INVENTORY-MANIFEST.json regenerated Tests: - tests/feat-3347-graphify-auto-update-config.test.cjs (8 tests): isValidConfigKey accepts graphify.auto_update, CANONICAL_CONFIG_DEFAULTS default false, config-set round-trip, sibling key preservation - tests/feat-3347-graphify-auto-update-hook.test.cjs (18 tests): all bail paths (non-Bash, non-HEAD-advancing, enabled=false, auto_update=false, CI=true, non-default-branch, missing graphify bin, live-PID lock), dispatch path with mock graphify bin (sync running status + detached transition to ok/failed), stale-PID lock, all five HEAD-advancing command matchers, git.base_branch override Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
7827e1ddee |
fix(#3129): replace bypassed bash regex with token-walk git-cmd.js classifier (#3141)
* fix(#3129): replace bypassed bash regex with token-walk git-cmd.js classifier Root cause: gsd-validate-commit.sh used: if [[ "$CMD" =~ ^git[[:space:]]+commit ]] This regex silently bypasses Conventional Commits enforcement for: git -C /path commit -m ... (working-directory prefix) GIT_AUTHOR_NAME=x git commit (env-var prefix) /usr/bin/git commit -m ... (full-path executable) Fix: introduces hooks/lib/git-cmd.js with isGitSubcommand(cmd, sub) — a token-walk classifier that handles all four forms by: 1. Skipping leading VAR=VALUE env assignments 2. Validating the git executable (basename check for full-path support) 3. Consuming git global options (-C <path>, --git-dir=, -p, etc.) 4. Checking the subcommand token The hook delegates to this classifier via node shell-out. node is already called twice in this hook (config check + JSON parse), so no new runtime dependency. This becomes the single source of truth for all hooks that gate on git subcommands (pre-commit-review-gate, post-push-verify, etc.). Regression test: 27 assertions — tokenize correctness, 12 must-match cases (including all 3 bypass forms), 8 must-not-match cases, 3 source checks. All are real behavioral tests, not string comparisons. Suite: 7035/7035. Closes #3129. * fix(lint+hook+changeset): allow-test-rule, fix HOOK_DIR quote injection, fix changeset pr+typo |