Commit Graph

15 Commits

Author SHA1 Message Date
Jakub Zych
6cfa0c55d2 refactor: drop 12 runtimes, keep Claude, Codex, OpenCode, Cursor, ZCode, Antigravity
Removes kilo, kimi, kimi-code, copilot, windsurf, augment, trae, qwen, hermes,
cline, codebuddy and pi end to end: capability descriptors, installer branches
and converters (bin/install.js 14.9k -> 11.2k lines), TypeScript converters,
hook surfaces and runtime homes, review lanes qwen/kimi-code, the two pi
migrations, Kimi payload normalization in the hook guards, dead hostBehaviors
vocabulary, launcher home probes, fixtures, runtime-specific tests and the
prose that presented them as supported.

Installer output for the six kept runtimes is byte-identical to before the
prune. The Kimi tool-vocabulary tests in workflow-guard, read-guard and
read-injection-scanner are left in place pending a decision.
2026-10-06 20:02:40 +02:00
Jakub Zych
a9a7a328e6 refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD
across contents and paths, upstream package/repo coordinates -> @golem15/msd-core
and golem15com/msd-core. Deep links into upstream history, sibling upstream
packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is.

Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line,
package/plugin identity, regenerated lockfile, install-tree fixtures, derived
registries and benchmark baseline; migration checksum baseline re-locked
(MSD keeps its own install state, so no install had applied the old sums);
sort-order and regex-escaped expectations in tests adjusted.
2026-10-06 01:47:40 +02:00
Tom Boucher
c0b2a05d2f fix(#4594): one canonical dispatch-identity owner — the emitted format and the parser that reads it back (#4693)
* fix(#4594): give dispatch identity one owner for the emitted format and its parser

The isolation guards decided whether a run-scoped sentinel applied to a
dispatch by regex-scraping model-authored prose. The scrape returned values in
a different namespace from the ones the sentinel records, so the comparison
could never succeed:

  sentinel  { phase: "03", plan: "03-02-hardening" }   <- $PHASE_NUMBER / $plan_id
  prose     "Execute plan 02 of phase 03-auth."
  scraped   { phase: "03-auth.", plan: "02" }          <- greedy (\S+), both wrong

#4594 reports only the phase half. Measured against a real phase-plan-index
run, plans[].id is phase-prefixed, plan-numbered AND slugged, while the prose
carries a bare in-phase plan number — so the plan field mismatches too, and the
Claude path is dead rather than latent. A fresh sentinel was therefore
discarded on every executor dispatch and every legitimate ISOLATION=none
degrade was denied, leaving the work unrun.

hooks/lib/dispatch-identity.js is now the single owner of both halves. The two
prompt-body producers emit a canonical marker carrying the same shell values
the sentinel records, so producer and consumer agree by construction. The prose
frame stays as a fallback, bounded by the phase-token grammar ADR-2121 owns and
deliberately reporting no plan — an absent identifier means "cannot compare"
and is safe; a wrong one is a false mismatch and is not.

The prose sentence itself is byte-identical: the executor agent reads it too,
so the marker is purely additive (Hyrum's Law).

An inapplicable sentinel is now named in the guards' deny reason instead of
being dropped silently — the silence is why this survived three producers and
two consumers unnoticed. Interpolated values come from a sentinel file and from
prompt text, so both are length-bounded and stripped of control characters.

ADR-4630 locks the seam and maps the epic's three phases.

Refs #4630
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4594): resolve eight review findings across the dispatch-identity seam

Three orthogonal review engines ran on 43418af144 — the code-review skill's
Standards and Spec axes, and an isolated adversarial security pass — plus a
self-review of the committed diff. Every finding is fixed here; none deferred.

F1 (major, reproduced). A keyless or unknown-key-only marker — the literal
"[gsd:dispatch]" or "[gsd:dispatch run=..]" — matched the marker grammar and
returned source:'marker' with both fields null, suppressing the prose fallback
entirely. Any prompt text containing that literal silently disabled identity
narrowing, so a fresh sentinel applied to a dispatch it was never scoped to,
defeating #3045 SECURITY F2. Prompt text is attacker-influenceable. A marker
that yields neither recognized key is no longer a marker: the scan continues to
later markers, then later texts, then prose. Forward-compatible tolerance of
unknown keys is unchanged.

F2/F3 (major). The first cut duplicated sanitizeForReason,
describeSentinelDiscard and REASON_INTERPOLATION_MAX_LEN byte-for-byte across
both guards — the exact defect class this epic exists to delete, and with no
cold-load justification, since both hooks already require hooks/lib/. They now
live in hooks/lib/isolation-deny-reason.js, and buildSentinelDiscard lives in
isolation-sentinel.js beside the comparison it mirrors, returning the nested
{sentinel:{phase,plan}, dispatch:{phase,plan}} shape instead of a bespoke
four-field bag that renamed the pairs already flowing through the seam.

F4 (hard violation). The visibility test asserted on the deny reason's prose.
CONTRIBUTING.md prohibits raw text matching on hook output, which is why every
deny carries a stable reason_code. The discard is now a structured
sentinel_discarded field on each hook's stdout JSON, and the test asserts that;
the sentence stays for the operator but is no longer the contract.

F5 (hard violation). The 64-character truncation limit had no boundary
coverage. 63/64/65 are now exercised against the single consolidated helper.

F6 (minor). sanitizeForReason stripped C0/C1 controls but not U+2028/U+2029 or
the bidi overrides, so a crafted value could still reflow or reverse the
message. Both classes are stripped, with a test each.

F7 (major). The producer/template parity test was vacuous — it rendered a
marker and re-parsed its own output, and would have passed with both templates
deleted. It now reads the two workflow templates, extracts each marker line,
substitutes the measured values and asserts the owner's parser returns them.
Proven red by deleting one template's marker line before being proven green.

F8 (doc). ADR-4630 and the design notes claimed the marker is guaranteed on the
orchestrator-worktree path because that prompt is built in shell. It is not:
executor-isolation-dispatch.md:131 says plainly that those are template
placeholders, not shell variables, so {plan_id} is model-substituted there too.
A false guarantee in a design lock is worse than a stated limit. Both documents
now say the marker is model-substituted on both paths and that the prose
fallback is the real floor everywhere. The "3 workflow templates" count was
also wrong — 3 prose sites across 2 files, 2 of which carry the marker.

Refs #4630
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4594): refresh the compact-content baseline and acknowledge execute-phase.md growth

Refs #4630.

The dispatch-identity marker and its substitution note grew
gsd-core/workflows/execute-phase.md by 525 bytes (91846 -> 92371), which drifts
two real-tree guards that lint:ci does not run:

- tests/benchmark-compact-content.test.cjs asserts the committed baseline is
  "up to date"; the split for execute-phase.md moved off 25827 -> 25952 and on
  23576 -> 23701, taking its compaction reduction 8.72% -> 8.67%. Baseline
  regenerated with scripts/benchmark-compact-content.cjs --write.
- tests/emitted-attribution.test.cjs requires a growth acknowledgment trailer
  for any emitted file that grows, keyed on the bare filename. Added below.

The growth is two additions and no rewrites: the [gsd:dispatch ...] marker line
inside the Agent() prompt's <objective>, and the note telling the orchestrator
to substitute {plan_id} with the plan's id verbatim. Both are load-bearing --
the marker is what lets a guard hook match a dispatch to the sentinel the
per-plan gate wrote, and without the note the orchestrator has no instruction
telling it the value must not be paraphrased.

Emitted-Drift-Ack-Growth: execute-phase.md — adds the canonical [gsd:dispatch] identity marker and its {plan_id} substitution note, which the isolation guards compare verbatim against the run-scoped sentinel (#4594)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4594): set changeset fragment pr to 4693

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 15:36:08 -04:00
Tom Boucher
241646a43a fix(#4651): classify .env names by final extension, and close the trailing-dot alias bypass — Phase 1 of #4636 (#4659)
* test(#4651): failing-first coverage for final-extension classification

Phase 1 of epic #4636, absorbing #4580. Tests only; no fix. These MUST fail.

The guard classifies a name by comparing everything after `.env.` as one
token against a set whose members are FINAL EXTENSIONS. So `.env.local.example`
yields suffix `local.example`, which is not a member, and a committed
secret-free template is refused. That is a category error, not strictness.

Two arms are covered because the same classification is hand-rolled twice in
one file: `isSecretBasename` for Read/Bash, and `globAltSelectsSecret`
(`lit.startsWith('.env.')`) for Grep globs. Fixing one alone would ship a
guard that allows `cat .env.local.example` while refusing
`Grep --glob '.env.local.example'` — the same file, the same hook, opposite
answers. A cross-arm parity loop over one shared list asserts the two cannot
drift.

Rows that exist because they are the ones nobody enumerates:

- `.env.example.local` must stay BLOCKED. Final extension is `local`; this is
  dotenv's documented local-override convention and a real secret. Any fix
  shaped as "contains example" admits it.
- `.env.local.` must stay BLOCKED — empty final extension is not a member.
- `.env.` must stay ALLOWED. Note #4580's proposed patch adds
  `if (suffix === '') return true;`, which flips it to blocked; that breaks the
  existing `allows` assertion in this suite and broadens the protected set,
  which epic #4636's non-goals forbid. Not applied.
- `.env.local.exam*` (partial glob literal) must stay BLOCKED — it can select
  `.env.local`, and a partial literal cannot be classified.
- `*.example` and `*` must stay ALLOWED — regression protection on the arm
  that already works.

Local behavioral repro of the current guard, confirming the tests fail for the
right reason rather than by construction:

  .env.local.example  rc=2 (blocked)   <- the defect
  .env.example        rc=0 (allowed)
  .env.local          rc=2 (blocked)
  .env.example.local  rc=2 (blocked)
  .env.               rc=0 (allowed)
  glob .env.local.example  rc=2        <- the second arm

Regressions are folded into the owning module's suite rather than a new
tests/fix-NNNN-*.test.cjs file, per scripts/lint-regression-test-names.cjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): classify by final extension so .env.<name>.example is readable

Phase 1 of epic #4636, absorbing #4580. Implements ADR-4650 decision 5.

The guard compared everything after `.env.` as ONE token against a set whose
members are FINAL EXTENSIONS. `.env.local.example` yielded `local.example`,
which is not a member, so a committed, secret-free template was refused — the
guard blocked the one file that exists so nobody has to open the real `.env`.

That is a category error, not strictness. The fix is not "add local.example to
the set"; it is to compare the right token. hooks/lib/filename-classification.js
now owns that distinction and is the only place it is expressed.

Both arms are fixed, because the same classification was hand-rolled twice in
this one file:

  - isSecretBasename (Read/Bash) now tests finalExtension(suffix).
  - globAltSelectsSecret (Grep --glob) split its first branch. With no
    wildcard the alternative IS a whole filename, so it is classified exactly
    via isSecretBasename. With a wildcard present the literal is only a
    PARTIAL prefix (`.env.local.exam*` can still select `.env.local`) and
    cannot be classified, so the original conservative rule stays.

Fixing only the first would have shipped a self-contradicting guard: `cat
.env.local.example` allowed while `Grep --glob '.env.local.example'` refused —
same file, same hook, opposite answers. A cross-arm parity loop over one shared
list now asserts the two cannot drift.

Two deliberate departures from #4580's suggested patch, both verified:

  - Its `if (suffix === '') return true;` is NOT applied. That flips `.env.`
    from allowed to blocked, breaking an existing assertion in this suite and
    broadening the protected set, which epic #4636's non-goals forbid.
  - `fullSuffix` was drafted alongside finalExtension and removed before
    commit: zero production consumers, and none planned (Phases 2-4 are
    containment, duplicate draining and the path-join ratchet, none of which
    classify filenames). A zero-caller export is dead code. The distinction is
    pinned instead by a test asserting finalExtension('local.example') is
    'example' and explicitly NOT 'local.example'.

The protected set is unchanged. `.env.example.local` stays BLOCKED — its final
extension is `local`, dotenv's local-override convention and a real secret;
any fix shaped as "contains example" admits it.

Scoped out by measurement, not assumption: src/validate.cts:395 and
src/phase.cts:1674 also hand-roll lastIndexOf('.'), but both parse phase
identifiers (`3.2` -> parent `3`), owned by the phase-id.cts seam. Folding
them in would repeat this same category error in the opposite direction.

Checkpoint 1 (prove RED) on the tests-only commit 91d3d6e1: outcome=failed,
26 failures / 45330, all 26 in the two new test files, zero pre-existing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): document the widened template exemption and cover the Bash arm

Two findings from the isolated adversarial review, both fixed in place.

1. The header's "Stated cost" passage named only the four literal template
   names, but since this change the exemption keys on the FINAL EXTENSION, so
   the trusted set is `.env.<anything>.{example,sample,template,dist}` — an
   unbounded family. The reviewer demonstrated it: `.env.prod-real-secrets.example`
   is allowed. That is the deliberate and necessary cost of fixing #4580, but
   it was materially larger than what the header disclosed, and a silent
   expansion of a security guard's trusted set is not acceptable. The passage
   now states the family, the concrete bypass, and that it applies across
   Read, Grep and Bash alike.

2. The cross-arm parity loop asserted Read and the exact-literal Grep glob but
   not Bash, whose `namesSecret` -> `isSecretBasename` path is genuinely
   distinct. The Bash arm was covered only by two one-off tests outside the
   shared table, so the table could not have caught a drift there. The loop now
   drives all three arms from the same TEMPLATES/SECRETS arrays.

No classification logic changed. The Read-arm behavioral table is byte-identical
before and after: rc=0 for .env.local.example / .env.example / .env. ; rc=2 for
.env.local / .env.example.local / .env / .secrets.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): treat trailing dots and spaces as aliases of the protected file

Closes a Windows path-alias bypass surfaced by the isolated adversarial review
of this phase. Maintainer-approved as in scope.

Win32 strips trailing dots and spaces from every path component, so `.env.`,
`.env..`, `.env `, `.env. `, `.env .`, `.secrets.` and `.secrets ` all resolve
to the real `.env` / `.secrets` on Windows. The guard allowed every one of them
— a bypass of a file it already protects, reachable from Read, Grep and Bash
alike. `isSecretBasename` now normalizes the basename before classifying.

The whole class is fixed, not the reported name. `.env.` alone would have left
`.secrets.` and the trailing-space forms open, which is the same
one-cause-explains-every-failure trap this epic exists to close.

Two consequences, both measured rather than assumed:

  - `.env.example.` flips blocked -> ALLOWED. It aliases the already-trusted
    `.env.example` template, so this is correct; it was previously blocked only
    because the trailing dot broke final-extension parsing.
  - A Bash token that is exactly `.env` plus trailing whitespace flips
    allowed -> BLOCKED. Verified this is CONSISTENCY, not a new false-positive
    class: the bare `.env` token was ALREADY blocked as an operand in the same
    position before this change, so the alias now simply behaves like the thing
    it aliases.

The header's "No whitespace trimming" guarantee is preserved and now stated
precisely: leading and interior whitespace is still never trimmed, so prose
like a commit message mentioning `.env` in a sentence stays prose and stays
allowed. Only TRAILING dots and spaces are stripped. Two tests pin that.

This lands at the same behavior #4580's proposed `if (suffix === '') return
true;` would have produced for `.env.`, which this phase earlier rejected. The
rejection was correct on its stated grounds — that line broadens the protected
set, which epic #4636's non-goals forbid. The Windows framing is different:
normalizing an alias of an already-protected file is not a broadening, and the
fix is reached by normalization rather than by special-casing an empty suffix,
so it generalizes to `.secrets.` and the space forms.

Cannot be reproduced on this host — the remote matrix is Linux-only and Windows
coverage arrives from CI — so this ships on the Win32 path-normalization
contract plus the CI lane, and that limitation is stated rather than implied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): one owner for path segmentation, closing a Read/Grep divergence

Four findings from the two-axis review, all fixed in place.

The real one: the guard had TWO path-segmentation rules. `lastSegment` (used
by Read and Bash via `namesSecret`) splits on both `/` and `\`, while
`classifyGrepGlob` hand-rolled its own on `/` only. Measured:

  Read  of `config\.env`      rc=2  BLOCKED
  Grep  --glob 'config\.env'  rc=0  ALLOWED

Same logical file, opposite answers — precisely the divergence this epic
exists to remove, sitting inside the file this phase was already fixing.
`lastSegment` now lives in hooks/lib/filename-classification.js and both arms
call it. All five path-bearing cases (both separators) now agree.

Note on how this was nearly missed: the first measurement of it reported
"both allow", which looked like the reviewer was wrong. That reading was a
measurement artifact — `config\.env` inside a printf'd JSON payload is an
invalid escape, so the hook fails open at rc=0 and the test was observing
JSON breakage rather than the predicate. Re-measured with correct escaping,
the divergence is real. The tests added here use properly escaped literals
and were verified by running, not by reasoning about the escaping.

Also fixed:

  - Both fast-check properties were satisfied by a degenerate
    always-return-'' implementation: "never contains a dot / is a suffix" and
    "never ends with dot-or-space / is a prefix" are both trivially true of
    the empty string. They now additionally pin content preservation — the
    removed tail must match /^[. ]*$/, and a name with nothing to strip must
    come back unchanged.
  - The cross-arm parity loop used only bare basenames, so it could not have
    caught the divergence above. It now covers path-bearing names with both
    separators.
  - That loop's description overclaimed: Read and Bash BOTH route through
    `namesSecret`, so they are not independent paths; only the Grep glob arm
    is genuinely separate. The description now says so rather than implying
    three-way independence.
  - `normalizeWindowsBasename` runs on every platform, not only Windows. Its
    doc now states that explicitly: the guard must answer identically
    everywhere, and a name is judged by what Win32 would resolve it to.

No classification logic changed; the 12-name regression sweep is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4651): regenerate install-tree goldens, correct the guard's user-facing docs

Three things, all consequences of the fix rather than new behavior.

1. Install-tree goldens. `hooks/lib/filename-classification.js` is a SHIPPED
   file — package.json `files` includes `hooks` — so every per-runtime install
   tree gains a path. Checkpoint 2 failed on exactly this: 11 failures, all in
   tests/golden-install-tree.test.cjs, against 45356 passing. Regenerated via
   scripts/gen-install-tree-fixtures.cjs; 11 goldens changed, matching the 11
   failures one-for-one.

   This ripple was identified at design time and then not acted on. Fleet's
   impact preview named golden-install-tree.test.cjs before any code was
   written, and 40-design.md records it under "Ripples identified". Writing a
   risk down is not the same as discharging it, and a full matrix run was spent
   discovering something already known.

2. docs/USER-GUIDE.md made a precise and now-false claim about the guard's
   protected set: it named `.env.example` / `.sample` / `.template` / `.dist`
   as the four exempt names. The exemption keys on the FINAL EXTENSION, so the
   exempt set is the unbounded family `.env.<anything>.{example,sample,template,dist}`.
   The page now states that family, the widened residual, that order matters
   and only the last segment counts (`.env.example.local` is a secret), and
   that trailing dots and spaces are stripped because Windows resolves them to
   the protected file. A wrong user-facing model of what a security guard
   protects is worth correcting even though Fixed/Security changesets are
   exempt from the required-docs rule.

   docs/ARCHITECTURE.md and docs/INVENTORY.md say "templates such as
   `.env.example` exempt" — non-exhaustive, still true, deliberately left
   alone. Same for the ja-JP / zh-CN / ko-KR / pt-BR rows, which carry the same
   hedged phrasing; hand-translating a security description unreviewed is not
   something to do silently.

3. Two changeset fragments, not one. A refusal corrected is `Fixed`; a bypass
   closed is `Security`. Folding the second into the first would under-report
   it in the release notes. Both carry `pr: 0` for backfill once the PR exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4651): backfill changeset PR number to 4659

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 08:42:43 -04:00
allcounter
723ea08dc2 fix(#4016): imperative-override injection patterns tolerate filler words (#4061)
* fix(#4016): imperative-override patterns tolerate filler words

The narrow imperative-override family tolerates no filler between the
verb and the noun, so a planted "Forget all of your instructions"
(measured in a real public transcript) matched none of the 14 patterns
and both consuming hooks stayed silent.

One combined filler-tolerant pattern is appended; the narrow four stay
untouched to keep the change merge-friendly. Known trade-offs, disclosed
in #4016: linter-doc prose like "ignore rules on a single line" now
trips a LOW advisory, and the overlap with the narrow patterns means one
sentence can count twice toward severity thresholds.

Regression tests assert the previously-missed phrasings fire in BOTH
consuming hooks (gsd-prompt-guard and gsd-read-injection-scanner), not
just in the raw pattern list, per the agent brief in #4016.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015aGr6fvmDMLznT7TvTXrsb

* chore(#4016): changeset fragment for PR #4061

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015aGr6fvmDMLznT7TvTXrsb

* test(#4016): pin the disclosed linter-doc FP as single-pattern LOW, never blocking

Review follow-up on PR #4061: the combined filler-tolerant pattern's
disclosed false-positive class (linter-doc prose such as "use
eslint-disable-next-line to ignore rules on a single line") was
documented in prose only. Two tests now pin it:

- the prose matches exactly ONE shared pattern (the #4016 combined
  pattern, not a narrow one), so it cannot silently start double-counting
  toward the 3+ HIGH threshold;
- through the real gsd-read-injection-scanner subprocess with
  security.injection_blocking=true, the prose yields a single-finding
  LOW advisory and no block decision — with an in-test positive control
  proving a 3+-pattern payload DOES block in the same directory, so the
  non-blocking assertion cannot pass vacuously.

Samples are fragment-built like the existing SAMPLES rows so this file's
own diff does not trip the CI injection scanner.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#4016): replace the five narrow imperative-override patterns with one superset

The first cut appended a filler-tolerant combined pattern next to the five
narrow verb patterns. Both consumers count one finding per matching pattern
toward the severity threshold, so the overlap made one sentence count twice:
"Ignore previous instructions. Forget your instructions." scored 2 (LOW) on
next and 3 (HIGH, blockable) on the branch. It also left `override` out of
the combined pattern.

Replace the narrow family (ignore x2, disregard, forget, override) with ONE
superset pattern over ignore|disregard|forget|discard|override. At least one
filler (all|of|the|your|my|system|previous|prior|above|earlier) must sit
between verb and noun, enforced by a lookahead with no repetition; the two
noun-less/bare forms the old list accepted (`disregard (all) previous`,
`forget instructions`) are kept as explicit tails so the new pattern is a
strict superset. Bare "override rules" / "ignore instructions" are ordinary
repo prose (6 measured hits across docs and source) and stay unmatched.

Corpus measurement over 3019 .md/.js/.cjs/.mjs files (injection-sample tests
excluded): the old family hit 2 lines, the new pattern hits 3, the only new
one being a documented injection example in planner-reversibility.md that
the old family missed (the issue's own class).

Tests: SAMPLES reshaped to the 10-entry list; superset proof table (17 legacy
phrasings, each matching exactly one pattern); five issue phrasings including
`override all of your previous instructions` counted exactly once through
both hook subprocesses; double-count regression (1 finding, LOW); design pin
that bare verb+noun matches nothing; linter-doc FP pin split into bare
(silent) and determined (single LOW, never blocks). All fragment-built; the
CI prompt-injection scanner reports 0 findings on every touched file.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012rvqL1wk6s8dQEXsFm4RB1

* chore(#4016): changeset body in the canonical bold-lead format

.changeset/README.md Format: a leading bold change sentence, then an em-dash
explanation. Also drops the verbatim planted phrase from the body so the
rendered CHANGELOG line does not trip the pattern it describes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012rvqL1wk6s8dQEXsFm4RB1

* fix(#4016): render a bounded pattern label in the prompt-guard advisory, pin plural prompts

Review round 4 of PR #4061 left two nits open.

1. gsd-prompt-guard.js pushed `pattern.source` verbatim into its typed
   finding and, through renderFinding, into the user-facing advisory. With
   the #4016 superset pattern that source is 300 characters, so a genuine hit
   surfaced an advisory dominated by a raw regex dump. The read scanner has
   trimmed its equivalent since #3523 (`\s+` -> `-`, strip `()\`, cut at 50).
   That transform is hoisted into hooks/lib/injection-patterns.js as
   `describePattern` and used by BOTH hooks, so one finding renders the same
   label everywhere. Byte-identical to the scanner's old inline output for
   all 10 patterns (measured). No new staging dependency: both hooks already
   require this module.

2. The noun alternation `prompts?` had no positive coverage for the plural
   branch. One filler-regression row now exercises `... previous prompts ...`
   and runs through the existing once-per-hook, exactly-one-pattern loops.

The parity test's prompt-guard count assertion moves off substring-matching
the advisory prose onto the typed `findings` surface added in #3546, per
CONTRIBUTING's raw-text-matching prohibition. New test: the superset source
exceeds the bound (positive control), the prompt guard never embeds it, and
both hooks carry the identical label in `findings[0].match`.

Tests: parity, read-scanner, kimi field-shadowing, prompt-injection-scan,
hooks-crash-policy, dead-exports: 206 run, 196 pass, 0 fail, 10 pre-existing
platform skips. eslint clean; changeset lint ok; hooks runtime-build-seam lint
ok; the CI prompt-injection scanner reports 0 findings on the PR diff.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016GGp8kEB5zCDmJ6TYHP1Nj

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 06:58:45 -04:00
Behruz Nassre Esfahani
4499933807 fix(#3802): resolve the heredoc body before validating the commit subject (#3816)
* fix(#3802): resolve the heredoc body before validating the commit subject

With hooks.community: true, gsd-validate-commit.sh blocked EVERY heredoc-form
commit with CONVENTIONAL_COMMITS_VIOLATION regardless of the message, including
Claude Code's own documented idiom:

    git commit -m "$(cat <<'EOF'
    feat(auth): add login flow
    EOF
    )"

Reproduced before changing anything: conforming heredoc -> exit 2; plain
-m "feat(auth): add login flow" -> exit 0.

Root cause is the extraction regex `-m[[:space:]]+"([^"]+)"`. Bash `[^"]`
matches newlines, so the capture ran from the quote after -m to the FINAL quote
at `)"`, swallowing the whole span. `head -1` then returned the literal
`$(cat <<'EOF'` as the subject, which can never satisfy Conventional Commits.

Fixed by not answering a regex bug with another regex. hooks/lib/git-cmd.js
already exists because "a naive regex misses all three" invocation forms, and
extractBranchArgument is the established precedent for pulling an argument off a
git command line. extractCommitSubject joins it on the same tokenizeShellLike
seam — which, checked first, already returns the entire heredoc span as ONE
token, leaving only "resolve the body to its first line" as new logic.

Because the walk starts at the subcommand, `git -C <path> commit` and
env-prefixed invocations now extract correctly too — forms the raw string scan
never handled.

Deliberately unchanged, and pinned as such: a glued `-mfeat: x` and
`--message=...` still yield no message, exactly as the regex left them. The fix
stays scoped to the reported defect rather than widening on a true observation.

Two things I got wrong and corrected by measuring rather than reasoning:

  - I expected `git commit -m ""` to be blocked. Checked against the ORIGINAL
    hook: allowed before, allowed now, identical. The scanner drops the empty
    token so it takes the null path. My expectation was wrong, not the code.
  - That exposed a false comment I had just written, claiming the exit-status
    split prevents silently allowing `-m ""`. It does not. The split IS
    load-bearing, but for a heredoc whose body's first line is blank, which
    resolves to an empty subject and is correctly blocked. The comment now names
    the real case and records that `-m ""` is not it.

Tests at both layers: 9 unit rows on extractCommitSubject beside its sibling in
tests/worktree-safety.test.cjs, and 5 behavioral rows piping real PreToolUse
payloads through the hook in tests/hooks-opt-in.test.cjs. Replacing
firstLineOfMessageArg with a plain first-line return reds 8 of them across both
files. (A first mutation attempt silently no-opped and reported green — the
mutated body is echoed in the transcript for the run that counted.)

Out of scope, per the issue: the hooks.commit_types config surface, split off by
the maintainer as #3811 and explicitly sequenced after this.

Verified: `npm run lint:ci` exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3802): confine the fix to heredoc resolution, closing four regressions

Codex review of the first attempt. It was right, and the finding is one my own
rules already name: a true observation is not a licence to widen the diff.

The first attempt replaced the shell's `-m` extraction with a token walk. That
looked like the better abstraction — this module exists precisely because a
naive regex misses invocation forms — but selecting WHICH argument is the
message was never the defect, and changing it regressed four forms that
upstream allowed, plus opened a bypass:

  - `git commit -- -m WIP`             -- introduces pathspecs; `-m` is a path
  - `git commit --amend && echo -m WIP` a later command's flag became the message
  - `git commit -m "" --allow-empty-message`  the shared scanner drops empty
                                        tokens, so the next flag became the
                                        message
  - `git commit -m WIP`                unquoted argument
  - `-m "WIP notes <<EOF\nfix: smuggled subject"` was ALLOWED — the opener was
    recognised unanchored, so validation skipped past the real, non-conforming
    subject. An enforcement bypass, not a misclassification.

Now confined to the actual defect. The shell's `-m` capture is restored byte for
byte, and only the subject-from-message step is delegated, to a PURE STRING
helper `resolveCommitSubject()` that never tokenizes. Verified as a differential
against the upstream hook run inside the real tree: the only behaviours that
change are the two intended heredoc rows (2 -> 0); all four forms above read
identical, and the bypass case blocks.

That differential also corrected my own control. An earlier comparison ran the
upstream hook from a scratch directory, where its `lib/` could not resolve
`../../gsd-core/bin/lib/token-scanner.cjs`, so the classifier failed open and
reported exit 0 for everything. That made a real regression look pre-existing.
Re-run inside the tree, `<<-"TAG"` (a double-quoted tag nested in the
double-quoted argument) is genuinely pre-existing — the capture truncates — and
is now recorded as a known limitation rather than silently "fixed".

Also fixed from the review:
  - `<<-` strips leading TABS from body lines; returning the raw line blocked a
    conforming message.
  - a non-identifier tag such as `END-MSG` is a valid bash word and was rejected.
  - an immediately-following terminator is an EMPTY message, not a subject.
  - a node/library failure now falls back to the previous `head -1` instead of
    skipping validation, so a broken extractor degrades to old behaviour rather
    than becoming a new silent-allow path.

Tests strengthened per the review: the opener-spelling rows now assert BOTH
directions per spelling, since "conforming passes" alone would also pass if the
resolver returned an empty subject for a spelling it failed to parse. Added
differential rows pinning the five previously-allowed forms, and a row for the
bypass. Dropped two rows whose comments claimed the raw scan could not handle
`-C`/env-prefix invocations — it could; the claim was wrong.

Replacing resolveCommitSubject with a plain first-line return reds 9 rows across
both files. (Mutant body echoed in the transcript; an earlier mutation attempt
on this branch silently no-opped and reported green.)

Verified: `npm run lint:ci` exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3802): keep the installed hook runtime-neutral

`hooks/lib/git-cmd.js` ships into every runtime, including hermes and qwen,
where tests/install.test.cjs enforces that no Claude reference leaks into the
installed tree. My JSDoc named the idiom after the runtime that documents it.

Reworded to describe the SHAPE rather than the vendor; the runtime is still
named in the changeset, which feeds CHANGELOG.md where such references are
allowed, and in the tests, which are not installed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#3802): backfill changeset pr number

The fragment shipped with the documented `pr: 0` placeholder, which the
changeset lint treats as always-silent, because the number does not exist until
the PR is opened. Backfilled to 3816 now that it does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3802): close the truncated-capture hole, add the required test artifacts

Review round 1. Major 3 was the one that mattered, and it disproved a claim I
had stated in falsifiable form — the PR body said only two behaviours change;
the differential found five.

Major 3 — an embedded `"` truncates the `-m` capture, so the resolver received a
PREFIX of the real subject and the length gate measured the wrong string. Before
this fix the whole form was blocked outright, so the gate was unreachable; the
fix opened the path and then mismeasured it. A new enforcement hole, so it is
CLOSED here rather than declared.

Closed precisely rather than bluntly. A first attempt refused to resolve any body
with no terminator, which also blocked commits whose SUBJECT was intact and whose
quote sat further down the body — a false positive of its own. Truncation is only
fatal to the line it lands IN, and a captured line is complete exactly when
another line follows it, because the capture kept its newline. So an unterminated
body whose subject line is followed by more text stays measurable; only a subject
line running to the end of a truncated capture falls back to the opener, which
fails the format gate exactly as this form did before the fix.

Major 1 — fast-check property rows for the new parser, via the shared seeded
setup helper rather than requiring fast-check directly, per repo convention:
totality (a security property here, since an exception on this path fails OPEN),
idempotency, and that the result is always a single line drawn from the input —
the third catches a resolver that concatenated or trimmed while satisfying the
first two.

Major 2 — the 72-char gate is now exercised at {71, 72, 73} on the RESOLVED
heredoc subject, with the fixture length asserted so a mis-built fixture cannot
silently pass. 92 chars did not show which side of `> 72` the code sits on.

Minor 1 — leading blank body lines are skipped, as git's cleanup=whitespace does.
A conforming commit written that way was still blocked, which is the same defect
class #3802 reports.

Nit 1 — a backslash-escaped delimiter (`<<\EOF`) is now the same delimiter rather
than failing closed on a delimiter that includes the backslash.

Nit 5 — changeset trimmed from 2,208 chars of design note to the user-visible
change.

Mutation discipline, including a correction to my own: dropping the truncation
guard reds the unit rows, and the pre-review naive shape reds the hook-level row
too. My first mutant did NOT distinguish the hook row — removing the guard made
an empty slice and blocked for an unrelated reason, so the row passed and looked
proven. Only mutating to the actual pre-review shape showed it discriminates.

Verified: `npm run lint:ci` exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3802): measure the subject as git does — strip trailing whitespace, split CRLF

git's cleanup=whitespace strips whitespace at BOTH ends of a line; the
resolver handled only the leading direction, so a 72-char subject with
trailing spaces measured 75 and stayed blocked — the defect class #3802
reports, surviving one round further (review of #3816, Major 2). The
resolved subject now drops trailing spaces and tabs; the plain non-
heredoc path is untouched, keeping the fix confined to heredoc
resolution. The length-gate boundary rows gain dirty fixtures: 72+3
trailing spaces passes, 73+1 stays blocked on LENGTH.

split('\n') left \r on every body line, so on CRLF input the delimiter
never matched: the truncation guard was inert, an empty message resolved
to 'EOF\r', and a real 72-char subject measured 73. Split on /\r?\n/
(Minor 3).

The three property tests never reached the parser — the pinned-seed
fc.string corpus contained no newline and no opener, so every property
reduced to f(s) === s (Major 1). The generator now constructs heredoc-
shaped input (all opener spellings, <<- tabs, optional terminator, CRLF)
and each property asserts a floor on inputs its corpus actually resolved.
All new rows proved failing-first against the pre-fix resolver.

Also records the unquoted-delimiter expansion limit as one JSDoc
sentence (Informational 5).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3802): close two recognition bypasses, pin the dquoted-delimiter limit

Codex whole-PR review found two enforcement bypasses in the resolver:

- The opener's path prefix was \S*, which accepted `id;/bin/cat` — the
  resolver then validated the heredoc BODY while bash runs `id` first
  and git's real subject is id's OUTPUT. The prefix is now a
  path-character class; any shell metacharacter fails recognition and
  the form falls back to the opener line and the format gate.

- The blank-line skip used JavaScript trim(), whose Unicode whitespace
  class skips lines git KEEPS: a NBSP first body line resolved to the
  SECOND line while git's real subject is the NBSP line (verified
  against git stripspace — the c2a0 bytes survive). Blank is now git's
  ASCII space/tab only; a Unicode-blank line is returned and fails the
  format gate, the same fail-closed direction git takes.

Both proven failing-first at resolver AND hook level. Also: the
<<"TAG" spelling is recorded as a documented limit — the -m capture
stops at the delimiter's own quote so the caller can never deliver it
(fail closed; widening the capture would change every embedded-quote
case) — with a hook-level row pinning the limit; and the derivation
property no longer accepts '' unconditionally, only for heredoc-shaped
input, so a conditional constant-'' regression can't satisfy the corpus
floor unnoticed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3802): recognition whitespace is ASCII, and '' answers to the generator

Codex round 2: the opener's \s accepted Unicode whitespace bash does
not split on — $(<NBSP>/bin/cat was recognized here while bash reads
<NBSP>/bin/cat as the executable NAME, so recognition claimed a
substitution that does not run cat. Every whitespace position in the
recognition is now [ \t], the same ASCII rule as the blank-line skip,
proven failing-first.

The derivation property's ''-acceptance now consults GENERATION-TIME
metadata: the heredoc generator records whether it built an empty
message (terminator reachable, all scanned lines ASCII-blank, <<- tab
stripping accounted for), and '' is accepted exactly then — a resolver
conditionally degrading to '' on non-empty heredocs now fails, closing
the residual round-1 permissiveness without re-deriving resolver logic.

The changeset no longer overstates the opener spellings: it names the
capture-deliverable set and the documented <<"EOF" limit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3802): nothing after the terminator escapes measurement

Round-3 BLOCKER: everything after the heredoc terminator was silently
discarded, so `-m "$(cat <<'EOF'\nfeat: ok\nEOF\n) <200 a's>"` —
one 200+ char real subject once bash substitutes — measured 8 chars and
dodged COMMIT_SUBJECT_TOO_LONG, a hole the base did not have. The
canonical idiom's tail is exactly one closing-paren line; any other tail
now falls back to the opener line and the format gate, the pre-fix
behaviour for the whole form. Proven failing-first at resolver and hook
level, including the glued-text and second-substitution variants.

Also from round 3: `cat<<'EOF'` (no space) is legal bash and now
resolves — the token before << is still literally cat; the env-prefixed
and option-terminated spellings join the JSDoc KNOWN LIMIT list instead
(fail closed, modelling bash prefix words is cost with no reported
user); the changeset states the embedded-quote truncation limit for the
message body, not just the <<"EOF" spelling; the dquoted unit and hook
rows now cross-reference each other; and the fast-check setup helper's
docstring no longer claims property-file exclusivity.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3802): glued text outside the closing quote must not shrink the measurement

Codex on the round-3 guard: bash concatenates -m "$(…)"suffix into ONE
argument, but the capture holds only the quoted part — so the resolver
measured the heredoc body (8 chars) for a 200+ char real subject, a
net-new length-gate bypass the base did not have (base measured the
opener and blocked). When the closing quote is followed by anything but
whitespace or end-of-command, the hook now skips the resolver and keeps
the pre-fix first-line subject: the heredoc form fails the format gate
exactly as on base, and the plain single-line form keeps base behavior
unchanged — both pinned as differential rows, the glued-suffix row
proven failing-first against the unguarded script.

The property generator's ''-oracle now models the post-terminator guard
it previously predated: expectEmpty requires the FIRST reachable
terminator to be followed by the one canonical closing-paren line, so a
resolver regressing to '' on a non-canonical tail (e.g. a body line that
doubles as an early terminator) fails the derivation property instead of
being blessed by stale metadata.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: retrigger CI — the previous wave never started (Actions queue stall)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#3802): only resolve a heredoc whose body bash does not rewrite

Round-4 review found two net-new enforcement bypasses: commands the base
hook blocked (exit 2) that this branch allowed (exit 0). Both reproduced as
a base-vs-head differential against the real hook, not inferred.

The predicate "may I resolve this?" was computed from the resolver's input
string alone, while two of its determinants live outside that string:

  1. WHICH -m quote arm produced the input. Inside -m '...' bash performs no
     command substitution, so $(cat <<'EOF' is literal text and git's real
     subject is the opener line. The resolver ran on both arms, so all four
     delimiter spellings went 2 -> 0 on the sq arm — reachable by the
     ordinary slip of typing ' for ". The hook now records MSG_QUOTE and
     gates the resolver on dq; sq keeps head -1, exact base parity.

  2. WHETHER the delimiter suppresses expansion. Only <<'D', <<"D" and <<\D
     do; a bare <<D is expanded by bash before git sees it. Resolving the
     literal dodged the format gate (feat: $UNSET_VAR reaches git as feat:)
     and the length gate (feat: ${LONG} reaches it at any length). The
     opener regex now separates the backslash-quoted and bare alternatives
     and refuses the bare one — the same fail-closed rule the metacharacter,
     truncation and post-terminator guards already follow.

A test row asserted exit 0 for a bare-delimiter body, so the suite defended
the second bypass and the fix could not land without editing a test that
read as intentional. That row and its two unit counterparts now assert the
block, per RULESET.TESTS.delete-bad-tests. Two unrelated rows used <<-EOF
to exercise tab stripping; they move to <<-'EOF' so each tests what it names.

Scoping the adjacency guard to the matched arm — required by the fix above —
also removes a spurious block (round-4 Minor 1): a double-quoted heredoc
whose body mentioned a glued single-quoted token tripped the sq arm.

The JSDoc claimed <<"EOF" was unreachable through the caller and that the
bare-delimiter gap was pre-existing. Round 4 disproved both; both corrected
here, along with the matching changeset sentence.

Verified: 7 bypass commands now block at head (was allow), the #3802 fix and
plain-form parity are unchanged across 8 control commands, hooks-opt-in 44/44,
worktree-safety 401/401, property-test non-vacuity 73/200 against a floor of
20, lint:ci exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG

* fix(#3802): resolve only where the captured text is provably git's subject

Codex review of the full PR found two more inputs where the validated text
is not the subject git receives, both net-new bypasses (base 2 -> head 0),
plus one escalation of round-4 Minor 2. All reproduced here against the real
hook and confirmed against real commits before fixing.

BLOCKER — the matched -m need not be git's message. The capture is a search
over the whole command and the double-quoted arm runs first, so it could
select a -m that is not the subject at all. git concatenates multiple -m
values and takes the FIRST as the subject, so

    git commit -m 'WIP first' -m "$(cat <<'EOF' … )"

commits the subject `WIP first` while the hook validated the heredoc. Same
for an unquoted earlier -m, for a heredoc after `--` (a pathspec, not a
message), and for one belonging to a later `&& echo`. The mis-selection is
pre-existing; resolving it is what made it a bypass. The hook now resolves
only when nothing before the matched -m could have been an earlier message,
an end-of-options marker, or another command.

BLOCKER — cleanup mode is part of the predicate. The resolver skips leading
blank lines and strips trailing whitespace because git's DEFAULT
cleanup=whitespace does. Under --cleanup=verbatim git does neither, so a
72-char subject plus three trailing spaces is committed at 75 bytes while
the hook measured 72 — COMMIT_SUBJECT_TOO_LONG dodged. This one hides from
`git log --pretty=%s`, which strips trailing whitespace in its own output;
the raw commit object shows 75 vs 72. Any named mode other than whitespace,
in either the --cleanup= or -c commit.cleanup= form, now refuses to resolve.

MAJOR — recognition trusted any path ending in /cat, so a planted
`../evil/cat` printing `WIP injected` had its heredoc body validated while
git's real subject was `WIP injected`. Only a bare `cat` or an absolute path
is recognised now. A bare `cat` shadowed on PATH is a documented residual and
is not fixable from a string — nor a meaningful boundary, since planting an
executable already allows running git directly.

The changeset and the JSDoc both asserted that a `"` anywhere in the message
blocks. Measured false: a `"` on a later body line resolves fine, because the
subject completes before the truncation point; only a `"` in the subject line
blocks. The changeset also listed <<"EOF" as covered when it measures 2/2.
Both rewritten to claim only what is measured, and the residual false
positives are now named.

Verified: 4 + 2 + 3 new bypass commands now block, with non-vacuity controls
proving the default path still resolves; all round-4 maintainer blockers stay
closed; the #3802 fix and plain-form parity unchanged across 7 controls;
hooks-opt-in 47/47, worktree-safety 402/402, property non-vacuity 73/200
against a floor of 20, lint:ci exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EAbQy7n4mLMB7h3TnZ8GdG

* fix(#3802): scope the cleanup-mode guard to the command outside the message

The guard scanned the whole $CMD for `--cleanup=` / `commit.cleanup=`,
and the heredoc BODY sits verbatim inside $CMD, so any conforming
message that merely MENTIONED the token was refused, fell back to the
opener line, and was blocked with CONVENTIONAL_COMMITS_VIOLATION. These
are ordinary English in this repository, whose own hooks and docs
discuss cleanup modes constantly. Reproduced against the real hook:
`fix: document commit.cleanup=strip behavior` blocked, the same message
without the token allowed (review of #3816, round 5 — BLOCKER).

Scoping to $MSG_PREFIX alone, as prescribed, would have reopened the
round-4 length-gate bypass the guard exists for: git accepts the flag on
EITHER side of -m, and `git commit -m "<heredoc>" --cleanup=verbatim` is
caught today only because the scan is command-wide. Measured, not
assumed. The scan now covers MSG_PREFIX + MSG_SUFFIX — the whole command
minus the one span that is message text — joined with a space so a token
cannot be forged across the seam.

Swept the guard class rather than the reported instance. The adjacency
guard does not share the defect: an in-body `-m "foo"bar` is refused by
the already-documented embedded-quote capture limit (any `"` in the
subject line truncates the capture), and an in-body `-m ` without quotes
resolves and is allowed. Deliberately untouched.

Both directions pinned failing-first: the three false-positive rows red
against the unscoped guard, and the trailing-flag row reds against
prefix-only scoping. Each mutation was echoed back to prove it landed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XogDtuuuGQEfsWaLSaZCLB

* fix(#3802): read commit options the way bash hands them to git

Round 6 reported the adjacency guard scanning all of $CMD for a glued
`-m "..."`, so a glued -m belonging to a chained-after command refused a
heredoc that was never truncated. Glue is a property of the ONE character
following the matched span, so that character is now the whole window.
Separators and redirections are excluded because bash does not
concatenate across them: in `-m "msg"&& echo hi` the argument ends at the
quote, so there is no truncated capture to defend against.

An independent full-PR pass then found three accept-direction defects
this PR had introduced in earlier rounds, each measured against a real
commit by reading the raw commit object — `git log --pretty=%s` strips
the trailing whitespace that makes the length wrong and hides it:

  --cle=verbatim         git accepts any unambiguous prefix of a long
                         option, so the mode was set by a token that is
                         not the literal --cleanup. 75-byte subject
                         recorded, 72 measured.
  -am 'WIP first'        git reads this as -a -m, so the real subject is
                         `WIP first` and the heredoc is only the second
                         message. The scan looked for a standalone -m.
  --clean""up=verbatim   bash removes quotes before git sees the
  -""m                   argument, so a spliced spelling is the same
                         option and matched no literal.

The two option-name scans now read their window with quote characters
removed, which is what bash does to it, and the cleanup class covers
git's abbreviations. The adjacency test deliberately keeps the raw text:
it asks about a literal character position, not an option name.

Narrowing the cleanup window to git's own command segment was tried and
reverted. `;`, `&` and `|` end a command only outside quotes, and this is
a substring scan, not a parse: an unconditional trim cut the window short
on `--author "a&b"`, and a quote-aware trim still cut it on `--author
a\&b`. Each hid a real trailing --cleanup=verbatim and accepted a 75-byte
subject. The resulting false positive — a --cleanup carried by a chained
command refuses the commit — is documented and pinned instead. Refusing a
commit git would take is recoverable; accepting an over-long subject is
not.

Sixteen rows in tests/hooks-opt-in.test.cjs. Seven mutations, including
both reverted narrowings, so no dead end can be reintroduced silently.

* fix(#3802): close six accept-direction bypasses in the resolve guards

Round 7's FIRST-MESSAGE GUARD Major does not reproduce. Measured against the
real hook in a complete tree at the reviewed head: the classifier gate runs
before any guard, so `git add -A && git commit …` (git->add stops on a
non-commit subcommand) and `cd dir && git commit …` (the first executable is
not git) exit 0 without a guard being evaluated. The control is the proof — a
subject the bare form blocks with CONVENTIONAL_COMMITS_VIOLATION exits 0 in
both chained forms, so the hook never validated them and cannot be
over-blocking them. The guard is unchanged; scoping this scan to $MSG_PREFIX
alone is what reopened the round-4 trailing-flag bypass.

The class was real, though, one shape further out: `FOO=bar; git commit …` IS
classified and then refused, because assignment detection is prefix-anchored
and the tokenizer does not split operators. Pinned as a counterexample and
disclosed rather than generalised away; narrowing it means changing
isGitSubcommand, the shared git-commit detector every gating hook uses, and it
fails closed.

Six accept-direction bypasses are fixed. Each let the hook resolve and ALLOW a
commit whose real subject the rules refuse; the three that turn on git's
recorded subject were confirmed against the RAW COMMIT OBJECT, since
`git log --pretty=%s` strips trailing whitespace and hid two of them:

  --cleanup=whitespace -m <72+spaces> --cleanup=verbatim  git kept 75 bytes
  -mWIP -m <heredoc>                                      git recorded `WIP`
  --mes=WIP -m <heredoc>                                  git recorded `WIP`
  -\m WIP -m <heredoc>                                    git recorded `WIP`
  git commit --amend --no-edit \n echo -m <heredoc>        echo's argument read
  --squash=HEAD -m <heredoc>                              `squash! …`

Causes: one BASH_REMATCH inspected only the FIRST cleanup directive while git
applies the last, so multiplicity now refuses rather than guesses at an
argument order a substring scan cannot recover; the option scan required a
trailing space or `=`, missing attached values and long-option abbreviations;
dequoting removed quotes but not the syntactic backslashes bash also removes;
the separator scan omitted newline; and --squash/--fixup have git compose the
subject, so the supplied message is not the subject at all. Every fix widens
refusal, the direction this file documents as recoverable.

The multiplicity count first broke the hook outright: the script runs under
`set -euo pipefail` and grep exits 1 when it matches nothing, which is the
common case, so every ordinary commit died at exit 1 with no verdict. Guarded,
and only caught because the probe runs the real hook rather than the scan.

Five new rows, all five proven red against the pre-fix hook, each carrying a
non-vacuity assertion that the canonical single-`-m` heredoc still resolves.
Changeset corrected on three counts: "all fail-closed" was wrong (persistent
commit.cleanup fails OPEN, as do the -C/-c/-F/-t message sources), "global
options are all walked through" was too broad, and the chained-before claim
now states what is measured.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018FUAVz49BghqxoJgwt7EW9

* fix(#3802): stop the separator and glue classes matching a literal backslash

Round 8's Major, with two corrections to its account.

`;`, `&` and `|` are metacharacters inside `[[ ]]`, so an inline bracket
class must escape each one. POSIX bracket expressions have no escape
mechanism of their own, so on bash 3.2 -- the system /bin/bash on macOS,
already a supported target here per the `declare -A` ban in
tests/install.test.cjs -- those backslashes reach the regex engine and add
a literal `\` to the class. bash 4+ consumes them, which is why this is
invisible on a modern bash. The hazard is specific to bracket
expressions: `\(` outside one is made literal correctly on every version,
and the subject validator and the `-m` capture classes were checked and
are unaffected.

The prescribed fix is not taken, because it does not parse. Inline
`[;&|]` is a bash SYNTAX ERROR on 3.2 and on 5.3 alike -- the backslashes
exist to get the metacharacters past the `[[ ]]` parser, so removing them
leaves an unparseable script. Each class is held in a variable and
expanded unquoted on the right of `=~` instead, which is a plain regex on
both versions.

One root cause, consequences in BOTH directions. The reported half is the
separator scan over-blocking. The half not reported is the accept
direction, and it is the more serious: the glue class is NEGATED, so on
bash 3.2 a backslash-glued suffix fell inside the exclusion and the hook
RESOLVED a heredoc it should have declined -- measured exit 0 on 3.2
against the unfixed hook, exit 2 everywhere else, with a letter-glued
control refused in all four cells.

The reported repro is not actually fixed by this, and the changeset says
so. A `\`-newline line continuation carries a literal newline, which the
round-7 separator guard refuses on every bash, so that shape stays
blocked with or without this change. Narrowing the newline guard is not
attempted: telling a continuation from a separator by substring scan is
the class that was tried twice in earlier rounds and reverted both times,
and an escaped backslash sitting immediately before a real newline is
indistinguishable from a continuation. Disclosed as a known fail-closed
limit instead.

Every new row runs under each bash on the machine. Against the unfixed
hook both bash 3.2 rows go red while all four bash 5.3 rows stay green --
written the ordinary way these rows would run under PATH bash, pass
against the broken hook, and prove nothing. Two non-vacuity controls per
interpreter prove the validator is reached rather than passing
everything. All 8 rows of the established differential harness are
byte-identical before and after on both versions: no regression, no new
refusal.

* fix(#3802): remove the $ of a dollar-quote from the option-name scans

Independent round-8 review, accept direction.

The option-name windows are dequoted so they match "the command as bash
hands it to git" -- round 6 removed quote characters, round 7 removed
syntactic backslashes. Both passes missed that bash has two further
quoting forms whose introducer is a `$`: `$'...'` and `$"..."`. Removing
the quote characters alone left that `$` stranded INSIDE the option name,
so `-$"m"` dequoted to `-$m` and matched no literal, while bash passed a
real `-m` to git.

Measured on bash 3.2.57 and 5.3.15 against a real repository: the hook
allowed

    git commit --allow-empty -$"m" WIP -m "$(cat <<'EOF'
    fix: a perfectly ordinary conforming subject
    EOF
    )"

with exit 0, and `git cat-file -p HEAD` recorded the subject `WIP`.

The comparison that establishes this is HEAD-internal, not a differential:
the same command spelled `-m WIP` is refused (exit 2). The merge-base
refuses EVERY heredoc form, including a perfectly conforming one, so its
exit 2 on this input says nothing about whether any guard fired -- it is
the absence of the feature, not a working check. The same miss covered
`$'m'`, spliced `--message`, `--cleanup`, `--squash` and `--fixup`.

An option NAME finished by a command substitution -- `--clean$(printf
up)=verbatim` -- is a different problem and gets its own guard: bash runs
a program to complete the name, so the argv git receives is not derivable
from this string at all, and resolution is refused rather than guessed.
The guard is scoped to the NAME: the class is a `-`-leading token whose
characters up to the substitution contain no `=`. A substitution
supplying a VALUE -- the ordinary `--author="$(git config user.name)"`,
spaced or glued, in either window -- is untouched and still resolves,
pinned in both directions. It is a SHAPE, not a segmentation of the
command line; segmenting was tried twice in earlier rounds and reverted
both times, and that reasoning stands.

Both new rows fail against the unfixed tree with their own assertions,
proven in a complete worktree at the previous head rather than a hook
copied out of its tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy

* fix(#3802): recognise a canonical cat, not any absolute path ending in /cat

Independent round-8 review, accept direction.

Round 4 restricted heredoc-opener recognition to an absolute path, after a
relative `./cat` was measured being trusted to echo its stdin. It stopped
at "absolute", so any absolute path ENDING in `/cat` was still trusted --
the same claim the round-4 reasoning had rejected one spelling earlier.

Measured on bash 3.2.57 and 5.3.15 against a real commit: with an
executable at `/.../fake-cat/cat` printing `WIP injected`, the hook
validated the conforming heredoc body and allowed the commit (exit 0)
while `git cat-file -p HEAD` recorded the subject `WIP injected`. The
same command through `./cat` was already refused, which is the control
that shows this is the round-4 class one spelling out rather than a new
one.

Recognition is now the canonical system locations -- bare `cat`,
`/bin/cat`, `/usr/bin/cat` -- which is the only identity claim a string
can support. `/usr/local/bin` is deliberately excluded: it is
user-writable on ordinary machines, which is the plantable case this
guard exists for. Anything else falls back to the opener line and the
format gate: fail closed, exactly the pre-fix behaviour for the form.

The pre-existing residual is unchanged and still documented: a bare `cat`
shadowed earlier on PATH is indistinguishable here, and is not a
meaningful boundary -- anyone able to plant an executable on PATH can run
`git commit` directly. This hook stays an authoring guard, not a security
control.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy

* fix(#3802): an option name carrying a shell expansion is unresolvable

Independent review, round 9, accept direction. Four more spellings, and a
change of strategy that is the actual point of this commit.

Rounds 6, 7 and 8 each tried to EMULATE what bash does to an argument
before git sees it -- round 6 removed quote characters, round 7 syntactic
backslashes, round 8 the `$` that introduces a dollar-quote -- and each
round review found another transform that had been missed. Round 9 found
four more. All measured on bash 3.2.57 and 5.3.15 against a real
repository, each with the plain spelling of the same command as its
control (refused, exit 2) and `git cat-file -p HEAD` for the subject git
actually recorded:

    -$'\155' WIP        hook 0, real subject `WIP`   ANSI-C octal -> m
    -$'\x6d' WIP        hook 0, real subject `WIP`   ANSI-C hex   -> m
    -`printf m` WIP     hook 0, real subject `WIP`   backtick substitution
    x= … -${x}m WIP     hook 0, real subject `WIP`   parameter expansion
    -? WIP              hook 0, real subject `WIP`   pathname expansion

and the same class through the cleanup guard, where git recorded a
75-character subject the length gate had measured as 72:

    --cle$'\141'nup=verbatim, --clean`printf up`=verbatim, --cle?nup=verbatim

The last two settle it. An option name finished by a PARAMETER expansion
depends on a variable's value at run time; one finished by a PATHNAME
expansion depends on the contents of the working directory. Neither is
derivable from the command string at any level of effort, so emulation
cannot be completed -- not "has not been completed yet". A fifth patch in
that direction would have the same shape as the previous four.

The rule is therefore no longer "normalise it and match the literal". It
is: an option NAME carrying a shell expansion or quoting construct is
UNRESOLVABLE, and unresolvable refuses. One rule covers every spelling
above and every spelling nobody has thought of yet, in the fail-closed
direction. The dequoting passes are kept rather than replaced: they still
normalise the deterministic removals, so the guards RECOGNISE
`--clean""up=` and `-\m` as the options they are instead of merely
refusing them, which keeps the existing rows meaningful.

Scope is unchanged and still pinned in both directions: the class is a
`-`-leading token whose characters up to the construct contain no `=`, so
a construct supplying a VALUE -- `--author="$(git config user.name)"`,
the backtick spelling, `--date="${NOW}"`, a glob character inside an
author string, a pathspec after `--` -- still resolves. Nine such forms
are asserted to pass beside the seven that must refuse.

The class is bracket-only and holds no backslash, per round 8: a POSIX
bracket expression has no escape mechanism, and a backslash written
inside one becomes a literal member on bash 3.2.

The new rows fail against the previous head with their own assertion
message, in a complete worktree with the lib built, not a copied hook.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy

* docs(#3802): disclose and pin the two spellings the round-9 class over-blocks

A scoped review of the round-9 class asked one question -- does it refuse a
conforming heredoc commit that the previous head accepted -- and found two
spellings that it does. Both measured on bash 3.2.57 and 5.3.15, previous
head 518d97b64 exit 0, current head exit 2:

    git commit -S$SIGNING_KEY -m <conforming heredoc>
    git commit -m <conforming heredoc> -- -*.txt

Disclosed and pinned rather than narrowed, for two reasons.

Narrowing is not available cheaply. Dropping the bare `$` member reopens
`-$xm`: with `xm=m` bash hands git a real `-m`, which is the parameter
expansion bypass the round-9 commit exists to close. Skipping tokens after
`--` means deciding where git's options end from a substring scan, which
is the class this file has already reverted twice for opening
accept-direction holes -- a `--` inside a quoted value (`--author "a -- b"`)
would truncate the window and hide a real trailing directive.

And the limits are narrower than they look, because in both cases the
spelling a developer actually reaches for still resolves:

    -S "$KEY"  and  --gpg-sign="$KEY"        resolve
    '-*.txt', "-*.txt", ':(exclude)-*.txt'   resolve

The pathspec one is worth stating precisely: a glob only reaches git AS a
pathspec when it is quoted, because an unquoted one is expanded by the
shell before git is executed. So the refused spelling is not passing a
glob to git at all, and the spellings that do are unaffected.

Refusing a commit git would take is the recoverable direction; accepting a
non-conforming subject is not. That is the trade this file already makes
everywhere else, and it is made explicitly here.

Nine rows pin the working spellings beside the three that refuse, so a
later narrowing cannot silently drop the cases that must keep working.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy

* fix(#3802): join backslash-newline continuations before the resolve guards

Round 9's Major, with a correction to its diagnosis.

The cited bracket classes at :223 and :260 no longer exist -- round 8
moved both into SEP_CLASS and GLUE_CLASS, and a lone backslash before -m
resolves (exit 0) at the reviewed head on both bash 3.2.57 and 5.3.15.
What refuses the repro is the NEWLINE a `\`-continuation carries: round
7's separator guard reads any newline in a window as a command boundary,
and `git commit \` newline `  -m "$(cat <<'EOF' …` was refused for that
reason. Round 8 disclosed it as a fail-closed limit; round 9 calls the
idiom common and the limit a Major, and it is fixed here.

It was left as a limit because "is this newline a continuation" looked
like the segmentation question this file has reverted twice. It is not:
bash's rule is local and character-level. A newline preceded by an ODD
run of backslashes is a continuation and bash removes both; an EVEN run
(`\\` then newline) is a literal backslash followed by a real newline,
which IS a separator. Both scan windows are joined that way immediately
after they are cut from the command and before any dequote copy is
derived, in three bash-3.2-safe parameter expansions: every `\\` pair is
parked on \x01, any backslash-newline that remains is a lone one and is
removed, then the pairs are restored.

Measured on both bashes, both directions:

    git commit \<nl>  -m <heredoc>                 2 -> 0   the fix
    git commit \\<nl>  -m <heredoc>                2 -> 2   literal \ + real separator
    git commit<nl>  -m <heredoc>                   2 -> 2   bare newline
    -m <heredoc>\<nl>suffix                        2 -> 2   bash glues it; the glue guard sees it glued
    git commit … \<nl>  --allow-empty<nl>echo -m … 2 -> 2   the REAL newline still separates

The prescribed `[\;&|]` is not taken: a backslash written inside a
bracket expression becomes a literal member on bash 3.2, which is the
round-8 defect from the other side.

Rows run under each bash on the machine. The fix row fails against the
previous head in a complete worktree with the lib built; the four control
rows were measured against that same head and were already refused, so
they pin existing behaviour rather than the change.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TadqrpTE2m6gCB7CaNNLcy

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 03:51:43 +00:00
Tom Boucher
d24e22b156 enhance(#3912): gsd-tools declares outcomes, pinned at v1 (#3983)
* enhance(#3912): gsd-tools declares outcomes, pinned at v1

ADR-3889 §4. Phase 6 already moved error()'s terminator onto the seam, so what
remained was the declaration — and the pin that makes it invisible today.

The census corrected two documented figures before any code changed.
ERROR_REASON has exactly 25 members (the ADR and epic were right; an earlier
note of mine claiming 23 was wrong and is corrected). And output({error}) is
**64 sites across 9 files, not the 60 ADR-2980 ratified** — the module shape
holds but the total drifted +4: frontmatter 7 not 6, phase 4 not 2, roadmap 3
not 2. That matters because this phase's criterion demands the pin be asserted
over the enumerated population rather than sampled; asserting over a stale 60
would leave four sites unpinned while claiming full coverage, which is the
shape of failure this epic exists to remove.

The issue does not state the fact that shapes the design: output() never
touches the exit code. Confirmed by reading it — it writes fd 1 and returns.
So a declared outcome for those 64 sites had nowhere to be READ. The mapping
was never the work; wiring somewhere for the declaration to land was.

The seam already existed twice over. cli-exit.cts holds two globalThis-Symbol
cells, each because the module is emitted to three locations and a module-level
`let` would let instances disagree, and runMain already maps a code returned by
main(). A third cell inherits that solution. output() records DEGRADED for any
{error} payload — key-order agnostic, which is exactly why the "42 sites"
figure undercounts — and runMain projects the cell only when main() returns
nothing, so an explicit return still wins.

error() maps its reason through a table over the closed 25-member enum, leaving
all 278 call sites untouched; 226 of them pass no reason at all. The version
gate lives in error(), NOT in projectOutcome: registered names are
version-invariant there, so mapping a reason straight through would make USAGE
project to 64 under v1 and break the pin on its first line. projectOutcome is
left exactly as Phase 2 shipped it, DEGRADED's 0/80 asymmetry included.

Proven rather than asserted. v1 is byte-identical across three real CLI paths —
config-get plain, config-get --json-errors, and an output({error}) path —
matching exit code and exact bytes against the pre-change build. Under
GSD_EXIT_CONTRACT=v2 the same commands now exit 66 (CONFIG_KEY_NOT_FOUND ->
NO_INPUT) and 80 (DEGRADED), both looked up through the registry. An
anti-vacuity test pins that v1 and v2 genuinely differ for at least one reason,
because without it a mapping where everything projects to 1 under both versions
would satisfy every other assertion and the declaration would be theatre.

A1 iterates all 25 enum members and A3 asserts over the measured 64-site
population, so a 26th reason or a 65th site fails until it is given a mapping —
the drift guard this phase needs, given ADR-2980's own count had drifted +4
unnoticed.

Verification runs on the remote runner.

Refs #3912

* fix(#3912): the outcome cell must never lower an exit code

The remote run caught a fail-open that this phase introduced, in the phase
whose entire purpose is removing fail-opens.

`state validate --strict` on a missing STATE.md exited **0** where it must exit
1. Mechanism: `runMain` projected the pending outcome whenever `main()` returned
void, and under v1 DEGRADED projects to 0 — so a `process.exitCode` already set
non-zero by the command was clobbered down to success. Confirmed live against a
fixture, before and after.

This refutes a review conclusion recorded earlier in this phase, that the cell
was "fail-closed and can never mask a failure as success". It could, and did.
Recording that plainly so the assumption is not repeated: the cell's danger was
never only that it might add a failure — it was that projecting it
unconditionally overwrites whatever decision came before.

Projection is now guarded: it may set a code only when none is set, and an
already-non-zero exit code always wins. The full precedence — explicit `main()`
return, then an existing non-zero exitCode, then the declared outcome — is
written at the projection site. A regression test drives a void return with a
pre-set non-zero code and a pending DEGRADED, and fails against the pre-fix
build.

The second failure was my test encoding the wrong contract, not a code defect.
It asserted `output({found:false, error: undefined})` records DEGRADED because
the KEY is present. `JSON.stringify` drops undefined, so the payload the user
receives is `{"found":false}` — carrying no error at all, and calling that
degraded would hand back exit 80 under v2 for output that reads as clean. The
discriminator is a serializable error VALUE, not key presence. The test now
pins `{error: undefined}` as explicitly NOT degraded, and the design doc's
wording is tightened to match.

Verification runs on the remote runner.

Refs #3912

* docs(#3912): the versioned exit contract, and a flag defect the docs found

Diataxis pass for Phase 8, plus a real fix that only surfaced because writing
the how-to meant running its own examples.

The docs. ADR-2980's "Revisit if" clause asked for exactly the versioned
projection this phase provides, so it gets an amendment naming #3912 /
ADR-3889 section 4 as that boundary: v1 stays 0 byte-for-byte, v2 projects
DEGRADED to 80. The amendment also records the count drift rather than
restating a stale figure — the ADR ratified 60 output({error}) sites in 9
modules; the AST-measured population is 64 across the same 9 (frontmatter 7
not 6, phase 4 not 2, roadmap 3 not 2). The pin is asserted over the
enumerated 64. json-errors.md gains the outcome-declaration reference,
including the precedence order a review pass got wrong and the suite refuted:
an explicit main() return, then an already-set non-zero process.exitCode, then
the declared outcome. Projection may only ever set a code, never lower one.

A how-to is owed here and is written, not skipped. Under v1 nothing changes,
so the audience is an operator opting into v2 and needing to know what the
codes mean for a CI gate — a migration, which is how-to shaped. It covers
turning v2 on, the code table, why 80 is "ran and reported a condition" rather
than a crash, and how to split a gate that treats any non-zero as fatal. No
tutorial: there is no new entry point to learn, and under the default contract
a reader would be walked through observing nothing.

The defect. Running the how-to's own Step 1 example returned

    $ gsd-tools --exit-contract=v2 state validate --strict
    Error: Unknown command: --exit-contract=v2          (exit 64)

while the same flag trailing the subcommand worked and exited 80. The flag
half-worked, by argv position. resolveContractVersion scans argv
non-destructively, so the token survived into the dispatcher, which treats
argv[2] as the command name. --json-errors had already solved precisely this
at gsd-tools.cjs:4455, under a comment naming the hazard verbatim: "The argv
splice must happen here too, otherwise the dispatcher below sees
--json-errors as an unknown command." The later flag never got the same
treatment.

Fixed rather than documented around: the version is resolved first — which
memoizes the cell and makes an invalid value throw early — and then every
--exit-contract= token is spliced out of the dispatcher's argv copy.
--exit-contract is now listed in TOP_LEVEL_USAGE, where it never was. The
regression test pins leading position, trailing position, agreement between
the two, and a loud failure on v3 rather than a silent fall back to v1.

Neither review engine would have caught this: the defect is invisible in the
diff, because the diff does not touch argv handling. It surfaced only from
running the documentation's own example. Writing a how-to is an execution pass.

Verification runs on the remote runner.

Refs #3912

* fix(#3912): the flag splice has to run before the run-with-timeout return

An isolated review of the previous commit found that the fix did not deliver
what it claimed, and that two of its own tests were weak. All three findings
reproduced by execution before any change was made.

The fix was placed below a return. main() intercepts `run-with-timeout` at
gsd-tools.cjs:4436 and returns from there — above both the --json-errors block
and the --exit-contract splice added in the previous commit. So the flag still
died in leading position for that one command:

    $ gsd-tools --exit-contract=v2 run-with-timeout 5 -- node -e "..."
    Error: Unknown command: run-with-timeout        (exit 64, child never ran)

The previous commit message and the test's describe-block both claimed
position-independence unconditionally. That was an overclaim, not a gap left
open, and it is the part worth naming: the fix was verified by hand on the
commands I happened to think of, and `run-with-timeout` returns before the
code I was verifying.

Both global-flag blocks now run above the interception, with a comment naming
it so a later edit cannot slide them back down. Moving --json-errors up fixes
the identical pre-existing bug for that flag, verified failing beforehand
(exit 1, sdk_unknown_command). Fixing the sibling is deliberate: same defect,
same block, and a known-broken twin next to a fixed one is not a resting state.

Two tests were not pulling their weight. The invalid-value test was vacuous —
it passed against the pre-fix build, because `--exit-contract=v3` already
exited 1 there and already printed the resolve error lazily through
error() -> getContractVersion. Both its assertions held before the fix, so it
pinned nothing. The real discriminator is that the pre-fix build emits BOTH
"Unknown command: --exit-contract=v3" and the resolve error, while the fixed
build emits only the latter; the test now asserts that absence.

The leading-position and leading==trailing tests asserted proxies — "not 64",
"no Unknown command", "the two agree" — none of which pin a value, and all of
which would survive both positions being identically broken. With a .planning
directory and no STATE.md, state-snapshot exits exactly 80 under v2 and 0
under v1 in both positions. Those numbers are pinned now. The multi-token case
the descending splice loop exists for is covered too, and run-with-timeout has
regression tests for both flags.

The lesson is narrower than "test more". Hand-verifying the production
behavior does not verify that the test would have caught its absence. The
pre-fix binary has to be run against the test's own assertions.

Investigated and deliberately not changed: splicing before --cwd parsing
degrades one diagnostic from "Missing value for --cwd" to "Invalid --cwd:
<path>", but that is pre-existing — verified on the pre-fix build via
--json-errors, which already did it. This change joins the pattern rather than
creating it, and both forms exit 64 on malformed input either way.

Verification runs on the remote runner.

Refs #3912

* chore(#3912): backfill changeset pr numbers to 3983

* test(#3912): pin the reason-table invariant as set equality, not a count

A graph-backed review flagged the unchecked lookup in
expectedErrorCode3912. Investigated by execution: the drift guard DOES
hold — for an unmapped reason under v2 the production error() yields 1
while the table yields undefined, so the assertion fails. Not a
correctness defect, and deliberately NOT made tolerant, since a tolerant
lookup would destroy the guard.

Two real problems remained. The guard asserted the wrong invariant: it
counted the TABLE's keys at 25 rather than checking they match the
ENUM's values, so a renamed member keeps the count at 25 and slips past,
and a 26th member leaves the table at 25 and slips past too. Both were
then caught only indirectly, by an undefined mismatch producing 'must
exit undefined'. It is now a sorted set equality, so the failure names
the specific missing or extra reason.

And the comment above it described a '?? FAIL' fallback that does not
exist anywhere in the function. It now states what the code actually
does, verified by running it rather than by reading it.

Refs #3912

---------

Co-authored-by: sim <sim@local>
2026-08-28 08:09:05 -04:00
Tom Boucher
2ea5efc151 enhance(#3911): hooks declare their crash policy (#3960)
* enhance(#3911): give hooks an exit seam that needs no build

ADR-3889 Phase 7 foundation. The 19 shipped enforcement hooks hold 91 of the
epic's 128 terminators and cannot reach `terminateNow` today.

The obvious route — requiring `gsd-core/bin/lib/cli-exit.cjs`, as
gsd-agent-isolation-guard.js already does for two other modules — is rejected.
That precedent carries its own warning (#3582): those files are tsc output,
gitignored and absent on a raw plugin-marketplace or git-clone install, so the
hook must first call ensureRuntimeBuild() to self-heal. Making the module a
hook needs IN ORDER TO TERMINATE depend on a build inverts the dependency, and
its failure mode is precisely the fail-open this phase exists to remove: a
guard that cannot terminate cannot deny. `lint-hooks-runtime-build-seam`
already encodes that concern, and Design B would have had to add an
ensureRuntimeBuild() call to all 19 hooks to satisfy it.

So `hooks/lib/` becomes a third emit location for cli-exit and a fifth for the
registry, preserving the invariant `src/cli-exit.cts`'s own header states: it
imports nothing but node:fs and its sibling registry, and the generator
dual-emits that sibling alongside each copy so a relative require resolves next
to whichever copy loaded it. Shipping needed no change — build-hooks.js already
declares HOOKS_SUBDIRS_TO_COPY = ['lib'].

Proven, not asserted: the two files are copied into an otherwise-empty tmpdir
and a child process requires them and terminates — PASS exits 0, HOOK_DENY
exits 2 with the payload on both stdout and stderr. That test fails the moment
the hooks copy gains a require reaching outside hooks/lib/.

Also fixed inline: the registry's fifth target let any `--write` test overwrite
the real committed hooks/lib/exit-code-registry.js, because the test helper
derived only three of the other output paths. It now redirects all five, and a
regression test asserts every committed artifact is byte-identical after a
redirected write.

Install-tree goldens pick up the two new shipped paths across 11 runtimes —
insertions only, no removals. lint:ci was green while they were stale, so this
was found by regenerating rather than by a gate.

Verification runs on the remote runner.

Refs #3911

* enhance(#3911): declare a crash policy, and migrate the write guard

Adds `hooks/lib/hook-exit.js` — the hook-facing vocabulary over `terminateNow`,
hand-written because the cli-exit copy beside it is generated:

  allow(payload)          exit 0
  deny(payload, stderr?)  exit 2
  crash(onCrash, payload) whichever the hook DECLARED

`crash()` takes the policy as a required argument with no default, which is the
whole mechanism: fail-open by accident stops being expressible. A hook must
name ALLOW or DENY at the call site, and an unrecognized value terminates
INTERNAL rather than guessing. Fail-open stays legal; fail-open by omission
does not.

`gsd-write-guard.js` is the first hook migrated, all 12 sites, and it exposed a
gap in the seam. `terminateNow`'s doc comment justified its fd-2 write by
citing this hook's `emitBlock` — but modeled it as sending the same bytes to
both streams, when `emitBlock` actually sends full JSON to stdout and only the
bare `reason` string to stderr, because Kimi's hook bus feeds stderr verbatim
back to the model. Migrating as written would have turned a readable sentence
into a JSON blob for Kimi-backed agents.

#3911 requires both "all 19 hooks terminate through terminateNow" and "no
hook's effective default changes". Those are jointly satisfiable only by
teaching the seam to carry a distinct stderr payload, so `terminateNow` gains
an optional third argument: omitted, behavior is byte-for-byte what it was; a
string is written raw, which is exactly the Kimi case. The doc comment's
inaccurate claim about emitBlock is corrected in place.

Proven rather than asserted: the pre-migration file is reconstructed from HEAD
and driven with the same catastrophic-shrink payload as the migrated one —
exit code, stdout and stderr all byte-identical.

Verification runs on the remote runner.

Refs #3911

* enhance(#3911): all 19 hooks terminate through the seam

Migrates the remaining 18 enforcement hooks onto allow/deny/crash. An AST walk
now reports zero `process.exit(` call sites across every `hooks/*.js` — down
from the 91 the census measured.

Each hook with an outer catch declares its policy once, at module top, with the
reason that policy is right for that specific guard: a read guard that cannot
scan must not block the read; a statusline that renders every prompt must
degrade rather than crash; an injection scanner must not retroactively block a
result already returned. Those sentences are the deliverable — they are what
turns fail-open-by-accident into fail-open-on-purpose. No hook's effective
default changed.

Wiring exposed two defects, both fixed here rather than noted.

A SECOND stdout/stderr-splitting site turned up in `gsd-workflow-guard.js`'s
`emitForceAddBlock`, matching the pattern already known from the write guard —
full JSON to stdout, bare reason to stderr for the Kimi bus. It uses the
`stderrPayload` argument added in the previous commit, which is now carrying
its second real caller rather than one special case.

More seriously, `terminateNow` emitted both streams inside ONE try, so a
payload that failed to serialize aborted before the stderr write ever ran. The
two windsurf guards write nothing to stdout on a block and only a reason string
to stderr, so `deny(undefined, reason)` exited 2 with EMPTY stderr — a deny
that silently loses its reason, which is the exact "fails with success" class
this epic exists to close. The streams are now emitted independently, each with
its own guard, and `undefined` means "nothing to write for this stream" rather
than an error. Regression tests inject a throwing write on one fd and assert
the other still receives its payload; they fail against the single-try version.

Byte-identity was proven per hook, not assumed: each pre-change file is
reconstructed from HEAD and driven side by side with the migrated one across
its normal path, its deny path, malformed stdin and empty stdin — exit code,
stdout and stderr compared.

Verification runs on the remote runner.

Refs #3911

* enhance(#3911): harden the three shell hooks, and pin every hook's policy

`gsd-phase-boundary.sh`, `gsd-session-state.sh` and `gsd-validate-commit.sh`
gain `set -euo pipefail`.

The expected hazard did not materialize, and that is worth recording: every
intentionally-non-zero command in all three is already the condition of an
`if`/`elif`, which `set -e` never fires on, and none of them reads a
possibly-unset variable or pipes through a grep that may legitimately match
nothing. No `|| true` guards were needed. Each hook was still checked
command-by-command before the flags went in rather than after.

Twenty-one before/after cases across the three hooks — disabled and enabled,
planning and non-planning, missing STATE.md, malformed JSON, the Kimi payload
shape, quoted and unquoted `-m`, valid and over-long Conventional Commits —
all match on exit code, stdout and stderr.

The hardening is shown to actually fire, not merely added: with a stubbed
`node` that fails at the JSON-emit step, phase-boundary and session-state go
from silently exiting 0 with empty stdout to failing visibly with the error
surfaced. No such case could be constructed for `gsd-validate-commit.sh`,
whose every statement already sits inside an if-condition — recorded as
unproven rather than claimed.

`tests/hooks-crash-policy.test.cjs` adds the per-hook coverage the issue asks
for, table-driven over all 19 hooks rather than 76 hand-written cases: normal
allow, deny where a deny path exists, crash-honors-the-declared-policy, and an
unclosed-stdin case — the one `process.exitCode` structurally cannot serve. The
deny assertions encode each hook's ACTUAL stream split rather than a uniform
shape, since four of the six deliberately differ. A drift guard enumerates
`hooks/*.js` and fails if a terminating hook is ever added without a row.

Writing those tests surfaced two hooks that emit a block decision in their JSON
body and exit 0. Both were checked rather than assumed, and neither is a
fails-with-success: `gsd-read-injection-scanner.js` is PostToolUse, where the
tool has already run and exit 2 has no meaning, and `gsd-cursor-subagent-start.js`
follows Cursor's JSON-body protocol. They are deliberately left alone — a
mechanical sweep to `deny()` would have broken exactly these two.

Verification runs on the remote runner.

Refs #3911

* fix(#3838): the commit validator says when it could not validate

#3911 claims to subsume #3838. Measurement said otherwise, so this closes it
for real rather than by assertion.

`set -euo pipefail`, added earlier on this branch, does NOT fix #3838: bash
exempts a command used as an `if` condition from `set -e`, and all three of the
hook's swallow-and-pass sites are exactly that shape. Verified against the
hardened hook with a node shim that fails only the classifier call — a
non-conforming commit still exited 0 with empty stdout AND empty stderr,
indistinguishable from "your commit conforms". That is the defect verbatim.

All three sites named in #3838 now capture the real exit status instead of
consuming it as a condition, and each distinguishes its genuine negative from
"could not run":

- the classifier: 0 = is a git commit, 1 = genuinely not one, anything else =
  could not classify. Its `node -e` now wraps the require and the call in
  try/catch and exits 3 on a throw, so a broken require chain can never be
  mistaken for `isGitSubcommand` legitimately returning false — which is the
  arm that matters, since `token-scanner.cjs` is a gitignored build artifact
  and a fresh checkout lands there.
- the opt-in config read and the JSON command extraction get the same
  treatment.

On "could not run" the hook emits a diagnostic to stderr naming which check
failed and why, then exits 0. The issue confirms this is safe — it is a
PreToolUse hook, so stderr does not disturb the JSON protocol — and ranks it
the smallest sufficient fix. The gate still fails open, but it can no longer do
so silently, which is the whole complaint: a validator that disables itself
quietly costs more than one that is absent, because it is trusted.

Both controls are unchanged and pinned by tests: a conforming commit still
passes silently, a non-conforming one still exits 2 with its existing block
payload. The defect test asserts stderr is non-empty and names the failure; it
fails against the pre-fix hook.

Verification runs on the remote runner.

Refs #3911, #3838

* docs(#3911): document the hook crash-policy contract

Reference and Explanation via a new docs/features fragment (FEATURES.md is
generated from it), INVENTORY rows for the three new hooks/lib files, and an
ARCHITECTURE note on the hooks section.

How-To: docs/how-to/declare-a-hook-crash-policy.md, indexed from docs/README.md
— a hook author now has to choose and declare a crash policy, which is more
than one step and crosses into which harness protocol their hook speaks. It
covers allow/deny/crash, writing an ON_CRASH reason that is actually useful,
when a deny needs a distinct stderr payload, the two hooks whose harness reads
a JSON-body decision and must NOT use deny(), and what to do when a check
cannot run at all — with #3838 as the worked example.

Refs #3911

* test(#3911): prove the seam actually ships, and stop hand-rolling temp cleanup

Two review findings.

The acceptance criterion 'hooks/dist/** stays in parity via the build seam
(lint:hooks-runtime-build-seam)' was misstated and unmet: that lint checks
something else — that a hook requiring a compiled gsd-core/bin/lib module also
calls ensureRuntimeBuild(). Nothing exercised that the three new hooks/lib
files reach hooks/dist/lib at all. That gap is not theoretical: #770 is a
recorded ship-blocking bug where a new hook never shipped because a copy list
missed it. The suite now builds dist through the repo's own ensureBuiltHooks(),
byte-compares each shipped copy against its source, and spawns a child that
requires the SHIPPED dist copy and denies — which is what catches a copy that
exists but cannot resolve its sibling registry.

gsd-validate-commit.sh hand-duplicated mktemp/run/rm three times; one idempotent
trap on EXIT replaces them, guarded so cleanup cannot alter the exit status.
Behavior-neutral across five cases, with temp-file counts taken before and
after each run.

Refs #3911

* fix(#3911): stage transitive hook lib requires, not just one level

The remote run returned 7 failures across 3 real causes.

The important one is a PRODUCTION bug this phase exposed rather than caused.
`writeCursorHooksJson` scanned each hook script for `./lib/X` requires exactly
one level deep and never re-scanned the lib files it staged for their own
sibling requires. Nothing had a transitive lib dependency before, so the gap
was invisible. Adding hook-exit.js -> cli-exit.js -> exit-code-registry.js
made real Cursor installs ship a bundle that dies at require time with
MODULE_NOT_FOUND. It now walks to a fixed point, and a real installed Cursor
hook runs to completion.

The staging harness in shared-hooks-dir-resolution hand-copied its fixture, so
the injection scanner crashed at require time and its exit-1 was being read as
a policy decision. Migrated to copyScriptWithDeps, which walks the require
graph — the repo's recorded rule for this class, since adding another
copyFileSync keeps it alive for the next person.

The missing-lib-source test in cursor-hook-workspace-roots hardcoded which lib
file it expected to be named in the abort message; the same throw now fires for
a different file first. Its assertion is unchanged in substance — staging still
must abort rather than ship a broken hook — only the name is no longer pinned.

The last one was my own test asserting an uppercase reason code. Measured
against origin/next: the pre-change hook emits the same lowercase
'config_unreadable', so the test was wrong, not the migration. Corrected to the
real value rather than making the code match the test.

Verification runs on the remote runner.

Refs #3911

* chore(#3911): regenerate the cursor install-tree golden

The staging fix means a Cursor install now correctly carries the two
transitive lib files it was silently missing. Additive only — no path was
removed. The golden diff is the evidence the packaging defect was real.

Refs #3911

* chore(#3911): backfill the changeset PR number

Refs #3911

* fix(#3911): a git probe that timed out is not a negative

A macOS CI lane failed three deny cases at 2084ms, 2112ms and 2177ms — just
past the 2000ms budget these hooks give their git probes. The three that passed
took 72ms, 595ms and 651ms. Under shard contention `git rev-parse` overruns,
the hook reads the non-zero result as "not a git repo", and allows with exit 0
and empty stdout AND empty stderr. Under load, the guards silently stop
guarding. That is ADR-3889's thesis exactly, sitting inside the security hooks
this phase is about.

The repo had already recognized the class in one place — gsd-cursor-subagent-start.js
fail-closed-denies on `git_timed_out` (#3045) — but nowhere else.

`hooks/lib/git-probe.js` classifies a probe's outcome, distinguishing a real
non-zero exit from ETIMEDOUT, a signal kill, and a spawn failure, rather than
folding all four into `status !== 0`. Three guards route their eight git probes
through it.

The resolution is the same shape #3838 took, and the same one that issue
endorsed as smallest-sufficient: fail open, but loudly. **No exit code changes
on any path** — a developer on a loaded machine is still not blocked, which
keeps #3911's declaration-pass contract intact for exit codes. What changes is
that the hook now says on stderr which probe could not answer, instead of
presenting silence as a clean verdict.

Scope was checked across every hooks/*.js, not just the three that failed:
gsd-agent-isolation-guard spawns no git; gsd-statusline's two probes gate only
a cosmetic display segment, not an allow/deny decision, and are left alone.

The C2 deny assertion was a real-race test — it demanded exit 2 while a slow
git legitimately yields 0. It now requires the hook to either deny, or allow
with a diagnostic naming the probe that could not run; a silent allow still
fails, so the assertion is not vacuous. A deterministic regression stubs git on
PATH to sleep past the budget rather than waiting for load to reproduce it.

Verification runs on the remote runner.

Refs #3911

* test(#3911): a PATH shim cannot intercept the hooks' git spawn on Windows

The deterministic timeout regression stubbed git on PATH and asserted the
guard reports rather than silently allows. It passes on Linux and macOS and
failed on Windows in 83ms and 176ms — the stub was never invoked at all.

Mechanism: the hooks call spawnSync('git', args) with no shell:true, so on
Windows CreateProcess resolves git.exe only and never a PATH .cmd shim. The
git.cmd branch could not have worked and is removed rather than left implying
a Windows path that does. Adding shell:true to the hooks to serve a test would
change product behavior and widen an injection surface, so the case is skipped
on win32 only, with the mechanism written into the skip reason so a future
reader does not 'fix' it that way.

Linux and macOS keep the coverage, and macOS is where the underlying fail-open
was actually caught.

Refs #3911

---------

Co-authored-by: sim <sim@local>
2026-08-27 22:21:10 -04:00
Tom Boucher
bf2332e67c fix(#3582): route every hook's compiled-module require through the self-heal build seam (#3629)
* test(3582): failing-first cold-tree coverage and the seam drift lint

On a plugin-channel install the compiled gsd-core/bin/lib/*.cjs are legitimately
absent (ADR-457 build-at-publish; the npm package builds before publishing, a raw
tree materialization never does). gsd-tools.cjs calls ensureRuntimeBuild() before
requiring ./lib; no hook does, so the isolation guard's Cannot-find-module lands in
its fail-closed catch and is misreported as an unreadable dispatch-isolation
configuration, blocking every executor dispatch.

These tests fail on that: cold-tree runs of the isolation guard, statusline, cursor
guard and update worker, plus the seam's actionable build error surfacing instead of
the generic misreport.

Also adds the drift lint the acceptance criteria require, with a fixture proving it
CAN fail — a guard never shown to fail is worthless. It is red here by design: it
flags today's unfixed hooks, which is exactly the defect.

* fix(3582): route every hook's compiled-module require through the self-heal seam

RED proven at 5b174b0d: 11 failures — the cold-tree runs for the isolation guard,
cursor guard and update worker, the fail-closed-with-actionable-message assertion, and
the lint's own real-tree check.

The compiled runtime library is produced by build:lib and gitignored (ADR-457,
build-at-publish). The npm package builds before publishing; a plugin-marketplace or
git-clone install materializes the raw tree and never does, so on that channel those
modules are legitimately absent. The self-heal seam added by #2002 exists to heal exactly
this, and the CLI entrypoint already calls it — no hook did. The isolation guard's
Cannot-find-module therefore landed in its fail-closed catch and was reported as
'could not read or resolve dispatch-isolation configuration', so an ARTIFACT ABSENCE was
misdiagnosed as an unreadable project config and every executor dispatch was blocked.

All SEVEN affected files now call the seam before their first compiled require. The issue
named four; a scan found six; implementing it surfaced a seventh — the shared isolation
sentinel helper, used by BOTH guards, which requires two compiled modules itself and
would have defeated the guards' own fix on a genuinely cold tree. Same defect class, so
fixed here rather than left as a known-broken remainder.

Failure posture is deliberately split by hook kind:
- Gates (agent isolation guard, cursor subagent start) surface the seam's actionable
  build error distinctly instead of swallowing it into the generic text, and stay
  fail-closed — a genuinely unreadable project config still DENIES exactly as before.
- Cosmetic and detached hooks (statusline, update worker, update check, update banner)
  DEGRADE rather than crash: the statusline draws on every render and the worker is a
  detached process, so a build failure there must not take down the prompt.

The npm path is untouched: the seam's already-built fast path returns immediately, so
prebuilt installs pay nothing and behave bit-for-bit as before.

Adds a drift lint, wired into the CI lint chain, so the invariant is enforced rather than
remembered — without it the next hook to add a compiled require reintroduces the class
silently. It is proven able to fail: a fixture hook requiring a compiled module without
the seam is flagged, and one that uses the seam is not. Verified directly — on the
unfixed tree it named all seven offenders; with the fix it passes.

While writing the lint's comment stripper, a naive whole-text block-comment regex ate its
own fixture, because this repo's comments legitimately spell the compiled-lib glob whose
star-slash reads as a comment opener. Rewritten as a line-based scanner with a regression
test pinning that case.

* fix(3582): test the three untested seam call sites and assert typed reason codes

Two independent reviews converged on the same major gap: the fix wired the seam into
seven files but only four had cold-tree tests. The adversarial pass put it plainly —
deleting the shared isolation-sentinel helper's seam call would not have failed any test
in the diff. That file was my own addition beyond the issue's four, so it shipped
untested; that is now closed.

- Shared isolation-sentinel helper: its seam call is only reached when .planning is NOT
  directly under cwd, and every existing cold-tree fixture puts it there, so the early
  return always fired first. Now covered, and proven load-bearing by mutation: with the
  call removed the spy records zero seam invocations and the test fails.
- update-check hook and update-banner hook: cold-tree tests added asserting the DEGRADED
  VERDICT — the fallback cache filename, and silent suppression when the package name
  degrades to null — rather than merely 'did not throw'. The banner hook previously had
  no test file at all.

Standards violation fixed: two tests asserted on free-form prose via assert.match against
a JSON reason string, which CONTRIBUTING bans by name — its own BAD example is exactly
that. The ESLint rule only covers readFileSync/spawnSync text, so tooling did not catch
it. Both isolation guards now emit a machine-readable reason_code from a frozen enum,
following the repo's existing REASON convention, and the tests assert that instead. The
human-readable message is unchanged for operators; only the assertion target moved.

The duplicated degrade boilerplate across the three cosmetic hooks was deliberately NOT
extracted, and the reason is recorded at each site: both viable shapes — a
path-parameterized helper, or a ceremony-only wrapper — defeat the drift lint's per-file
literal co-occurrence check, so extracting would require the lint to special-case its own
helper. Triplication is the lesser evil while the lint stays a co-occurrence scan.

The lint's header now states what it does and does not catch (literal quoted requires
only; hooks/ scan root), so a future reader does not over-trust a guard that a
concatenated path or a require inside a non-hooks helper would evade.

* chore(3582): regenerate the committed install-tree fixtures

Adding a new shipped hook helper changed the install tree, and those fixtures are
committed-and-derived (regen:derived / gen:install-tree), so 12 'install tree — <runtime>'
tests failed on 541a1913. Regenerated rather than hand-edited.

The delta across all 15 runtime fixtures is exactly two lines — the new helper under both
its hooks/ and gsd-hooks/ install paths — and nothing else, so the regeneration pulled in
no unrelated drift.

This is the bookkeeping ripple a new file under hooks/ carries; it was not visible from
lint:ci, which passed both before and after.

* chore(3582): backfill changeset PR number (#3629)

---------

Co-authored-by: sim <sim@local>
2026-08-18 14:11:23 -04:00
Tom Boucher
268ca7e32d fix(#3504): harden hook injection patterns and force-add guard (#3510)
* test(#3504): add failing-first parity, fail-closed, and bypass suites

* fix(#3504): harden hook injection patterns and force-add guard

* test(#3504): stage the scanner lib dependency in shared-hooks fixture

* chore(#3504): backfill changeset pr number

* test(#3504): build parity samples from fragments for the ci scan

---------

Co-authored-by: sim <sim@local>
2026-08-14 21:19:35 -04:00
Tom Boucher
470389f3a2 chore(#3212): tokenizer-first for stateful grammars — a shared scanner — Phase 3 (#3424)
* feat(#3414): promote git-cmd.js token-walk into a shared scanner, fix #3169

Phase 3 of epic #3212 (ADR-3212 §4). New src/token-scanner.cts generalizes
hooks/lib/git-cmd.js's proven token-walk (#3129 — "has not re-opened"):
tokenizeShellLike (quote-aware shell tokenizer, byte-identical port) and
indentWidth (bullet-nesting depth).

git-cmd.js migrates onto tokenizeShellLike with zero behavior change
(parity-asserted against every existing #3129 fixture in
tests/worktree-safety.test.cjs's folded block); isGitSubcommand's phases
1-3 (env-prefix skip, executable check, global-option consume) extracted
into skipToSubcommand, shared with the new extractBranchArgument (git
checkout -b / git branch <name>) — a new capability exercising the seam
on the domain the ADR names, not a migration of existing duplicated logic
(none existed).

Fixes #3169: src/decisions.cts's parseDecisionLines couldn't distinguish
a cross-reference bullet nested under an open decision from a fresh
malformed declaration attempt. An earlier bold-run-content-classification
design was tried and disproven against the repo's own existing FIX-B
fixtures (D-02, "no colon no dash") before being adopted — both have
identical shape under any content-only rule. Nesting depth (via
indentWidth) is the actual distinguishing signal: a bullet indented
deeper than the currently-open decision's own bullet is elaboration,
folded into its text like a continuation line, never tested against the
parse-miss guard. A bullet at the same-or-shallower indent is unchanged.

Scope-narrowing disclosed, not silent: of the ADR's four named bugs
(#3197, #3169, #2570, #2528), three no longer need this phase's work.
were independently fixed and closed since the ADR was authored — #2570's
fix is already a correctly-bounded regex per the ADR's own decidability
test (no scanner needed); #2528's fix is a deliberate, twice-reviewed
non-scanner design (its own code comment records a scanner-based attempt
that regressed a symmetric case and was reverted) that this phase does
not disturb. Only #3169 required new work.

get_impact: isGitSubcommand CRITICAL/196 affected symbols,
parseDecisionLines CRITICAL/164 affected symbols (ADR §6 due diligence).

Six-gate ripple: .gitignore, eslint.config.mjs, docs/INVENTORY.md,
docs/INVENTORY-MANIFEST.json (regenerated), CONTEXT.md glossary.

Design: .gsd/phase/chore-3414-tokenizer-first-seam/40-design.md
Test matrix: .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3414): add required fast-check property tests per code review

TESTING-STANDARDS.md:169 requires at least one fast-check property test
for any module that implements parsing — src/token-scanner.cts had none,
an orthogonal Standards-axis review finding. Adds two seeded property
tests (mirroring Phase 1/2's fast-check-setup.cjs convention):
indentWidth counts exactly a generated leading-space run; tokenizeShellLike
round-trips a generated array of whitespace/quote-free words joined with
single spaces.

The design doc's own "no property test needed" rationale was wrong — it
argued no algebraic law applied, but the standard is unconditional for
parsing modules regardless of whether one "feels" applicable. Corrected
in .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md.

Also fixes two Spec-axis wording drifts the same review found between
the design doc and the shipped code (doc-only, no behavior change):
extractBranchArgument's documented signature dropped an unused
subVariants parameter that was never implemented, and the #3169
fail-first fixture description corrected from "15-decision plan via
cmdDecisionCoverageVerify" to the actual compact 3-decision analog via
the real blocking gate, check.decision-coverage-plan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3414): add changeset for #3169 fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3414): backfill changeset pr number to 3424

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-13 23:08:34 -04:00
Tom Boucher
8f75e27554 fix(#3045): fail closed when an executor dispatch drops its resolved isolation (#3069)
* feat(#3045): deny an executor dispatch that drops its isolation flag

Every isolation gate already resolved correctly. The resolved value then reached
the executor through a prose instruction telling the model to substitute it into
a call the model composes itself, and nothing verified the substitution. When it
was dropped, the executor edited and committed in the user's primary checkout
with no consent and no warning.

A prose backstop would be the same class of artifact as the defect, so this is a
shipped PreToolUse hook on the Agent tool. It fires at the instant of the call
rather than being read once at the top of a workflow, which is the only placement
the model cannot skip.

The guard is inert unless it can positively establish that this is a GSD project,
that the project resolves to harness isolation, and that the dispatch targets an
executor. A non-GSD repo has no invariant to enforce. Where it cannot read the
configuration at all, it denies rather than assuming, with its own reason -- a
guard that cannot verify must not answer safe. A malformed payload allows rather
than throwing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#3045): extend the isolation guard to Cursor

Cursor is the second of only two runtimes that resolve harness isolation, so
shipping the guard for Claude alone left half the exposed surface unguarded while
the changeset implied it was covered.

The two runtimes fail differently. On Claude the harness flag is a per-dispatch
kwarg the model must copy into a call it composes, and the defect is that it can
be dropped. On Cursor the flag is --worktree, which applies to the whole session,
and the subagent-start payload carries no isolation field at all. There is no
flag to check, so the guard verifies the effective state instead: whether the
workspace is genuinely running outside the user's primary checkout. That is a
stronger check than the Claude one because it tests reality rather than intent,
and it is commented so nobody later rewrites it into a flag check.

Isolation is established two ways, either sufficient: the workspace resolves to a
linked git worktree, or it sits under the worktree root Cursor manages. The
second matters because a directory Cursor placed there is a legitimate isolated
session even before it becomes a distinct git worktree, where linkage alone would
report no repository.

Detecting linkage required a new primitive rather than the existing context
resolver. That resolver short-circuits on finding a local .planning directory
before it ever compares the git directory to the common one -- and an isolation
worktree normally has its own checked-out .planning. Reusing it would have read a
correctly isolated session as unisolated and denied it, which is the failure
direction that gets a guard switched off. The comparison is now its own
shortcut-free function that the resolver delegates to after its own shortcut, so
existing behavior is unchanged, and the case that would have broken is pinned.

The subagent type is checked before any configuration is read, so an unreadable
config cannot deny a dispatch this guard would never have enforced against.

The input-schema comment on the Cursor hook documented only the fields common to
every event and omitted the ones specific to this one. That omission cost a
halt during this work; it now documents both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3045): enforce the resolved dispatch decision, not the host capability

The guard keyed on the registry's dispatch.isolation, which says only that a
runtime is CAPABLE of harness worktrees. The decision that actually governs a
dispatch is the one the workflow resolves after gating, and that legitimately
comes out as sequential in three documented cases: a project setting
use_worktrees false, a per-plan submodule intersection, and the base-check
auto-degrade. The workflow tells the model to omit the flag in exactly those
cases, and the guard was denying every one of them.

The third case matters most. The preceding fix made the base-check degrade on
git timeouts and a missing git binary, where it had previously answered "safe".
That correction is right, and it means a transient hang now degrades to
sequential far more often than before -- so the two changes composed into a trap
where the workflow behaved exactly as designed and the guard blocked it.

The workflow already resolves isolation in shell, deterministically, which is
what makes it a trustworthy source in a way the model-authored call is not. It
now records that resolved value through a dedicated verb, and both guards read
it first. A fresh record is authoritative, so sequential dispatches pass
untouched. Absent or stale, the guards fall back to the capability check
combined with the project's use_worktrees setting, which still covers the case
that never reaches the workflow.

Also widened the matcher to accept Task alongside Agent, since a host that names
the tool Task would otherwise leave the guard silently inert while implying
coverage; stopped assuming Claude when no runtime is declared, which is the
shipped default and would have demanded a Claude-only argument elsewhere; and
made a non-git project inert rather than denied, since advising a worktree
session is not actionable without a repository.

The original diagnosis never modeled sequential mode as legitimate. That
omission is what let this through, and it is now recorded there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3045): record at resolution and bind the record to its dispatch

Two independent reviews converged on the same failure: the guard was fail-open in
a default install, so it did not catch the defect it exists to catch. A shipped
project carries no runtime key, which made "runtime not confidently known" the
common case rather than a corner one. A record asserting that isolation was
required but carrying no flag then fell through to a capability lookup that
answered "none", and the dispatch was allowed. The flag itself only arrived from
a second shell block -- the same block a model dropping the argument would also
skip. A test had pinned that behavior as intended.

The record is now written by the resolver, as an unavoidable consequence of
asking for the value, rather than by a step the model is told in prose to go and
run. A guard against a prose-carried value cannot itself depend on prose. Mode,
flag and identifiers are written together and atomically, so the flagless window
is gone, and a record asserting isolation with no resolvable flag now denies
instead of degrading. Runtime is also resolved from the installer's own recorded
default, which makes confident resolution the normal case.

The per-plan submodule gate degrades after the phase-level decision and never
re-recorded, so a plan that legitimately ran sequentially was denied against a
still-fresh phase record. It now records its own, scoped to the plan.

A record also authorized any dispatch for four hours. One phase degrading to
sequential could silently license an unisolated dispatch in the next. Records
now carry phase and plan, the guards require them to match, and the window is
minutes rather than hours -- the resolver rewrites it before every dispatch, so
a long window bought nothing and only widened the hole.

The flag validator rejected any value beginning with two dashes, which is exactly
the form Cursor and Windsurf declare, so their real value could never have been
stored. Writer and reader also derived the record path differently and diverged
inside a linked worktree without local planning state.

The predictable path remains a way to silence the control without leaving a trace
in the diff. It grants no access an agent with shell does not already have, so it
is documented as accepted rather than redesigned around.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3045): correct the staleness boundary and unmask a vacuous parity test

The remote runner returned twenty failures. One was a real production defect the
boundary case existed to catch: a record whose age exactly equalled the staleness
window was treated as fresh, so it stayed authoritative for one tick past its own
expiry. Freshness is now strictly inside the window.

The parity test meant to stop the two guards' executor lists from drifting could
never have failed. Its project fixture was a bare directory rather than a
repository, so the non-git inert branch answered before the executor list was
ever consulted. It asserted agreement it never actually measured. The fixture is
now a real repository, like every sibling in the file.

A test also asserted that Windsurf declares the worktree flag. It does not --
Windsurf resolves to no isolation by design, having no named concurrent dispatch
to isolate. The test claimed a registry fact that was never true, and a comment
in the resolver repeated it. Both corrected, and the test now proves what it
should have all along: that the parser accepts any bare flag value, rather than
one runtime's supposed value.

The new guard was missing from the bundled-hook whitelist, which is the surface
that decides what actually ships, and the per-plan gate had gained calls to the
launcher without the preamble those calls require. The changeset carried
parenthetical product descriptions the purity rule forbids.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3045): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3045): make the guard tests hold on Windows

Two tests redirect HOME to control where the installer-persisted runtime default
is read from. Node resolves the home directory from USERPROFILE on Windows and
never consults HOME, so both silently read the real runner profile, found no
recorded runtime, and asserted against a project the hook had not recognised. The
production code was already correct in asking the platform rather than the
variable; only the tests were wrong to assume one variable answers everywhere.
The helpers now mirror the override onto both.

The symlink spoofing test also created a directory symlink unconditionally, which
needs elevated privileges on Windows. It survived on this runner, but it would
fail on any host without them, so the creation is now attempted and the test
skips explicitly when it cannot be done -- a bare return would have counted as a
pass and hidden the gap.

Skipping alone would have left the platform uncovered, so the behaviour it proves
is now also driven in-process through an injected realpath, following the seam
already used for the clock. That case no longer depends on privileges at all, and
the end-to-end test keeps its original assertions wherever symlinks work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 23:42:16 -04:00
Tom Boucher
c7c2fe3c2b fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd (#2680)
* fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd

gsd-cursor-session-start.js and gsd-cursor-stop.js both resolved the project as
path.join(process.cwd(), '.planning', 'STATE.md'). Under the cursor-agent CLI,
hooks are invoked with cwd set to the Cursor config dir (~/.cursor), not the
workspace — so the lookup always missed. sessionStart could only ever emit the
"no .planning/ workflow found" nudge and stop's verify-work reminder could never
fire, even with .planning/STATE.md sitting in the workspace. Slash commands were
unaffected, which is why only the hook layer looked blind.

Both hooks already buffered stdin into `raw` and never parsed it; the payload's
workspace_roots carries the real path.

Multi-root was left open in the report ("first root vs any root"). Resolved
forward: prefer the first root that actually carries .planning/STATE.md, so a
workspace whose GSD project is not the first root still resolves — strictly
better than first-root-only and identical to it in the single-root CLI case.
Falls back to roots[0], then to cwd, keeping IDE behavior unchanged if the IDE
ever invokes hooks from the workspace.

The resolver is duplicated verbatim across the two scripts rather than shared via
hooks/lib/: these hooks ship standalone, and a new hooks/lib/ file must be
registered in the GENERATED installer's GSD_HOOK_LIB_FILES allowlist — the
installer-omits-shipped-file class that yields MODULE_NOT_FOUND at runtime. Per
CLAUDE.md "Generative Fix Divergence", the duplication carries a parity assertion
so the copies cannot drift.

Failing-first, demonstrated by direct invocation with cwd != workspace:
  pre-fix  sessionStart -> "no .planning/ workflow found"   stop -> {}
  post-fix sessionStart -> ".planning/STATE.md is present"  stop -> reminder

tests/fix-2587-cursor-hook-workspace-roots.test.cjs spawns the real scripts as
child processes with a cwd lacking .planning/ and workspace_roots pointing at it.
Boundary coverage on the roots array (0 / 1 / 2 entries), plus malformed-JSON
fail-open, junk-entry filtering, the parity assertion, and a guard that neither
script resolves .planning from cwd again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* fix(#2587): extend workspace_roots fix to subagentStart; keep cwd a candidate

Three findings from the isolated review, all fixed.

1. MISSED SITE (high). gsd-cursor-subagent-start.js carried the identical
   defect at line 43 — its own header documents workspace_roots in the input
   schema, but it resolved .planning/ from process.cwd() anyway. Under the
   cursor-agent CLI that meant every Cursor subagent (planner, executor,
   verifier) started with "no .planning/ workflow found" and no phase context.
   The report named only sessionStart and stop; the defect class was wider.
   Verified pre-fix vs post-fix by direct invocation with cwd != workspace.

2. SEMANTIC NARROWING (medium). The first cut searched only workspace_roots and
   fell back to cwd solely when the array was EMPTY. So when roots were supplied
   but none carried .planning/ while cwd did, the hook reported absent — where
   the pre-fix code, which always used cwd, reported present. That contradicted
   the fallback's own stated intent of preserving IDE behavior. cwd is now a
   CANDIDATE in the search (`[...roots, process.cwd()]`), so the fix is a strict
   superset of both the old behavior and the CLI fix, never a narrowing.

3. STALE GOLDEN FIXTURES (high, would have failed CI). The golden-install-parity
   fixtures store a content hash per installed file; these three hooks appear in
   13 of the 19 runtime fixtures. Regenerated via `npm run gen:golden` — the
   diff is exactly the three hook hashes in exactly those 13 runtimes.

Tests extended: subagentStart resolution via workspace_roots; the stop hook's
absent branch (previously only session-start's was covered); an explicit
regression guard that a project at cwd is still found when roots miss; parity now
asserts all THREE copies byte-identical; and the cwd guard sweeps the whole
RESOLVING_HOOKS list so a future hook in this family cannot be left on cwd.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* refactor(#2587): extract cursor workspace resolution to a shared hooks/lib module

The duplicate-plus-parity-test approach was the wrong call. The reported issue
named two hooks; a third (subagentStart) had the identical defect. That is the
signature of a systemic problem, and three copies of a resolver guarded by a
parity assertion is a divergence risk maintained by hand rather than a fix.

hooks/lib/cursor-workspace.js is now the single implementation. All three Cursor
hooks require it; none defines a local copy. Divergence is prevented
structurally instead of by asserting three copies stay byte-identical.

The reason duplication looked necessary was real, and is fixed properly here
rather than worked around: Cursor sets hostBehaviors.skipSharedHooksInstall
(#2089), so it never reaches the installer's bulk hooks/lib copy — it was the
ONE runtime shipping these hooks WITHOUT hooks/lib (verified against all 19
golden fixtures: cursor had the hook scripts, no lib). A naive require would
have thrown MODULE_NOT_FOUND at load, BEFORE each hook's own try/catch, wedging
every session on precisely the runtime this bug is about.

writeCursorHooksJson (src/runtime-hooks-surface.cts) now stages the hooks/lib
helpers the staged scripts actually require, discovered by scanning their
require('./lib/…') calls rather than a hardcoded name — so a future helper
cannot be silently omitted. This is narrower than flipping
skipSharedHooksInstall, which would wrongly pull in every shared hook.
cursor-workspace.js is also added to GSD_HOOK_LIB_FILES so uninstall and the
manifest manage it for the runtimes that do receive hooks/lib.

Verified against a REAL install (runMinimalInstall, cursor/global): the helper
is staged, and all three INSTALLED hooks resolve the workspace end-to-end from a
cwd that is not the project.

Also closes the review gap that the stop hook was excluded from the
cwd-candidate regression loop — it now sweeps RESOLVING_HOOKS. The byte-parity
test is replaced by a structural guard (every hook requires the shared module,
none redefines it) plus a new install test asserting the helper is staged and
the installed hook actually loads against it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* fix(#2587): fail loud on a missing hook lib source; drop unsubstituted version marker

Two findings from the installer-focused review.

H1 — the staging step's `if (!fs.existsSync(libSrc)) continue;` silently defeated
the very guarantee it was added for. Reproduced: delete hooks/lib/cursor-workspace.js
from source, run the cursor install — it exits 0, prints "Done!", and ships the
three hook scripts with an EMPTY hooks/lib/. The installed hook then throws
`Cannot find module './lib/cursor-workspace.js'` at load, before its own
try/catch, wedging every session — and nothing surfaces until a user hits it.
The scan protected against a required-but-UNLISTED helper while leaving
required-but-MISSING wide open (typo, bad rebase, an accidental delete).
It now throws: a missing helper source is a packaging bug and aborts the install.

M1 — hooks/lib/cursor-workspace.js carried a `gsd-hook-version: <placeholder>`
marker that NOTHING substitutes: copyLibDir stamps .sh files only, and
writeCursorHooksJson's staging applies just the colon-to-dash rewrite. Verified
the literal was reaching disk on both the bulk (--claude) and Cursor
(--cursor) paths. hooks/lib/git-cmd.js — the only pre-existing hooks/lib/*.js —
carries no such marker, so this was newly introduced, not inherited. Marker
removed, matching that precedent, with a note on why. (The explanatory comment
deliberately does not spell the token out, or it would reintroduce the literal.)

M2 — the require-scan regex demanded the exact compact form, so
`require( "./lib/x.js" )` would silently fail to stage its helper and compound
H1. Now tolerant of interior whitespace and either quote style.

Regression test added for H1 — the reviewer confirmed the invariant had zero
coverage repo-wide: a source tree carrying the hooks but no hooks/lib/ must make
writeCursorHooksJson throw rather than produce a broken install.

Re-verified end to end: the missing-source case throws, no unsubstituted literal
ships, and the installed hook still resolves the workspace from a foreign cwd.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* chore(#2587): backfill changeset pr number (#2680)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 19:57:58 -04:00
Tom Boucher
dacc23137a feat(3347): opt-in auto-update of knowledge graph after main HEAD advances
Closes #3347

Config:
- Add graphify.auto_update (default false) to manifests:
  sdk/shared/config-defaults.manifest.json, config-schema.manifest.json

Hook:
- hooks/gsd-graphify-update.sh — PostToolUse Bash matcher
  - Gates: tool_name=Bash, HEAD-advancing git op, CI=unset, in git repo,
    current branch == default branch (git.base_branch override or main/
    master/trunk fallback), graphify.enabled && graphify.auto_update both
    true, graphify on PATH, no live PID lock
  - Writes .planning/graphs/.last-build-status.json with status=running
    synchronously, then detaches hooks/lib/gsd-graphify-rebuild.sh
- hooks/lib/gsd-graphify-rebuild.sh — detached rebuild runner
  - PID-lock acquire + trap-on-exit cleanup
  - graphify update . then cp graphify-out/* → .planning/graphs/
  - Status file rewritten to status=ok|failed with exit_code, duration_ms,
    head_at_build
- Portable detach (subshell + disown, no setsid dependency)

Installer:
- bin/install.js: register hook as PostToolUse Bash matcher (5s timeout)
- Add to gsdHooks uninstall list and expectedShHooks warning list

Planner / researcher status surface (issue #3347 reviewer must-have AC):
- agents/gsd-planner.md and agents/gsd-phase-researcher.md
  load_graph_context steps now read .last-build-status.json and surface:
  running → "rebuild in flight"; failed → "auto-rebuild FAILED at {ts},
  context is from prior build"; ok with stale head_at_build → "HEAD has
  advanced since last build"

Settings:
- get-shit-done/workflows/settings.md adds "Graph auto-update" question
  with No-Recommended default; bullets and update_config block updated

Inventory:
- docs/INVENTORY.md hook count 12 → 13 with new row
- docs/INVENTORY-MANIFEST.json regenerated

Tests:
- tests/feat-3347-graphify-auto-update-config.test.cjs (8 tests):
  isValidConfigKey accepts graphify.auto_update, CANONICAL_CONFIG_DEFAULTS
  default false, config-set round-trip, sibling key preservation
- tests/feat-3347-graphify-auto-update-hook.test.cjs (18 tests):
  all bail paths (non-Bash, non-HEAD-advancing, enabled=false,
  auto_update=false, CI=true, non-default-branch, missing graphify bin,
  live-PID lock), dispatch path with mock graphify bin (sync running
  status + detached transition to ok/failed), stale-PID lock, all five
  HEAD-advancing command matchers, git.base_branch override

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-15 11:59:43 -04:00
Tom Boucher
7827e1ddee fix(#3129): replace bypassed bash regex with token-walk git-cmd.js classifier (#3141)
* fix(#3129): replace bypassed bash regex with token-walk git-cmd.js classifier

Root cause: gsd-validate-commit.sh used:
  if [[ "$CMD" =~ ^git[[:space:]]+commit ]]
This regex silently bypasses Conventional Commits enforcement for:
  git -C /path commit -m ...     (working-directory prefix)
  GIT_AUTHOR_NAME=x git commit   (env-var prefix)
  /usr/bin/git commit -m ...     (full-path executable)

Fix: introduces hooks/lib/git-cmd.js with isGitSubcommand(cmd, sub) —
a token-walk classifier that handles all four forms by:
  1. Skipping leading VAR=VALUE env assignments
  2. Validating the git executable (basename check for full-path support)
  3. Consuming git global options (-C <path>, --git-dir=, -p, etc.)
  4. Checking the subcommand token

The hook delegates to this classifier via node shell-out. node is
already called twice in this hook (config check + JSON parse), so no
new runtime dependency.

This becomes the single source of truth for all hooks that gate on
git subcommands (pre-commit-review-gate, post-push-verify, etc.).

Regression test: 27 assertions — tokenize correctness, 12 must-match
cases (including all 3 bypass forms), 8 must-not-match cases, 3 source
checks. All are real behavioral tests, not string comparisons.
Suite: 7035/7035. Closes #3129.

* fix(lint+hook+changeset): allow-test-rule, fix HOOK_DIR quote injection, fix changeset pr+typo
2026-05-05 15:02:15 -04:00