Commit Graph

8 Commits

Author SHA1 Message Date
Tom Boucher
7e905aa137 feat(#2505): Phase 0 — Kimi PreToolUse guard vocabulary normalization (precondition; carries PR #2326 forward) (#2518)
* fix(#2304): normalize Kimi tool vocabulary in PreToolUse guard payload checks

The Kimi [[hooks]] registrations translate the matcher to Kimi's tool
vocabulary (WriteFile|StrReplaceFile) but the guard scripts early-exit
unless the payload's tool_name is a Claude name (Write/Edit/MultiEdit),
so every guard was dormant on Kimi: the matcher fired, the script saw
WriteFile, and exit(0)'d.

Normalize the payload's tool_name at the top of each guard
(WriteFile -> Write, StrReplaceFile -> Edit; bare or module-qualified
kimi_cli.tools.file:* forms) before the check. Inlined per guard rather
than a hooks/lib/ helper because hook scripts are staged as standalone
files on every hook surface, and a sibling require is a staging
dependency that can fail silently.

Regression tests pipe Kimi-vocabulary payloads at each guard and assert
it engages (typed fields: exit status, decision, hookSpecificOutput) —
verified red against the pre-fix scripts, green after.

* fix(#2304): normalize Kimi tool_input fields and route block reasons to stderr

Cross-AI review of the initial fix, verified against kimi-cli source,
found the tool_name normalization alone leaves the guards dormant on a
real Kimi runtime: kimi-cli forwards tool_input verbatim
(src/kimi_cli/hooks/events.py), and its tool schemas
(src/kimi_cli/tools/file/{write,replace}.py) use path/content and
edit.old/edit.new (single Edit or list) — not Claude's
file_path/old_string/new_string. The guards read file_path, got '',
and exited 0 past the now-open tool_name gate.

Extend the per-guard normalization to the payload fields
(path -> file_path, edit -> old_string/new_string with list flattening),
and write the worktree guard's block reason to stderr as well as the
stdout JSON — Kimi feeds stderr, not stdout, back to the model on
exit 2 (docs/en/customization/hooks.md exit-code table).

Regression tests rewritten to Kimi's actual payload shapes (plus an
edit-list case and a stderr-reason assertion) — verified red against
the name-only fix, green after.

* fix(#2304): join all edit[] entries into old_string, matching new_string

Review nit on #2326: old_string took only edits[0].old while new_string
joined the whole list. Symmetric join removes the latent trap for any
future consumer sizing before/after content (e.g. the #2255 write guard).

* fix(#2304): normalize Kimi ReadFile vocabulary in read-injection scanner

Review Major 2 on #2326: gsd-read-injection-scanner.js had the identical
dormancy — its Kimi matcher fires on 'ReadFile' but the SCANNED_TOOLS
check only knew 'Read', so injected content in read files was never
flagged on Kimi installs.

Folds the same inlined normalization block into the scanner and extends
the shared KIMI_TOOL_NAMES map with ReadFile:'Read' in all four copies so
they stay byte-identical. Harmless in the three write guards: a
normalized 'Read' falls out of their Write/Edit allowlist exactly as the
unmapped name did. Field mapping verified against kimi-cli upstream
(src/kimi_cli/tools/file/read.py Params.path); the existing
path->file_path copy covers the scanner's file_path read.

* test(#2304): parity test binding the four inlined Kimi normalization copies

Review Major 1 on #2326: KIMI_TOOL_NAMES + normalizeKimiPayload is
deliberately inlined in four hook scripts (staging-dependency rationale,
unchanged), with the inverse table in bin/install.js — five
hand-maintained surfaces and nothing binding them.

Static binding, zero runtime coupling:
- the four inlined blocks must be byte-identical;
- each guard-map entry must be the value-inverse of
  convertKimiToolName() for its Claude name;
- every guard-relevant Claude tool (Write/Edit/MultiEdit/Read) must have
  a reverse entry — a vocabulary rename or extension that updates the
  installer without updating the guards now fails in CI instead of
  leaving a guard silently dormant (the #2304 recurrence door).

Negative-controlled: diverging one copy or dropping a map entry fails
the suite against the fixed code.

* test(#2304): regenerate golden parity fixtures for guard hook changes

CI red on #2326: all 10 golden-parity failures were the staged guard
hooks drifting from their fixtures. Regenerated with npm run gen:golden
(after npm run build) under throwaway HOME/CLAUDE_CONFIG_DIR; diff
verified to change exactly the four PR-touched guard entries per
surface, nothing else.

* test(#2304): regression tests for Kimi ReadFile engaging the scanner

Mirrors the per-guard Kimi vocabulary tests the PR added for the three
write guards: bare and module-qualified ReadFile produce the advisory,
path exclusions still apply post-normalization, unknown Kimi names stay
fail-open. Negative-controlled against the pre-fold scanner (the two
positive cases fail there; exclusion/fall-through correctly pass on
both sides).

* fix(#2304): normalize Kimi Shell vocabulary in workflow guard

Withdraws the disclosed out-of-scope split: verification showed the
Bash->Shell case needs NO different mapping — kimi-cli's Shell.Params
names its field `command` (src/kimi_cli/tools/shell/__init__.py), same
as Claude's Bash — and the guard's write branch (Write/Edit/MultiEdit
allowlist) was ALSO dormant on Kimi under its Shell|WriteFile|
StrReplaceFile matcher. Same defect class as the other four hooks.

Folds the identical inlined block into gsd-workflow-guard.js and
extends the shared map with Shell:'Bash' in all five copies (harmless
outside the workflow guard: a normalized Bash falls out of the other
guards' checks as before). Parity test now binds five copies and adds
Bash to the dormancy alarm. New workflow-guard test file exercises the
observable block (force-add on a worktree-agent branch): Shell bare and
module-qualified block with WORKTREE_AGENT_FORCE_ADD_FORBIDDEN, benign
Shell passes, Claude Bash unchanged — negative-controlled against the
pre-fold guard (the two Kimi cases fail there). Golden parity fixtures
regenerated; diff verified to change exactly the five guard entries per
surface.

* fix(#2304): map Kimi tool_output and route workflow-guard block to stderr

Third-party review (cross-AI verifier) caught two gaps in the revision:

1. Kimi PostToolUse events carry `tool_output`, not `tool_response`
   (kimi-cli src/kimi_cli/hooks/events.py post_tool_use()), so the
   read-injection scanner — which reads data.tool_response — was STILL
   dormant on real Kimi payloads; the earlier tests passed because they
   sent Claude-shaped payloads. The shared normalization block now maps
   tool_output -> tool_response (inert in PreToolUse guards, where the
   field is absent), and the scanner's Kimi tests send the real shape.

2. The workflow guard's force-add block wrote its reason to stdout only.
   Kimi's exit-2 protocol feeds stderr back to the model — the exact
   fix this PR already applied to the other blocking guard — so the
   newly-awakened block would have been a silent denial. Reason now
   also routed to stderr, asserted in the test.

Also: the scanner's "unknown name" test now uses a genuinely unmapped
name (FetchURL) — Shell stopped qualifying when it entered the map —
and the workflow guard's write branch (WriteFile advisory,
StrReplaceFile .planning pass) gains behavioral coverage. All five
copies stay byte-identical (parity test green); golden fixtures
regenerated, diff verified to the five guard entries per surface.
Negative-controlled: 3 new assertions fail against the pre-fix hooks.

* docs(#2304): update changeset to cover the full five-guard fix

Review round 2 (2026-07-18) flagged the changeset as stale: it was
written for the first commit and still described only the three guards
named in the issue. The shipped diff grew to five guards plus two
payload dimensions the original body never mentioned. The body now
names gsd-read-injection-scanner and gsd-workflow-guard, the ReadFile
and Shell vocabulary entries, the tool_output -> tool_response mapping,
and the workflow guard's stderr block-reason routing.

* test(#2304): regenerate kilo golden fixture after #2305 landed on next

The branch's fixture sweep predates 50efae13 (fix(#2305), PR #2327),
which made Kilo ship the five shared guard hooks. Rebased onto next and
re-ran the full generator sweep (gen:golden, size:baseline, and the
four registry/contract generators); the only delta across all of them
is kilo.json's five guard-hook hashes, matching this PR's hook edits.

* fix(#2304): fold Kimi normalization into the two shell hooks

The 2026-07-19 review found the last two guards with the #2304 dormancy:

- hooks/gsd-graphify-update.sh gated on tool_name == "Bash" but is
  registered on Kimi with matcher 'Shell' — Gate 1 never matched and the
  auto-rebuild was silently dormant. kimi-cli's Shell.Params names its
  field `command` (src/kimi_cli/tools/shell/__init__.py), same as Claude
  Bash, so only the name needs mapping: strip the module-path prefix,
  map Shell -> Bash.
- hooks/gsd-phase-boundary.sh read only tool_input.file_path, but Kimi's
  file tools name the field `path` (src/kimi_cli/tools/file/write.py +
  replace.py) — the hook read '' and .planning/ writes went undetected.
  Falls back to tool_input.path when file_path is absent, mirroring
  normalizeKimiPayload's precedence in the JS guards.

The normalization is reimplemented in shell — a byte-identity assertion
cannot span the JS<->shell boundary, so the parity test gains a
shell-guard vocabulary block that pins both scripts' mapping facts to
convertKimiToolName's live vocabulary instead of faking a byte binding.
Behavior is covered by negative-controlled tests beside each hook's
existing suite (verified red against the pre-fix scripts): Kimi Shell
dispatch (bare + module-qualified) with a WriteFile negative control in
graphify-auto-update.slow.test.cjs, and Kimi path detection, file_path
precedence, and a non-.planning negative control in hooks-opt-in.test.cjs.

Changeset updated to name all seven guards; golden install-parity
fixtures regenerated (diff is exactly the two hook entries per runtime;
size baselines unchanged).

* fix(#2304): use a Map for KIMI_TOOL_NAMES so prototype keys cannot pass the guard fall-through

A bare bracket lookup on an object literal resolves 'constructor',
'__proto__', 'toString', 'valueOf' and 'hasOwnProperty' through
Object.prototype to truthy functions/objects, so `if (!mapped)` failed
to short-circuit and data.tool_name was assigned a non-string. Map.get
returns undefined for those keys — the same shape the repo already uses
in canonicalizeRuntimeName (src/runtime-name-policy.cts). Applied
identically to all five inlined copies (review M1, PR #2326).

No new bypass class: unrecognized strings already fail open by design;
this fixes the lookup being wrong, not the posture.

* test(#2304): enumerate normalized guards by scanning hooks/, not a hardcoded list

The parity test's file list was a literal five-entry array — a sixth guard
with its own copy-pasted normalization block would be silently uncovered,
the exact divergence mode the test exists to prevent (review M2). Now the
list is a scan of hooks/*.js for the KIMI_TOOL_NAMES marker, with a floor
assertion so a scan that finds nothing fails instead of passing vacuously.
Also parses the Map declaration introduced by the M1 fix, and carries the
allow-test-rule annotation documenting the source-text scanning (review m4).

* test(#2304): parse hook JSON output instead of substring-matching raw stdout

workflow-guard.test.cjs asserted on unparsed stdout while read-guard.test.cjs
in the same PR parses the JSON envelope first — match the better pattern at
all four assertion sites (review m5).

* test(#2304): regenerate golden parity fixtures after Map conversion in the five guards

* docs(#2304): reset changeset pr:0 placeholder for Phase 0 PR (#2507)

The closed PR #2326's changeset carried pr:2326. Phase 0 of epic #2505
re-lands this fix on a fresh branch; the pr: field will be backfilled
to the real Phase 0 PR number immediately after gh pr create returns.

* docs(changeset): backfill PR #2518 for Phase 0 (#2507)

---------

Co-authored-by: 0xdhx <darkhawkx@gmail.com>
2026-07-21 23:42:02 -04:00
Tom Boucher
a28dcec981 chore(#597): replace count-based ratchet guards with AST lint + named-set allowlists (#603)
The windows-test-parity ratchet greps test source for fs.rmSync-without-
maxRetries (and six other Windows-portability anti-patterns), failing when an
integer offender COUNT exceeds a frozen baseline (rmSync: 95). A count ratchet
is a Goodhart metric: fixing one offender and adding another keeps the count
constant, so a new defect slips through green. Replace it — and every other
count ratchet in the repo — with a layered, masking-proof design.

Behavioral seam test
- tests/helpers-cleanup.test.cjs proves helpers.cleanup() carries the Windows
  EBUSY retry budget. cleanup() delegates retries to Node's fs.rmSync via
  maxRetries (it owns no loop), so the test asserts the option contract
  (recursive/force/maxRetries>0/retryDelay>0) + real-FS removal + the cwd-guard,
  rather than a loop that does not exist. The EBUSY risk is now tested ONCE at
  the helper, not approximated textually at every call site.

Write-time ESLint rule (AST-accurate, replaces the grep)
- eslint-rules/no-raw-rmsync-in-tests.cjs (error in tests/**/*.test.cjs) bans
  raw fs.rmSync, steering to cleanup(). Catches member, computed (fs['rmSync']),
  destructured and aliased forms; escape hatch is inline
  `// eslint-disable-next-line local/no-raw-rmsync-in-tests -- <reason>` only.
- Migrated 336 raw fs.rmSync teardown calls across ~116 test files to cleanup().
  ~18 genuinely load-bearing sites (mid-test SUT/fault-injection removals,
  error-swallowing or name-colliding local teardown helpers) keep the raw call
  with an inline eslint-disable + reason.

Shared anti-ratchet primitive
- scripts/lib/allowlist-ratchet.cjs:
  - assertWithinAllowlist: fails on NOVEL ids (new offender introduced) AND on
    STALE ids (a known offender was fixed but not pruned) — identity, not count,
    and a ratchet DOWN toward zero.
  - assertTightCeiling: a size/length budget whose ceiling must stay within a
    grace band of the high-water mark, so budgets may only tighten, never creep.

Ratchets converted onto the primitive
- windows-test-parity-guard.test.cjs: rmSync rule deleted (now ESLint-enforced);
  the remaining six patterns moved from integer baselines to named-set
  allowlists with ratchet-down.
- scripts/lint-test-file-count.{cjs,allowlist.json}: per-module integer counts →
  named filename sets (closes the swap-a-file-keep-the-count blind spot); a
  module dropping under cap now FAILS to force pruning its allowlist entry.
- enh-2790 skill-count `<= 63` → named skill allowlist (ratchets toward ~58).

Size budgets hardened (tighten-only)
- agent-size / workflow-size / feat-3039 help-tiered: ceilings lowered to the
  current high-water mark and an assertTightCeiling anti-creep check added per
  tier. Fixed external-contract limits (description ≤100 chars, agent ≤100 KB)
  are intentionally left as-is — they are not grandfathered creeping budgets.

No user-facing behavior change (tests + tooling only); no USER_FACING_PREFIXES
touched, so no changeset fragment is required.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 22:43:49 -04:00
Tom Boucher
b8c33647d8 refactor(tests): retire output-grep & source-grep via typed surfaces (finish #2974) (#462)
* refactor(#455): implement typed surfaces to retire grep tests

Production surfaces added:
- hooks/managed-hooks-registry.cjs: new CJS module exporting MANAGED_HOOKS
  as a typed array; gsd-check-update-worker.js now requires it instead of
  declaring an inline array
- bin/install.js: elevate inline gsdHooks to module-level GSD_UNINSTALL_HOOKS,
  export it alongside runtimeMap/allRuntimes (already exported)
- scripts/build-hooks.js: export HOOKS_TO_COPY; guard build() behind
  require.main===module so tests can require the file without triggering a build
- get-shit-done/bin/lib/init.cjs: add --json mode to agent-skills command,
  emitting typed IR { agent_type, block, skills_count } for test assertions
- get-shit-done/bin/gsd-tools.cjs: wire --json flag for agent-skills dispatch

Category-B source-grep migrations:
- tests/managed-hooks.test.cjs: require MANAGED_HOOKS from registry, drop fs.readFileSync+regex
- tests/orphaned-hooks.test.cjs: require MANAGED_HOOKS+HOOKS_TO_COPY as typed exports
- tests/hooks-opt-in.test.cjs: replace gsdHooks regex-parse with GSD_UNINSTALL_HOOKS import
- tests/install-minimal-hooks.test.cjs: replace gsdHooks regex-parse with GSD_UNINSTALL_HOOKS
- tests/copilot-install.test.cjs: replace src.includes() checks with typed
  assertions on runtimeMap, allRuntimes, parseRuntimeInput, buildRuntimePromptText
- tests/agent-skills.test.cjs: migrate to --json typed IR assertions

pending-migration-to-typed-ir token cleared (87 of 87 files):
- 78 files already had source-text-is-the-product; removed duplicate token
- 5 files already used typed assertions; reclassified or annotated
- 4 files required individual reclassification to source-text-is-the-product
  or architectural-invariant

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#455): update workflow-guard test to typed GSD_UNINSTALL_HOOKS import; isolate HOME in runtime-launcher (D) test

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#455): guard install.js main() behind require.main===module so the typed export is require-safe

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#455): document --json typed surfaces for agent-skills, progress, validate context

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#455): add changeset fragment for new --json surfaces

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#455): complete grep migration for files flagged by lint-tests

The branch commit 4e630d99 stripped `allow-test-rule: pending-migration-to-typed-ir`
from ~80 test files without replacing their assertions or adding the correct
exemption annotation. The files were NOT source-grep tests — they read .md
workflow/agent/command/reference files (source-text-is-the-product) or hook
source files for structural invariants (structural-regression-guard). No
assertion logic was changed; only the correct allow-test-rule annotation was
added to each file per CONTRIBUTING.md exception matrix.

73 files: `source-text-is-the-product` — workflow/agent/command/reference .md
7 files:  `structural-regression-guard` — hook .js / bin/install.js structural checks

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 11:39:52 -04:00
Tom Boucher
f55069ecbf test(#2974): migrate 8 test files to typed-IR assertions (#3016)
* test(#2974): migrate 8 test files to typed-IR assertions

Replaces raw stdout/stderr substring matching with structured-field
assertions per CONTRIBUTING.md "Prohibited: Raw Text Matching on Test
Outputs". Adds shared infrastructure for typed error emission so this
pattern is the easy path going forward.

Shared infrastructure:
- core.cjs: ERROR_REASON frozen enum + setJsonErrorMode/getJsonErrorMode
- gsd-tools.cjs: --json-errors CLI flag, parsed before subcommand dispatch
- config.cjs: typed reasons at all 7 error sites
- graphify.cjs: GRAPHIFY_REASON enum + reason/timeout_ms in execGraphify result
- bin/install.js: pure buildSdkFailFastReport() IR builder + renderer
- hooks/gsd-session-state.sh, gsd-phase-boundary.sh: emit Claude Code
  hookSpecificOutput JSON envelope with typed state_present/config_mode/
  planning_modified/file_path fields (no-op when hooks.community is off)

Test migrations (all pass, 171 tests across the 8 files):
- bug-2649-sdk-fail-fast: assert on ir.reason / ir.context / ir.fix_command
- bug-2687-config-read-warning-parity: assert.equal stderr === ''
- bug-2796-arg-parsing-regression: assert on result.json.updated/.phase
- bug-2838-summary-rescue: parse rescue footer, assert mtime invariant
- bug-2943-config-get-context-window: parse JSON, assert ERROR_REASON.CONFIG_KEY_NOT_FOUND
- graphify: assert reason === GRAPHIFY_REASON.ENOENT/TIMEOUT
- hooks-opt-in: parse hookSpecificOutput, assert typed fields
- security-scan: reclassified as source-text-is-the-product (scan label
  output and CI workflow YAML ARE the deployed contract)

Verification: lint-no-source-grep clean (0 violations), full suite
6741/6741 pass.

Closes #2974

* test(#2974): address CR feedback — typed code field, robust idempotency

Two CodeRabbit findings on #3016 addressed:

1. tests/hooks-opt-in.test.cjs:355 (Minor, inline) —
   parsed.reason.includes('Conventional Commits') was still substring
   matching after the typed-IR migration. Fixed at the source: the
   gsd-validate-commit hook now emits a typed `code` field
   ('CONVENTIONAL_COMMITS_VIOLATION', 'COMMIT_SUBJECT_TOO_LONG')
   alongside the human-readable `reason`. Test asserts strictEqual
   on the code; the prose copy is no longer part of the test contract.

2. tests/bug-2838-summary-rescue-gitignored-planning.test.cjs:224-250
   (Outside-diff) — mtimeMs alone can stay unchanged on coarse-grained
   filesystems (HFS+, FAT) when two rewrites land within the same
   timestamp tick, falsely passing the idempotency assertion.
   Replaced with a full snapshot (mtimeMs, ctimeMs, size, ino, sha256
   of contents) compared via assert.deepStrictEqual — the hash
   catches any rewrite the timestamp would miss.

Verification: 30/30 pass on the two affected files; lint-no-source-grep
clean (0 violations across 368 test files).
2026-05-02 09:27:23 -04:00
Tom Boucher
ef43f5161f fix(#2969): deterministic Step 5 verification gate for /gsd-reapply-patches (#2972)
* fix(#2969): deterministic Step 5 verification gate for /gsd-reapply-patches

The prior Step 5 "Hunk Verification Gate" was prescribed correctly in the
workflow text — but executed laxly by the LLM, which filled in `verified: yes`
without actually checking content presence. The reporter observed three
distinct files (skills/gsd-discuss-phase/SKILL.md, skills/gsd-autonomous/
SKILL.md, get-shit-done/workflows/new-project.md) where archives contained
substantive user-added blocks that did not survive into the merged result, yet
the gate reported clean.

Move verification from LLM-driven prose into a deterministic Node script the
workflow calls. The script can't be shortcut.

Changes:

- scripts/verify-reapply-patches.cjs (new): pure Node, no external deps.
  For each file in the patches dir, computes user-added significant lines as
  the line-set diff between backup and pristine baseline (when available;
  falls back to "every significant backup line" when no pristine — over-broad
  but the safe direction for this bug class). Asserts each line appears
  literally in the merged installed file via String.prototype.includes.
  Filters trivial lines (length < 12 chars, pure punctuation, decorative
  comments) so harmless drift doesn't trigger false failures. Exits 0 on
  pass, 1 on any miss with per-file diagnostic, 2 on usage error.
  Supports --json for workflow consumption.

- get-shit-done/workflows/reapply-patches.md: rewrite Step 5 to call the
  script and parse its JSON output. The Step 4 Hunk Verification Table
  remains as advisory Claude-readable summary, but the gate is now the
  script's exit code.

- tests/bug-2969-verify-reapply-patches.test.cjs (new): 6 tests covering
  (a) pass when every line survives, (b) fail when a line is missing,
  (c) fail when the merged file is deleted entirely, (d) --json structured
  report shape, (e) backup-meta.json is correctly skipped as metadata,
  (f) no-pristine-dir fallback exercises the safe over-broad path. All pass.

Out of scope: the manifest-baseline tightening described in #2969 Failure 1
(saveLocalPatches comparing against the wrong baseline so prior silent wipes
poison subsequent updates). That's a separate, bigger architectural change
involving pristine-content infrastructure; this PR addresses the gate fidelity
half so users at least see the diagnostic when content goes missing.

Closes #2969 (partial — Failure 2 only)

* fix(#2969): preserve #1999 Hunk Verification Table assertions alongside new script gate

CI failure on PR #2972 surfaced that tests/reapply-patches.test.cjs (the
#1999 contract) asserts Step 5 references:
  - "Hunk Verification Table"
  - `verified: no` failure condition
  - explicit STOP/halt/abort directive
  - "table absent / missing" halt path

My initial Step 5 rewrite for #2969 substituted the deterministic script
for the table-based gate entirely, stripping those references. The script
is the strictly stronger gate, but the existing #1999 test enforces the
table-based safety net as a defense-in-depth contract.

Restore both gates as a layered Step 5:

  - 5a (binding): deterministic verifier script — script gate, exits
    non-zero on any miss, cannot be shortcut by the LLM
  - 5b (advisory): Hunk Verification Table review — preserved as
    redundant safety net for the case where the script has a bug or the
    pristine baseline is unavailable

Both gates must pass. Verified: tests/reapply-patches.test.cjs (5 tests
in the #1999 suite) and tests/bug-2969-verify-reapply-patches.test.cjs
(6 tests in the #2969 suite) all pass — 21/21 total in this fixture.

* fix(#2969): address CodeRabbit findings on workflow + script

Five CR findings on PR #2972, all valid; addressed in this commit:

1. (Major) Stderr was merged into VERIFY_OUTPUT via `2>&1`, so any Node
   warning, deprecation notice, or stack trace would corrupt the JSON
   parse downstream. Capture stdout only; stderr remains on the
   controlling terminal for operator visibility.

2. (Major) verifyFile() crashed with EISDIR/EACCES instead of producing
   a structured diagnostic when the installed path was a directory or
   unreadable. Wrap statSync/readFileSync in try/catch and emit a
   per-file fail row; the whole-run gate continues with structured
   output. Added test case asserting the directory-at-installed-path
   case fails with `not a regular file` diagnostic instead of crashing.

3. (Minor) PRISTINE_FLAG built as a single string + unquoted expansion
   would split paths with spaces. Switched to a bash array (VERIFY_ARGS)
   that preserves whitespace through expansion.

4. (Minor) Fenced code block missing language tag (markdownlint MD040).
   Added `text` tag to the error message block.

5. (Minor) Usage comment said pristine fallback was "backup-meta lookup"
   but the actual code path falls back to significant-line checks from
   backup content. Corrected the comment to match implementation.

Verified all 21 tests in tests/reapply-patches.test.cjs (#1999 contract)
+ tests/bug-2969-verify-reapply-patches.test.cjs (now 7 tests with the
new directory case) pass.

* test(#2969): structured JSON assertions, no substring matching on script output

Replace every assert.match(r.stdout, /pattern/) call with structured
assertions on the parsed JSON report from the script's own --json mode.
The script's --json contract IS the structured shape we test against —
the test author should never depend on the human-readable formatter
output, just as no test should depend on substring presence in source.

Changes:

  - All 7 tests now run the verifier with --json (via a runVerifier()
    helper) and parse the resulting JSON document into { status, report,
    stderr }. Diagnostic stderr is preserved as a separate channel for
    debug output but is not used for assertions.
  - Each previously substring-matched diagnostic ("Failures: 1",
    "not a regular file", "installed file missing after merge",
    file path, dropped line) is now a deepEqual / equal / Array.includes
    against typed report fields: report.failures, report.results[i].status,
    report.results[i].reason, report.results[i].file,
    report.results[i].missing[].
  - Added an explicit "documented shape" test asserting the JSON output
    has exactly the keys { file, missing, reason, status } per result —
    locks the public contract of the --json mode.
  - DRY'd up fixture reset into a resetFixture() helper since every test
    starts with a fresh patches/installed/pristine triple.

Linter: scripts/lint-no-source-grep.cjs reports 0 violations across 348
test files. Combined run of bug-2969-...test.cjs (7 tests) +
reapply-patches.test.cjs (5 tests in the #1999 suite) all pass —
22/22 in the relevant fixture.

* fix(#2969): typed REASON enum + raw-text-matching rule shipped repo-wide

This commit closes the loop on the no-source-grep discipline:

1. scripts/verify-reapply-patches.cjs:
   - Frozen REASON enum exposes the diagnostic surface as stable codes:
     OK_NO_USER_LINES_VS_PRISTINE, OK_NO_SIGNIFICANT_BACKUP_LINES,
     FAIL_INSTALLED_MISSING, FAIL_INSTALLED_NOT_REGULAR_FILE,
     FAIL_READ_ERROR, FAIL_USER_LINES_MISSING.
   - Each result.reason is now a code from this enum, not free text.
     Tests assert via REASON.X equality, not regex on prose.
   - REASON exported from module.exports.

2. tests/bug-2969-verify-reapply-patches.test.cjs:
   - Full rewrite. Every assertion on typed structured fields:
     report.results[0].status === 'fail',
     report.results[0].reason === REASON.FAIL_INSTALLED_NOT_REGULAR_FILE,
     report.results[0].missing.includes(droppedLine) (Array set membership,
     not String substring).
   - Locks the REASON enum surface via Object.keys(REASON).sort() deepEqual.
   - Locks the JSON report shape via Object.keys(report).sort() deepEqual.
   - Zero regex, zero String#includes, zero startsWith/endsWith on text.

3. CONTRIBUTING.md:
   - New section "Prohibited: Raw Text Matching on Test Outputs" with
     concrete BAD/GOOD examples (substring on file content; assert.match
     on stdout; "structured parser" hiding string ops; regex on free-form
     reason fields).
   - The rule statement: "Tests assert on typed structured values. If
     the code under test produces text, the code under test must also
     expose a structured intermediate representation, and the test must
     assert on that IR — never on the rendered text."
   - Required structured-surface table: file IR, --json mode, frozen
     enum, fs facts.
   - "Hiding grep behind a function is still grep" callout — the
     parser-wrapper anti-pattern.
   - New `pre-existing-text-matching` exemption category for the 8
     grandfathered files. Marked Transitional; new tests cannot use it.

4. scripts/lint-no-source-grep.cjs:
   - Three new patterns enforced (in addition to the existing .cjs-source
     readFileSync rule):
     - assert.match/doesNotMatch on .stdout/.stderr
     - .stdout/.stderr.<includes|startsWith|endsWith>(
     - readFileSync(...).<includes|startsWith|endsWith>(
   - Aggregated violations per file (multiple findings now report together).
   - Updated diagnostic message references both CONTRIBUTING.md sections.

5. 8 pre-existing tests annotated with `// allow-test-rule:
   pre-existing-text-matching` so the lint passes on this commit; each
   carries the prose "Tracked for migration to typed-IR assertions; do
   not copy this pattern." Files: bug-2649, bug-2687, bug-2796, bug-2838,
   bug-2943, graphify, hooks-opt-in, security-scan.

Verification: lint 0 violations across 348 test files; full suite passes.

* fix(#2969): rename exemption category to pending-migration-to-typed-ir + cite tracking issue

Per maintainer feedback:
1. "Grandfathered" / "legacy" framing is wrong — both terms imply
   permanent or condoned exemption. The 8 files are tracked for
   correction, not exempted.
2. Each annotated file must cite the tracking issue so the migration
   work is auditable.

Changes:
- CONTRIBUTING.md: rename exemption category from
  `pre-existing-text-matching` to `pending-migration-to-typed-ir`. Update
  prose to "Tracked for correction, not exempted" and require each
  annotation to cite the open migration issue (e.g.
  `// allow-test-rule: pending-migration-to-typed-ir [#NNNN]`).
- 8 test files: update annotation to cite #2974 (the tracking issue
  opened for migrating these files to typed-IR assertions).
2026-05-01 16:14:39 -04:00
Tom Boucher
2703422be8 refactor(tests): standardize to node:assert/strict and t.after() per CONTRIBUTING.md (#1675)
* refactor(tests): standardize to node:assert/strict and t.after() per CONTRIBUTING.md

- Replace require('node:assert') with require('node:assert/strict') across
  all 73 test files to enforce strict equality (no type coercion)
- Replace try/finally cleanup blocks with t.after() hooks in core.test.cjs
  and hooks-opt-in.test.cjs per the test lifecycle standards
- Utility functions in codex-config and security-scan retain try/finally
  as that is appropriate for per-function resource guards, not lifecycle hooks

Closes #1674

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(tests): add --test-concurrency=4 to test runner for parallel file execution

Node.js --test-concurrency controls how many test files run as parallel child
processes. Set to 4 by default, configurable via TEST_CONCURRENCY env var.
Fixes tests at a known level rather than inheriting os.availableParallelism()
which varies across CI environments.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(security): allowlist verify.test.cjs in prompt-injection scanner

tests/verify.test.cjs uses <human>...</human> as GSD phase task-type
XML (meaning "a human should verify this step"), which matches the
scanner's fake-message-boundary pattern for LLM APIs. This is a
false positive — add it to the allowlist alongside the other test files
that legitimately contain injection-adjacent patterns.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-04 14:29:03 -04:00
Tom Boucher
ca6a273685 fix: remove marketing text from runtime prompt, fix #1656 and #1657 (#1672)
* chore: ignore .worktrees directory

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(install): remove marketing taglines from runtime selection prompt

Closes #1654

The runtime selection menu had promotional copy appended to some
entries ("open source, the #1 AI coding platform on OpenRouter",
"open source, free models"). Replaced with just the name and path.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(kilo): update test to assert marketing tagline is removed

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(tests): use process.execPath so tests pass in shells without node on PATH

Three test patterns called bare `node` via shell, which fails in Claude Code
sessions where `node` is not on PATH:

- helpers.cjs string branch: execSync(`node ...`) → execFileSync(process.execPath)
  with a shell-style tokenizer that handles quoted args and inner-quote stripping
- hooks-opt-in.test.cjs: spawnSync('bash', ...) for hooks that call `node`
  internally → spawnHook() wrapper that injects process.execPath dir into PATH
- concurrency-safety.test.cjs: exec(`node ...`) for concurrent patch test
  → exec(`"${process.execPath}" ...`)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: resolve #1656 and #1657 — bash hooks missing from dist, SDK install prompt

#1656: Community bash hooks (gsd-session-state.sh, gsd-validate-commit.sh,
gsd-phase-boundary.sh) were never included in HOOKS_TO_COPY in build-hooks.js,
so hooks/dist/ never contained them and the installer could not copy them to
user machines. Fixed by adding the three .sh files to the copy array with
chmod +x preservation and skipping JS syntax validation for shell scripts.

#1657: promptSdk() called installSdk() which ran `npm install -g @gsd-build/sdk`
— a package that does not exist on npm, causing visible errors during interactive
installs. Removed promptSdk(), installSdk(), --sdk flag, and all call sites.

Regression tests in tests/bugs-1656-1657.test.cjs guard both fixes.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: sort runtime list alphabetically after Claude Code

- Claude Code stays pinned at position 1
- Remaining 10 runtimes sorted A-Z: Antigravity(2), Augment(3), Codex(4),
  Copilot(5), Cursor(6), Gemini(7), Kilo(8), OpenCode(9), Trae(10), Windsurf(11)
- Updated runtimeMap, allRuntimes, and prompt display in promptRuntime()
- Updated multi-runtime-select, kilo-install, copilot-install tests to match

Also fix #1656 regression test: run build-hooks.js in before() hook so
hooks/dist/ is populated on CI (directory is gitignored; build runs via
prepublishOnly before publish, not during npm ci).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-04 14:15:30 -04:00
Tibsfox
4157c7f20a feat(hooks): add opt-in community hooks for GSD projects
Port 3 community hooks from gsd-skill-creator, gated behind hooks.community config flag. All hooks are registered on install but are no-ops unless the project config has hooks: { community: true }.

gsd-session-state.sh (SessionStart): outputs STATE.md head for orientation. gsd-validate-commit.sh (PreToolUse/Bash): blocks non-Conventional-Commits messages. gsd-phase-boundary.sh (PostToolUse/Write|Edit): warns when .planning/ files are modified.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 04:23:44 -07:00