Commit Graph

3169 Commits

Author SHA1 Message Date
Tom Boucher
28cf6b444f test(#489): stop docs-parity tokenizer matching repo path open-gsd/gsd-core as a command (#491)
extractCommandTokens regexes matched the /gsd-core substring inside the
org/repo path open-gsd/gsd-core#22. Add a negative lookbehind so only
actual invocations (BOL / space / backtick / paren) match, not path or
word-embedded segments. Add a regression test covering the repo-path
false positive while proving real broken command refs are still caught.

Refs #489

Co-authored-by: CI Rebase Check <ci@gsd-redux>
2026-05-29 18:08:27 -04:00
Tom Boucher
b9ea06fa8b ci(#483): resolve transitive dependencies in affected-test selection + zero-dependent widen backstop (#485)
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-29 17:12:58 -04:00
Tom Boucher
48824f258c fix(#444): resolver preamble checks repo-local .claude install path (#476)
The gsd_run resolver preamble now probes
<repo-root>/.claude/get-shit-done/bin/${_GSD_SHIM_NAME} as the second
check — immediately after the existing get-shit-done/bin/ check and
before command -v / $HOME/.claude fallbacks. This covers the install
layout produced by npx @opengsd/get-shit-done-redux@latest --claude --local.

A _GSD_RUNTIME_ROOT variable is introduced to bind the repo-root
expression once and reuse it for both checks without repeating the
git rev-parse subshell.

76 workflow files regenerated via node scripts/sync-runtime-launcher.cjs.
All parity, size-budget, and new regression tests pass.

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 17:12:54 -04:00
Tom Boucher
d7dafafd03 fix(#442): --config-dir= no longer truncates paths containing equals signs (#475)
Extract a pure `parseConfigDirFromArgs(argsArray)` seam from the
closure-based `parseConfigDirArg()` and fix the equals-form parser to
use `slice(indexOf('=') + 1)` instead of `split('=')[1]`, so that
paths like `/tmp/gsd=a` or `/tmp/a=b=c` are preserved in full.

Both `--config-dir=<path>` and `-c=<path>` are fixed.  The pure seam
is exported via `module.exports` so the 12-case unit test can assert
on typed return values without spawning a child process.

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 17:12:50 -04:00
Tom Boucher
6687087627 fix(#481): remove residual dead silent-expiry poll helper + Atomics.wait afterEach retry in bug-1974 (racy caller already removed by #453) (#482)
The 45s detached-subprocess poll caller was removed by #453; this deletes the
residual dead waitForStateMatch helper (silent-expiry anti-pattern) and the
redundant Atomics.wait-based afterEach retry loop (cleanup() already retries via
fs.rmSync maxRetries:20). Net deletion; deterministic tests untouched.

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 17:12:45 -04:00
Tom Boucher
5f18379da5 fix(#478): delete wall-clock elapsed-time assertions per ADR 456 (keep correctness invariants) (#480)
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-29 17:12:41 -04:00
Tom Boucher
d2ff4ac092 docs(#22): add ADR for plan-vs-codebase drift guard (defaults + resolver seam) (#484)
Consolidated decision record: source-grounding verification default-on
(plan_review.source_grounding), intel.enabled stays opt-in, and the
three-valued symbol-resolver seam with a climbable adapter ladder.

Refs #22

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 17:12:37 -04:00
Tom Boucher
d3eaf6aec1 docs(#23): document changeset extract CLI contract (#479)
Add scripts/changeset/README.md specifying the cli.cjs extract
subcommand: invocation, flags, version validation, exit-code table
(0/1/2), and output shapes for text and --json modes.

Corrects two inaccuracies from the triage table against the source:
v-prefixed versions ARE accepted (stripped), and it is pre-release/
build suffixes that are rejected — not the v prefix. Also documents
the full exit-1 surface (missing flags, invalid semver, missing
changelog) and the exit-2 overlap (empty range vs malformed argv) so
external callers do not conflate "no releases in range" with failure.

Closes #23

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 17:12:33 -04:00
Tom Boucher
e22596be04 fix(#471): make perf-407 lock-buffer-alloc test deterministic via clock-seam; remove real-worker race (#472)
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 13:20:32 -04:00
Tom Boucher
b735f270c5 chore(#469): clear residual warn-level lint in effort test files (#470)
Remove unused `os` import left by #463's effort-API conversion and fix
two no-useless-escape chars in codex-config test description string.
Part of ESLint harness cleanup effort (#452).

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 12:25:27 -04:00
Tom Boucher
c7e5a88353 enh(#466): refresh opus-tier model IDs to current GA (Opus 4.8 / codex gpt-5.5) (#467)
* enh: bump opus-tier model IDs to current GA (Opus 4.8 / codex gpt-5.5)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#466): changeset for opus-tier model-ID refresh

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-29 11:52:47 -04:00
Tom Boucher
b8c33647d8 refactor(tests): retire output-grep & source-grep via typed surfaces (finish #2974) (#462)
* refactor(#455): implement typed surfaces to retire grep tests

Production surfaces added:
- hooks/managed-hooks-registry.cjs: new CJS module exporting MANAGED_HOOKS
  as a typed array; gsd-check-update-worker.js now requires it instead of
  declaring an inline array
- bin/install.js: elevate inline gsdHooks to module-level GSD_UNINSTALL_HOOKS,
  export it alongside runtimeMap/allRuntimes (already exported)
- scripts/build-hooks.js: export HOOKS_TO_COPY; guard build() behind
  require.main===module so tests can require the file without triggering a build
- get-shit-done/bin/lib/init.cjs: add --json mode to agent-skills command,
  emitting typed IR { agent_type, block, skills_count } for test assertions
- get-shit-done/bin/gsd-tools.cjs: wire --json flag for agent-skills dispatch

Category-B source-grep migrations:
- tests/managed-hooks.test.cjs: require MANAGED_HOOKS from registry, drop fs.readFileSync+regex
- tests/orphaned-hooks.test.cjs: require MANAGED_HOOKS+HOOKS_TO_COPY as typed exports
- tests/hooks-opt-in.test.cjs: replace gsdHooks regex-parse with GSD_UNINSTALL_HOOKS import
- tests/install-minimal-hooks.test.cjs: replace gsdHooks regex-parse with GSD_UNINSTALL_HOOKS
- tests/copilot-install.test.cjs: replace src.includes() checks with typed
  assertions on runtimeMap, allRuntimes, parseRuntimeInput, buildRuntimePromptText
- tests/agent-skills.test.cjs: migrate to --json typed IR assertions

pending-migration-to-typed-ir token cleared (87 of 87 files):
- 78 files already had source-text-is-the-product; removed duplicate token
- 5 files already used typed assertions; reclassified or annotated
- 4 files required individual reclassification to source-text-is-the-product
  or architectural-invariant

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#455): update workflow-guard test to typed GSD_UNINSTALL_HOOKS import; isolate HOME in runtime-launcher (D) test

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#455): guard install.js main() behind require.main===module so the typed export is require-safe

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#455): document --json typed surfaces for agent-skills, progress, validate context

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#455): add changeset fragment for new --json surfaces

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#455): complete grep migration for files flagged by lint-tests

The branch commit 4e630d99 stripped `allow-test-rule: pending-migration-to-typed-ir`
from ~80 test files without replacing their assertions or adding the correct
exemption annotation. The files were NOT source-grep tests — they read .md
workflow/agent/command/reference files (source-text-is-the-product) or hook
source files for structural invariants (structural-regression-guard). No
assertion logic was changed; only the correct allow-test-rule annotation was
added to each file per CONTRIBUTING.md exception matrix.

73 files: `source-text-is-the-product` — workflow/agent/command/reference .md
7 files:  `structural-regression-guard` — hook .js / bin/install.js structural checks

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 11:39:52 -04:00
Tom Boucher
bf68ad4d93 fix(tests): deterministic concurrency via injectable clock seam; delete flaky racing tests (#459)
* feat(#453): add deterministic clock seam to lock modules

Introduces get-shit-done/bin/lib/clock.cjs exporting realClock with
now() (Date.now) and sleep() (Atomics.wait). acquireStateLock,
writeStateMd, and readModifyWriteStateMd in state.cjs each accept an
optional trailing clock param (default: realClock). withPlanningLock in
planning-workspace.cjs gains the same seam. No production behavior
change — all callers that omit the param continue to use realClock.

Adds tests/helpers/clock.cjs (makeFakeClock) and tests/clock-seam.test.cjs
with 20 deterministic in-process tests covering: lock serialization,
timeout throw at maxWaitMs boundary, stale-lock takeover, lock released
on error path, withPlanningLock timeout recovery, exit-cleanup integration,
readModifyWriteStateMd call-site coverage (7 cmd*), and roadmap analyze
behavioral assertion (50 phases, no elapsed-time gate).

Deletes/converts per research verdicts: removes 11 source-grep/elapsed-time/
non-deterministic-concurrent tests across concurrency-safety.test.cjs,
locking-bugs-1909-1916-1925-1927.test.cjs, and bug-1974-context-exhaustion-
record.test.cjs. All deleted tests have deterministic replacements in
clock-seam.test.cjs or surviving barrier-based tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#453): update module inventory for clock.cjs; make EEXIST-retry assertion behavioral

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#453): satisfy lint-tests — allow-test-rule annotation on readFileSync/includes runtime output check

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 11:18:10 -04:00
Tom Boucher
4b908e6bcd test(tests): antagonistic tier — fast-check property tests + Stryker mutation testing (PR-gated) (#461)
* feat(#454): add antagonistic tier — fast-check property tests + Stryker mutation config

Adds property-based testing (fast-check v4) and mutation testing scaffolding
(Stryker v8) as the antagonistic validation tier for lib/*.cjs pure logic.

## Files added

### Shared setup
- tests/helpers/fast-check-setup.cjs — configureGlobal({ numRuns:200, seed:42 })
  for deterministic CI; override locally with GSD_FC_SEED

### Property test suites (node:test + fast-check)
- tests/context-utilization.property.test.cjs — boundary at 60%/70% thresholds
  (exact Math.ceil boundary, not Math.floor), TypeError on all invalid inputs,
  overflow clamping to 100%/critical, shape invariants (7 tests, all pass)
- tests/prompt-budget.property.test.cjs — estimateTokens monotonicity + ceil(len/4)
  exactness; applyBudget shape invariant, instructions/roadmap verbatim, budget
  envelope, omit tracking (11 tests, all pass)
- tests/frontmatter.property.test.cjs — extractFrontmatter/reconstructFrontmatter/
  spliceFrontmatter never-throw + type shape + splice→extract round-trip (9 tests)
- tests/adr-parser.property.test.cjs — shouldRejectAdrStatus boundary (3 statuses
  only), parseAdrMarkdown shape + title trim invariant (discovered: parser trims
  trailing whitespace) (9 tests, all pass)
- tests/config-schema.property.test.cjs — isValidConfigKey never throws, returns
  boolean, accepts all VALID/RUNTIME_STATE_KEYS, rejects empty/null/unknown (8 tests)

### Stryker mutation config
- stryker.config.mjs — testRunner:'command', mutate bin/lib/**/*.cjs minus 13
  generated files, coverageAnalysis:'off', thresholds {high:80,low:60,break:50},
  incremental:true, reporters html+clear-text+progress

### CI workflow
- .github/workflows/mutation.yml — PR-gating job (pull_request + workflow_dispatch),
  runs stryker --incremental --since origin/next (changed files only), uploads
  HTML artifact; SINCE_REF via env not interpolation (injection-safe)

### Package config
- package.json: +test:mutation, +test:mutation:since scripts
- .gitignore: +.stryker-tmp/, +.stryker-incremental.json, +reports/mutation/
- package-lock.json: fast-check@4.8.0, @stryker-mutator/core@9.6.1

## Verified
node --test on all 5 property test files: 44 tests, 0 failures.
Stryker NOT run (slow; reserved for CI).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#454): pin upload-artifact to v7.0.1 SHA used across repo (bad SHA ea165f8d)

The SHA ea165f8d65b6e75b540449d3ec4f5dde0c5a4e1 (labeled v4.6.2) does not
resolve on GitHub Actions. All other workflows in this repo pin
043fb46d1a93c77aae656e7c1c64a875d1fc6a0a (v7.0.1) — align mutation.yml.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#454): mutation workflow — replace invalid --since flag with changed-core-files --mutate scoping

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 11:07:07 -04:00
Tom Boucher
33afb4f6eb chore(#452): add ESLint 9 flat-config harness with three custom AST rules (#460)
Install eslint@9 + typescript-eslint@8 + globals@16 + eslint-plugin-n@17 +
eslint-plugin-no-only-tests@3 + typescript as devDependencies.

eslint.config.mjs (flat config):
- Global ignores: node_modules, dist, .worktrees, .claude, coverage, the
  12 generated get-shit-done/bin/lib/*.cjs files
- Block for get-shit-done/bin/**/*.cjs + scripts/**/*.cjs: js.recommended +
  eslint-plugin-n + local plugin; generic quality rules (no-var, prefer-const,
  no-unused-vars, no-empty, n/no-process-exit)
- Block for tests/**/*.test.cjs: no-only-tests (error), local timing rules,
  no-restricted-syntax timing bans

eslint-rules/ local plugin (three AST rules, all at warn pending cleanup):
- no-source-grep: flag readFileSync on source .cjs/.js/.ts + text methods
- no-magic-sleep-in-tests: flag Atomics.wait and await-new-Promise(setTimeout)
- no-elapsed-assertion: flag assert*() on timing props (elapsed/duration/took/ms)

tsconfig.lint.json: allowJs + checkJs + noEmit for future type-aware passes.

tests/eslint-rules.test.cjs: 15 RuleTester unit tests (all pass, 0 fail).

package.json: add lint/lint:fix scripts; remove lint:tests (subsumed by ESLint
local/no-source-grep). Rules that produced pre-existing errors downgraded to
warn: no-useless-escape, no-unsafe-finally, no-regex-spaces, no-control-regex,
no-irregular-whitespace. ESLint exits 0 (warnings ok).

.github/workflows/test.yml lint-tests job: add npm ci + ESLint step; remove
"Lint — no source-grep tests" step (now covered by ESLint); bump timeout 3→5
min. .gitignore: add node_modules/.cache/eslint/ entry.

eslint --fix auto-cleaned: no-regex-spaces in tests, prefer-const in state.cjs,
redundant eslint-disable-next-line comments.

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 10:54:49 -04:00
Tom Boucher
ed1c20061b docs(#456): add testing standards, ADRs 452/456/457, and CONTEXT.md entries (#458)
- docs/adr/452-eslint-lint-harness.md (Accepted): adopt ESLint flat config
  with typescript-eslint, eslint-plugin-n, eslint-plugin-no-only-tests, and
  local AST-rule plugin; retire homegrown scripts/lint-*.cjs regex checkers;
  three test-rigor rules ship at warn, promoted to error after #453 cleanup
- docs/adr/456-test-rigor-architecture.md (Accepted): deterministic-over-racing
  via injectable clock seam + node:test mock.timers; antagonistic tier with
  fast-check + Stryker at 80% threshold; typed-surface mandate; delete-bad-tests
  policy with no-permanent-quarantine
- docs/adr/457-generated-cjs-single-source.md (Proposed): future direction to
  collapse ~59 hand-written bin/lib/*.cjs to TS-generated single source;
  eliminates tsconfig.lint.json stopgap; marked Proposed / not yet executed
- TESTING-STANDARDS.md: orients to existing docs; codifies six test-rigor
  contracts; adds new policies (no-timing-assertion, clock-seam, property-based,
  mutation-score, delete-bad-tests); pairs each with exact ESLint rule names;
  markdownlint-clean (MD040 fences, MD056 table columns)
- CONTEXT.md: adds six RULESET.TESTS.* predicates (no-timing-assertion,
  clock-seam, property-based-testing, mutation-score, delete-bad-tests,
  eslint-harness) and five glossary terms (clock seam, deterministic scheduler,
  property-based test, mutation testing/score, ESLint harness)
- docs/adr/README.md: adds index rows for ADRs 452, 456, 457

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 10:33:36 -04:00
Tom Boucher
5ca646f015 feat(#443): unified cross-provider effort controls + fast-mode-aware routing (#463)
* test(#443): RED unified effort + fast_mode + resolve-execution

All 68 tests failing as expected — no implementation yet.
Covers: effort cascade (tier defaults, overrides, invalid fallthrough),
fast_mode cascade (boolean-only, tier defaults), resolveEffortForTier
escalation, renderEffortForRuntime clamping, resolve-execution CLI,
config schema new keys, QA hostile-input matrix.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(#443): unified cross-provider effort + fast_mode knobs and resolve-execution query

Adds config-driven effort control (universal ladder: minimal<low<medium<high<xhigh<max)
and fast_mode propagation knobs, with per-runtime rendering that clamps the unique
tail values (max=Anthropic-only clamps to xhigh on Codex; minimal=Codex-only clamps
to low on Claude).

Key changes:
- config-schema.manifest.json: add effort.default, fast_mode.enabled as validKeys;
  add 4 dynamicKeyPatterns for effort.routing_tier_defaults, effort.agent_overrides,
  fast_mode.routing_tier_defaults, fast_mode.agent_overrides; fix stale _comment
- config-defaults.manifest.json: add effort and fast_mode blocks with tier defaults
- model-catalog.cjs: add EFFORT_RENDERING map, renderEffortForRuntime(), RUNTIMES_WITH_FAST_MODE
- model-profiles.cjs: re-export new catalog exports
- core.cjs: add resolveEffortInternal, resolveFastModeInternal, resolveEffortForTier,
  VALID_EFFORTS, EFFORT_SET, nextEffort; pass effort/fast_mode through loadConfig
- commands.cjs: replace reasoning_effort in cmdResolveModel with unified effort;
  add cmdResolveExecution (superset command with effort_rendered, effort_param,
  effort_propagation, fast_mode, fast_mode_supported)
- gsd-tools.cjs: add resolve-execution case with --effort/--fast-mode/--attempt flags
- tests/feat-443: 69 tests covering cascade, rendering, escalation, CLI, schema, QA matrix
- tests/commands.test.cjs: convert 3 reasoning_effort assertions to unified effort
- docs/CONFIGURATION.md: document effort + fast_mode + resolve-execution sections
- settings-advanced.md: list new effort/fast_mode keys in confirmation table

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#443): remove dead catalog effort lane; unify codex effort through renderEffortForRuntime

- Remove resolveReasoningEffortInternal (catalog-driven effort function) from
  core.cjs and its export; remove from commands.cjs destructure import
- Convert tests/issue-2517-runtime-aware-profiles.test.cjs: all 11 effort
  assertions now use resolveEffortInternal + renderEffortForRuntime; Claude
  effort is first-class (output_config.effort); unknown runtimes assert param===null
- Convert tests/feat-3023-model-phase-types.test.cjs: replace the entire
  resolveReasoningEffortInternal describe with unified effort assertions;
  effort derives from AGENT_DEFAULT_TIERS routing tier, not phase-type tier

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#443): ADR for unified cross-provider effort + fast-mode routing

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* test(#443): architecture-level QA invariants + test-strategy doc

Add 48-test integration suite (feat-443-effort-fast-mode.integration.test.cjs)
covering 8 architectural invariants: cross-provider validity (never emit a value
the real API would 400 on), param/channel contract stability, resolve-execution
JSON contract (all 8 keys + correct types), totality across the full 33-agent
registry, fast-mode honesty (claude always fast_mode_supported=false), precedence
first-valid-wins matrix for both effort and fast_mode cascades, dynamic-routing
composition (effort escalation independent of model tier), and config-set round-trip
for all new effort/* and fast_mode/* key namespaces. Append test-strategy section
with invariant rationale and E2E gap documentation to docs/TESTING-SUITES.md.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#443): add failing install-wiring tests for effort per-runtime injection (RED)

TDD RED: 10 failing tests covering:
- Claude .md gets effort: injected per tier (planner=xhigh, mapper=low, executor=high)
- Gemini .md does NOT get effort: (already passing — Gemini-safe)
- Codex .toml gets model_reasoning_effort via unified resolver
- Config-driven: effort.agent_overrides drives both Claude .md and Codex .toml
- Source purity: agents/*.md have no effort: key (already passing)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(#443): wire effort per-runtime at install (Claude .md frontmatter + Codex .toml unified)

- Import AGENT_DEFAULT_TIERS and renderEffortForRuntime from model-catalog.cjs
- Add readGsdEffectiveEffortConfig(targetDir): reads merged effort config from
  .planning/config.json (per-project wins) + ~/.gsd/defaults.json (global fallback),
  same probe pattern as readGsdRuntimeProfileResolver
- Add resolveInstallTimeEffort(effortCfg, agentName): pure function matching
  resolveEffortInternal() precedence (agent_overrides > routing_tier_defaults > default > 'high')
  without loadConfig side-effects (no sub-repo detection, no migration writes)
- Claude agent copy loop: inject `effort: <value>` into frontmatter ONLY for
  runtime === 'claude'; all other .md runtimes (Gemini, Qwen, Hermes, etc.) stay
  effort-free (Gemini-safe source contract preserved in agents/*.md)
- generateCodexAgentToml: add effortCfg param; emit model_reasoning_effort from
  unified resolver (replaces old catalog entry.reasoning_effort); Codex clamps
  max → xhigh via renderEffortForRuntime('codex', ...)
- installCodexConfig: pass readGsdEffectiveEffortConfig(targetDir) to
  generateCodexAgentToml so per-project config wins for Codex .toml too
- Update failing tests to GREEN: 12/12 pass; all 17 install tests pass;
  2847/2848 unit tests pass (1 pre-existing failure: policy-shell-pinning)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#443): source install effort defaults from manifest (kill drift) + guard test

Replace hardcoded _GSD_EFFORT_MANIFEST_TIER_DEFAULTS and the 'high' fallback in
resolveInstallTimeEffort with values read from config-defaults.manifest.json at
module init, using the same __dirname-relative path install.js already uses for
all shared manifests. Add feat-443-effort-defaults-drift.test.cjs to assert
equality between install.js's runtime constants and the manifest on every CI run.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): reconcile Codex TOML tests with unified effort design

The #443 unified effort resolver makes generateCodexAgentToml always emit
model_reasoning_effort (driven by resolveInstallTimeEffort, not model_profile_overrides).
The test 'generated TOML omits reasoning_effort when runtime has none' had an
obsolete premise — model_profile_overrides.reasoning_effort:'' no longer suppresses
unified effort. Convert it to assert the new invariant: Codex TOML always carries a
valid model_reasoning_effort from the agent's routing tier (xhigh for gsd-planner,
a heavy-tier agent), while model_profile_overrides model override is still respected.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): make install.js effort resolution lazy (no load-time side effects breaking launcher-parity)

Replace module-load-time IIFE + hard throw (config-defaults.manifest.json read)
and top-level require of model-catalog.cjs with a lazy _getGsdEffortCatalog()
getter that initialises on first call from resolveInstallTimeEffort /
generateCodexAgentToml / Claude .md effort injection.  Requiring install.js in
unrelated test contexts (e.g. runtime-launcher-parity) no longer triggers
manifest IO or throws, eliminating the load-time side effect that changed
subprocess exit codes / stderr on the bench.

Drift-guard exports (_GSD_EFFORT_MANIFEST_TIER_DEFAULTS / _GSD_EFFORT_MANIFEST_DEFAULT)
preserved as lazy getter properties on module.exports so feat-443-effort-defaults-drift
still validates them without forcing eager load.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): isolate install-wiring test HOME to stop \$HOME/.claude pollution breaking launcher-parity

runGlobalInstall() now redirects HOME to a per-call isolated tmpdir in addition
to the existing runtime-specific env-var redirects (CLAUDE_CONFIG_DIR,
GEMINI_CONFIG_DIR, CODEX_HOME). This ensures install.js code that uses
os.homedir() directly — including the ~/.cache/gsd update-check deletion,
~/.gsd/defaults.json reads, and any HOME-relative npm subprocess writes —
never touches the real \$HOME during the test.

Without the HOME isolation the install test (which is new to this branch and
is now picked up by Docker's raw \`tests/*.test.cjs\` glob) could write or
delete files under the real \$HOME, causing runtime-launcher-parity test (D)
to fail: (D) asserts a loud non-zero exit when \$RUNTIME_DIR/gsd-tools.cjs is
absent and gsd-tools is not on PATH, but the launcher's \$HOME/.claude fallback
arm succeeds if \$HOME/.claude/get-shit-done/bin/gsd-tools.cjs exists.

Also sets GSD_SKIP_STALE_SDK_CHECK=1 to suppress the \`npm ls -g\` subprocess
that the global installer spawns — irrelevant to effort-wiring assertions,
slow, and potentially writes to ~/.npm cache.

All 12 feat-443 install-wiring assertions preserved. Drift-guard 5/5. Unit
suite 2848/2850 (pre-existing policy-shell-pinning.test.cjs failure on next).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#443): add changeset fragment for effort + fast-mode routing

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(#443): set GSD_TEST_MODE before requiring install.js in drift-guard test to prevent HOME leak

Without GSD_TEST_MODE=1, require('bin/install.js') runs the module's main
install block (guarded by !GSD_TEST_MODE), performing a real global Claude
install into $HOME/.claude/. On CI ubuntu where node is on standard PATH,
the launcher's $HOME/.claude fallback arm then finds gsd-tools.cjs, causing
runtime-launcher-parity test (D) to exit zero when it must exit non-zero.

Root cause: feat-443-effort-defaults-drift.test.cjs (unit suite) runs
alphabetically before runtime-launcher-parity.test.cjs in the same node
--test invocation. Each runs in a separate worker process but shares the
same HOME. The drift test's install leaks gsd-tools.cjs into that HOME,
then the launcher test's bash subprocess finds it via the $HOME/.claude arm.

Fix: add process.env.GSD_TEST_MODE = '1' at the top of the drift-guard
test, before the require(installPath) call. This matches the pattern used
by feat-443-effort-fast-mode.test.cjs and feat-443-effort-install-wiring
.install.test.cjs.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): deterministic resolve-execution arg parsing + validate install-time effort (Codex adversarial findings)

Finding 1: resolve-execution --effort low gsd-planner misrouted 'low' as the agent.
Replace find(non-dash) with a proper flag-consuming loop that collects a single
positional; validate missing/extra positionals and malformed --attempt values.

Finding 2: resolveInstallTimeEffort returned unvalidated effort strings (e.g. "ultra")
verbatim. Each precedence layer now checks GSD_EFFORT_SET (imported once from
core.cjs) before accepting a value, mirroring resolveEffortInternal exactly.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): newline-agnostic effort frontmatter injection (Windows CRLF) + CRLF-safe assertions

Extracts injectEffortFrontmatter(content, effortValue) pure helper that detects
EOL (LF vs CRLF) from the opening '---' line and inserts 'effort: <value>'
before the closing '---' delimiter using the same EOL as the surrounding
frontmatter. Regex now uses /^---\r?\n([\s\S]*?)^---\r?$/m instead of the
LF-only /^(---\n[\s\S]*?)(---)(\n|$)/ that silently skipped CRLF files on
Windows (git core.autocrlf=true checkout).

Also adds 7 unit tests covering LF, CRLF, idempotency, no-frontmatter, and
complex frontmatter cases. Exports injectEffortFrontmatter from module.exports.

Fixes 6 CI failures in tests/feat-443-effort-install-wiring.install.test.cjs
on windows-latest runners (lines 138, 145, 152, 261, 345, 356).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-29 10:32:58 -04:00
Tom Boucher
199251c0f6 test(#432): make perf-316/perf-407 lock-race tests deterministic (#450)
The perf-316 and perf-407 regression tests had two race-driven failure modes:

1. False-fail: setTimeout(80) before assert.ok(fs.existsSync(lockPath)) could
   fire before Worker A finished writing the lock file under CI load. The
   handoff noted this fired on origin/next's coverage job.

2. False-pass (latent): if Worker B raced past Worker A's release before
   retrying, sabCount === 1 from the no-retry success path matched what
   the PRE-FIX buggy code produced (one SAB for the single successful
   open/write). The existing test had no witness for retry-path coverage,
   violating test-rigor Contract 4 (exercise the path you claim to cover).

Fix (test-only):

- Replace setTimeout(80) with await-{pid}-message synchronization. Worker A
  posts {pid} AFTER fs.writeFileSync/openSync returns (single-thread source
  order within the worker), so by the time the parent receives it the lock
  file exists on disk. The MessagePort buffers messages posted before the
  parent attaches its listener, so there is no listener-race.
  Ref: https://nodejs.org/api/worker_threads.html#event-message_1

- Add a 5000ms safety timeout on the lock-written signal so a hung Holder
  worker surfaces as a specific error, not a generic test-timeout.

- Add a fs.openSync/fs.writeFileSync stub in Worker B that counts atomic-
  create attempts (O_CREAT|O_EXCL for perf-316 / { flag: 'wx' } for perf-407).
  Assert lockAttempts >= 2 to prove the SUT entered the retry loop. This
  closes the Contract-4 hole: sabCount === 1 now discriminates pre-fix
  (sabCount === lockAttempts) from post-fix (sabCount === 1, hoisted).

- Bump holdMs from 400ms to 1000ms to guarantee >=4 retries (perf-316,
  200ms+jitter delay) or >=9 retries (perf-407, 100ms delay) on the
  slowest CI worker. Test wall time grows ~600ms; well under the existing
  8000ms timeout.

Closes #432

Co-authored-by: CI Rebase Check <ci@open-gsd.dev>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-28 20:10:08 -04:00
Tom Boucher
381be435fe fix(#445): upgrade Windows checkout v4→v5.0.1 + remove FORCE_JAVASCRIPT_ACTIONS_TO_NODE24 (node20 EOL) (#446)
Problem: Node 20 deprecation warnings fired on every CI run. Despite
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24=true in test.yml (added in PR #350),
the warning is not silenced — the env var only changes which warning fires
(both paths call context.Warning()). The only silent path is for the action
itself to declare `using: node24`.
Ref: https://github.com/actions/runner/blob/main/src/Runner.Common/Util/NodeUtil.cs
Ref: https://github.com/actions/runner/blob/main/src/Runner.Worker/JobExtension.cs

Three locations used @v4 actions (node20 runtime):

1. test.yml — Windows lanes (test + test-full jobs) pinned to
   actions/checkout@11bd71901b (v4.2.2).
   Original pin rationale (PR #162 / issue #161): v6 uses includeIf.gitdir:
   for auth injection, which is unreliable on Windows git 2.54.
   v5 never used includeIf, so it is safe for Windows.
   Upstream confirmation: https://github.com/actions/checkout/pull/2425
   v5 action.yml declares `using: node24`: https://raw.githubusercontent.com/actions/checkout/v5/action.yml

2. changeset-required.yml — floating actions/checkout@v4 + actions/setup-node@v4

3. docs-required.yml — same floating @v4 pattern

Changes:
- test.yml: both Windows checkout steps v4.2.2 → v5.0.1
  (SHA 11bd71901bbe5b1630ceea73d27597364c9af683 → 93cb6efe18208431cddfb8368fd83d5badbf9bfd)
- test.yml: update comment on Windows checkout to reflect v5 rationale
- test.yml: no FORCE_JAVASCRIPT_ACTIONS_TO_NODE24 was present on origin/next
  (already absent; the env block was on the main-branch version only)
- changeset-required.yml: actions/checkout@v4 → @93cb6efe18208431cddfb8368fd83d5badbf9bfd (v5.0.1)
- changeset-required.yml: actions/setup-node@v4 → @a0853c24544627f65ddf259abe73b1d18a591444 (v5.0.0)
- docs-required.yml: same as changeset-required.yml

SHAs resolved from upstream tags:
- checkout v5.0.1: gh api repos/actions/checkout/git/ref/tags/v5.0.1 → 93cb6efe18208431cddfb8368fd83d5badbf9bfd
- setup-node v5.0.0: gh api repos/actions/setup-node/git/ref/tags/v5.0.0 → a0853c24544627f65ddf259abe73b1d18a591444

Closes #445

Co-authored-by: CI Rebase Check <ci@gsd-redux>
2026-05-28 18:53:41 -04:00
Tom Boucher
d83e58eea0 fix(#437,#439,#440): restore defaults.run.shell + 'zsh {0}' format + Windows .cmd shell:true (PR #434 fallout) (#438)
* fix(#437): restore defaults.run.shell at job level (step-level matrix expr rejected by GHA)

Per actions/runner workflow-v1.0.json schema, `jobs.<job_id>.defaults.run.shell`
allows `matrix` context (job-defaults-run has context:[matrix,...]); step-level
`shell:` does not (run-step's shell field is plain string with no context array).
PR #434 used step-level shell:${{matrix.shell}}, which GHA's parser rejects with
"Unrecognized named-value: 'matrix'" — blocking every push to next and every
release.yml dispatch.

This commit:
- Removes step-level `shell: ${{ matrix.shell }}` from test-full (test.yml)
  and smoke (install-smoke.yml) jobs (17 directives).
- Adds `defaults.run.shell: ${{ matrix.shell }}` at job level in those two jobs.
- Fixes pre-existing shellcheck SC2129 in test.yml (individual >> redirects →
  grouped brace form) and SC2010 in install-smoke.yml (ls|grep → glob loop).

Verified locally with actionlint 1.7.12 (exit 0). Policy linter still 0 violations
(matrix.shell now resolves via job.defaults.run.shell which the linter already
handles per workflow-policy.cjs:effectiveShell).

Refs: actions/runner#444 (open since 2020), GHA contexts page section "Context availability".

* fix(#439): inline ci-smoke-skip back to shell (Node port required pre-checkout file resolution)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#440): use platform-correct npm.cmd on Windows for spawn (and surface-check other Node ports)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#437): use 'zsh {0}' format string in matrix.shell for macOS (zsh not in GHA built-ins)

Per https://docs.github.com/en/actions/using-workflows/workflow-syntax-for-github-actions
(jobs.<job_id>.defaults.run.shell section):

  "You can use built-in shell keywords like bash, pwsh, python, sh, cmd, and
  powershell, or define a custom set of shell options."

zsh is not in the built-ins list. GHA accepts custom shells via a format string
containing '{0}', which it replaces with the temporary script file path at
runtime (same pattern as the perl {0} example in the docs).

Bare `shell: zsh` triggers: "Invalid shell option. Shell must be a valid
built-in or a format string containing '{0}'".

Precursor: 514cb429 introduced the matrix shell-pinning pattern; this completes
it by switching the macOS rows from the bare value to the required format string.

Also updates scripts/workflow-policy.cjs to normalise 'zsh {0}' to 'zsh' before
the policy comparison, so the repo-baseline test continues to pass (the linter
was correctly treating 'zsh {0}' as a distinct value from the policy 'zsh').

Affects:
- .github/workflows/test.yml: test-full matrix (node 22 + node 24 macOS rows)
- .github/workflows/install-smoke.yml: smoke matrix (macOS node 24 row)
- scripts/workflow-policy.cjs: detectViolation strips ' {0}' format suffix

* fix(#440): add shell:true to spawnSync on Windows for .cmd files (Node docs requirement)

Per https://nodejs.org/docs/latest-v22.x/api/child_process.html:

  ".bat and .cmd files require a terminal to run and cannot be launched
  directly with execFile(). To run these scripts on Windows, use
  child_process.spawn() with the shell option, child_process.exec(), or
  spawn cmd.exe with the script as an argument."

  "On Windows, .bat and .cmd files require a shell to execute. Use
  child_process.exec() or child_process.spawn() with the shell: true option."

On Windows, npm is installed as npm.cmd (a batch wrapper). Without
shell: true, spawnSync resolves the binary directly and fails with
ENOENT / "npm binary not found on PATH" because the OS cannot execute
a .cmd file without cmd.exe as the intermediary.

The fix uses `shell: process.platform === 'win32'` so the shell spawning
is only activated on Windows; macOS/Linux continue to resolve the plain
npm binary directly with shell: false, preserving the existing behaviour
on non-Windows platforms.

Updated both spawnSync(npmCmd, ...) call sites:
- npm --version check (line 182)
- npm ci --dry-run lockfile-sync check (line 215)

* fix(#437): bug-410 defaults test — set USERPROFILE for Windows os.homedir() redirect

On Windows, os.homedir() reads USERPROFILE (not HOME), so the test's
process.env.HOME = FAKE_HOME redirect was silently ignored. finishInstall's
path.join(os.homedir(), '.gsd') resolved to the real user home and the
defaults.json write either failed (permissions) or landed outside the temp
dir, causing the existsSync assertion to return false.

Fix: also set process.env.USERPROFILE = FAKE_HOME so os.homedir() returns
the sandboxed directory on Windows. Node.js docs (os.homedir):
https://nodejs.org/docs/latest-v22.x/api/os.html#oshomedir

Refs: #437 (fix/437-restore-defaults-run-shell), Windows pwsh compat

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#437): precommit-alias-drift hook test — use path.delimiter for PATH

Hardcoded ':' PATH separator breaks Windows where process.env.PATH uses ';'.
The malformed PATH passed to bash caused the mock git/npm stubs in binDir
to be invisible to the hook script; npm was never called and the marker
file never written.

Fix: replace ':' with path.delimiter in both PATH constructions so the
env var is well-formed on Windows (';') and POSIX (':') alike.

Refs: #437 (fix/437-restore-defaults-run-shell), Windows pwsh compat

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#437): prepush-enterprise-email hook test — use path.delimiter for PATH

Same root cause as precommit-alias-drift: hardcoded ':' PATH separator is
invalid on Windows (';' required). The malformed PATH meant bash ran the
real git binary instead of the mock stub, which rejected the placeholder
SHAs 'refs-local-sha' / 'refs-remote-sha' with a fatal ambiguous-argument
error rather than returning the fixture commit list.

Fix: replace ':' with path.delimiter in both execFileSync PATH env values.

Refs: #437 (fix/437-restore-defaults-run-shell), Windows pwsh compat

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#437): set MSYS2_PATH_TYPE=inherit so mock stubs take precedence in Git Bash PATH

Root cause: Git Bash (MSYS2) on Windows prepends its own system directories
(/mingw64/bin, /usr/bin, /bin) to the PATH at process startup before the
user-supplied Windows PATH entries. This placed the real git/npm binaries
ahead of the mock stubs in binDir even though binDir was first in the Windows
PATH passed to execFileSync. The path.delimiter fix (0042fe0d) made the PATH
syntactically correct for Windows (semicolons) but did not change the MSYS2
system-dir prepend order.

The real git rejected placeholder SHAs (refs-local-sha, refs-remote-sha) with
"fatal: ambiguous argument", producing the observed Windows CI failure. For the
pre-commit test, the real git output nothing (no staged files on a fresh
checkout), so the grep match failed and npm was never called.

Fix: set MSYS2_PATH_TYPE=inherit in the env passed to both bash spawns.
With inherit, MSYS2 uses only the converted Windows PATH without prepending
system directories, so binDir (converted from Windows to POSIX) is first in
the search path and the mock stubs are found.

grep/tr/printf remain available: the GHA Windows runner PATH includes
C:\Program Files\Git\usr\bin which contains these utilities; MSYS2 converts
that Windows entry to a POSIX path on startup. The /usr/bin/env shebang in
mock stubs resolves through MSYS2's virtual filesystem mount (not via PATH)
and is always accessible regardless of MSYS2_PATH_TYPE.

On macOS/Linux this variable is ignored; no behaviour change on those platforms.

Source: https://www.msys2.org/wiki/MSYS2-introduction/#path
(MSYS2_PATH_TYPE controls whether system dirs are prepended to converted PATH)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#437): hook test mocks — use cmd-shim pattern for Windows bin resolution

On Windows, bash (Git Bash / MSYS2) resolves PATH commands by scanning for
extensionless files, but cmd.exe and Win32 process creation resolve via
PATHEXT (.CMD, .BAT, .EXE). When execFileSync('bash', [hookPath]) runs a
hook that calls `git` or `npm`, both resolution paths may fire. The previous
approach set MSYS2_PATH_TYPE=inherit in the child env, but that variable is
only read in /etc/profile (login-shell path) — bash launched without --login
never sources /etc/profile, so the variable had no effect:
https://github.com/msys2/MSYS2-packages/blob/master/filesystem/profile

Fix: adopt the cmd-shim three-file pattern used by npm itself:
https://github.com/npm/cmd-shim
For each mock binary, write:
  <name>          extensionless bash script (bash PATH scan)
  <name>.cmd      batch wrapper delegating to bash (PATHEXT / cmd.exe)
  <name>.ps1      PowerShell wrapper (completeness)

This is the same approach used by stevemao/mock-bin for test mocking with
Windows CI green on AppVeyor:
https://github.com/stevemao/mock-bin

The .cmd and .ps1 files are only written on process.platform === 'win32'.
MSYS2_PATH_TYPE is removed from the child env — it was ineffective and is
no longer needed with the shim files in place.

* fix(#437): tarball-smoke — raise CHILD_TIMEOUT_MS on Windows to 600 s

The CI failure showed a test duration of 120003.1812 ms — matching the
previous CHILD_TIMEOUT_MS = 120_000 exactly. When spawnSync hits its
timeout, it sends SIGTERM and returns { status: null, stdout: '', stderr: '' }
per the Node.js docs:
https://nodejs.org/docs/latest-v22.x/api/child_process.html
  "status: <number> | <null> — The exit code of the subprocess, or null if
   the subprocess terminated due to a signal."

The installResult check is `status !== 0`; null !== 0 is true, so the
timeout fired the INSTALL_FAILED path with empty stdout/stderr, which made
the root cause invisible in CI logs.

GitHub-hosted Windows runners are slower than Linux/macOS for
filesystem-heavy operations (npm install -g of a 1499-file tarball):
https://docs.github.com/en/actions/using-github-hosted-runners/about-github-hosted-runners/about-github-hosted-runners#standard-github-hosted-runners-for-public-repositories

Fix: use 600_000 ms (10 min) on Windows, keeping 120_000 ms on POSIX.
600 s matches the SLOW_HOST_TIMEOUT already used in the test before() helper
for the pack + install fixture step.

Also expose `signal` and `installError` in the INSTALL_FAILED details object
so a future timeout (status=null, signal='SIGTERM', stdout='') is immediately
diagnosable in CI logs without guesswork.

* fix(#437): chmod +x via bash on Windows for hook test mocks (root cause: fs.writeFileSync mode=0o755 no-op on NTFS)

Root cause: Node's fs.writeFileSync mode=0o755 is a no-op for the execute
bit on Windows NTFS. Per https://nodejs.org/docs/latest-v22.x/api/fs.html:
"on Windows only the write permission can be changed." Bash's access(X_OK)
therefore skips the mock file; the real git/npm binary is found later in PATH
and the hook runs against real state instead of the test double.

Fix: after writeFileSync, invoke Git Bash's chmod via the POSIX emulation
layer (Cygwin/MSYS2), which sets the NTFS execute ACL that Node cannot reach:

    const posixPath = filePath.replace(/\\/g, '/');
    execFileSync('bash', ['-c', `chmod +x "${posixPath}"`], { stdio: 'pipe' });

execFileSync('bash', ...) works because Git for Windows ships bash on PATH in
all GHA Windows runners. Forward-slash conversion is required because MSYS2
bash auto-converts /c/foo paths but not mixed-separator paths.

Why prior approaches didn't take effect:
- MSYS2_PATH_TYPE=inherit: only read in /etc/profile (login-shell path);
  execFileSync('bash', ...) launches non-interactively without --login, so
  /etc/profile is never sourced.
  Ref: https://github.com/msys2/MSYS2-packages/blob/master/filesystem/profile
- .cmd/.ps1 cmd-shim wrappers: bash does POSIX command resolution and does
  not honor PATHEXT, so wrappers are not found by bash's own PATH scan.
  They are not wrong (kept for non-bash callers) but do not fix bash's X_OK.

Files changed: tests/precommit-alias-drift-hook.test.cjs,
               tests/prepush-enterprise-email-hook.test.cjs

* refactor(#437): hooks use GIT_OVERRIDE/NPM_OVERRIDE env-var DI; tests drop PATH-mocking

Four prior rounds (path.delimiter join, MSYS2_PATH_TYPE=inherit, cmd-shim
.cmd/.ps1 wrappers, chmod-via-bash post-write) all failed to make MSYS2
bash's PATH-lookup find the mock executables. The root cause is that none
of those approaches can reliably override bash's own command-resolution
on NTFS without fighting NTFS execute-ACLs or login-shell profile sourcing.

The simplest robust solution is to bypass PATH entirely:

Hooks: each hook now binds GIT_CMD="${GIT_OVERRIDE:-git}" (and NPM_CMD for
pre-commit) at the top. When env vars are unset the hooks invoke bare
`git`/`npm` exactly as before — zero behavior change for users.

Tests: writeMockBin/binDir/PATH manipulation replaced by writeMock(), which
writes a .sh mock to a tmpDir and passes its absolute path via GIT_OVERRIDE
/ NPM_OVERRIDE in the execFileSync env. Bash inside the hook executes the
path directly via the seam — no PATH scan, no NTFS ACL check, no MSYS2
profile dependency.

Test-rigor principle: the new seam (env-var injection) is platform-
independent and doesn't rely on bash's command-resolution mechanism on
the host OS.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-28 18:06:19 -04:00
Tom Boucher
48b1e35187 fix(#431): enforce H1 shell policy (linux=bash, macOS=zsh, windows=pwsh) across PR + release gates (#434)
* test(#431): policy-shell-pinning linter — RED baseline (37 violations on origin/next)

Adds scripts/workflow-policy.cjs: H1 shell-policy linter with POLICY map,
VIOLATION enum, matrix expansion, effective-shell resolution order, and
runPolicyLint({ workflowsDir }) entry point.

Adds tests/policy-shell-pinning.test.cjs: 8 tests (baseline + 6 synthetic
counter-tests). Synthetic tests 2–7 pass; baseline test is intentionally RED
(37 violations: 28 in test.yml, 9 in install-smoke.yml — all macos/windows
lanes using shell: bash instead of native zsh/pwsh).

Adds js-yaml@4.1.1 as devDependency for YAML parsing.

* fix(#431): switch ubuntu/windows lanes to native shells; extract bash-isms to Node

Remove all explicit shell: bash pins from ubuntu-only jobs (changes, lint-tests,
coverage, required-tests, smoke-unpacked) — ubuntu runner default is bash, which
is both H1-compliant and the runner default, making the pin redundant.

For the test and test-full mixed-OS jobs (ubuntu+windows, windows+macos):
- Move bash-ism steps to shell-agnostic Node scripts:
    scripts/ci-guard-runner.cjs       — RUNNER_ENVIRONMENT check
    scripts/ci-rebase-check.cjs       — git fetch+merge PR base branch
    scripts/check-npm-integrity.cjs   — Node port of check-npm-integrity.sh
    scripts/ci-prepare-test-scope.cjs — write .ci-selected-tests.txt
    scripts/ci-smoke-skip.cjs         — set skip= output for full-only matrix entries
- Remove shell: bash from simple npm/node command steps (runner default applies)

This brings Windows violations from 19 to 0. Remaining 17 violations are all
MACOS_MISSING_EXPLICIT_ZSH in mixed-OS matrix jobs (test-full: windows+macos,
install-smoke smoke: ubuntu+macos) — these require job splitting to fix; see
BLOCKER in PR description.

* fix(#431): update workflow-shell-pinning test for H1 policy

The old test required all Windows-targeting npm steps to pin shell: bash
(to prevent pwsh stderr-swallow). Under H1, Windows runners must use
pwsh (native, no pin needed) — shell: bash on Windows is now the
violation, not the fix.

Update findViolations() to flag npm steps with effectiveShell === 'bash'
(rather than effectiveShell === null). Update synthetic tests to verify
the H1-inverted semantics: defaults.run.shell: bash on Windows is now 2
violations, not 0. Update test name and assertion messages to describe
the H1 constraint rather than the old missing-pin constraint.

* fix(#431): extend policy linter to resolve matrix.shell expressions

- expandRunsOn now captures all matrix.include row keys as realization
  context (os, node-version, shell, full_only, etc.) instead of only os
- effectiveShell now accepts a realizationContext and resolves
  ${{ matrix.<key> }} expressions against it before checking policy
- Unresolvable matrix key in shell expression emits UNRESOLVABLE_MATRIX
- Add 3 new tests: positive (zsh+pwsh per row → 0 violations),
  counter (bash in macOS row → WRONG_SHELL_FOR_OS), counter (missing
  shell key → UNRESOLVABLE_MATRIX)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#431): apply matrix.shell pattern to test-full and smoke jobs (clears BLOCKER)

test-full job (test.yml):
- Add shell: pwsh/zsh per matrix.include row (windows-latest→pwsh,
  macos-latest→zsh)
- Add job-level defaults.run.shell: ${{ matrix.shell }}
- No step-level shell pins existed to remove

smoke job (install-smoke.yml):
- Add shell: bash/zsh per matrix.include row (ubuntu→bash, macos→zsh)
- Add job-level defaults.run.shell: ${{ matrix.shell }}
- No step-level shell pins existed to remove

Policy linter now reports 0 violations across all workflow files.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#431): migrate .sh check scripts to .cjs; remove .sh originals

- Add scripts/check-env.cjs: Node.js port of check-env.sh with
  identical exit codes (0/1/2), human-readable and --json output,
  --help flag, and all 5 checks (node-version, npm-version,
  lockfile-present, lockfile-sync, version-manager-pin)
- Migrate all callers:
  - package.json check:env → node scripts/check-env.cjs
  - package.json check:integrity → node scripts/check-npm-integrity.cjs
  - scripts/ci-test-scope.cjs path strings → .cjs equivalents
  - .github/workflows/release.yml rc+finalize jobs → node .cjs (drop chmod+x)
  - .github/workflows/security-scan.yml → node .cjs (drop chmod+x)
  - tests/check-env.test.cjs → spawn node process.execPath [.cjs]
  - tests/npm-integrity-gate.test.cjs → spawn node process.execPath [.cjs]
- Delete scripts/check-env.sh and scripts/check-npm-integrity.sh

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#431): update doc references from .sh to .cjs

Update SECURITY.md and docs/contributing/bootstrap.md to reference the
canonical Node invocation instead of the removed bash scripts.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#431): use per-step shell:matrix.shell instead of defaults.run.shell (GHA compat)

GHA does not reliably resolve matrix expressions inside defaults.run.shell.
Per-step shell: always resolves correctly. Removed the defaults.run.shell block
from the test-full job (test.yml) and the smoke job (install-smoke.yml), and
added shell: \${{ matrix.shell }} directly on every run: step in both jobs.

Codex finding: defaults.run.shell with matrix expressions is not a
GHA-supported pattern; per-step shell: is the safe form.

* fix(#431): policy linter validates every matrix.include row independently

Removed runner-label-only dedup from expandRunsOn() in workflow-policy.cjs.
The prior guard (if !realizations.find(r => r.runner === runner)) collapsed
two macos-latest rows with different node-version/shell contexts into one,
hiding the second row's policy violation.

Each matrix.include row is a distinct CI realization with its own context;
validating it twice is harmless but skipping it causes false negatives.

Added counter-test (Test 8) in tests/policy-shell-pinning.test.cjs:
two macos-latest rows (shell:zsh compliant + shell:bash violation) must
produce exactly one WRONG_SHELL_FOR_OS violation on the second row.

* fix(#431): remove dedup-by-runner in Cartesian matrix.<key> expansion (Codex round 3)

The base-list path in expandRunsOn (matrix.<key> arrays, e.g. matrix.os)
previously guarded each push with `if (!realizations.find(r => r.runner === runner))`,
collapsing duplicate runner values into a single realization and hiding policy
violations on later rows of a Cartesian matrix.

Remove the guard unconditionally; each entry in the base-list array now produces
its own realization, matching the same fix already applied to the matrix.include path.

Add counter-test "Cartesian matrix os × shell — dedup must not collapse rows by
runner alone": matrix.os: [macos-latest, macos-latest] + shell: ${{ matrix.shell }}
now yields 2 realizations (not 1). Documents that Cartesian cross-product expansion
(carrying all keys into realization context) is a separate follow-up; current violations
are UNRESOLVABLE_MATRIX pending that work.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#431): remove 60s timeout regression on npm ci --dry-run (parity with check-env.sh)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#431): ci-rebase-check.cjs — return truthy sentinel on success (Codex round 4)

run() used execFileSync with stdio:'inherit', which returns null on success.
Caller checked `result !== null`, always false → every successful fetch fell
through to "failed after 3 attempts" exit-1 path.

Fix: run() now returns true on success, false on failure.
Update caller from `result !== null` to `if (result)`.

Adds tests/ci-rebase-check.test.cjs (5 tests) covering the sentinel contract
and a local-bare-remote integration smoke that verifies the full fetch+merge
path exits 0 when fetch succeeds.

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
2026-05-28 09:23:59 -04:00
Tom Boucher
a5eceb1faf fix(#211): add ~/.claude/get-shit-done/bin fallback to gsd_run launcher (closes #394) (#427)
Extends the canonical runtime-launcher snippet with a third resolution arm
that probes $HOME/.claude/get-shit-done/bin/${_GSD_SHIM_NAME} between the
PATH check and the hard-error exit. Global Claude-Code installs (--claude
without --local) with no PATH wiring and no RUNTIME_DIR no longer hit the
hard-error path.

Resolution order: local/RUNTIME_DIR -> PATH -> ~/.claude/... -> hard error.

Propagated to 76 workflow .md files via sync-runtime-launcher.cjs. Parity
test (G) and bug-211 regression test (4 assertions) added.
2026-05-27 23:17:59 -04:00
Tom Boucher
92cd7e03dd fix(#338): write local Claude install hook wiring to settings.local.json (+ one-shot migration) (#426)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-27 23:14:43 -04:00
Tom Boucher
6f33b18f69 fix(#376): rewrite /gsd: → /gsd- in Claude-installed hook .js files (#424)
* fix(#376): rewrite /gsd: → /gsd- in Claude-installed hook .js files

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#376): preserve .sh branch + {{GSD_VERSION}} stamp in restructured hook-copy loop

Trim the .js branch comment/whitespace so the `else {` and
`entry.endsWith('.sh')` fall within the 1500/2000-char assertion windows
anchored on `configDirReplacement` in the regression tests for #1834 and
#2136. The .sh read+substitute+chmod path is intact; the new #376 hyphen-
namespace rewrite for .js/.cjs files is also preserved.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-27 23:12:15 -04:00
Tom Boucher
6cba62b39d fix(#397): preserve executor-authored STATE.md fields (template-default-only replacement) (#422)
Introduces KNOWN_TEMPLATE_DEFAULTS and KNOWN_STATUS_PATTERNS in state-document.cjs
to enumerate every string a GSD handler writes.  stateReplaceFieldIfTemplate
consults this table and only replaces the field when the existing value is a known
template default (or absent) — executor-authored values are left untouched.

Wire-in:
- record-session: Resume File now only overwritten when caller passes --resume-file
  OR existing value is 'None'.  Router no longer defaults resume_file to 'None'
  before calling the handler.
- advance-plan (both branches): Status and Last Activity guarded via
  stateReplaceFieldIfTemplate.
- updateCurrentPositionFields: Status and Last activity in the Current Position
  section guarded; bare ISO date shape is the trigger for replacement, prose
  narrative is preserved.
- planned-phase: Status and Last Activity guarded the same way.

Regression tests (7 cases) in tests/bug-397-state-preserve-executor-authored.test.cjs
cover each data-loss shape; all 106 existing state.test.cjs tests still pass.

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-27 23:03:15 -04:00
Tom Boucher
978823cd06 fix(#410): guard ~/.gsd/defaults.json write with GSD_TEST_MODE (#130 sibling) (#421)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-27 23:01:15 -04:00
Tom Boucher
f5f51b5b49 fix(#408): align ci-test-scope smoke handling with #395 changeset (drop unconditional injection; unit fallback) (#420)
- Remove `DEFAULT_SMOKE_TESTS` and `WINDOWS_SMOKE_TESTS` constants (now dead after the unconditional injection block is dropped)
- Drop the `addAll(targeted, DEFAULT_SMOKE_TESTS)` / `addAll(windows, WINDOWS_SMOKE_TESTS)` block from the `codeChanged` branch
- When `codeChanged && targetedTests.length === 0`, push `'unit'` as the fallback suite token
- Two new regression tests in `tests/ci-test-scope.test.cjs` covering the no-injection and unit-fallback contracts (bug #408)
2026-05-27 22:59:26 -04:00
Tom Boucher
e4aca8dab0 fix(#416): return null when active milestone has no archive; tighten **Milestone:** regex (#419)
* fix(#416): return null when active milestone has no archive (no fall-through to prior milestone's archive)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#416): handle bold-formatted Milestone: field in archive dir resolver

The STATE.md regex in getActiveMilestoneArchiveDir failed to extract the
version from **Milestone:** vX.Y format (bold wraps the label+colon).
The old pattern captured '**' instead of the version, causing the
milestone→archive lookup to produce a false candidate path, then return
null (post-fix behavior) instead of falling through to the version-sort
fallback — breaking the #3164 consistency scanner tests.

Fix: extend the regex to skip optional trailing '**' after the colon so
both 'milestone: vX.Y' and '**Milestone:** vX.Y' parse correctly.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-27 22:57:42 -04:00
Tom Boucher
ec0a32ea07 fix(#407): hoist sleep SharedArrayBuffer out of withPlanningLock retry loop (#418) 2026-05-27 22:48:31 -04:00
Tom Boucher
1067a0f3fd docs(#415): ADR — prevent stale-base reintroduction of retired runtime tokens (#417) 2026-05-27 21:38:04 -04:00
Tom Boucher
bd98e568f6 fix(#308): bound websearch fetch with timeout and retry (#387)
cmdWebsearch called fetch() with no timeout and no retry, so a hung
connection blocked indefinitely and transient 429/5xx/network failures
were not recovered. Add AbortSignal.timeout (configurable via
GSD_WEBSEARCH_TIMEOUT_MS, default 10s) and a bounded retry loop
(max 2 retries, exponential backoff + jitter) for 429/5xx/network
errors, honoring Retry-After on 429 (capped at 60s). Non-429 4xx fail
immediately (no wasted retries). Transient-exhausted failures report an
`attempts` count. Worst-case time is bounded by timeout*(1+retries)+backoff.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 21:32:06 -04:00
Tom Boucher
7ea6a06645 fix(#378): poll scoped package name in update check (#414)
The update worker queried the unscoped 'get-shit-done-redux' via
`npm view`, which returns E404 — so `latest` stayed null and
`update_available` could never become true. Now derives the name from
package.json (`require('../package.json').name`) so it always matches
the actual published scoped name (@opengsd/get-shit-done-redux).

Fixes #378.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 21:10:07 -04:00
Tom Boucher
485ea1bd3d fix(#411): restore gsd_run launcher in next.md and align policy-160 test (#412)
* fix(#411): restore gsd_run launcher in next.md (re-run sync after #406 regression)

* fix(#411): update policy-160 route0 test to expect gsd_run canonical resolver

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-27 21:03:38 -04:00
Tom Boucher
21dcf58750 fix(#138): add --default true to nyquist_validation config-get in validate-phase/audit-milestone (#405)
config-get calls lacked --default, causing stderr noise ("Key not found") and a fragile empty-variable fallback when workflow.nyquist_validation was absent; added --default true (matching schema default). Fixes #138.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:27:33 -04:00
Tom Boucher
af36416d14 fix(#160): resume partially-executed phases before current_phase routing (#406)
When a session dies mid-execution (hang, token exhaustion, API drop),
STATE.md's current_phase can be advanced past a phase that still has
PLAN.md files without matching SUMMARY.md files. Without this fix,
/gsd-next and /gsd-progress would route by current_phase and silently
skip the partially-executed phase, producing a data-loss-shape outcome.

Adds a Route 0 cross-phase incomplete-execution scan to both
next.md and progress.md. Before any current_phase-based routing,
the scan finds the lowest-numbered phase where plans outnumber
summaries and routes to /gsd-execute-phase <that-phase> to resume it.
Opt out with --no-resume to fall back to the prior-phase defer prompt;
--force bypasses all gates as before.

Rework (codex review):
- Route 0 now ordered AFTER Gates 1-3 (repo/state validity always run)
  but BEFORE the prior-phase completeness-scan defer prompt — eliminating
  the double-decision where the default path would both prompt the user
  (C/S/F) AND resume the phase anyway. Prior-phase defer prompt moved to
  a new prior_phase_completeness step; only reached via --no-resume.
- --force flow made coherent across all three steps: safety_gates jumps
  directly to determine_next_action, skipping Gates, Route 0, AND
  prior_phase_completeness. resume_incomplete_phase and prior_phase_completeness
  now correctly state --force never reaches them. success_criteria entry
  updated to reflect --force → determine_next_action (not prior_phase_completeness).
- Scan uses $GSD_SDK (canonical resolver form) throughout next.md, matching
  the file's existing convention. progress.md uses $ROADMAP already loaded
  by analyze_roadmap. Neither file uses bare gsd-sdk.
- Errors are surfaced rather than suppressed: removed 2>/dev/null on the
  main roadmap.analyze call; added explicit WARNING emission when the scan
  cannot run, so the invariant fails closed instead of failing open.
- Predicate aligned to plans-without-summaries (plans.length > summaries.length)
  in both files, consistent with determine_next_action Route 4.
- command references use canonical /gsd: namespace form

Fixes #160

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:22:14 -04:00
Tom Boucher
fd061fccc5 fix(#130): skip opencode permission config under GSD_TEST_MODE (#404)
finishInstall called configureOpencodePermissions unconditionally, causing
fs.mkdirSync + fs.writeFileSync to run even under GSD_TEST_MODE='1', violating
the side-effect-free contract. Guarded the call with !process.env.GSD_TEST_MODE.

Fixes #130

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:22:09 -04:00
Tom Boucher
616b387f12 perf(#320): hoist By-Phase state table regex to module scope (#403)
Static regex literal `byPhaseTablePattern` was recompiled on every call to `updatePerformanceMetricsSection`; hoisted to module scope (compiled once; stateless /i used with .match → safe to share across calls). `phaseRowPattern` uses dynamic interpolation and stays in-function. `Fixes #320`.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:22:06 -04:00
Tom Boucher
b058c5861f ci(#319): shallow checkout for docs/changeset lint workflows (#402)
Lint workflows used fetch-depth:0 (full clone); switched to depth 50 +
explicit base-ref fetch so the three-dot diff (origin/${base}...HEAD) has
its merge-base; fails closed if merge-base is deeper than 50. Added
policy test asserting fetch-depth:50 on both workflows. Fixes #319.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:22:03 -04:00
Tom Boucher
636a9463e6 ci(#318): drop runtime npm self-upgrade from release lanes (#401)
Release jobs ran `npm install -g npm@latest` before each publish step,
adding ~30 s and version drift risk on every run; removed both occurrences,
relying on Node 24's bundled npm pinned via setup-node. Added a policy test
(tests/policy-release-no-npm-self-upgrade.test.cjs) that will fail RED if
the antipattern is re-introduced in release.yml or hotfix.yml. Fixes #318.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:21:59 -04:00
Tom Boucher
ee820442e6 perf(#317): collapse redundant existsSync+readFileSync in context-monitor hook (#400)
Per-PostToolUse hot path did stat-then-read ×3 (config.json, metrics bridge,
warn sentinel); collapsed to read-with-ENOENT-catch (fewer blocking syscalls,
no TOCTOU), behavior identical. Fixes #317.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:21:55 -04:00
Tom Boucher
1868947088 perf(#316): hoist state-lock sleep buffer out of retry loop (#399)
acquireStateLock was allocating a fresh SharedArrayBuffer on every retry
iteration via Atomics.wait(new Int32Array(new SharedArrayBuffer(4)), ...).
The buffer is never mutated and never escapes, so hoisting it before the loop
is a provably-equivalent transformation — Atomics.wait always sees value 0
whether the buffer is fresh or reused.

Fixes #316

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:21:52 -04:00
Tom Boucher
a393e28b34 perf(#315): memoize subrepo detection within loadConfig (3 scans → 1) (#398)
loadConfig called detectSubRepos(cwd) at up to 3 sites per invocation (root-config
requiresFilesystem migration, workstream-config requiresFilesystem migration, and the
planning.sub_repos filesystem re-sync) — all with the same cwd, yielding identical
results. Introduce a per-call lazy memo (getDetectedSubRepos) so the directory scan
runs at most once per loadConfig call while preserving all conditional logic. Fixes #315.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:21:50 -04:00
Tom Boucher
4361d83279 perf(#314): index roadmap plan lookups by id (O(lines×plans) → O(lines+plans)) (#396)
Hot path in cmdRoadmapAnnotateDependencies called planData.find() on every
checklist line. Replaced with a first-wins Map built once before the loop so
each line resolves in O(1); first-wins preserves exact .find() semantics and
null-on-miss → wave-1 default is unchanged.

Fixes #314

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:21:46 -04:00
Tom Boucher
8c8887f00e fix(#370): scope affected-tests runner to PR suites, exclude push-only install/slow (#395)
Root cause: the affected-tests runner called runAllSuites() on critical-path
changes (running every suite including install/slow on all matrix cells including
Windows), and pickAffectedTests injected DEFAULT_SMOKE_TESTS (an install test)
as the empty-selection fallback — causing install suite tests to run on PR lanes
where they are push-only per docs/TESTING-SUITES.md.

Fix: PR_EXCLUDED_SUITES filter at the pickAffectedTests chokepoint strips
install/slow from every selection path (direct-change, reverse-index, stem-match).
Empty selection now returns [] and the caller runs the unit suite as smoke.
Critical-path fallback replaces runAllSuites with PR_FULL_SUITES
(unit, integration, security). suiteOf exported from run-tests.cjs
(with require.main guard) so affected-tests-lib reuses canonical detection.

Fixes #370

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:21:43 -04:00
Tom Boucher
269ed3e3b5 perf(#313): dedupe intel export extraction with a Set (#393)
intelExtractExports deduped export names with `if (!arr.includes(x)) arr.push(x)`
across ~8 extraction loops (one doubly-nested over an export block) — O(n^2).
Accumulate into Sets (add/has/size) and materialize to an array once at return.
Set dedups by value and preserves insertion order, so the returned export list
and its first-seen order are identical. Adds behavior-lock tests for dedup + order.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:21:39 -04:00
Tom Boucher
d77170a25d perf(#312): index argv once in parseNamedArgs (#392)
parseNamedArgs re-scanned argv with indexOf/includes once per flag —
O(flags * argv) — on the command-dispatch hot path (24 call sites across
gsd-tools + init/state/validate routers). Build a first-index Map of argv
tokens in a single pass and use it for the flag lookups, dropping it to
O(argv + flags). Semantics are identical: firstIndex.get(t)??-1 === indexOf(t),
firstIndex.has(t) === includes(t); first-occurrence-wins and the value-token
rejection are preserved. Adds the first behavior-lock tests for the module.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:21:36 -04:00
Tom Boucher
25d24219bf perf(#311): index subrepo routing by first path segment (#390)
cmdCommitToSubrepo routed each changed file to a sub-repo via
subRepos.find(...) inside the file loop — O(files * repos). Extract a pure
groupFilesBySubrepo() that buckets sub-repos by first path segment and scans
only the matching bucket, dropping it to expected O(files + repos). First-
match-in-array-order semantics (incl. multi-segment sub-repos) are preserved
exactly. Adds the first behavior-lock test for the routing path.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:21:29 -04:00
Tom Boucher
09123e395e test(#382): make graphify-auto-update hook tests deterministic on Docker (#389)
The hook detaches a background rebuild; the tests raced it two ways on
slow/contended Docker (Mac passed): (1) a tight ~2.2s status poll budget
expired before the detached rebuild wrote its terminal status -> "1 subtest
failed"; (2) cleanupHookRepo's rmSync threw EBUSY/ENOTEMPTY while the child
was still writing -> "failed running after hook".

Replace the ad-hoc poll budgets with a shared waitForBuildStatus() that
waits for the real terminal status ('ok'/'failed') under a generous 30s
deadline (all assertions are outcome-based, so this is deterministic, not a
timing assertion), and make teardown best-effort so a residual temp dir can
never fail a passing test. No production/hook code changed; no assertion
weakened.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:21:25 -04:00
Tom Boucher
24b62cbe71 ci(#310): scope release and hotfix gates to unit coverage (#388)
release.yml (rc + finalize) and hotfix.yml (finalize) ran the full
`npm run test:coverage` suite on the release path, redundantly re-running
the integration/install/security/slow suites that already passed on the
PR lanes into next. Switch those three sites to `npm run test:coverage:unit`
(same c8 config, unit suite only) to cut release latency. Full-suite
coverage remains available via the dedicated lanes / `test:coverage:all`.

Adds a workflow-contract regression test asserting the release/hotfix gates
invoke the unit coverage command (exact-line match, not substring).

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:21:22 -04:00
Tom Boucher
3599cb0c3e perf(#306): build learnings dedupe index once per bulk import (#386)
learningsCopyFromProject called learningsWrite K times in one process,
and each call re-scanned the entire learnings store to dedupe — O(K*N).
Build the content_hash -> id index once at the start of the bulk import
and thread it through; single-write behavior and the return contract are
unchanged (no caller reads `id` on the created:false branch). O(K*N) ->
O(N+K). Adds a regression test asserting store scan count is independent
of import size.

The larger persistent on-disk index (atomic updates, corruption rebuild,
cross-process dedupe) is deferred — needs design decisions and is not
required to resolve the bulk-import scan this issue reports.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:21:18 -04:00