Commit Graph

19 Commits

Author SHA1 Message Date
Tom Boucher
4056d830bc refactor(#651): consolidate verification-status routing into one queryable seam (#755)
* refactor(#651): consolidate verification-status routing into one queryable seam

The passed/gaps_found/human_needed verification status was re-encoded as
bare strings across three prose surfaces (gsd-verifier emits, execute-phase
routes, ship gates), each independently deciding the per-status next action
with no parity coupling — the DEFECT.GENERATIVE-FIX class.

Give the enum one home: src/verification.cts (-> bin/lib/verification.cjs)
exposing `gsd_run query verification.status <phaseDir>` returning a typed
{status, next_action, next_command}. ship.md and execute-phase.md now consume
the query instead of re-deriving the routing in prose; gsd-verifier.md points
at the shared vocabulary as the single emitter (values unchanged).

Also fixes the latent broad-grep status misread (DEFECT.FRONTMATTER-SCALAR-
BROAD-GREP): execute-phase.md read `grep "^status:"` over the whole report, so
a body `status:` line could misroute a valid phase. Extraction is now
frontmatter-scoped in one place. A parity test fails if a verifier status
gains no route. Lands the two CONTEXT.md DEFECT entries captured on the issue.

Closes #651

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#651): set changeset pr to 755

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 01:26:56 -04:00
Tom Boucher
f7e902f1cf feat(#52): add agent_skills_security.trusted_global_roots allowlist for global skills (#754)
* feat(#52): add agent_skills_security.trusted_global_roots allowlist

Opt-in allowlist so a global: agent skill whose SKILL.md realpath resolves
outside the default global skills base (e.g. ~/.claude/skills) is accepted
when its real target lies under a user-declared trusted root. Default [] is
byte-identical to prior behavior; the symlink-escape guard is preserved and
simply re-applied against each declared root.

- src/security.cts: loadTrustedGlobalRoots — tilde-expand (~ and ~/), reject
  project-relative and dangerously broad roots (filesystem/UNC root, homedir),
  realpath-canonicalize each root every run and drop non-existent ones.
- src/init.cts: on base-check failure the guard consults the trusted roots
  (hoisted out of the loop); emits a stderr NOTE when a skill is accepted via
  a trusted root so the widened boundary is visible.
- src/core.cts: thread agent_skills_security through loadConfig.
- config-schema.manifest.json: allow the new key path.
- docs/CONFIGURATION.md: document the option and its security model.
- tests/agent-skills.test.cjs: unit + end-to-end CLI coverage (regression,
  feature, negative, broad-root hardening, stderr NOTE).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#52): add changeset fragment for trusted_global_roots (#754)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 01:06:41 -04:00
Tom Boucher
7e76f1a736 feat(#703): add --granularity override flag to /gsd:plan-phase (#750)
* feat(#703): add --granularity override flag to /gsd:plan-phase

Add a `--granularity <coarse|standard|fine>` flag to /gsd:plan-phase that
overrides the configured planning granularity for a single invocation.

The override is a new highest-priority tier above the existing precedence
chain (granularities[phaseType] -> granularity -> planning.granularity ->
'standard') in resolveGranularityInternal; when the flag is absent, resolution
is byte-for-byte unchanged. cmdInitPlanPhase now resolves with phaseType
'planning' so granularities.planning participates, and emits the resolved
value in the init JSON, which the plan-phase workflow forwards to the planner
prompt. Invalid values are rejected at the CLI boundary via a shared
assertValidGranularityOverride helper.

Closes #703

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#703): set changeset pr to 750

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 23:56:53 -04:00
Tom Boucher
cf8bd3cd5e fix(#683): auto-degrade phase execution to sequential on worktree base mismatch (#749)
* fix(#683): auto-degrade phase execution to sequential on worktree base mismatch

Claude Code forks worktree-isolated executors off the repository default
branch (origin/HEAD), not the orchestrator's HEAD. Running /gsd-execute-phase
on a branch diverged from the default (unmerged milestone/feature branch) left
every executor without the phase's plan files and tripped the
worktree-branch-check guard with `exit 42` — 100% reproducible, all OSes.

- New module src/worktree-base-ref.cts: HEAD-vs-fork-base drift detection
  (origin/HEAD with symbolic-ref fallback) and no-clobber worktree.baseRef
  management, exposed as `worktree base-check` / `worktree set-baseref`.
- execute-phase.md: pre-dispatch, for Claude Code with worktrees enabled,
  auto-degrades the run to sequential on the main tree when a base mismatch
  is detected, recommending worktree.baseRef:"head". The exit-42 guard stays
  as a backstop.
- Installer: fresh local Claude installs set worktree.baseRef:"head" in
  .claude/settings.local.json (no-clobber, respecting an explicit shared
  settings.json value); upgrades print an opt-in notice pointing at
  `gsd-tools worktree set-baseref`.
- Docs: how-to guide, CLI/config reference, planning-config cross-ref.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#683): auto-apply worktree.baseRef on upgrade; gate fresh+upgrade on use_worktrees

Per maintainer direction: on a local Claude Code UPGRADE, set
worktree.baseRef:"head" automatically (no opt-in notice) when the project's
workflow.use_worktrees is enabled, instead of merely printing a remediation
notice. For consistency the FRESH path is now gated the same way: both paths
compute worktrees-enabled once (bounded walk-up read of .planning/config.json,
default enabled unless workflow.use_worktrees === false) and apply the
no-clobber baseRef only when enabled — never overwriting an explicit value in
settings.local.json or a shared settings.json. gsd-tools worktree set-baseref
remains for manual use. Docs + changeset updated; tests hardened (file-exists
assertions, fresh+disabled case, upgrade idempotency).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#683): measure workflow byte-budget on LF, fixing Windows-only CI failure

The workflow-size-budget test failed only on Windows: git checks out the .md
files as CRLF (no eol=lf in .gitattributes) and byteCount used
fs.statSync().size (raw on-disk bytes), counting an extra \r per line. That
inflated execute-phase.md — the XL high-water-mark file pinned near its ceiling
by the tighten-only ratchet — from 88492 LF bytes to ~90245 on Windows, over
the 90000 XL ceiling, while passing on the LF-checkout Mac/Linux runners.

The ceilings are explicitly "calibrated against raw `wc -c`" on an LF checkout,
so the measurement should be LF-based on every platform. byteCount now reads the
file and counts Buffer.byteLength after stripping CR, making the budget
platform-independent (a no-op on LF checkouts; verified statSync === normalized
for all 88 workflow files). No ceilings changed. Added a regression test
asserting CRLF and LF content of the same file count identically.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#683): make worktree-base-ref test path mocks Windows-safe (path.join)

tests/worktree-base-ref.test.cjs keyed its injected readFile/writeFile mocks
(and a few expected `file` values) with forward-slash template literals like
`${claudeDir}/settings.local.json`. The module composes those paths with
path.join(), which emits backslashes on Windows, so the mock keys never matched
the module's lookup → readFile returned null → resolveEffectiveBaseRef /
cmdWorktreeBaseCheck / cmdWorktreeSetBaseRef (and the JSONC variants) failed on
the Windows full-test runner only (they passed on Mac/Linux, and the install
tests passed because they use the real filesystem). The module is correct;
only the test fixtures hardcoded '/'.

All mock keys and path assertions now use path.join(base, ...) mirroring the
module, so they match on every platform (no-op on POSIX). 19 path references
across 16 lines.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 23:40:24 -04:00
Tom Boucher
1bea220d58 refactor(#720): lazy-load MVP-only reference bodies on non-MVP runs (#746)
* refactor(#720): lazy-load MVP-only reference bodies (eager @-import → gated Read)

Convert eager @-imports of MVP-only reference bodies into lazy "Read" instructions
gated on MVP_MODE / WALKING_SKELETON / MVP+TDD, so non-MVP planning/execution runs
no longer pull MVP guidance into context. Covers both the workflow files and the
planner/executor agent definitions (the dominant context-cost path):

- workflows/plan-phase.md: planner-mvp-mode.md + skeleton-template.md (L146/936/937/941)
- workflows/execute-phase.md: execute-mvp-tdd.md halt-report ref, now gated on gate-trip (L191)
- agents/gsd-planner.md: planner-mvp-mode.md, user-story-template.md, skeleton-template.md
- agents/gsd-executor.md: execute-mvp-tdd.md

The dedicated always-MVP mvp-phase workflow keeps its eager imports (intentional).
Behaviour is unchanged; non-MVP runs simply carry less loaded context. Adds a
regression guard mirroring the discuss-phase lazy-load test, and documents the
conformance in docs/ARCHITECTURE.md.

Refs #720

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#720): add changeset fragment (pr #746)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 18:23:17 -04:00
Tom Boucher
b79b767629 refactor(bin): replace process.exit() in CLI entrypoints + ratchet rule to error (#738) (#741)
Part 2 of 2 of the n/no-process-exit cleanup (completes umbrella #738; part 1
was #739/scripts). Converts the 20 flagged process.exit() calls in the three
hand-written gsd-core/bin CLI entrypoints and flips n/no-process-exit to error.

- New src/cli-exit.cts -> gsd-core/bin/lib/cli-exit.cjs (ExitError + runMain),
  the gsd-core-side equivalent of scripts/lib/cli-exit.cjs; registered in
  .gitignore, eslint ignores, and the inventory manifest like its siblings.
- gsd-tools.cjs: 13 apply-prompt-budget exits -> throw ExitError; main()->runMain.
- verify-reapply-patches.cjs: 6 exits -> throw ExitError / return verdict; runMain.
- check-latest-version.cjs: 1 exit -> return verdict; runMain.
- eslint.config.mjs: n/no-process-exit warn -> error.

Scope note: the gsd-core/bin/lib/*.cjs modules (core, state, profile-pipeline,
roadmap-command-router, adr-parser, ui-safety-gate) are tsc-generated and
eslint-ignored (ADR-457), so their process.exit calls were never flagged and are
intentionally left untouched. Only the linted hand-written entrypoints are in scope.

Exit codes verified unchanged for all three entrypoints.

Closes #738

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 16:27:05 -04:00
Tom Boucher
ba231ecbfc chore: clean up clear-cut ESLint warnings (#732) (#734)
Pay down pre-existing error→warn lint debt. Removes dead imports/vars, unused functions, redundant regex/string escapes, and stale eslint-disable directives; converts unused `catch (_e)` to optional catch binding (src/*.cts).

No behavior change. Lint 345→125 warnings (0 errors); deferred categories (n/no-process-exit, test-sleeps, control-regex) tracked in #732 for follow-up. Full test suite green (0 failures); code-review verified all removals unused and all escape fixes semantics-preserving.

Closes #732

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 11:24:48 -04:00
Tom Boucher
42b74100f1 feat(#159): auto-use existing RESEARCH.md in /gsd:plan-phase --research-phase (#718)
* feat(#159): auto-use existing RESEARCH.md in /gsd:plan-phase --research-phase

When RESEARCH.md already exists in research-only mode and neither --research
nor --view is passed, emit a one-line notice and exit cleanly instead of
prompting update/view/skip. This matches the promptless auto-use of standard
/gsd:plan-phase <N> (§5.1) and removes the §5.0/§5.1 inconsistency, making
AI-agent and CLI invocations non-interactive in the common case. The two
explicit-flag escape hatches (--research to refresh, --view to print) cover
any deviation.

Closes #159

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#159): point changeset fragment at PR #718

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#159): tighten research-phase reference register (Diataxis)

Make the 'no modifier' research-phase entries descriptive rather than
imperative and drop the trailing 'pass --research/--view' clauses, which
duplicated the adjacent --research/--view documentation. Reference docs
describe; the recovery flags are documented in their own entries. The
emitted runtime notice in the workflow keeps naming the flags (in-band
recovery), unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 11:09:54 -04:00
Tom Boucher
11afca2968 feat(#656): Research module — content-addressed cache + provider seam + registry-API legitimacy (#664)
* feat(#656): add Research Store module (content-addressed cache, TTL staleness)

Content-addressed research cache behind a clock seam: researchKey (sha256, deterministic), putResearch/getResearch ({hit,stale}, never throws), ttlForSource (curated HIGH 30d / MED 7d / web LOW 1d), two-tier resolveStorePath (curated -> ~/.gsd/research-cache, web/synthesis -> project .planning/research/.cache). 28 behavioral + property tests; boundary coverage at ttl-1/ttl/ttl+1.

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#656): add Research Provider module (waterfall + confidence + plan)

Single source of truth for the Balanced provider waterfall (docs Context7->Ref->Jina, web Exa+Tavily, fallback Perplexity/Brave, Firecrawl scrape-only). classifyConfidence stamps HIGH|MEDIUM|LOW by provider (never throws). providerAvailability maps config flags to usable providers. planResearch checks the Research Store (injected seam) and returns cache-hits + a per-question fetch plan, falling through the waterfall to the always-available websearch terminal. 22 behavioral + property tests.

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#656): add Package Legitimacy module (registry-API verdicts, slopcheck optional)

Replaces the pip-install-or-degrade slopcheck prose gate with code: classifyPackage (pure, never throws) computes OK|SUS|SLOP from tunable thresholds (minAgeDays 30, minWeeklyDownloads 1000, requireRepo). checkPackages queries injectable npm/PyPI/crates registry adapters (real https with 5s timeout, degraded-not-thrown on failure); slopcheck is one optional adapter that can only escalate severity, never degrade to [ASSUMED]. 34 behavioral + property tests; boundary coverage on age and downloads (limit-1/limit/limit+1).

Known follow-up: real npm adapter must add api.npmjs.org last-week downloads fetch (currently null -> unknown-downloads). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#656): detect Tavily/Ref/Perplexity/Jina provider keys; complete npm downloads adapter

config: add tavily_search/ref_search/perplexity/jina availability flags (env var or ~/.gsd/<x>_api_key), mirroring brave_search/exa_search/firecrawl, so the Research Provider waterfall can gate them. package-legitimacy: real npm adapter now fetches api.npmjs.org last-week downloads (bounded, degraded-not-thrown) so weeklyDownloads is populated. +12 config tests; 34 legitimacy tests unchanged.

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#656): expose Research seam via gsd-tools query (research-plan, research-store, package-legitimacy)

Routes the L2-hybrid surface so agents reach it as CLI: 'query research-store get/put' (cache, HOME-sandboxable), 'query research-plan --input' (cache-hits + fetch plan from planResearch), 'query package-legitimacy check --ecosystem' (async registry verdicts). Commands skip .planning root resolution and appear in top-level usage. 5 behavioral runGsdTools tests; command-contract unchanged (335).

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#656): document Research module (CONTEXT predicates, ADR-0656, architecture, changeset)

Adds GSD-RESEARCH.* + DEFECT.RESEARCH-PROVIDER-PROSE-DRIFT predicates to CONTEXT.md, ADR-0656 recording the L2-hybrid seam decision, a docs/ARCHITECTURE.md Research Module subsection, and an Added changeset fragment (pr:0, backfill on PR). Notes the #657 deferrals (agent collapse + install.js MCP mapping).

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#656): sync inventory for research modules

Regenerate INVENTORY-MANIFEST.json and bump docs/INVENTORY.md CLI Modules count 82->85 with rows for research-store/research-provider/package-legitimacy (DEFECT.INVENTORY-DRIFT).

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#656): eslint-ignore generated research .cjs artifacts (ADR-457)

research-store/research-provider/package-legitimacy .cjs are tsc-generated from src/*.cts, so they belong in the ESLint ignore block (lint the .cts source, not the emitted .cjs). Fixes tests/551-eslint-bin-lib-coverage.

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#656): backfill changeset pr number to #664

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#656): satisfy eslint lint-tests gate

Fix 20 eslint errors in the new research files: use helpers.cleanup() instead of raw fs.rmSync() in tests (local/no-raw-rmsync-in-tests, Windows-EBUSY retry budget); drop redundant '| string' union members and unnecessary type assertions; deterministic object normalization in researchKey (no-base-to-string). Logic unchanged; 6180 tests still green.

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): harden package legitimacy per review (W1/W2/I3/I4)

W1: httpsGet now reads statusCode; npm/PyPI/crates map 404 -> exists:false -> SLOP (registry-existence is the #1 slopsquatting defense; previously only npm caught it). Transport made injectable (_setHttpGet) for hermetic 404 tests. W2: suspicious-postinstall is now terminal SLOP independent of the optional slopcheck adapter, and the regex drops the bare https?:// arm (over-fired on esbuild/sharp/node-gyp) for shell-exec/download-exec signatures only. I3: checkPackages now threads version to registry.lookup and adapters verify that specific version exists. I4: moreServerVerdict -> moreSevereVerdict. +11 regression tests (all RED-first); 45 total green.

Addresses review by @davesienkowski on #664. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): research-store tier coherence + freshness + version TTL (W4/I1/I2/I4)

I1: tier now derives from source (curated -> user ~/.gsd, else -> project .planning), not kind, so put-tier and get-tier can't diverge; kind is a key component only. W4: getResearch searches both tiers and returns the freshest (non-stale preferred), never letting a stale curated entry shadow a fresh web one; blank version caps TTL at 1 day (no 30d on version-blind keys). I2: atomic platformWriteSync instead of raw fs.writeFileSync on the shared global path. I4: dropped the dead ttlForSource arm. CLI get now searches both tiers. +5 RED-first regression tests; 38 green.

Addresses review by @davesienkowski on #664. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): expose classifyConfidence as a CLI route, killing dead code (W3)

Adds 'gsd-tools query classify-confidence --provider X [--verified]' so research agents get the confidence tier FROM CODE (provider waterfall + verification lever) instead of asserting it in prose. classifyConfidence previously had no runtime caller. HIGH means 'trusted provider'; --verified raises web results to MEDIUM (verification semantics documented in ADR-0656). +4 behavioral tests.

Addresses review by @davesienkowski on #664 (W3). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): close Codex adversarial-review findings (path-traversal, version-age, malformed-cache)

HIGH: research key must be 64-hex sha256 (isValidResearchKey) + resolved-path containment check in put/get + CLI validation -> blocks '../../x' arbitrary-file-write. HIGH: package legitimacy now derives publishedAt from the REQUESTED version (npm time[version], PyPI releases[version] upload_time, crates versions[].created_at) so a new malicious version of an old package can't inherit old age and evade 'too-new'. MEDIUM: getResearch validates entry shape (finite fetched_at + positive ttl + required fields) -> malformed cache entry is a miss, not fresh-forever. +regression tests (RED-first); 111 green.

Codex adversarial review (required pre-PR gate). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): close code-review correctness findings

(1) package-legitimacy CLI now rejects unknown --flags instead of silently consuming the following package as a flag value; only --ecosystem takes a value. (2) crates recent_downloads (90-day) normalized to a weekly figure before the minWeeklyDownloads threshold (was ~13x too lenient). (3) research-plan --input validates parsed JSON is an object with an Array questions before destructuring -> clean usage error instead of an uncaught TypeError on null/bad input. (4) research-store put rejects a flag value that is itself a --flag (no more storing '--source' as content). (5) planResearch skips questions whose text is not a non-empty string instead of emitting question:undefined. +13 RED-first regression tests; 143 green.

Code-review gate. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(#657): extract researcher documentation_lookup to shared @-reference

6 researcher agents carried a near-duplicate <documentation_lookup> block; consolidate into gsd-core/references/research-documentation-lookup.md (@-included). Unifies the ctx7 CLI fallback to the safer 'command -v ctx7' guard (drops silent 'npx --yes ctx7@latest' execution in 5 agents). Behavior-preserving dedup; inventory 63->64 references. Phase A of the agent collapse.

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(#657): extract researcher philosophy + verification-protocol to shared @-references

philosophy and the pitfalls+pre-submission-checklist common-core were near-duplicated in project/phase researchers; consolidate into gsd-core/references/research-{philosophy,verification-protocol}.md (@-included). phase-researcher keeps its 3 extra checklist items inline. Pre-submission domains checklist made agent-agnostic so project-researcher doesn't lose features/architecture coverage. Write-contract intentionally left inline (bug-214 tests assert it verbatim). Inventory 64->66 refs. Behavior-preserving. Phase A.

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#657): wire gsd-phase-researcher to the Research seam (Phase B / S1)

The phase researcher now CALLS the code seam instead of carrying inline mechanics: provider waterfall -> 'gsd-tools query research-plan' (+ research-store put to cache digests); confidence-tier prose -> 'gsd-tools query classify-confidence'; slopcheck pip-install protocol -> 'gsd-tools query package-legitimacy check'. This makes the Research module a real runtime consumer (validates the seam end-to-end, addresses reviewer S1) and removes the duplicated waterfall/confidence/slopcheck prose. RESEARCH.md output contract, commit step, structured returns, and Phase-A @-includes unchanged. package-legitimacy-gate.test.cjs rewritten prose-grep -> behavioral (asserts the seam invocation).

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#657): wire gsd-project-researcher to the seam + add tavily/ref/jina MCP tools (Phase C.1)

project-researcher now calls gsd-tools query research-plan / classify-confidence (+ research-store put) instead of the inline provider waterfall + confidence-tier prose (mirrors the phase-researcher rewire; no package-legitimacy — phase-only). Output contract (STACK/FEATURES/ARCHITECTURE/PITFALLS/SUMMARY.md + sections, no-commit, structured returns, Phase-A @-includes) unchanged. Adds mcp__tavily/ref/jina__* to the project/phase/ui researcher tools frontmatter (Balanced provider set) so install.js MCP mapping (C.2) has a consumer.

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#657): cover tavily/ref/jina MCP install handling + frontmatter parity guard (Phase C.2)

Investigation: exa/firecrawl have no explicit per-runtime tool-mapping — every mcp__<server>__* except context7 rides the generic passthrough (Copilot lowercases; OpenCode/Cursor/Windsurf/Augment keep as-is; Gemini auto-discovers). tavily/ref/jina are handled identically, no install path broken. Added 12 copilot-install passthrough tests + a mcp-tool-inheritance parity guard (tavily co-declared with exa, jina with firecrawl, ref present across the 3 web researchers) so the MCP set can't drift. No io.github registry ids invented (none sourceable in-repo); documented as a follow-up. 488 tests green.

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#657): profiles as source of truth for researcher agents + drift-guard (Phase C.3)

scripts/research-profiles.cjs declares each of the 7 researcher agents' identity + contract (name, description, color, tools, required @-includes, required gsd-tools seam calls, output-contract markers). scripts/gen-research-agents.cjs --check validates every committed agent against its profile; --write regenerates ONLY the frontmatter from profiles (body untouched) and is a verified no-op against the current agents (zero diff = fidelity). tests/research-agent-profiles.test.cjs is the DEFECT.GENERATIVE-FIX drift guard. Design note: profiles govern the generatable/contract surface rather than destructively regenerating the disparate operational prose bodies (those were deduped via @-includes in Phase A). scripts/ is not inventoried (no inventory change).

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#657): complete agent provider-dispatch + parity guard; align legitimacy field; validate profiles

Adversarial-review findings: (HIGH) the seam-wired agents' Step-C dispatch only mapped 6 providers, so a planResearch result of jina/ref/perplexity/brave (reachable via the waterfall fallbacks) had no handling -> agent stall; completed both agents' dispatch to all 9 PROVIDER_WATERFALL ids + a catch-all, and added a parity test asserting agent dispatch stays in sync with research-provider PROVIDER_WATERFALL (DEFECT.GENERATIVE-FIX). (MEDIUM) phase-researcher package-legitimacy JSON example used 'package' but the module returns 'name' -> aligned. (LOW) gen-research-agents checkAgent now returns a clear failure for a malformed profile instead of throwing. +parity/validation tests (RED-first).

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): make classifyConfidence verification-evidence-driven (W3)

Confidence conflated provider authority with claim verification — context7/ref
stamped HIGH purely by provider identity, and the only verification lever was a
self-set --verified flag. Split into two axes: provider authority (static) +
verification evidence (code-computed). HIGH now requires ground-truth
corroboration (legitimacyVerdict OK), independent of provider; authority alone
caps at MEDIUM; SLOP caps at LOW; the self-reported --verified is demoted to a
MEDIUM-only web lever. HIGH = corroborated-against-authoritative-source, not a
correctness guarantee. Adds --legitimacy-verdict to the classify-confidence CLI;
updates CONTEXT.md predicate + ADR-0656 (tier set unchanged, ADR-consistent).

Addresses davesienkowski's W3 review on #664.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#656): bind classify-confidence verdict to code, closing CLI self-grading

Adversarial review found the new --legitimacy-verdict flag was caller-supplied,
so an agent could self-assert OK->HIGH without any real legitimacy check —
reintroducing the exact self-grading hole W3 closes. Remove the free flag; the
CLI now computes the verdict via checkPackages only when --package/--ecosystem
is given (code-computed, not agent-asserted). Update the stale CLI test
(context7 alone -> MEDIUM) and extend the property test to vary legitimacyVerdict.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 17:58:48 -04:00
Tom Boucher
ecc42cefc6 fix(#669): /gsd-review --cursor actually invokes cursor-agent (#686)
* fix(#669): /gsd-review --cursor actually invokes cursor-agent

The Cursor reviewer branch in review.md never ran the agent:
- detection probed `cursor` (the IDE launcher) instead of the headless
  `cursor-agent` binary
- the invocation used the two-token `cursor agent` (the IDE treats `agent`
  as a file-path argument, so the agent never starts)
- the prompt was piped via stdin, but `cursor-agent -p` reads the prompt
  from a command-line argument, and `2>/dev/null` hid the empty result

Probe `cursor-agent`; invoke `cursor-agent -p --mode ask --trust
--output-format text` with the prompt passed as a file-path-reference
argument (avoids the OS arg-length limit on large prompts); capture stderr
so failures are diagnosable. Invert tests/cursor-reviewer.test.cjs to assert
the corrected contract, with negative guards against the two-token form and
the stdin pipe. The sibling `agy` reviewer already used the argument form.

Closes #669

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#669): set changeset pr number to 686

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 10:36:19 -04:00
Tom Boucher
afe61f015c fix(#687): bound agy print mode with its native --print-timeout (#689)
`/gsd-review --agy` hung indefinitely on large prompts. agy's print mode runs the
full tool-enabled agent, and on a big, file-path-rich prompt its agentic Cascade
loops on the code_search/grep tool and never converges; the transcript fallback
only runs after agy exits, so it can't recover a run that never exits.

The agy CLI exposes no per-tool deny (that lives in the Antigravity SDK), but it
does expose --print-timeout — agy's native print-mode cap. Pass it explicitly so a
stalled run self-terminates through the tool's own mechanism; a non-zero exit
discards any partial output so the existing transcript fallback / "review failed"
stub take over. Adds a regression test.

Closes #687

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 10:35:27 -04:00
Tom Boucher
0e259a589c fix(#705): route hardcoded $HOME gsd-tools invocations in agents/commands through gsd_run (#707)
* fix(#705): route hardcoded $HOME gsd-tools invocations in agents/commands through gsd_run

The hardcoded `node "$HOME/.claude/gsd-core/bin/gsd-tools.cjs" <cmd>` form
(fixed for workflows in #621/#637) survived in agent/command surfaces and
misresolves on global/shim-only installs. Route every agent-executed
invocation through the resolved `gsd_run` launcher in gsd-phase-researcher,
gsd-planner (load_graph_context extracted to a shared reference to stay under
the planner size budget), import, and graphify. Add a regression guard over
agents/ + commands/ + gsd-core/references/ bash blocks. User-facing display
messages and docs are intentionally left untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#705): use repo changeset fragment format (type: Fixed, pr: 707)

The hand-written fragment used the standard changesets package format
(package: bump) which lacks the type:/pr: frontmatter the repo's
docs-required lint consumes (fail_malformed_fragment / missing_type).
Regenerated via scripts/changeset/new.cjs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 09:26:29 -04:00
Joe
0fbce0fbc7 fix(#637): route 3 more workflows through gsd_run launcher (hardcoded $HOME sweep) (#642)
* fix(#637): route 3 more workflows through gsd_run launcher (hardcoded $HOME sweep)

The hardcoded `node "$HOME/.claude/gsd-core/bin/gsd-tools.cjs"` invocation form
fixed in plan-phase.md (#621) survived in three more workflows. Same bug class:
on a global/shim-only install with no project-local runtime, the hardcoded path
can miss a working install, so the step reports the tool "not found" instead of
resolving it via the launcher. #3668 introduced gsd_run resolution; these sites
were missed.

- plan-review-convergence.md: convert the 3 hardcoded invocations (init,
  roadmap get-phase, state planned-phase) to gsd_run. File already carried the
  canonical preamble (first gsd_run is the earlier convergence-enabled check).
- ingest-docs.md, spec-phase.md: convert their hardcoded invocations to gsd_run
  and inject the canonical launcher preamble via
  `node scripts/sync-runtime-launcher.cjs` (these files previously had no
  gsd_run and no preamble). The injected preamble is byte-equal to
  _runtime-launcher.snippet.sh and precedes the first gsd_run call, per
  runtime-launcher-parity invariant (B).
- Add tests/bug-637-workflow-no-hardcoded-home-tool.test.cjs: repo-wide
  regression guard asserting NO workflow .md invokes gsd-tools via a hardcoded
  $HOME path. Generalizes the plan-phase-only guard from #621 — the parity test
  guards retired $GSD_SDK / bare /gsd-tools tokens but not this form, which is
  how it survived across four files. Fails on the pre-fix files, passes after.

runtime-launcher-parity 7/7; full unit suite green (3477 pass / 0 fail).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#637): add changeset fragment for PR #642

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#637): update stale bug-2801 assertion to expect gsd_run

bug-2801 pinned ingest-docs.md to the hardcoded node "$HOME/.../gsd-tools.cjs" init form, which #637 replaces with the gsd_run launcher. Flip the assertion to expect gsd_run init ingest-docs; the bare-gsd-tools rejection and CLI-handler tests are unchanged, and bug-637's repo-wide guard now owns the no-hardcoded-$HOME invariant.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-05 08:57:20 -04:00
Tom Boucher
3042b79178 fix(#690): promote CHANGELOG 1.3.x + gate release-notes promotion (#694)
* fix(#690): promote CHANGELOG 1.3.x + gate release-notes promotion

/gsd:update showed an empty "What's New" preview after updating to 1.3.1
because CHANGELOG.md's 1.3.x content was never promoted out of [Unreleased]
into dated sections, so `scripts/changeset/cli.cjs extract` returned exit 2
("no releases in range").

- CHANGELOG.md: split [Unreleased] into dated [1.3.0] and [1.3.1] sections
  (1.3.1 = hono advisory bump + installer-migration checksum self-heal, #670;
  1.3.0 = the feature release), restoring an empty [Unreleased].
- scripts/changeset/cli.cjs: new `verify` subcommand that exits non-zero when
  CHANGELOG has no dated `## [x.y.z]` heading for a version; hoist shared
  stripV/resolveChangelogPath helpers used by extract + verify.
- .github/workflows/release.yml: gate the finalize job on `verify` (after the
  build, before tag/publish) so an unpromoted CHANGELOG can never ship again.
- gsd-core/workflows/update.md: move `rm -f $CHANGELOG_TMP` after the
  human-readable extract re-run so the preview no longer degrades to
  "(changelog unavailable)".
- tests: regression guard for the 1.3.x headings + extract range + verify
  command coverage (present/absent/undated/v-prefixed/--json/prerelease).

Closes #690

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#690): add changeset fragment for #694

Fixed-type fragment for the user-facing /gsd:update preview fix and the
release-notes promotion gate.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 22:48:21 -04:00
Tom Boucher
0d97532a57 fix(#586): make ship PHASE_VERIFICATION_INCOMPLETE actionable, drop dead pass status branch (#650)
* fix(#586): make ship PHASE_VERIFICATION_INCOMPLETE actionable, drop dead `pass` arm

The ship preflight gate blocked with PHASE_VERIFICATION_INCOMPLETE but named no
next step, and accepted a `pass` status the verifier never emits. Capture the
verification status and route per value (gaps_found / human_needed / missing),
mirroring execute-phase's status table; accept only `passed`.

Closes #586

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#586): backfill changeset PR number 650

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#586): scope ship verification status to frontmatter only

Codex adversarial review of PR #650 flagged that the status gate grepped
`^status:` over the entire VERIFICATION.md, so a `status:` line in the report
body (a code block / copied artifact) concatenates into a non-matching value and
blocks a genuinely-passed phase with the wrong next action. Restrict extraction
to the leading YAML frontmatter block, first match only. Adds a behavioral
regression test that runs the gate's own bash pipeline against a passing report
whose body contains decoy `status:` lines.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#586): drop manual PR ref from changeset body

The changelog renderer auto-appends `(#<pr>)` from the fragment's pr: field
(scripts/changeset/serialize.cjs, github-release-notes.cjs). The manual trailing
`(#586)` produced a double, mismatched ref (issue #586 + auto PR #650); remove it
to match the sibling-fragment convention.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#586): make ship-586 bash-fence regex Windows-safe (CRLF)

The behavioral test extracted the gate's bash block with /```bash\n.../ — a
literal \n that fails to match Windows CRLF checkouts and trips the
windows-test-parity-guard (fenceRegexLiteralNewline). Use ```bash\r?\n and
normalize the captured block to LF before running it. Full unit suite: 0 fail.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#586): run ship-586 bash-pipeline tests on POSIX only

On Windows CI the behavioral tests failed: git-bash is present (so the old
hasBash guard ran them) but receives a Windows-style tmpdir path it cannot glob,
so extraction returned empty. The extraction logic is platform-independent and
the gate's bash only runs in a POSIX workflow context, so skip the pipeline
execution on win32. POSIX (macOS/Linux) still runs and asserts it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 20:51:53 -04:00
Tom Boucher
b38ae41243 fix(#619): resolve gsd-tools via runtime shim in codebase-drift-gate (#645)
* fix(#619): resolve gsd-tools via runtime shim in codebase-drift-gate

The post-execution drift check ran the bare PATH binary
`gsd-tools verify codebase-drift`. On a shim-only install (gsd-tools.cjs
present, `gsd-tools` not on PATH) that exits 127, `2>/dev/null` hides it,
and the `|| echo` fallback marks the gate skipped — so codebase-drift
detection silently never runs. Non-blocking by contract, so nothing
surfaced; it just quietly stopped working.

Resolve gsd-tools through the runtime shim launcher (gsd_run) instead.
The canonical launcher preamble is now defined once in the always-run
drift-check block (the file's first gsd_run block); the conditional
auto-remap block reuses gsd_run from the workflow's shared shell scope,
keeping the file compliant with the single-canonical-preamble parity
invariant (tests/runtime-launcher-parity.test.cjs). This is the same
single-preamble pattern established by discuss-phase (#614). Non-blocking
is preserved for the drift command's internal failures via the unchanged
`|| echo '{"skipped":...}'` fallback.

Scope decision (the issue's open question): workflow step-file bash blocks
share one shell scope, so the preamble is defined once before the first
gsd_run call — matching discuss-phase and enforced by the parity test.

Regression test (bug-619-...): contract assertions (gsd_run not bare
gsd-tools; single preamble in the drift block; fallback intact) plus a
behavioral proof that runs the shipped drift-check block against a
shim-only topology and asserts the shim actually executes where the old
bare-binary form would have skipped. Red→green verified.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#619): add changeset for codebase-drift-gate shim fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-03 08:08:34 -04:00
Tom Boucher
c89197972e fix(#630): pin wave-cleanup to orchestrator root via manifest, not worktree-list first-entry (#643)
* fix(#630): pin wave-cleanup to orchestrator root via manifest, not list first-entry

Follow-up to #590. #590 fixed the dispatch-side orchestrator cwd anchor
(ORCHESTRATOR_WT via git rev-parse --show-toplevel) but the two
wave-cleanup guards still resolved PRIMARY_WT from `git worktree list
--porcelain`'s first entry — always the main checkout. An orchestrator
running from a non-primary (per-phase lane) worktree was therefore cd'd
off its own lane at cleanup, tripping the #3174 branch-drift assertion
(ORCH_BRANCH != EXPECTED_BRANCH) and refusing merge-back — the same
failure #590 set out to fix, surviving on the cleanup side.

Persist the dispatch-time orchestrator root (show-toplevel, captured from
the lane the orchestrator dispatches from) into WAVE_WORKTREE_MANIFEST as
`orchestrator_root`, and resolve PRIMARY_WT from it at both cleanup sites.
The git-worktree-list first entry survives only as a guarded fallback for
pre-#630 manifests. Byte-identical for a primary orchestrator (its root
IS the first entry); unblocks the non-primary-orchestrator topology.

Regression test (bug-630-...): behaviorally proves the pivot by running
the shipped manifest-reader one-liner against a real non-primary-worktree
git topology — it resolves to the lane while first-entry resolves to main
— plus contract assertions. Updates the #3425 worktree-cleanup contract
tests to the new manifest-based resolution.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#630): add changeset for wave-cleanup orchestrator-root fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#630): canonicalize paths with realpathSync.native for Windows 8.3 parity

On Windows the CI runner's os.tmpdir() yields an 8.3 short name (RUNNER~1)
while `git worktree list` reports the long form (runneradmin); plain
realpathSync preserved each input's form, so the first-entry/main sanity
comparison mismatched. Canonicalize both sides (and the reader output)
via fs.realpathSync.native, which reconciles 8.3 and long forms. Test-only;
the shipped manifest reader is unaffected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-03 08:07:16 -04:00
Tom Boucher
2815720aad fix(#621): route plan-phase post-planning-gaps through gsd_run launcher (#635)
* fix(#621): route plan-phase post-planning-gaps through gsd_run launcher

The post-planning-gaps step in gsd-core/workflows/plan-phase.md invoked
gsd-tools via a hardcoded `node "$HOME/.claude/gsd-core/bin/gsd-tools.cjs"`
path twice on one line — for the gap-analysis call and its nested
init.plan-phase phase_req_ids query — bypassing the gsd_run launcher that
every other call in the workflow uses. On non-default install/runtime layouts
(relocated/global installs, non-Claude runtimes) the hardcoded path does not
resolve, so the assistant reported the gap-analysis tool as "not found" and
fell back to a frontmatter-only coverage check even when a working install
existed. #3668 fixed this class earlier in the file but missed this block.

Route both invocations through gsd_run, matching the rest of the workflow.
No hardcoded $HOME gsd-tools path remains in plan-phase.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#621): add changeset for plan-phase gsd_run gap-analysis fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#621): update bug-2851 §13e guard for the gsd_run gap-analysis form

The #621 fix migrates plan-phase.md's post-planning-gaps gap-analysis call
from the hardcoded node "$HOME/.claude/gsd-core/bin/gsd-tools.cjs" form to the
gsd_run launcher (the canonical resolvable form every other call in the file
uses). bug-2851's §13e subtest pinned that line to the absolute-$HOME form and
now asserts stale behavior.

Update the §13e assertion to require `gsd_run gap-analysis` (still rejecting a
regression to the hardcoded $HOME path), retitle it, and note the migration in
the file header. The generic bare-`gsd-tools` sweeper test is unchanged —
gsd_run is a resolvable launcher, not a bare gsd-tools call.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 23:02:25 -04:00
Tom Boucher
463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00