Commit Graph

114 Commits

Author SHA1 Message Date
Tom Boucher
8f2ebbe9bf feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity

Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap).

--gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): backfill changeset PR number (#1996)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): drop Gemini CLI from issue templates (review nit)

Removes the sunset Gemini CLI runtime from the two GitHub issue-template
runtime lists that the removal PR missed, per @davesienkowski's review nit:
- feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise
  request a feature for a runtime GSD no longer supports)
- bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json
  retrieval-help line

Leaves the post-removal templates fully consistent with the Antigravity redirect.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 13:32:51 -04:00
Jeremy McSpadden
e5ef323b15 feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command

Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.

* feat: add /gsd-start smart-entry command

State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.

- src/smart-entry.cts: deterministic situation classifier (no-project,
  paused, blocked, verify-failed, needs-first-phase, planning, executing,
  verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
  ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire  case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
  dispatcher presenting an AskUserQuestion menu (with --text fallback for
  non-Claude runtimes) and dispatching to existing commands. Falls back
  to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
  situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
  (markdown-layer invariants + every emitted command resolves to a real
  slash command).

Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.

* refactor: rename smart-entry command to /gsd:next

Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.

All affected tests (188) pass; lint:ci clean.

* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)

Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.

- detectSignals now reads total_phases + percent from nested progress{}
  first, then scalar fm, then body; current_phase falls back to the
  body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
  body Phase field) covering verify-pending + executing situations.

Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.

* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)

Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.

Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
   layer is down — read .planning/STATE.md directly with the Read tool
   and synthesize a minimal situation + actions menu so /gsd:next stays
   useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
   /gsd:progress as before.

Matches the direct-read resilience the live agent already did by hand.

* docs: add gsd-next skill surface

* chore: trigger no-mistakes validation

* no-mistakes(review): Fix smart-entry phase ordering

* no-mistakes(review): Fix decimal smart-entry phase ordering

* no-mistakes(test): Fix smart-entry next test contracts

* no-mistakes(document): Docs synced for smart entry

* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: shorten next.md description and update golden install parity fixtures

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: trigger no-mistakes validation

* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files

Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.

* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: regenerate golden install parity fixtures for /gsd:next

Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.

* Fix smart-entry verify-failed phase scoping and empty resolve shim step

Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.

* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: recapture all 16 golden fixtures with updated smart-entry.md hash

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: regenerate fixtures + inventory manifest after rebase onto next

Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next

Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).

Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)

The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs

Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo

Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
  recommended action, 1-4 unique-id /gsd:* actions (previously the
  one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout

Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:

- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
  zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
  kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
  consolidation file doing dozens of real installs — 41s even on a fast Mac
  (much worse on the slow Windows I/O path), plus an install-heavy cluster.

Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.

Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 12:18:25 -04:00
Tom Boucher
3c13903dcd feat(#1866): agent-side self-load of configured agent_skills
Each of the 22 consumer agents now self-loads its configured agent_skills
in its mandatory init step, so .planning/config.json agent_skills.<type>
reaches the agent on every runtime — including Cursor and /gsd-autonomous,
where Skill()-delegated workflow bash init did not reliably execute.

- gsd-core/references/agent-skills-bootstrap.md: shared contract
  (query + Read + dedup guard that skips when <agent_skills> is already
  in the prompt, so Claude's orchestrator-side injection never doubles)
- 22 agents/gsd-*.md: one self-load line naming the agent's own type
- gsd-core/workflows/autonomous.md: note that delegated agents self-load
- tests/agent-skills-bootstrap.test.cjs: regression + parity (CONSUMER_AGENTS
  bijection + fast-check property) — Generative-Fix-Divergence guard
- docs: ADR-1866, CONFIGURATION dual-injection How It Works, INVENTORY
  row, Changed changeset

Closes #1866
2026-07-01 20:09:01 -04:00
Tom Boucher
da37986cd0 fix(#1847): resolve standard tier to claude sonnet 5
Point the sonnet/standard tier at Claude Sonnet 5 (`claude-sonnet-5`,
GA 2026-06-30) across the Anthropic-backed runtimes and provider presets,
replacing the superseded `claude-sonnet-4-6`. Mirrors the change into the
CONFIGURATION.md and settings-advanced.md runtime-defaults tables (the
#3229 catalog↔docs parity gate) plus the pt-BR/zh-CN translations, and
updates the tests that pin the old ID. Regenerates the workflow size
baseline for the (smaller) settings-advanced.md.

Scope is Sonnet only — opus/haiku IDs are untouched. Prepared as a 1.6.1
hotfix off the v1.6.0 tag.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 33260555b4)
2026-06-30 22:21:18 -04:00
Tom Boucher
fd576528a7 fix(#1747): register four search-provider keys in the config schema (#1814)
* fix(#1747): register four search-provider keys in the config schema

buildNewProjectConfig emits seven search-provider availability flags and
research-provider.cts providerAvailability() consumes all seven, but only
three were registered in VALID_CONFIG_KEYS (config-schema.manifest.json).
config-loader.cts then printed an 'unknown config key(s)' warning for the
four unregistered keys (tavily_search, ref_search, perplexity, jina) on
every freshly generated .planning/config.json.

Register the four missing keys in the schema manifest and document them
alongside brave/exa/firecrawl in CONFIGURATION.md. Add a regression test
plus a structural drift guard that requires every config-driven
research-provider flag to be in VALID_CONFIG_KEYS, so a future provider
addition cannot silently reintroduce the drift.

* fix(#1747): move regression into owning test file + add changeset

lint-regression-test-names rejects new bug-NNNN-*.test.cjs files; move the
#1747 regression (four provider keys in VALID_CONFIG_KEYS + provider-flag
drift guard) into tests/bug-2530-valid-config-keys.test.cjs, the canonical
home for VALID_CONFIG_KEYS regressions, and delete the standalone file.

Add the missing .changeset fragment — config-schema.manifest.json lives
under gsd-core/ (user-facing), so changeset-lint requires a fragment.

* test(#1747): regenerate golden-install-parity fixtures for schema change

Adding four provider keys to config-schema.manifest.json shifts its shipped
content hash (65dea848 -> 7d398e94); recapture all 16 runtime fixtures via
UPDATE_GOLDEN=1. Each fixture changes exactly one line — the manifest hash.
2026-06-28 22:42:36 -04:00
Tom Boucher
4ced0a64cc feat(#1561): assumption-delta advisory checkpoint (#1767)
* feat(#1561): assumption-delta advisory checkpoint

* chore(#1561): backfill changeset PR number (#1767)

---------

Co-authored-by: review-bot <review-bot@gsd>
2026-06-27 08:43:47 -04:00
Tom Boucher
b0d5ca3379 feat(#1517): support custom reviewer instances for /gsd:review (#1766)
* feat(#1517): support custom reviewer instances for /gsd:review

Add a bounded review.reviewer_instances config surface so one model-capable
adapter (e.g. opencode) can run as several independent reviewer identities in a
single /gsd:review pass. Instances participate only via review.default_reviewers,
expand before built-in slugs, are available iff their cli is detected, and a
non-matching entry is a hard error (typo must be loud). >=2 same-cli instances
emit a shared-adapter caveat in REVIEWS.md. Default path with no instances is
byte-for-byte unchanged.

Single-source instance->cli resolution lives in resolveReviewerSelection /
normalizeReviewerInstances (parity-locked in
tests/review-reviewer-instances.test.cjs). cli validated against
KNOWN_REVIEWER_SLUGS only (never arbitrary shell); model/agent opaque, never
shell-interpolated.

Closes #1517

* chore(#1517): backfill changeset pr:1766

---------

Co-authored-by: review-bot <review-bot@gsd>
2026-06-26 23:35:04 -04:00
Tom Boucher
2b215b4163 feat(#1688): warn on stale model bake for static-frontmatter runtimes (#1692)
* docs(#1650): fix stale opencode install-path claim in core settings

* feat(#1688): warn on stale model bake for static-frontmatter runtimes

* chore(#1688): backfill changeset pr field with real PR number

* test(#1688): make resolveAgentDir assertions use path.join for windows

* docs(#1688): codify windows path-literal-in-assert anti-pattern + align test
2026-06-25 09:12:40 -04:00
Alex V.
a63684c222 enhance(#1577): WebFetch/WebSearch injection isolation + opt-in blocking (#1585)
* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking

Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the
largest untrusted channel) in gsd-read-injection-scanner; shared
untrusted-input-boundary reference @-included by the 8 ingest agents
(randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring);
opt-in security.injection_blocking (default advisory — non-breaking).

arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth).

* fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized

- A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a
  circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already
  in the transcript. The prompt-level data/instruction boundary is the primary control.
- A2: registered security.injection_blocking in the config schema + defaults manifests (default
  false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads.
- A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention).
- A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale).
- A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input.
- Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) +
  drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as
  the noted pre-existing follow-up.

* fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate

The new reference quotes injection phrases ('ignore previous instructions',
'you are now…') as examples agents must NOT comply with, tripping the repo's
own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red
on HEAD). Allowlist it alongside the other security docs (security-model.md,
TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS
scanner test doesn't scan references/, so only the shell gate needed it.

Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15.

* fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer

trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so
no named web-ingress agent is uncovered, keeping the two justified additions
(gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10.
 - gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset.
 - gsd-assumptions-analyzer reads 5-15 codebase source files (external/source-
   document ingress per the boundary), though it has no web tools.
INGEST_AGENTS in the isolation test now asserts all 10; size baselines
regenerated (+60 bytes each, both well under the DEFAULT cap); changeset
reworded 8 -> 10.

Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39.

* docs(#1577): document security.injection_blocking + boundary seam

trek-e Major 2 + Minor:
 - docs/CONFIGURATION.md: add the top-level security.injection_blocking key to
   the Full Schema and a Security Settings subsection, distinguishing it from
   the workflow.security_* namespace; honest circuit-breaker-not-redactor
   framing matching ADR-1577 / security-model.
 - CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry.

Verified: lint:docs ok; config-field-docs + contributor-standards green.

* test(#1577): make read-injection property test git-text, not binary

trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as
degenerate-edge inputs. The NUL is what actually made git classify it binary
(git binary = NUL in first 8K). Replace both with text-safe escapes that keep
the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File
now diffs/blames line-by-line.

Verified: property test 2/2; no NUL/raw-noncharacter bytes remain.

* docs(#1577): align untrusted boundary docs

Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner.

* docs(#1577): align ADR ingest agent count

Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 17:07:23 -04:00
Tom Boucher
35478b615e refactor(#1646): route capability routers through Command Routing Hub per ADR-959 (#1647)
* refactor(#1646): route capability routers through Command Routing Hub per ADR-959

Phase 2 of parent #1641. Converts graphify, intel, and audit command
routers from hand-rolled if/else dispatch to routeHubCommandFamily,
implementing the ADR-959 §III(B) line 75 mandate. The three routers
now share the uniform dispatch shape with the 14 host routers.

src/cjs-command-router-adapter.cts
  * Imported ERROR_REASON from io.cjs.
  * UnknownCommand translation now passes ERROR_REASON.SDK_UNKNOWN_COMMAND
    as the second arg to error() — additive for host routers (their
    existing one-arg error callbacks ignore the second arg), required
    for capability routers whose tests assert reason === 'sdk_unknown_command'
    on the JSON-error envelope.

src/graphify-command-router.cts
  * Replaced 4-branch if/else with routeHubCommandFamily + handlers map.
  * Validation handlers (missing term, missing/invalid --budget) now
    return makeInvalidArgs(arg, reason, ERROR_REASON.USAGE) Results
    instead of calling error() directly (Q2=C, Q4=ii from grilling).
  * Success handlers keep direct output() calls.
  * Subcommands array is alphabetical for byte-identical 'Available:'
    text in the unknown-subcommand message.
  * The unknown-subcommand path is now owned by the Hub's manifest
    check (the adapter passes SDK_UNKNOWN_COMMAND).

src/intel-command-router.cts
  * Replaced 9-branch if/else with routeHubCommandFamily + handlers map.
  * Validation handlers (missing term, missing filePath for patch-meta
    and extract-exports) return makeInvalidArgs Results.
  * Preserved the timeAgo mutation in the non-raw status handler.
  * Preserved the lazy require('./intel.cjs') inside the route function.

src/audit-command-router.cts
  * routeAuditUat: routes through the Hub with a synthetic 'run'
    defaultSubcommand (no real subcommands). Gives uniform observability.
  * routeAuditOpen: captures --json in a closure, strips it from args
    before Hub dispatch (so it isn't mistaken for a subcommand by the
    manifest check), then branches on wantJson inside the handler to
    preserve the formatAuditReport success-path quirk.

docs/CONFIGURATION.md
  * Observability section: noted capability commands (graphify, intel,
    audit-uat, audit-open) now emit DispatchEvent records since #1646.

.changeset/capability-routers-via-hub.md
  * Changed fragment describing the user-visible audit-trail expansion.
    pr:0 placeholder will be backfilled after gh pr create returns the
    real PR number (DEFECT.CHANGESET-PR-FIELD-DRIFT).

Verification
  * graphify cutover tests: 119/119 pass (all unit, dispatch, behavior,
    error path, JSON-errors, and registry assertions)
  * intel cutover tests: 39/39 pass
  * audit cutover tests: 24/24 pass
  * bug-974-graphify-budget-missing-value regression test: pass
  * npm run test:unit (full suite): 2384 tests, 0 fail
  * gsd-test-summary on docker: outcome=passed, 0 failures
    (RULESET.PR-FLOW.docker-before-push)

JSON-error envelope parity verified byte-identical: reason values
('usage', 'sdk_unknown_command') and message texts are preserved
across all three routers' error paths.

* chore(#1646): backfill changeset pr: 1647 (DEFECT.CHANGESET-PR-FIELD-DRIFT)
2026-06-23 23:21:06 -04:00
Tom Boucher
207d8f1697 fix(#1626): make the security gate severity-aware via per-threat severity (#1635)
workflow.security_block_on was documented as the minimum threat severity
that blocks advancement, but threats carried no severity and the auditor's
threats_open count (the SECURITY.md gate field) counted every open threat
regardless of severity — so the threshold had no effect, and the auditor's
block_on vocabulary (open/unregistered/none) did not even match the config
enum (critical/high/medium/low/none).

- planner: add a Severity column to the STRIDE threat register; assign
  severity per threat.
- auditor: read severity; reconcile the <config> block_on domain to the
  severity enum; redefine threats_open as the count of OPEN threats whose
  severity is at or above block_on (none => 0). Below-threshold opens are
  reported as non-blocking and excluded from threats_open.
- SECURITY.md template + planning-config.md reconciled.

No gate-check site changed: threats_open == 0 stays the gate everywhere;
only its computation is now severity-filtered.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 19:04:13 -04:00
Dave
d1f7ba82f2 feat(#1592): add plan:pre codebase-drift pre-check before planner runs
Add a non-blocking, warn-only codebase-drift gate at plan:pre so a stale
STRUCTURE.md is surfaced before /gsd:plan-phase spawns the planner, instead
of being discovered mid-execution by the existing execute:wave:post gate.
Gated on a dedicated workflow.plan_drift_precheck toggle (default true),
independent of schema_drift_gate. Never blocks planning, never spawns the
mapper agent at plan time.

Review feedback (#1595):
- Use a documented conventional-commit type (feat, not enhance) per
  CONTRIBUTING.md / gsd-validate-commit.sh.
- Normalize the plan_drift_precheck command references to the colon prose
  form (/gsd:plan-phase, /gsd:map-codebase) to match plan-phase.md §5.65;
  registry regenerated from capability.json.
- Make the test temp dirs hermetic: drain mkdtemp dirs in an after() hook
  via the helpers.cleanup() budget (local/no-raw-rmsync-in-tests-compliant).

Closes #1592

Claude-Session: https://claude.ai/code/session_016JBiXEAofvB3prJim29XMS
2026-06-22 20:30:59 -04:00
Tom Boucher
2c718bf972 fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs (#1537)
* fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs

Generalizes the Codex-only #1515/#1519 fix to every non-Claude runtime, and
wires it into the real install path (where it was previously dead-on-arrival).

Root causes:
1. The runtime-default stamping lived only in `_applyRuntimeRewrites`, but the
   installer emits `gsd-core/workflows/*.md` via `copyWithPathReplacement`, which
   never calls it — so a real `--codex`/`--cursor`/etc. install emitted
   `--default claude` and worktrees-on. RUNTIME mis-resolved to claude and the
   workflow ran executors unisolated against the main checkout. (#1515/#1519 were
   also dead-on-arrival in real installs; this repairs them.)
2. Only `case 'codex'` was stamped; every other non-Claude runtime kept the
   Claude default.

Fix:
- New `_stampNonClaudeRuntimeDefaults(content, runtime)` (single shared helper)
  stamps `--default <runtime>` + `use_worktrees=false` for every `runtime !=
  claude`; called from both `_applyRuntimeRewrites` and, crucially,
  `copyWithPathReplacement` in bin/install.js (the real workflow emit path).
- Generalize the fail-closed worktree guard `= codex` -> `!= claude` in
  execute-phase/quick/diagnose-issues (worktree isolation is Claude-Code-only).
- Flip manager/autonomous inline-vs-background gating to `codex -> background,
  everything-else -> inline` (research: only Codex can background-nest the
  pipeline's subagents; all others run inline, which they support).

Worktree-capability determination is research-backed (official docs for all 14
non-Claude runtimes: none honor GSD's isolation="worktree" mechanism, only Codex
background-nests). New end-to-end real-install test asserts the EMITTED workflow
is stamped — the regression guard that would have caught the dead-on-arrival bug.

Closes #1521

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

* chore(#1521): backfill changeset PR number (#1537)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 13:48:47 -04:00
Tom Boucher
2436b76980 fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees (#1519)
* fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees

A Codex install with a runtime-neutral .planning/config.json resolved
RUNTIME=claude and enabled git worktree isolation, which Codex's
spawn_agent cannot honor. Two root causes:

1. Workflows read `config-get runtime` / `config-get workflow.use_worktrees`
   without `--raw`, so config-get's JSON-quoted output ("codex") was captured
   verbatim into the bash var and broke every `[ "$RUNTIME" = ... ]` check —
   the Codex fail-closed guard was dead even when runtime:codex was explicit,
   and Claude's own worktree degrade-check was dead too. Add `--raw` to those
   reads across execute-phase, autonomous, manager, diagnose-issues, quick.

2. The conversion engine emitted `--default claude` for every runtime. Stamp
   the codex-emitted workflows to `--default codex` (runtime) and
   `--default false` (use_worktrees) in _applyRuntimeRewrites case 'codex', so
   a neutral config on a Codex install resolves runtime=codex / worktrees off.

Also extend the Codex fail-closed worktree guard to quick.md and
diagnose-issues.md (they spawned isolation="worktree" with no runtime guard).

Regression test asserts source<->engine parity across all five workflows
(DEFECT.GENERATIVE-FIX) plus fast-check property coverage of the stamping.

Closes #1515

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

* chore(#1515): backfill changeset PR number (#1519)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 11:56:59 -04:00
Tom Boucher
a77f3c4b3a docs(#1452): document workflow.context_guard_mode in CONFIGURATION.md
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 18:12:03 -04:00
Tom Boucher
c330f70f65 feat(#1494): add workflow.mvp_mode to VALID_CONFIG_KEYS; document code_review_command and plan_chunked in planning-config.md (#1500)
* feat(#1494): add workflow.mvp_mode to VALID_CONFIG_KEYS; document code_review_command, plan_chunked, mvp_mode in planning-config.md

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: backfill PR number 1500 in changeset

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 15:32:07 -04:00
Tom Boucher
0d56f544d2 feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation (#1458)
* feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation

ADR-1244 Phase 6. Adds the capability matrix reference, generated FROM the committed registry so it
can never drift from the actual capability set:
- scripts/gen-capability-matrix.cjs (--write / --check); --check is a CI drift guard.
- tests/capability-matrix-sync.test.cjs (4 tests): drift guard, buildMatrix==committed, every cap
  present, no placeholders.
- docs/reference/capability-matrix.md regenerated from the registry (release-stable: shows engines.gsd,
  omits the lockstep per-cap version that would churn the file every release).
- Consolidated the duplicate trust-model doc: deleted docs/explanation/the-capability-trust-model.md,
  merged its content into capability-trust-model.md, redirected ~10 references; no stale links remain.
- Diataxis verification (now that gsd capability is a real command): corrected the matrix's third-party
  section — the matrix is the first-party catalogue; the overlay-aware view of installed third-party
  capabilities is 'gsd capability list', not this generated file.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1435): Added changeset for the capability matrix reference

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1435): address code-review — non-vacuous matrix test + generator polish

- capability-matrix-sync.test.cjs: assert the 'security registers a ship:pre gate' precondition
  unconditionally so the extension-point check can never degrade to a vacuous pass on registry drift.
- gen-capability-matrix.cjs: warn (stderr) on an unknown loop point at generation time; rename
  enginesOf -> fmtEngines for consistency with the other fmt* helpers (output unchanged).
- capability-trust-model.md: point the two how-to links at the real files
  (import-a-capability-from-a-url.md, version-a-capability.md) instead of the bare directory.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1435): backfill changeset PR number → #1458

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 12:45:26 -04:00
Tom Boucher
9219af3360 feat(#1433): capability trust gate + upgrade/compat (ADR-1244 Phase 4) (#1449)
ADR-1244 Phase 4 (D5 trust + D6 upgrade/compat). capability-trust.cjs (disclosure/consent, strict_known_registries, engines+compatVersions, reserved namespace) + capability-lifecycle.cjs (install/upgrade/remove/reconcile; ledger-as-commit-point _pending intent; atomic stage-then-swap; surgical marker-isolated shared-edit strip; owner-token lock) + capability-source promote/skipEnginesGate seams + loader pending-skip + config keys. No sandbox re-derived (consent+integrity+reversibility). 6 Codex adversarial rounds + /security-review (no HIGH) + /code-review; gsd-test green both platforms; CI green. Phase 5 (#1434) wires the CLI dispatch.

Closes #1433.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 21:40:37 -04:00
Tom Boucher
353f63d170 feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) (#1440)
* feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2)

Promote the registry from a frozen data file to loadRegistry({includeInstalled}),
composing the first-party registry with a validated installed overlay (ADR-1244 D2):

- Extract the conformance validator to a shared runtime-callable module
  (gsd-core/bin/lib/capability-validator.cjs); the generator re-exports it
  verbatim, guarded by a generative-parity test (no build-time/runtime drift).
- capability-loader.cts: loadRegistry({includeInstalled}) composes first-party
  ∪ validated overlay from $GSD_HOME/.gsd/capabilities (global) and
  <root>/.gsd/capabilities (project) via the canonical buildRegistry. First-party
  always wins (id/skill/agent/config/command-family + reserved gsd-/anthropic-
  prefixes); full merged-set cross-capability validation; engines.gsd load-time
  re-gate (skip-with-warning); gate-kind capabilities FAIL CLOSED; fragment-path
  escapes rejected.
- semverSatisfies (hand-written, no dep) for the engines.gsd gate, fail-closed.
- Wire surface/state + loop to the overlay; loop injects a blocking gate for each
  skipped gate-kind overlay (fail-closed).
- cwd-aware overlay config-key federation: config-loader _federatedConfigSchema(cwd)
  + config-schema isValidConfigKey(key, cwd) compose the overlay per loadConfig/
  config-set call (never eager at module load, never wrong-cwd); first-party path
  unchanged with no cwd.
- run-tests.cjs sandboxes GSD_HOME (idempotent — nested spawns reuse it) for test
  hermeticity; capability-loader.cjs git+eslint-ignored (tsc artifact);
  capability-validator.cjs stays linted (#551 migration coverage).

Closes #1431

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1431): add changeset for runtime capability registry overlay

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1431): kill config-schema cwd-aware federation mutants (Stryker ≥52)

The cwd-aware overlay config-key federation added to config-schema.cts
(_capabilityConfigSchema(cwd) + isCapabilityConfigKey/isValidConfigKey cwd
threading) introduced mutable surface uncovered by config-schema's mutation
test set, dropping its score to 39.58% (below the 52 break threshold). Add a
real-overlay-fixture describe block exercising every branch (cwd guard, overlay
loadRegistry, found-branch, first-party fallback, cwd threading); local Stryker
score 39.58% -> 77.08%.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 14:41:20 -04:00
Tom Boucher
137760a655 fix(#1296): align config docs/prompts/schema with consumers (#1299)
* fix(#1296): align config docs/prompts/schema with consumers

The user-facing config surface disagreed with what the consumers actually do
(subset of the #1216 audit). No runtime consumption behavior changes.

- workflow.subagent_timeout: settings-advanced.md prompt + docs/CONFIGURATION.md
  said "seconds (default 600)" but the consumer (map-codebase.md) uses
  milliseconds (default 300000). Relabeled all four spots in settings-advanced.md
  (prompt, parse-default list, example, confirmation table) + the CONFIGURATION.md
  row.
- review.models.<cli>: settings-integrations.md, docs/CONFIGURATION.md (Integration
  Settings), and docs/CLI-TOOLS.md documented a shell command, but review.md injects
  the value into a --model/-m flag. Relabeled to a bare model id and reconciled the
  contradictory CONFIGURATION.md sections.
- workflow.test_command + workflow.build_command: consumed via config-get
  (test_command in verify-phase/execute-phase/audit-fix/post-merge-gate;
  build_command in post-merge-gate) and documented, but absent from validKeys so
  `config set` rejected them. Registered both in config-schema.manifest.json and
  documented them in references/planning-config.md (overview + complete reference).

Regression tests: behavioral config-set tests (tests/config.test.cjs) + doc-parity
content guards (tests/config-field-docs.test.cjs).

Deferred to other #1216 clusters: security-gate wiring, search_gitignored wiring,
mvp_mode, source_grounding_authority labeling, and config-set enum enforcement.

Closes #1296
Refs #1216

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(changeset): Fixed fragment for #1296 config-surface alignment

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 19:57:11 -04:00
Tom Boucher
cf68841220 enh(#1243): consume Claude plugin-provided skills in agent_skills (epic #1258 Phase B) (#1261)
* feat(#1243): consume Claude plugin-provided skills via native Skill-tool directive + grant Skill to agent_skills-consumer agents

- Relax global skill name validation to accept namespaced form `^[A-Za-z0-9_-]+(:[A-Za-z0-9_-]+)*$`
- Namespaced names (containing colon) on claude runtime emit a Skill-tool load directive instead of a @-include line
- Namespaced names on non-claude runtimes are skipped with a warning
- Bare unresolved names retain existing warn-and-skip behavior (no promotion to directive)
- Grant `Skill` tool to all 22 agent_skills consumer agents; 5 generated agents updated via research-profiles.cjs + regen, 17 hand-authored agents edited directly
- Add 16 TDD tests in describe('bug #1243') covering happy/mixed/precedence/negative/cross-runtime/regression/grant cases

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#1243): document plugin-provided skills in agent_skills

Update the Agent Skills Injection reference in CONFIGURATION.md with
the three entry forms (project-relative, global:<name>,
global:<plugin>:<skill>), the Claude-only runtime behaviour of the
namespaced form and the warn-skip on other runtimes, the plugin
pre-install prerequisite, and the consumer-agent Skill tool grant.

Add docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md with a
step-by-step guide for installing the plugin, locating the namespaced
skill name, wiring it into agent_skills, and verifying injection.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1243): align agent_skills docs with emitted block format + mixed-block regression test (code-review)

- Replace two-section mixed-block example (bogus "Load these plugin-provided skills using the Skill tool:" header) with the actual single-section inline format in CONFIGURATION.md and docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md
- Fix quoted warning text in how-to doc to exactly match the emitted string: [agent-skills] WARNING: Plugin-namespaced skill "global:<name>" requires a Skill-tool-capable runtime (claude) — skipping on runtime "<runtime>"
- Replace phantom agent slugs (gsd-checker, gsd-researcher, gsd-advisor, gsd-synthesizer) in CONFIGURATION.md Supported Agent Types with real agents/gsd-*.md examples (gsd-plan-checker, gsd-phase-researcher, gsd-code-reviewer, gsd-ui-auditor, gsd-research-synthesizer)
- Add byte-identical mixed-block regression test: one path-resolvable global skill + one plugin-namespaced skill on claude runtime → asserts r.ir.block === single-section interleaved block, no secondary header

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1243): regenerate agent-size baseline for the Skill-tool grant

The 22 agent_skills-consumer agents each grew +7 bytes from adding `Skill`
to their tools list; refresh the committed per-agent size baseline (#1074 guard).

* chore(#1243): add Added changeset fragment

* fix(#1243): traceable allow-test-rule ref + separator-agnostic byte-identical tests (CI)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-15 00:49:12 -04:00
Tom Boucher
a375c4b354 feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) (#1201)
* feat(#956): add MemPalace memory capability (ADR-857 feature plug-in)

Adds an opt-in, default-resilient ADR-857 feature capability that wires
MemPalace (local-first memory: MCP server + CLI) into the GSD loop:
deliberate recall before discuss/plan and verbatim + temporal-KG capture
at phase boundaries. Three memory modes (augment default; kg_backend and
replace forward-declared). Master gate mempalace.enabled (default off);
every hook onError:skip, zero gates; absent/disabled MemPalace => loop
unchanged. Transport is rendered-markdown only — MemPalace runs
out-of-process, no third-party code in gsd-core (ADR-857 §7).

Capability: capabilities/mempalace/ (manifest + 2 fragments), skills
commands/gsd/mempalace-{recall,capture}.md, agent
agents/gsd-mempalace-curator.md. Registration: ns-context router,
utility cluster, KNOWN_SKILLS, help full.md, model-catalog, copilot
install list, size baselines; regenerated capability-registry +
inventory manifest. ship:post wired into ship.md (wire-on-demand).

HELD on #1196: this capability also declares hooks at discuss:pre and
discuss:post, which are structurally un-wireable until the host-loop
conformance model covers the discuss phase (discuss-phase.md is not in
HOST_LOOP_FILES). The phase6-capstone-conformance gate therefore fails
on exactly those two orphaned points by design — see #1196. Once #1196
lands, rebase onto next and the gate goes green with no further change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#956): backfill changeset PR number (#1201)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 02:09:10 -04:00
Tom Boucher
44024aa535 fix(#1133): honor model_policy on the claude runtime (forward-port to next) (#1144)
Forward-port of the #1133 hotfix (commit 22f237a5, 1.4.5 hotfix line) onto
next. The hotfix was authored against src/core.cts (v1.4.4); on next the
resolver logic lives in src/model-resolver.cts (ADR-457 extraction, #888),
so the patch is re-applied there rather than cherry-picked.

resolveModelInternal step 2.5 now honors model_policy on the claude runtime:
the policy-resolved full model ID is mapped back to a Claude Code agent alias
via CLAUDE_POLICY_ID_TO_ALIAS (reverse of MODEL_ALIAS_MAP + claude-fable-5 ->
fable). Bare aliases (opus/sonnet/haiku/fable) pass through; an ID with no
Claude alias warns once to stderr (deduped by agentType::policyModel::tier)
and falls back to the configured tier alias. Non-claude runtimes return full
IDs verbatim (unchanged). resolveModelForTier is intentionally unchanged.

The warn-dedupe cache lives in model-resolver.cts; core.cts composes the
exported _resetRuntimeWarningCacheForTests to clear both that cache and the
config-loader warning cache (config-loader cannot import model-resolver --
circular dependency).

Ports the 6 #1133 tests (rewriting the old claude-no-op test that asserted
the bug) plus one added test covering the MODEL_ALIAS_MAP reverse-map path.

Forward-port of #1133

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 19:56:54 -04:00
Tom Boucher
827011b865 fix(#1098): guard generate-claude-md against clobbering hand-crafted files; redirect default to .claude/CLAUDE.md (#1118)
/gsd-new-project wrote a repo-root CLAUDE.md full of broad project docs,
overwriting/diluting a hand-crafted instruction file. --force was parsed but
silently dropped, and nothing guarded an existing non-GSD file.

- Guard: an existing instruction file with no `<!-- GSD:<section>-start` markers
  (hand-crafted) is left untouched; report action:"skipped". --force (now wired
  through CmdGenerateClaudeMdOptions) overwrites intentionally. The marker check
  uses /<!-- GSD:[a-z]+-start/ so a file merely documenting GSD syntax is safe.
- Redirect: the Claude-family default output is now ./.claude/CLAUDE.md (a valid
  auto-loaded project-memory location) instead of repo-root ./CLAUDE.md. Aligned
  across the handler default, config-defaults.manifest.json, buildNewProjectConfig,
  the config template, new-project.md, and cmdGenerateClaudeProfile; advisory
  read-CLAUDE.md hints in plan-phase/quick/profile-user updated. Codex still
  writes AGENTS.md.

Closes #1098

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 16:03:51 -04:00
Tom Boucher
93c5ecd645 feat(#792): add devin-desktop runtime alias for windsurf (#1086)
windsurf now also answers to devin-desktop (CLI --devin-desktop) for the Windsurf→Devin Desktop rebrand; all paths unchanged. The .devin/skills/ workspace migration is split to #1085. Closes #792.
2026-06-11 22:32:49 -04:00
Tom Boucher
9e3b056b15 fix(#779): correct stale model-catalog model IDs verified against live providers (#1047)
Verify-first audit: gemini opus gemini-3-pro→gemini-3.1-pro-preview (undefined in gemini-cli source), codex sonnet gpt-5.3-codex→gpt-5.4 (deprecated per OpenAI); qwen3-coder-next verified valid, unchanged. Adds a regression guard + sourcing note. Closes #779.
2026-06-11 13:36:23 -04:00
Tom Boucher
e4dfa6b9ea fix(#1012): invoke fallow with its real CLI and wire the report normalizer (#1044)
* fix(#1012): invoke fallow with its real CLI and wire the report normalizer

The /gsd-code-review structural pre-pass invoked fallow with flags no published
fallow version accepts (--json, --profile, --stdin-files), so it failed on every
run and degraded silently per REQ-FALLOW-02 — the feature never delivered on any
fallow version. Three compounding defects:

1. Invalid flags. Real fallow audit uses --format json (not --json), -q/--quiet,
   --changed-since/--base for changed-files scoping (no file-list input), and
   --max-crap for thresholds. There is no --profile or --stdin-files.
2. Exit-code handling. fallow audit exits 1 when it FINDS issues (verdict=fail),
   0 when clean. The pre-pass treated any non-zero exit as a crash and discarded
   the output — i.e. it threw away exactly the findings it exists to surface.
   Success is now decided by whether a valid fallow JSON report was produced,
   not by the exit code.
3. Schema mismatch. normalizeFallowReport parsed a fictional top-level schema
   (unusedExports/duplicates/circularDependencies) fallow never shipped, and was
   dead code (the workflow embedded raw JSON; its tests asserted the fictional
   schema, one even calling a non-existent runFallowAudit and passing vacuously).

Fixes: align the invocation to fallow's documented agent-facing pattern; map the
profile preset (minimal/standard/strict) to --max-crap (50/30/15); scope phase
runs via --changed-since with a repo-scope fallback; rewrite the normalizer to
fallow's real schema (dead_code.unused_exports/unused_files/circular_dependencies
+ duplication.clone_groups) and wire it into the workflow so the reviewer
receives normalized findings; replace the fictional-schema fixtures and tests
with real-schema ones and delete the vacuous runFallowAudit test.

Closes #1012

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1012): backfill changeset PR number to 1044

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:36:50 -04:00
Colin Johnson
76f42ddb4b feat(#1014): add Claude Fable 5 model config (#1015) 2026-06-10 20:32:26 -04:00
Jeremy McSpadden
f61b97276e fix(#724): block convergence on actionable review findings (#728)
* fix(#724): block convergence on actionable review findings

* merge: integrate clean next (#936 inline) onto author tip + re-apply cursor fixes (Mode field, REVIEWS.md extraction) and review hardening (#724)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 14:54:15 -04:00
Tom Boucher
29c0a2f5a1 docs(#849): capture 1.4.0 release features across the docs base (#850)
Diataxis review of the 1.4.0 content (52 changesets, multi-runtime maturation
plus native packaging and new flags) against the existing docs base found most
per-feature docs already landed with their PRs. Fill the four remaining gaps,
each in its Diataxis quadrant:

- Reference: FEATURES.md Feature #36 (Multi-Runtime Support) updated in place
  with 1.4.0 additions — native skills emission (Cline/Kilo/OpenCode), new
  slash-command surfaces (CodeBuddy/Augment/Cursor), cross-runtime lifecycle
  hooks for context-headroom tracking, and the Gemini CLI extension package.
- Reference: CONFIGURATION.md gains a dedicated worktree.baseRef entry (values,
  .claude/settings.local.json location, auto-set-on-install behaviour).
- How-to: plan-a-phase.md gains an 'override planning granularity for one phase'
  section for the --granularity flag.
- Explanation: context-engineering.md gains a 'Lifecycle hooks and context
  headroom' section (the why of lifecycle hooks + forked context), cross-linked
  from multi-agent-orchestration.md.

Docs-only; documents already-shipped features, so no changeset required
(docs/ is not in the changeset-lint user-facing prefixes).

Closes #849

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 23:22:15 -04:00
Tom Boucher
f7e902f1cf feat(#52): add agent_skills_security.trusted_global_roots allowlist for global skills (#754)
* feat(#52): add agent_skills_security.trusted_global_roots allowlist

Opt-in allowlist so a global: agent skill whose SKILL.md realpath resolves
outside the default global skills base (e.g. ~/.claude/skills) is accepted
when its real target lies under a user-declared trusted root. Default [] is
byte-identical to prior behavior; the symlink-escape guard is preserved and
simply re-applied against each declared root.

- src/security.cts: loadTrustedGlobalRoots — tilde-expand (~ and ~/), reject
  project-relative and dangerously broad roots (filesystem/UNC root, homedir),
  realpath-canonicalize each root every run and drop non-existent ones.
- src/init.cts: on base-check failure the guard consults the trusted roots
  (hoisted out of the loop); emits a stderr NOTE when a skill is accepted via
  a trusted root so the widened boundary is visible.
- src/core.cts: thread agent_skills_security through loadConfig.
- config-schema.manifest.json: allow the new key path.
- docs/CONFIGURATION.md: document the option and its security model.
- tests/agent-skills.test.cjs: unit + end-to-end CLI coverage (regression,
  feature, negative, broad-root hardening, stderr NOTE).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#52): add changeset fragment for trusted_global_roots (#754)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 01:06:41 -04:00
Tom Boucher
cf8bd3cd5e fix(#683): auto-degrade phase execution to sequential on worktree base mismatch (#749)
* fix(#683): auto-degrade phase execution to sequential on worktree base mismatch

Claude Code forks worktree-isolated executors off the repository default
branch (origin/HEAD), not the orchestrator's HEAD. Running /gsd-execute-phase
on a branch diverged from the default (unmerged milestone/feature branch) left
every executor without the phase's plan files and tripped the
worktree-branch-check guard with `exit 42` — 100% reproducible, all OSes.

- New module src/worktree-base-ref.cts: HEAD-vs-fork-base drift detection
  (origin/HEAD with symbolic-ref fallback) and no-clobber worktree.baseRef
  management, exposed as `worktree base-check` / `worktree set-baseref`.
- execute-phase.md: pre-dispatch, for Claude Code with worktrees enabled,
  auto-degrades the run to sequential on the main tree when a base mismatch
  is detected, recommending worktree.baseRef:"head". The exit-42 guard stays
  as a backstop.
- Installer: fresh local Claude installs set worktree.baseRef:"head" in
  .claude/settings.local.json (no-clobber, respecting an explicit shared
  settings.json value); upgrades print an opt-in notice pointing at
  `gsd-tools worktree set-baseref`.
- Docs: how-to guide, CLI/config reference, planning-config cross-ref.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#683): auto-apply worktree.baseRef on upgrade; gate fresh+upgrade on use_worktrees

Per maintainer direction: on a local Claude Code UPGRADE, set
worktree.baseRef:"head" automatically (no opt-in notice) when the project's
workflow.use_worktrees is enabled, instead of merely printing a remediation
notice. For consistency the FRESH path is now gated the same way: both paths
compute worktrees-enabled once (bounded walk-up read of .planning/config.json,
default enabled unless workflow.use_worktrees === false) and apply the
no-clobber baseRef only when enabled — never overwriting an explicit value in
settings.local.json or a shared settings.json. gsd-tools worktree set-baseref
remains for manual use. Docs + changeset updated; tests hardened (file-exists
assertions, fresh+disabled case, upgrade idempotency).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#683): measure workflow byte-budget on LF, fixing Windows-only CI failure

The workflow-size-budget test failed only on Windows: git checks out the .md
files as CRLF (no eol=lf in .gitattributes) and byteCount used
fs.statSync().size (raw on-disk bytes), counting an extra \r per line. That
inflated execute-phase.md — the XL high-water-mark file pinned near its ceiling
by the tighten-only ratchet — from 88492 LF bytes to ~90245 on Windows, over
the 90000 XL ceiling, while passing on the LF-checkout Mac/Linux runners.

The ceilings are explicitly "calibrated against raw `wc -c`" on an LF checkout,
so the measurement should be LF-based on every platform. byteCount now reads the
file and counts Buffer.byteLength after stripping CR, making the budget
platform-independent (a no-op on LF checkouts; verified statSync === normalized
for all 88 workflow files). No ceilings changed. Added a regression test
asserting CRLF and LF content of the same file count identically.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#683): make worktree-base-ref test path mocks Windows-safe (path.join)

tests/worktree-base-ref.test.cjs keyed its injected readFile/writeFile mocks
(and a few expected `file` values) with forward-slash template literals like
`${claudeDir}/settings.local.json`. The module composes those paths with
path.join(), which emits backslashes on Windows, so the mock keys never matched
the module's lookup → readFile returned null → resolveEffectiveBaseRef /
cmdWorktreeBaseCheck / cmdWorktreeSetBaseRef (and the JSONC variants) failed on
the Windows full-test runner only (they passed on Mac/Linux, and the install
tests passed because they use the real filesystem). The module is correct;
only the test fixtures hardcoded '/'.

All mock keys and path assertions now use path.join(base, ...) mirroring the
module, so they match on every platform (no-op on POSIX). 19 path references
across 16 lines.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 23:40:24 -04:00
Tom Boucher
463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00
Tom Boucher
3bb2f8f1c5 docs: rebrand to GSD Core and restructure docs with Diataxis (#605)
* chore: wire docs/agents config into AGENTS.md Agent skills section

Add the `## Agent skills` discovery block pointing the engineering
skills at the existing docs/agents/{issue-tracker,triage-labels,domain}.md
files (issue tracker, triage label mapping, single-context domain docs).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: rebrand to GSD Core and restructure docs with Diataxis

Reorganise the root README and docs/ around the Diataxis framework
(tutorials, how-to guides, reference, explanation), add new how-to
guides and schema references (STATE.md / CONTEXT.md / PLAN.md /
planning artifacts), and cross-link the whole set. Update the lone
legacy gsd-build reference to open-gsd; keep internal get-shit-done/
filesystem paths unchanged (directory rename tracked separately in
open-gsd/gsd-core#604). Regenerate the ja-JP, ko-KR, pt-BR and zh-CN
localised trees to mirror the new structure.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: backfill changeset PR number (#605)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 08:13:09 -04:00
Tom Boucher
a11ba2dfcb feat(#68): per-phase granularity overrides (granularities.<phaseType>) (#595)
Closes #68. Per-phase-type granularity overrides via granularities.<phaseType>, mirroring models.<phaseType>. Includes maintainer-authorized sdk-seam reference cleanup.
2026-06-01 20:53:46 -04:00
Tom Boucher
9ffe45a7c3 feat(#163): tighten gsd-roadmapper granularity defaults to reduce thin-phase fragmentation (#591)
* feat(#163): tighten gsd-roadmapper granularity defaults to reduce thin-phase fragmentation

Tighten the Granularity Calibration buckets in gsd-roadmapper (Coarse 3-5->2-4,
Standard 5-8->4-6, Fine 8-12->6-10) and append inline Key guidance naming the
thin-phase failure pattern (single requirement / internal-quality goal /
task-shaped success criteria) with instruction to fold into a neighbor rather
than create a standalone phase. Implements the maintainer-approved proposal
verbatim.

Update the canonical English docs that hardcoded the old phase-count numbers:
docs/CONFIGURATION.md and docs/FEATURES.md. Translated docs are
community-maintained and are not updated per-PR (CONTRIBUTING.md language
policy).

Prompt/doc text only; no code, format, or downstream-consumer changes. Agent
size-budget and skills-awareness tests pass; full suite green.

Closes #163

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#163): add Changed changeset for roadmapper granularity tightening

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#163): lock tightened gsd-roadmapper granularity buckets

source-text-is-the-product test asserting the Granularity Calibration table
holds the tightened ranges (Coarse 2-4, Standard 4-6, Fine 6-10), that no row
maps to an old bucket, and that the Key paragraph carries the thin-phase
folding guidance. Would fail if the values regress.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 20:37:49 -04:00
Tom Boucher
2ba6b69d53 feat(#49): provider-neutral model policy presets
* feat(#49): provider-neutral model policy presets

Adds model_policy config surface with known-provider presets (openai/anthropic/google/qwen) and generic provider escape hatch. model_policy.runtime_tiers resolves before legacy model_profile_overrides. reasoning_effort is stripped for unsupported runtimes.

Closes #49

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#49): replace unregistered /gsd-settings-advanced token in docs

docs-parity-live-registry enforces every /token in docs/*.md maps to
a live command. /gsd-settings-advanced is a workflow filename, not a
registered command — use /gsd:settings instead.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#49): update INVENTORY.md count and manifest for config-types.cjs

inventory-counts and inventory-manifest-sync tests require the headline
count and INVENTORY-MANIFEST.json to reflect every file in bin/lib/.
config-types.cjs (new module added by feat(#49)) was missing from both.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-01 08:46:55 -04:00
Tom Boucher
0a12b06381 feat(#39): milestone-prefixed phase IDs (M-NN convention) + migration tool + validation (#565)
* feat(#39): milestone-prefixed phase IDs (M-NN convention) + migration tool + validation

- Add getMilestoneFromPhaseId() / getPhaseDirFromPhaseId() helpers to core.cjs
- Fix isDirInMilestone to match M-NN-style dirs (02-01-setup) against M-NN ROADMAP headings
- Extend heading regex to tolerate [bracket-token] scope prefix on phase headings
- Add W021 validation rule for milestone prefix mismatch
- Add gsd-tools roadmap validate + roadmap upgrade --convention milestone-prefixed
- Add phase_id_convention config field (null default, backwards-compatible)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#39): address 4 Codex review findings in milestone-prefixed phase ID implementation

- getMilestoneFromPhaseId: tighten regex to require a digit after the hyphen (rejects '1-' and '1-abc')
- isDirInMilestone: use convention-aware regex — only capture M-NN segments when ROADMAP itself uses hyphenated phase IDs, preventing legacy dirs like '01-02-setup' from being misread as phase '1-02'
- checkW021: add UNPREFIXED_PHASE_RE path so unprefixed headings (### Phase 1:) also fire W021 when convention is milestone-prefixed
- roadmap-upgrade: remove isMigratedDirName dir-name check (false-positive for legacy dirs); config + ROADMAP heading checks at lines 194 and 212 are sufficient

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: update changeset pr reference to #565

* fix(#39): restore phaseDirNameRe 2-digit minimum; add roadmap-upgrade to inventory

- validate.cjs: \d{1,} → \d{2,} to keep single-digit prefix rejection per W005 contract
- docs/INVENTORY.md: 79 → 80, add roadmap-upgrade.cjs row
- docs/INVENTORY-MANIFEST.json: regenerated (roadmap-upgrade.cjs entry)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-31 23:15:55 -04:00
Tom Boucher
0fbe1d899e chore(#191): retire the gsd-sdk shim — route everything at gsd-tools (#522)
* chore(#191): migrate gsd-sdk query call sites to gsd-tools query

Retiring the gsd-sdk shim. gsd-tools.cjs already accepts `query` as a
meta-prefix (gsd-tools query <command>), so this is a behavior-preserving 1:1
swap across the runtime reference prompts, the graphify hook's commit-detection
gate, and two bin/lib comment/message references.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#191): remove vestigial gsd-sdk shim code from installer + projection

The gsd-sdk shim was already not wired up (no gsd-sdk bin in package.json;
buildWindowsShimTriple had zero call sites). Remove the dead code:
- shell-command-projection.cjs: buildWindowsShimTriple + formatSdkPathDiagnostic
  (+ their now-unused PACKAGE_NAME import) and exports
- install.js: the re-export wrappers + imports, the #3406 stale-standalone-sdk
  detection (detectStaleStandaloneSdk/formatStaleStandaloneSdkWarning + its
  global-install call site), and the exports

Preserved (retained, not gsd-sdk): buildCodexHookWindowsShimIR (#3426) — only
its comments referenced the gsd-sdk pattern; reworded. Also kept the
homePathCoveredByRc 'reopen your shell' branch in maybeSuggestPathExport — its
logic is bin-dir-agnostic, only the message mentioned gsd-sdk; reworded to use
the actual bin dir.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#191): update tests for retired gsd-sdk shim

- bug-3441/bug-3442: drop the formatSdkPathDiagnostic / buildWindowsShimTriple
  assertions (functions removed); retained PATH-action + drift-guard tests stay
- bug-505: remove the 'still exported' assertions for detectStaleStandaloneSdk /
  formatStaleStandaloneSdkWarning / the shim contract surface (#505 kept them;
  #191 removes them)
- graphify-auto-update: migrate the hook-dispatch inputs gsd-sdk query commit ->
  gsd-tools query commit to match the migrated commit hook

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#191): point active docs at gsd-tools query (gsd-sdk shim retired)

Update the user/agent-facing docs (AGENTS, COMMANDS, CONFIGURATION, USER-GUIDE,
ship-pr-body-sections) that presented gsd-sdk query as a current command to
gsd-tools query. Historical docs (ADRs, PRDs, release notes) left untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#191): correct state.load vs state.json description for gsd-tools query

Adversarial-review (codex) finding: the migrated USER-GUIDE line claimed both
'gsd-tools query state.json' and 'state.load' resolve to the frontmatter-rebuild
handler. Verified they don't — state.load returns the CJS load shape
(config + state_raw + flags), state.json returns the frontmatter shape. Both are
available via gsd-tools query; corrected the text to say so.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#191): add changeset for gsd-sdk shim retirement

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-05-30 19:19:14 -04:00
Tom Boucher
79002a00cb chore(#518): rename npm package + bin to @opengsd/gsd-core (#519)
* chore: rename npm package + bin to @opengsd/gsd-core (functional)

- package.json: name @opengsd/get-shit-done-redux → @opengsd/gsd-core,
  bin key get-shit-done-redux → gsd-core, repository/homepage/bugs URLs
- package-lock.json: regenerated (npm install --package-lock-only)
- tests/**, scripts/**, bin/**, .github/**, agents/**, commands/**,
  get-shit-done/bin/**, get-shit-done/workflows/**:
  applied the 4-rule replacement (scoped npm ref, GitHub repo path,
  bin/clone invocations) per #505 single-source refactor

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: sweep live references to @opengsd/gsd-core

Update all live documentation (README.md + translations, docs/**,
CONTRIBUTING.md, VERSIONING.md, SECURITY.md, CONTEXT.md,
docs/CANARY.md) to reflect the renamed package and repository.

Rules applied:
- @opengsd/get-shit-done-redux → @opengsd/gsd-core (scoped npm name)
- open-gsd/get-shit-done-redux → open-gsd/gsd-core (GitHub repo)
- GSD-redux/get-shit-done-redux → open-gsd/gsd-core (stale badge org)
- bare bin/clone refs → gsd-core

CHANGELOG.md, docs/adr/**, docs/RELEASE-*.md, docs/research/**,
and .changeset/** are preserved byte-identical.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: add negative lookbehind to slash-command regex in bug-2954 test

The extractSlashReferences regex matched /gsd-core inside npm package
URLs (@opengsd/gsd-core), producing a false /gsd:core command reference.
Adding a negative lookbehind (?<![a-z]) excludes matches preceded by a
letter, so only standalone /gsd-<cmd> and /gsd:<cmd> tokens are found.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#518): add changeset for package rename

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#518): update package-identity expectations to the renamed coordinates

The rebase regenerated the seam to @opengsd/gsd-core (bin gsd-core, repo
open-gsd/gsd-core). The #498 seam tests assert deriveIdentity against the REAL
package.json, so their expected literals must follow the rename. The drift-lint
unit test is left as-is — its SEAM is a self-consistent fixture and its
stale-literal detection cases would shift if altered; the live-repo scan in it
already passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-05-30 17:25:02 -04:00
Tom Boucher
05cdec5f47 feat(#22): plan-vs-codebase drift guard (source-grounded reviewer + intel surface) (#487)
* feat(#22): add plan_review.source_grounding + _authority config keys

Two additive opt-out keys for the drift guard: source_grounding (bool,
default true) gates the source-grounded reviewer pass; _authority (enum
grep|intel|treesitter|lsp|scip, default grep) selects the resolver rung.
No existing default changed.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#22): add intel api-surface renderer + CLI subcommand

Renders .planning/intel/api-map.json into a human-readable API-SURFACE.md
for planner injection. Empty/missing map still writes a surface that
announces itself incomplete (absence = unknown, not 'does not exist').
Gated on intel.enabled like all intel functions.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#22): add source-grounding pass to plan-review-convergence

Default-on reviewer pass (plan_review.source_grounding) that enumerates
every symbol a plan cites, excludes declared new artifacts, resolves each
against source via the configured authority adapter, and records
three-valued verdicts. rung-0/1 MISSING is needs-acknowledgement, not a
hard block; UNCHECKABLE is logged in a REVIEWS.md coverage section.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#22): inject API-SURFACE.md into planner + require Artifacts section

When intel.enabled, plan-phase regenerates API-SURFACE.md and injects it
as a HINT (prefer, may be incomplete, absence = unknown), never a hard
rule. Every plan must now emit an 'Artifacts this phase produces' section
so the source-grounding reviewer can separate new symbols from references
to existing code.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#22): surface drift-guard in setup + settings, add docs

/gsd:new-project asks to enable plan_review.source_grounding (default Y);
/gsd:settings exposes the toggle and authority knob. Documents both config
keys in CONFIGURATION.md, the intel api-surface command in COMMANDS.md,
the drift guard in USER-GUIDE.md, and links ADR 22 from ARCHITECTURE.md.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#22): respect AskUserQuestion 4-option cap and plan-phase XL line budget

settings drift-guard toggle moved to its own 2-option question; #22
plan-phase additions condensed to bring the file back under the 1810-line
XL budget without dropping the intel gate, the incomplete-surface hint, or
the Artifacts-section requirement.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#22): use live slash-command forms in drift-guard docs

Doc-parity gate requires every slash-command token in docs/*.md to resolve
to a registered command. Corrected the command form(s) referenced in the
#22 drift-guard / api-surface documentation.

The unresolved token was /gsd-core, matched from the GitHub repo reference
"open-gsd/gsd-core#22" in docs/adr/22-plan-drift-guard.md. This is the
same pattern as the existing 'test-runner' exemption (open-gsd/gsd-test-runner).
Added 'core' to INTERNAL_COMPONENT_SLUGS with a matching explanatory comment.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#22): add changeset fragment for drift guard (PR #487)

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 17:08:11 -04:00
Tom Boucher
c7e5a88353 enh(#466): refresh opus-tier model IDs to current GA (Opus 4.8 / codex gpt-5.5) (#467)
* enh: bump opus-tier model IDs to current GA (Opus 4.8 / codex gpt-5.5)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#466): changeset for opus-tier model-ID refresh

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-29 11:52:47 -04:00
Tom Boucher
5ca646f015 feat(#443): unified cross-provider effort controls + fast-mode-aware routing (#463)
* test(#443): RED unified effort + fast_mode + resolve-execution

All 68 tests failing as expected — no implementation yet.
Covers: effort cascade (tier defaults, overrides, invalid fallthrough),
fast_mode cascade (boolean-only, tier defaults), resolveEffortForTier
escalation, renderEffortForRuntime clamping, resolve-execution CLI,
config schema new keys, QA hostile-input matrix.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(#443): unified cross-provider effort + fast_mode knobs and resolve-execution query

Adds config-driven effort control (universal ladder: minimal<low<medium<high<xhigh<max)
and fast_mode propagation knobs, with per-runtime rendering that clamps the unique
tail values (max=Anthropic-only clamps to xhigh on Codex; minimal=Codex-only clamps
to low on Claude).

Key changes:
- config-schema.manifest.json: add effort.default, fast_mode.enabled as validKeys;
  add 4 dynamicKeyPatterns for effort.routing_tier_defaults, effort.agent_overrides,
  fast_mode.routing_tier_defaults, fast_mode.agent_overrides; fix stale _comment
- config-defaults.manifest.json: add effort and fast_mode blocks with tier defaults
- model-catalog.cjs: add EFFORT_RENDERING map, renderEffortForRuntime(), RUNTIMES_WITH_FAST_MODE
- model-profiles.cjs: re-export new catalog exports
- core.cjs: add resolveEffortInternal, resolveFastModeInternal, resolveEffortForTier,
  VALID_EFFORTS, EFFORT_SET, nextEffort; pass effort/fast_mode through loadConfig
- commands.cjs: replace reasoning_effort in cmdResolveModel with unified effort;
  add cmdResolveExecution (superset command with effort_rendered, effort_param,
  effort_propagation, fast_mode, fast_mode_supported)
- gsd-tools.cjs: add resolve-execution case with --effort/--fast-mode/--attempt flags
- tests/feat-443: 69 tests covering cascade, rendering, escalation, CLI, schema, QA matrix
- tests/commands.test.cjs: convert 3 reasoning_effort assertions to unified effort
- docs/CONFIGURATION.md: document effort + fast_mode + resolve-execution sections
- settings-advanced.md: list new effort/fast_mode keys in confirmation table

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#443): remove dead catalog effort lane; unify codex effort through renderEffortForRuntime

- Remove resolveReasoningEffortInternal (catalog-driven effort function) from
  core.cjs and its export; remove from commands.cjs destructure import
- Convert tests/issue-2517-runtime-aware-profiles.test.cjs: all 11 effort
  assertions now use resolveEffortInternal + renderEffortForRuntime; Claude
  effort is first-class (output_config.effort); unknown runtimes assert param===null
- Convert tests/feat-3023-model-phase-types.test.cjs: replace the entire
  resolveReasoningEffortInternal describe with unified effort assertions;
  effort derives from AGENT_DEFAULT_TIERS routing tier, not phase-type tier

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#443): ADR for unified cross-provider effort + fast-mode routing

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* test(#443): architecture-level QA invariants + test-strategy doc

Add 48-test integration suite (feat-443-effort-fast-mode.integration.test.cjs)
covering 8 architectural invariants: cross-provider validity (never emit a value
the real API would 400 on), param/channel contract stability, resolve-execution
JSON contract (all 8 keys + correct types), totality across the full 33-agent
registry, fast-mode honesty (claude always fast_mode_supported=false), precedence
first-valid-wins matrix for both effort and fast_mode cascades, dynamic-routing
composition (effort escalation independent of model tier), and config-set round-trip
for all new effort/* and fast_mode/* key namespaces. Append test-strategy section
with invariant rationale and E2E gap documentation to docs/TESTING-SUITES.md.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#443): add failing install-wiring tests for effort per-runtime injection (RED)

TDD RED: 10 failing tests covering:
- Claude .md gets effort: injected per tier (planner=xhigh, mapper=low, executor=high)
- Gemini .md does NOT get effort: (already passing — Gemini-safe)
- Codex .toml gets model_reasoning_effort via unified resolver
- Config-driven: effort.agent_overrides drives both Claude .md and Codex .toml
- Source purity: agents/*.md have no effort: key (already passing)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(#443): wire effort per-runtime at install (Claude .md frontmatter + Codex .toml unified)

- Import AGENT_DEFAULT_TIERS and renderEffortForRuntime from model-catalog.cjs
- Add readGsdEffectiveEffortConfig(targetDir): reads merged effort config from
  .planning/config.json (per-project wins) + ~/.gsd/defaults.json (global fallback),
  same probe pattern as readGsdRuntimeProfileResolver
- Add resolveInstallTimeEffort(effortCfg, agentName): pure function matching
  resolveEffortInternal() precedence (agent_overrides > routing_tier_defaults > default > 'high')
  without loadConfig side-effects (no sub-repo detection, no migration writes)
- Claude agent copy loop: inject `effort: <value>` into frontmatter ONLY for
  runtime === 'claude'; all other .md runtimes (Gemini, Qwen, Hermes, etc.) stay
  effort-free (Gemini-safe source contract preserved in agents/*.md)
- generateCodexAgentToml: add effortCfg param; emit model_reasoning_effort from
  unified resolver (replaces old catalog entry.reasoning_effort); Codex clamps
  max → xhigh via renderEffortForRuntime('codex', ...)
- installCodexConfig: pass readGsdEffectiveEffortConfig(targetDir) to
  generateCodexAgentToml so per-project config wins for Codex .toml too
- Update failing tests to GREEN: 12/12 pass; all 17 install tests pass;
  2847/2848 unit tests pass (1 pre-existing failure: policy-shell-pinning)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#443): source install effort defaults from manifest (kill drift) + guard test

Replace hardcoded _GSD_EFFORT_MANIFEST_TIER_DEFAULTS and the 'high' fallback in
resolveInstallTimeEffort with values read from config-defaults.manifest.json at
module init, using the same __dirname-relative path install.js already uses for
all shared manifests. Add feat-443-effort-defaults-drift.test.cjs to assert
equality between install.js's runtime constants and the manifest on every CI run.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): reconcile Codex TOML tests with unified effort design

The #443 unified effort resolver makes generateCodexAgentToml always emit
model_reasoning_effort (driven by resolveInstallTimeEffort, not model_profile_overrides).
The test 'generated TOML omits reasoning_effort when runtime has none' had an
obsolete premise — model_profile_overrides.reasoning_effort:'' no longer suppresses
unified effort. Convert it to assert the new invariant: Codex TOML always carries a
valid model_reasoning_effort from the agent's routing tier (xhigh for gsd-planner,
a heavy-tier agent), while model_profile_overrides model override is still respected.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): make install.js effort resolution lazy (no load-time side effects breaking launcher-parity)

Replace module-load-time IIFE + hard throw (config-defaults.manifest.json read)
and top-level require of model-catalog.cjs with a lazy _getGsdEffortCatalog()
getter that initialises on first call from resolveInstallTimeEffort /
generateCodexAgentToml / Claude .md effort injection.  Requiring install.js in
unrelated test contexts (e.g. runtime-launcher-parity) no longer triggers
manifest IO or throws, eliminating the load-time side effect that changed
subprocess exit codes / stderr on the bench.

Drift-guard exports (_GSD_EFFORT_MANIFEST_TIER_DEFAULTS / _GSD_EFFORT_MANIFEST_DEFAULT)
preserved as lazy getter properties on module.exports so feat-443-effort-defaults-drift
still validates them without forcing eager load.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): isolate install-wiring test HOME to stop \$HOME/.claude pollution breaking launcher-parity

runGlobalInstall() now redirects HOME to a per-call isolated tmpdir in addition
to the existing runtime-specific env-var redirects (CLAUDE_CONFIG_DIR,
GEMINI_CONFIG_DIR, CODEX_HOME). This ensures install.js code that uses
os.homedir() directly — including the ~/.cache/gsd update-check deletion,
~/.gsd/defaults.json reads, and any HOME-relative npm subprocess writes —
never touches the real \$HOME during the test.

Without the HOME isolation the install test (which is new to this branch and
is now picked up by Docker's raw \`tests/*.test.cjs\` glob) could write or
delete files under the real \$HOME, causing runtime-launcher-parity test (D)
to fail: (D) asserts a loud non-zero exit when \$RUNTIME_DIR/gsd-tools.cjs is
absent and gsd-tools is not on PATH, but the launcher's \$HOME/.claude fallback
arm succeeds if \$HOME/.claude/get-shit-done/bin/gsd-tools.cjs exists.

Also sets GSD_SKIP_STALE_SDK_CHECK=1 to suppress the \`npm ls -g\` subprocess
that the global installer spawns — irrelevant to effort-wiring assertions,
slow, and potentially writes to ~/.npm cache.

All 12 feat-443 install-wiring assertions preserved. Drift-guard 5/5. Unit
suite 2848/2850 (pre-existing policy-shell-pinning.test.cjs failure on next).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#443): add changeset fragment for effort + fast-mode routing

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(#443): set GSD_TEST_MODE before requiring install.js in drift-guard test to prevent HOME leak

Without GSD_TEST_MODE=1, require('bin/install.js') runs the module's main
install block (guarded by !GSD_TEST_MODE), performing a real global Claude
install into $HOME/.claude/. On CI ubuntu where node is on standard PATH,
the launcher's $HOME/.claude fallback arm then finds gsd-tools.cjs, causing
runtime-launcher-parity test (D) to exit zero when it must exit non-zero.

Root cause: feat-443-effort-defaults-drift.test.cjs (unit suite) runs
alphabetically before runtime-launcher-parity.test.cjs in the same node
--test invocation. Each runs in a separate worker process but shares the
same HOME. The drift test's install leaks gsd-tools.cjs into that HOME,
then the launcher test's bash subprocess finds it via the $HOME/.claude arm.

Fix: add process.env.GSD_TEST_MODE = '1' at the top of the drift-guard
test, before the require(installPath) call. This matches the pattern used
by feat-443-effort-fast-mode.test.cjs and feat-443-effort-install-wiring
.install.test.cjs.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): deterministic resolve-execution arg parsing + validate install-time effort (Codex adversarial findings)

Finding 1: resolve-execution --effort low gsd-planner misrouted 'low' as the agent.
Replace find(non-dash) with a proper flag-consuming loop that collects a single
positional; validate missing/extra positionals and malformed --attempt values.

Finding 2: resolveInstallTimeEffort returned unvalidated effort strings (e.g. "ultra")
verbatim. Each precedence layer now checks GSD_EFFORT_SET (imported once from
core.cjs) before accepting a value, mirroring resolveEffortInternal exactly.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): newline-agnostic effort frontmatter injection (Windows CRLF) + CRLF-safe assertions

Extracts injectEffortFrontmatter(content, effortValue) pure helper that detects
EOL (LF vs CRLF) from the opening '---' line and inserts 'effort: <value>'
before the closing '---' delimiter using the same EOL as the surrounding
frontmatter. Regex now uses /^---\r?\n([\s\S]*?)^---\r?$/m instead of the
LF-only /^(---\n[\s\S]*?)(---)(\n|$)/ that silently skipped CRLF files on
Windows (git core.autocrlf=true checkout).

Also adds 7 unit tests covering LF, CRLF, idempotency, no-frontmatter, and
complex frontmatter cases. Exports injectEffortFrontmatter from module.exports.

Fixes 6 CI failures in tests/feat-443-effort-install-wiring.install.test.cjs
on windows-latest runners (lines 138, 145, 152, 261, 345, 356).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-29 10:32:58 -04:00
Tom Boucher
5f3eb42864 feat(observability): propagate parentTraceId on DispatchEvent — ADR-0174 SDK retirement Phase 1.4 (#178) (#225)
* test(#178): update DispatchEvent factory tests to propagate parentTraceId

P1.3 test 'parentTraceId is always undefined' replaced with four P1.4
contracts: absent → undefined, string → propagated, null → undefined,
non-string → undefined (defensive normalization policy).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(#178): propagate parentTraceId through DispatchEvent factory

Stop ignoring the parentTraceId parameter added as a forward-compat hook
in P1.3. Defensive normalization: only non-null strings are propagated;
null, non-string values, and absent callers all yield undefined, keeping
P1.3 behavior intact for all existing dispatch call sites.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#178): add Hub-level parentTraceId propagation tests

Four new assertions: req.parentTraceId propagates to event, absent →
undefined (P1.3 regression), shared parentTraceId across multiple
dispatches, and unique traceId invariant despite shared parentTraceId.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(#178): plumb parentTraceId through Hub dispatch and _notifyLogger

dispatch() now reads req.parentTraceId and passes it to _notifyLogger,
which forwards it to makeDispatchEvent. Backward-compatible: callers
that omit parentTraceId emit events with parentTraceId: undefined,
identical to P1.3 behavior.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#178): add trace correlation end-to-end test

Dispatches a root command then 3 children with parentTraceId=rootTraceId.
Reads the real .gsd-trace.jsonl audit file and verifies: 4 events total,
root has no parentTraceId, all children carry rootTraceId, all traceIds
unique, JS filter returns exactly the 3 children given the root's traceId.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#178): document traceId/parentTraceId in audit file

Update Observability section to note that audit events now carry both
traceId and parentTraceId, and explain the correlation filter pattern.
Note that leaf dispatches emit parentTraceId: undefined until the Phase 2
composer wires it automatically.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#178): add changeset for trace correlation seam

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#178): cover invalid parentTraceId values in DispatchEvent factory

Adds 9 new test cases for UUID v4 validation of parentTraceId:
empty string, whitespace, non-UUID, oversized, UUID v1, missing-hyphen,
extra-char (all dropped to undefined), plus UPPERCASE and lowercase v4
(both propagated). Tests are intentionally red until the implementation
commit that follows.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(#178): validate parentTraceId against UUID v4 before propagation

Adds UUID_V4_REGEX constant and isValidParentTraceId() helper to
event.cjs. makeDispatchEvent now silently coerces any parentTraceId that
fails the UUID v4 check (wrong version nibble, wrong variant, missing
hyphens, oversized, empty, etc.) to undefined. No stderr warn is emitted
— the factory remains pure and side-effect-free. Closes the correlation-
poisoning vector identified in the Codex adversarial review of PR #225.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#178): assert Hub silently drops invalid parentTraceId at the seam

Adds two tests to hub-logger-integration.test.cjs:
1. dispatch with 'junk' parentTraceId emits event with parentTraceId===undefined.
2. The logger-failure warn path is NOT triggered — the factory coerces the bad
   value before onEvent is called, confirmed by zero stderr output even when a
   logger that would throw on non-undefined parentTraceId is installed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#178): assert invalid parentTraceId does not poison correlation siblings

Adds one test to trace-correlation.test.cjs: dispatches a root, a valid
child (parentTraceId = rootTraceId), and an invalid child (parentTraceId =
'junk'). Asserts: valid child carries correct parentTraceId, invalid child
has parentTraceId dropped to undefined, filtering by rootTraceId yields
exactly 1 event (the valid child only), and all 3 events have unique
traceIds. Uses an isolated Hub + tmpdir to avoid shared fixture interference.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#178): document UUID v4 contract for parentTraceId

Appends one sentence to the Observability audit-trail paragraph in
CONFIGURATION.md: parentTraceId must be canonical UUID v4 (RFC 4122);
values that don't match are silently dropped from audit output. No section
restructuring — single sentence addition only.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 16:32:28 -04:00
Tom Boucher
2d14eb8873 feat(observability): add DispatchLogger seam — ADR-0174 SDK retirement Phase 1.3 (#177) (#223)
* test(#177): add DispatchEvent factory failing tests

Red tests for makeDispatchEvent shape, traceId UUID v4, uniqueness,
parentTraceId-always-undefined (P1.3), args redaction toggle, ISO 8601
timestamp, and all result variant passthrough.

* feat(#177): introduce DispatchEvent factory

makeDispatchEvent produces an immutable event record per dispatch:
- traceId: crypto.randomUUID() (UUID v4)
- parentTraceId: always undefined (P1.4 wires composer)
- command, result, timestamp (ISO 8601)
- args only included when includeArgs === true (default: omitted)

* test(#177): add arg redaction policy failing tests

Red tests for shouldIncludeArgs (GSD_AUDIT_ARGS env gating) and
redactEvent (strips args from frozen events, preserves all other
fields, returns a new object, never mutates the source).

* feat(#177): introduce arg redaction policy

shouldIncludeArgs(): only GSD_AUDIT_ARGS==='1' opts in; all other
values (unset, '', '0', 'true') default to omitting args.

redactEvent(event): returns a shallow copy of the event, dropping the
args field unless opted in. Never mutates the (frozen) source event.

* test(#177): add DispatchLogger interface failing tests

Red tests covering:
- no-op logger: silent on all events, never throws
- default logger: silent on ok, one flattened JSON line to stderr on error
- default logger: audit file creation + append-only + redaction + config gate
- GSD_AUDIT env var and config.audit.enabled config gate
- GSD_AUDIT_ARGS opt-in for args inclusion
All tests use real fs under os.tmpdir() — no mocked appendFileSync.

* feat(#177): introduce DispatchLogger with default and no-op implementations

createNoOpLogger(): silent on all events — Hub default when no logger injected.
createDefaultLogger({ cwd, config }):
  - Silent on ok result
  - Flattened JSON line to stderr on error: { kind, traceId, ...typedPayload }
  - Append-only audit at .planning/.gsd-trace.jsonl when GSD_AUDIT=1 or config.audit.enabled
  - Args redacted by default; GSD_AUDIT_ARGS=1 opts in
  - Logger errors caught internally; never break dispatch callers

* test(#177): add Hub+logger integration failing tests

Red tests verifying:
- onEvent called exactly once per dispatch (ok, error, handler-throw, unknown)
- DispatchEvent shape: traceId uniqueness, command, result.kind, parentTraceId
- Logger errors contained (dispatch still returns Result, warn line to stderr)
- Hub defaults to no-op when no logger injected
- End-to-end with createDefaultLogger: silent on success, stderr on error, audit file

* feat(#177): wire DispatchLogger into CommandRoutingHub

Add optional logger param to createHub({ ..., logger }).
Defaults to createNoOpLogger() — silent, no behaviour change for callers
that don't inject a logger.

After every dispatch (success and error):
- Normalises HubResult { ok } to DispatchEvent { kind: 'ok'|error-kind }
- Calls makeDispatchEvent({ command, args, result }) to mint the event
- Calls logger.onEvent(event) exactly once
- Wraps in try/catch: logger errors emit { level:'warn', source:'DispatchLogger' }
  to stderr but never propagate to dispatch callers

* chore(#177): gitignore .planning/.gsd-trace.jsonl audit file

The audit trail is local-only, append-only, and must never be committed.
Slotted under the existing "Local scratch + Claude-test artifacts" block.

* docs(#177): document GSD_AUDIT, GSD_AUDIT_ARGS, config.audit.enabled

New ## Observability section at end of CONFIGURATION.md covering:
- Default silent/stderr behaviour overview
- Stderr error JSON format
- Audit file opt-in (env var and config key)
- Args redaction policy and GSD_AUDIT_ARGS opt-in

Also slots GSD_AUDIT and GSD_AUDIT_ARGS into the existing
## Environment Variables table (alphabetical order).

* chore(#177): add changeset for observability seam

type: Added — new DispatchLogger seam with default silent/stderr/audit behaviour.
2026-05-24 15:22:36 -04:00
Tom Boucher
2a915c1b82 chore: migrate references from gsd-build to open-gsd/get-shit-done-redux (#120) (#121)
Security-motivated migration of all stale repository and npm-scope references.

Three categories of changes (58 files, 174 substitutions):

1. gsd-build → open-gsd (security-critical):
   - .github/workflows/release-sdk.yml — npm token comment, tarball filename pattern
   - .github/workflows/hotfix.yml — same
   - .changeset/fix-3406-detect-stale-sdk-shadow.md — @gsd-build/sdk → @open-gsd/sdk
   - .changeset/sharp-quails-leap.md — same
   - get-shit-done/workflows/update.md — CHANGELOG raw GitHub URL

2. GSD-redux org slug → open-gsd (canonical rename):
   - package.json + sdk/package.json — repository/homepage/bugs metadata
   - All README.*.md — live badge and link sections
   - CONTRIBUTING.md, CONTEXT.md, QUICK-WINS-CONFIRMED-BUGS.md
   - .coderabbit.yaml, .release-monitor.sh, scripts/sync-rulesets.sh
   - docs/** — all live agent/ADR/user-facing documentation
   - tests/** — repo slug assertions and test fixtures
   - scripts/changeset/cli.cjs + github-release-notes.cjs
   - .github/ISSUE_TEMPLATE/*, .github/pull_request_template.md
   - bin/install.js, get-shit-done/bin/lib/model-catalog.cjs
   - sdk/HANDOVER-*.md, sdk/src/*.test.ts

3. CLAUDE.md (gitignored local file — not in this commit):
   Updated separately outside git: --repo gsd-build/get-shit-done →
   --repo open-gsd/get-shit-done-redux with security warning.

Intentionally unchanged: CHANGELOG.md, docs/RELEASE-*.md,
.changeset/README.md, .changeset/build-hooks-atomic-write.md,
README.md migration table (historical fork record),
tests/changeset-serialize.test.cjs line 78 (serialization fixture).

The gsd-build/get-shit-done repo is compromised (rug-pull documented in
README.md). Do not push to or interact with that repo.

Closes #120
2026-05-22 12:28:16 -04:00
Tom Boucher
74cb493373 fix(3784): expose adaptive in model_profile settings flow (#91)
* fix(3784): expose adaptive in model_profile settings flow

Split the single 4-option model-profile AskUserQuestion into a two-question
flow: Q1 (Adaptive / Standard tier / Inherit) routes top-level intent; Q2
(Quality / Balanced / Budget) appears only when Standard tier is chosen.
Updates the confirm table and success_criteria to include adaptive.

Adds regression test asserting all five valid profiles are reachable
interactively via the settings UI.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* changeset: add Fixed entry for #3784 / PR #3795

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(3784): correct Q2-skip comment and remove duplicate brace in settings.md

Codex review followup:
- Replaced vague "preserve existing config" comment with accurate description:
  Q1 still writes model_profile on Adaptive/Inherit branches; only Q2 is skipped.
- Removed stray duplicate `{` line before the Spawn Plan Researcher question block
  (pseudocode had two consecutive `{` openers, one spurious).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(3784): address review — gate Q2 structurally, define cancel rule, harden tests

Addresses gsd-code-reviewer (M1/M2/m1/m2/m3/m4) and codex adversarial
(Q2 gating, save-mapping, Claude-only wording, step-of-2 wording).

- F1: Replace //comment-only Q2 gating with Conditional visibility block
  (mirrors code_review_depth / graphify.auto_update structural pattern)
- F2: Define model_profile cancel rule in update_config step (leave
  existing value unchanged when Q1="Standard tier…" but Q2 cancelled)
- F3: Fix Adaptive description — remove "Claude only" tail; describe
  heavy/light role tiers across all supported runtimes
- F4: Remove "step 1 of 2 for standard profiles" from Q1 question text
  (2-step nature now structurally documented by Conditional visibility)
- F5: Fix vacuously-true test disjunct (|| content.includes('Adaptive')
  always true — 6+ occurrences); assertion now requires role-based cost
  optimization + heavy roles wording
- F6: Add 4-option cap enforcement test (ASK_USER_QUESTION_OPTION_CAP=4
  named constant, counts per question object not per AskUserQuestion call)
  and brace-balance regression test (guards against bd53925f recurrence)

* docs(3784): list adaptive in model_profile reference docs

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-22 11:39:57 -04:00
Tom Boucher
dff176bfd2 chore: rebrand to GSD-redux/get-shit-done-redux
Mirror of code, issues, and PRs from the upstream gsd-build/get-shit-done,
which appears compromised or abandoned (maintainer unreachable since
2026-04-01; $GSD token linked to rug-pull).

- Adds rebrand notice block at top of English README
- Removes $GSD token badge and @gsd_foundation X badge (keeps Discord)
- Renames npm packages: get-shit-done-cc -> get-shit-done-redux,
  @gsd-build/sdk -> @gsd-redux/sdk
- Updates all repo URLs across docs, workflows, package.json, bin/
- Updates ci@gsd-build -> ci@gsd-redux in workflow git identities
- Leaves CHANGELOG and .changeset/* alone (historical, time-stamped)
2026-05-22 08:27:07 -04:00
Tom Boucher
6a5fa59129 feat(3081): auto-trim review prompts for small-context model reviewers (#3708)
* feat(3081): auto-trim review prompts for small-context model reviewers

Adds review.max_prompt_tokens and review.max_prompt_tokens_per_reviewer
config keys. When configured, the /gsd-review workflow deterministically
trims the assembled prompt before sending to each reviewer (drop CONTEXT
→ RESEARCH → REQUIREMENTS; head-shrink PROJECT.md; tail-truncate PLANs
proportionally; reserve disclosure-note tokens upfront). Trim metadata
is recorded in REVIEWS.md frontmatter. Reviewer is skipped with a
warning if even the minimum review set exceeds the budget.

Closes #3081

* fix(3081): register prompt-budget in SDK query registry and update inventory manifest

review.md references `gsd-sdk query prompt-budget` at three call sites, but the
command had no handler in the SDK registry — failing the registry-integration
drift-guard test on all 6 CI matrix legs. Added a native TypeScript SDK handler
(sdk/src/query/prompt-budget.ts) that ports the applyBudget logic from the CJS
module, registered it in DOMAIN_STATIC_CATALOG, and regenerated
docs/INVENTORY-MANIFEST.json to include the new cli_modules/prompt-budget.cjs entry.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(3081): bump ws to 8.20.1 and allowlist prompt-budget sibling pair

Two additional CI failures after the registry fix:

1. ws moderate CVE (GHSA-58qx-3vcg-4xpx, uninitialized memory disclosure):
   The advisory covers ws >=8.0.0 <8.20.1. Both root and sdk/package.json
   pinned ^8.20.0 which resolved to 8.20.0. Bumped both to 8.20.1 to clear
   the npm audit drift-guard test (bug-3588-npm-audit-clean.test.cjs).

2. lint-shared-module-handsync detected the new prompt-budget.ts / prompt-budget.cjs
   sibling pair without an allowlist entry. Added a cooperatingSiblings entry
   to scripts/shared-module-handsync-allowlist.json with classification and
   justification matching the established pattern.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(3081): align prompt-budget skip semantics across CJS and SDK dispatch paths

Replace brittle `[ $EXIT -eq 2 ]` guards with `[ $EXIT -ne 0 ]` in all three
local-reviewer blocks (Ollama, LM Studio, llama.cpp) in workflows/review.md.
Any non-zero exit from prompt-budget now triggers a skip with a descriptive
warning — exit 2/11 prints "budget too small", any other non-zero prints
"unexpected exit code". This ensures the SDK bridge dispatch path (exit 11
via GSDError(Blocked)) triggers the same skip as the CJS path (exit 2).

The SDK handler (sdk/src/query/prompt-budget.ts) already writes both metadata
and prompt files before throwing, so no change needed there.

The Ollama block also gains the missing OLLAMA_SKIP guard so the reviewer
invocation is actually skipped (previously the block only suppressed the
OLLAMA_PROMPT_FILE update but still ran the curl invocation).

SDK integration path (hardFailed via GSDError(Blocked) → exit 11) is covered
by handler unit tests in tests/prompt-budget.test.cjs; no gsd-sdk-*.test.cjs
exercising the full bridge dispatch for this command exists yet — that gap
remains and is documented here.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix prompt-budget trim ordering and review guard follow-ups

* perf: optimize prompt-budget and dedup reviewer trim workflow

* fix(3708): drop source-grep theater tests to satisfy lint-no-source-grep

All four test files added in commit 2df566ed were pure source-grep theater:
they read .cjs / .ts / .md source files and asserted that specific string
literals were present or absent. None exercised runtime behaviour.

Deleted:
- tests/gsd-tools-memory-optimizer.test.cjs   — 7 includes() on gsd-tools.cjs
- tests/prompt-budget-hotpath-optimizer.test.cjs — includes() on prompt-budget.cjs + .ts
- tests/prompt-budget-io-optimizer.test.cjs   — includes() on prompt-budget.ts + gsd-tools.cjs
- tests/review-workflow-budget-dedup.test.cjs — includes() on review.md

Behavioural coverage for the prompt-budget feature already exists in
tests/prompt-budget.test.cjs and tests/prompt-budget-cli.test.cjs (also
added by this PR). No replacement tests needed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(3708): correct budget-pressure threshold and minSet accounting

Two bugs in applyBudget caused premature trimming and false hard-fails:

1. UNNEEDED_TRIM: budgetUnderPressure compared baseTokens against
   effectiveBudget - NOTE_RESERVE_TOKENS, triggering trim pressure 80
   tokens before the budget was actually exceeded. Fix: compare against
   effectiveBudget directly; NOTE_RESERVE_TOKENS are still reserved in
   contentBudget once real pressure is confirmed.

2. FALSE_HARDFAIL: minSet included NOTE_RESERVE_TOKENS unconditionally,
   treating the note as mandatory even when no trim would occur and no
   note would be injected. Fix: exclude NOTE_RESERVE_TOKENS from minSet;
   a prompt that fits untrimmed needs no note and must not hard-fail.

Both fixes applied in CJS and TypeScript implementations. Two regression
tests added (cycles 11 and 12) that reproduce each case behaviorally.

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-18 23:13:09 -04:00
Tom Boucher
08848df839 docs(3562): pin minimum Codex CLI version (0.130.0) and explain the seam
Rationale for the version pin (the timeline that produced the oscillation):

  2026-05-08  Codex CLI 0.130.0 ships, dropping extra-skills-roots
              discovery via openai/codex#21485 (scans only ~/.codex/skills,
              cwd .codex/skills, and registered plugin roots).
  2026-05-14  GSD PR #3512 lands, removing ~/.codex/skills/gsd-* under the
              assumption Codex would auto-discover from extra roots.
              That assumption was already obsolete in shipped Codex.
  2026-05-15  #3562 filed — Codex CLI 0.130.0 users have zero $gsd-*
              commands after install.

The previous fix (#3427) was for Codex Desktop's official-skills surface,
which is a different product; that surface still exists on Desktop and
remains harmless duplication when both root scans see the gsd-* dirs.

Documents the supported version inline at the Codex sections of both
USER-GUIDE.md and CONFIGURATION.md, plus a one-line note in README's
Troubleshooting block. No runtime version-detection added — out of scope
and brittle against future Codex changes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-15 13:46:12 -04:00