Commit Graph

102 Commits

Author SHA1 Message Date
Tom Boucher
2c718bf972 fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs (#1537)
* fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs

Generalizes the Codex-only #1515/#1519 fix to every non-Claude runtime, and
wires it into the real install path (where it was previously dead-on-arrival).

Root causes:
1. The runtime-default stamping lived only in `_applyRuntimeRewrites`, but the
   installer emits `gsd-core/workflows/*.md` via `copyWithPathReplacement`, which
   never calls it — so a real `--codex`/`--cursor`/etc. install emitted
   `--default claude` and worktrees-on. RUNTIME mis-resolved to claude and the
   workflow ran executors unisolated against the main checkout. (#1515/#1519 were
   also dead-on-arrival in real installs; this repairs them.)
2. Only `case 'codex'` was stamped; every other non-Claude runtime kept the
   Claude default.

Fix:
- New `_stampNonClaudeRuntimeDefaults(content, runtime)` (single shared helper)
  stamps `--default <runtime>` + `use_worktrees=false` for every `runtime !=
  claude`; called from both `_applyRuntimeRewrites` and, crucially,
  `copyWithPathReplacement` in bin/install.js (the real workflow emit path).
- Generalize the fail-closed worktree guard `= codex` -> `!= claude` in
  execute-phase/quick/diagnose-issues (worktree isolation is Claude-Code-only).
- Flip manager/autonomous inline-vs-background gating to `codex -> background,
  everything-else -> inline` (research: only Codex can background-nest the
  pipeline's subagents; all others run inline, which they support).

Worktree-capability determination is research-backed (official docs for all 14
non-Claude runtimes: none honor GSD's isolation="worktree" mechanism, only Codex
background-nests). New end-to-end real-install test asserts the EMITTED workflow
is stamped — the regression guard that would have caught the dead-on-arrival bug.

Closes #1521

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

* chore(#1521): backfill changeset PR number (#1537)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 13:48:47 -04:00
Tom Boucher
2436b76980 fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees (#1519)
* fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees

A Codex install with a runtime-neutral .planning/config.json resolved
RUNTIME=claude and enabled git worktree isolation, which Codex's
spawn_agent cannot honor. Two root causes:

1. Workflows read `config-get runtime` / `config-get workflow.use_worktrees`
   without `--raw`, so config-get's JSON-quoted output ("codex") was captured
   verbatim into the bash var and broke every `[ "$RUNTIME" = ... ]` check —
   the Codex fail-closed guard was dead even when runtime:codex was explicit,
   and Claude's own worktree degrade-check was dead too. Add `--raw` to those
   reads across execute-phase, autonomous, manager, diagnose-issues, quick.

2. The conversion engine emitted `--default claude` for every runtime. Stamp
   the codex-emitted workflows to `--default codex` (runtime) and
   `--default false` (use_worktrees) in _applyRuntimeRewrites case 'codex', so
   a neutral config on a Codex install resolves runtime=codex / worktrees off.

Also extend the Codex fail-closed worktree guard to quick.md and
diagnose-issues.md (they spawned isolation="worktree" with no runtime guard).

Regression test asserts source<->engine parity across all five workflows
(DEFECT.GENERATIVE-FIX) plus fast-check property coverage of the stamping.

Closes #1515

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

* chore(#1515): backfill changeset PR number (#1519)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 11:56:59 -04:00
Tom Boucher
a77f3c4b3a docs(#1452): document workflow.context_guard_mode in CONFIGURATION.md
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 18:12:03 -04:00
Tom Boucher
c330f70f65 feat(#1494): add workflow.mvp_mode to VALID_CONFIG_KEYS; document code_review_command and plan_chunked in planning-config.md (#1500)
* feat(#1494): add workflow.mvp_mode to VALID_CONFIG_KEYS; document code_review_command, plan_chunked, mvp_mode in planning-config.md

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: backfill PR number 1500 in changeset

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 15:32:07 -04:00
Tom Boucher
0d56f544d2 feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation (#1458)
* feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation

ADR-1244 Phase 6. Adds the capability matrix reference, generated FROM the committed registry so it
can never drift from the actual capability set:
- scripts/gen-capability-matrix.cjs (--write / --check); --check is a CI drift guard.
- tests/capability-matrix-sync.test.cjs (4 tests): drift guard, buildMatrix==committed, every cap
  present, no placeholders.
- docs/reference/capability-matrix.md regenerated from the registry (release-stable: shows engines.gsd,
  omits the lockstep per-cap version that would churn the file every release).
- Consolidated the duplicate trust-model doc: deleted docs/explanation/the-capability-trust-model.md,
  merged its content into capability-trust-model.md, redirected ~10 references; no stale links remain.
- Diataxis verification (now that gsd capability is a real command): corrected the matrix's third-party
  section — the matrix is the first-party catalogue; the overlay-aware view of installed third-party
  capabilities is 'gsd capability list', not this generated file.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1435): Added changeset for the capability matrix reference

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1435): address code-review — non-vacuous matrix test + generator polish

- capability-matrix-sync.test.cjs: assert the 'security registers a ship:pre gate' precondition
  unconditionally so the extension-point check can never degrade to a vacuous pass on registry drift.
- gen-capability-matrix.cjs: warn (stderr) on an unknown loop point at generation time; rename
  enginesOf -> fmtEngines for consistency with the other fmt* helpers (output unchanged).
- capability-trust-model.md: point the two how-to links at the real files
  (import-a-capability-from-a-url.md, version-a-capability.md) instead of the bare directory.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1435): backfill changeset PR number → #1458

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 12:45:26 -04:00
Tom Boucher
9219af3360 feat(#1433): capability trust gate + upgrade/compat (ADR-1244 Phase 4) (#1449)
ADR-1244 Phase 4 (D5 trust + D6 upgrade/compat). capability-trust.cjs (disclosure/consent, strict_known_registries, engines+compatVersions, reserved namespace) + capability-lifecycle.cjs (install/upgrade/remove/reconcile; ledger-as-commit-point _pending intent; atomic stage-then-swap; surgical marker-isolated shared-edit strip; owner-token lock) + capability-source promote/skipEnginesGate seams + loader pending-skip + config keys. No sandbox re-derived (consent+integrity+reversibility). 6 Codex adversarial rounds + /security-review (no HIGH) + /code-review; gsd-test green both platforms; CI green. Phase 5 (#1434) wires the CLI dispatch.

Closes #1433.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 21:40:37 -04:00
Tom Boucher
353f63d170 feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) (#1440)
* feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2)

Promote the registry from a frozen data file to loadRegistry({includeInstalled}),
composing the first-party registry with a validated installed overlay (ADR-1244 D2):

- Extract the conformance validator to a shared runtime-callable module
  (gsd-core/bin/lib/capability-validator.cjs); the generator re-exports it
  verbatim, guarded by a generative-parity test (no build-time/runtime drift).
- capability-loader.cts: loadRegistry({includeInstalled}) composes first-party
  ∪ validated overlay from $GSD_HOME/.gsd/capabilities (global) and
  <root>/.gsd/capabilities (project) via the canonical buildRegistry. First-party
  always wins (id/skill/agent/config/command-family + reserved gsd-/anthropic-
  prefixes); full merged-set cross-capability validation; engines.gsd load-time
  re-gate (skip-with-warning); gate-kind capabilities FAIL CLOSED; fragment-path
  escapes rejected.
- semverSatisfies (hand-written, no dep) for the engines.gsd gate, fail-closed.
- Wire surface/state + loop to the overlay; loop injects a blocking gate for each
  skipped gate-kind overlay (fail-closed).
- cwd-aware overlay config-key federation: config-loader _federatedConfigSchema(cwd)
  + config-schema isValidConfigKey(key, cwd) compose the overlay per loadConfig/
  config-set call (never eager at module load, never wrong-cwd); first-party path
  unchanged with no cwd.
- run-tests.cjs sandboxes GSD_HOME (idempotent — nested spawns reuse it) for test
  hermeticity; capability-loader.cjs git+eslint-ignored (tsc artifact);
  capability-validator.cjs stays linted (#551 migration coverage).

Closes #1431

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1431): add changeset for runtime capability registry overlay

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1431): kill config-schema cwd-aware federation mutants (Stryker ≥52)

The cwd-aware overlay config-key federation added to config-schema.cts
(_capabilityConfigSchema(cwd) + isCapabilityConfigKey/isValidConfigKey cwd
threading) introduced mutable surface uncovered by config-schema's mutation
test set, dropping its score to 39.58% (below the 52 break threshold). Add a
real-overlay-fixture describe block exercising every branch (cwd guard, overlay
loadRegistry, found-branch, first-party fallback, cwd threading); local Stryker
score 39.58% -> 77.08%.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 14:41:20 -04:00
Tom Boucher
137760a655 fix(#1296): align config docs/prompts/schema with consumers (#1299)
* fix(#1296): align config docs/prompts/schema with consumers

The user-facing config surface disagreed with what the consumers actually do
(subset of the #1216 audit). No runtime consumption behavior changes.

- workflow.subagent_timeout: settings-advanced.md prompt + docs/CONFIGURATION.md
  said "seconds (default 600)" but the consumer (map-codebase.md) uses
  milliseconds (default 300000). Relabeled all four spots in settings-advanced.md
  (prompt, parse-default list, example, confirmation table) + the CONFIGURATION.md
  row.
- review.models.<cli>: settings-integrations.md, docs/CONFIGURATION.md (Integration
  Settings), and docs/CLI-TOOLS.md documented a shell command, but review.md injects
  the value into a --model/-m flag. Relabeled to a bare model id and reconciled the
  contradictory CONFIGURATION.md sections.
- workflow.test_command + workflow.build_command: consumed via config-get
  (test_command in verify-phase/execute-phase/audit-fix/post-merge-gate;
  build_command in post-merge-gate) and documented, but absent from validKeys so
  `config set` rejected them. Registered both in config-schema.manifest.json and
  documented them in references/planning-config.md (overview + complete reference).

Regression tests: behavioral config-set tests (tests/config.test.cjs) + doc-parity
content guards (tests/config-field-docs.test.cjs).

Deferred to other #1216 clusters: security-gate wiring, search_gitignored wiring,
mvp_mode, source_grounding_authority labeling, and config-set enum enforcement.

Closes #1296
Refs #1216

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(changeset): Fixed fragment for #1296 config-surface alignment

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 19:57:11 -04:00
Tom Boucher
cf68841220 enh(#1243): consume Claude plugin-provided skills in agent_skills (epic #1258 Phase B) (#1261)
* feat(#1243): consume Claude plugin-provided skills via native Skill-tool directive + grant Skill to agent_skills-consumer agents

- Relax global skill name validation to accept namespaced form `^[A-Za-z0-9_-]+(:[A-Za-z0-9_-]+)*$`
- Namespaced names (containing colon) on claude runtime emit a Skill-tool load directive instead of a @-include line
- Namespaced names on non-claude runtimes are skipped with a warning
- Bare unresolved names retain existing warn-and-skip behavior (no promotion to directive)
- Grant `Skill` tool to all 22 agent_skills consumer agents; 5 generated agents updated via research-profiles.cjs + regen, 17 hand-authored agents edited directly
- Add 16 TDD tests in describe('bug #1243') covering happy/mixed/precedence/negative/cross-runtime/regression/grant cases

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#1243): document plugin-provided skills in agent_skills

Update the Agent Skills Injection reference in CONFIGURATION.md with
the three entry forms (project-relative, global:<name>,
global:<plugin>:<skill>), the Claude-only runtime behaviour of the
namespaced form and the warn-skip on other runtimes, the plugin
pre-install prerequisite, and the consumer-agent Skill tool grant.

Add docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md with a
step-by-step guide for installing the plugin, locating the namespaced
skill name, wiring it into agent_skills, and verifying injection.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1243): align agent_skills docs with emitted block format + mixed-block regression test (code-review)

- Replace two-section mixed-block example (bogus "Load these plugin-provided skills using the Skill tool:" header) with the actual single-section inline format in CONFIGURATION.md and docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md
- Fix quoted warning text in how-to doc to exactly match the emitted string: [agent-skills] WARNING: Plugin-namespaced skill "global:<name>" requires a Skill-tool-capable runtime (claude) — skipping on runtime "<runtime>"
- Replace phantom agent slugs (gsd-checker, gsd-researcher, gsd-advisor, gsd-synthesizer) in CONFIGURATION.md Supported Agent Types with real agents/gsd-*.md examples (gsd-plan-checker, gsd-phase-researcher, gsd-code-reviewer, gsd-ui-auditor, gsd-research-synthesizer)
- Add byte-identical mixed-block regression test: one path-resolvable global skill + one plugin-namespaced skill on claude runtime → asserts r.ir.block === single-section interleaved block, no secondary header

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1243): regenerate agent-size baseline for the Skill-tool grant

The 22 agent_skills-consumer agents each grew +7 bytes from adding `Skill`
to their tools list; refresh the committed per-agent size baseline (#1074 guard).

* chore(#1243): add Added changeset fragment

* fix(#1243): traceable allow-test-rule ref + separator-agnostic byte-identical tests (CI)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-15 00:49:12 -04:00
Tom Boucher
a375c4b354 feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) (#1201)
* feat(#956): add MemPalace memory capability (ADR-857 feature plug-in)

Adds an opt-in, default-resilient ADR-857 feature capability that wires
MemPalace (local-first memory: MCP server + CLI) into the GSD loop:
deliberate recall before discuss/plan and verbatim + temporal-KG capture
at phase boundaries. Three memory modes (augment default; kg_backend and
replace forward-declared). Master gate mempalace.enabled (default off);
every hook onError:skip, zero gates; absent/disabled MemPalace => loop
unchanged. Transport is rendered-markdown only — MemPalace runs
out-of-process, no third-party code in gsd-core (ADR-857 §7).

Capability: capabilities/mempalace/ (manifest + 2 fragments), skills
commands/gsd/mempalace-{recall,capture}.md, agent
agents/gsd-mempalace-curator.md. Registration: ns-context router,
utility cluster, KNOWN_SKILLS, help full.md, model-catalog, copilot
install list, size baselines; regenerated capability-registry +
inventory manifest. ship:post wired into ship.md (wire-on-demand).

HELD on #1196: this capability also declares hooks at discuss:pre and
discuss:post, which are structurally un-wireable until the host-loop
conformance model covers the discuss phase (discuss-phase.md is not in
HOST_LOOP_FILES). The phase6-capstone-conformance gate therefore fails
on exactly those two orphaned points by design — see #1196. Once #1196
lands, rebase onto next and the gate goes green with no further change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#956): backfill changeset PR number (#1201)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 02:09:10 -04:00
Tom Boucher
44024aa535 fix(#1133): honor model_policy on the claude runtime (forward-port to next) (#1144)
Forward-port of the #1133 hotfix (commit 22f237a5, 1.4.5 hotfix line) onto
next. The hotfix was authored against src/core.cts (v1.4.4); on next the
resolver logic lives in src/model-resolver.cts (ADR-457 extraction, #888),
so the patch is re-applied there rather than cherry-picked.

resolveModelInternal step 2.5 now honors model_policy on the claude runtime:
the policy-resolved full model ID is mapped back to a Claude Code agent alias
via CLAUDE_POLICY_ID_TO_ALIAS (reverse of MODEL_ALIAS_MAP + claude-fable-5 ->
fable). Bare aliases (opus/sonnet/haiku/fable) pass through; an ID with no
Claude alias warns once to stderr (deduped by agentType::policyModel::tier)
and falls back to the configured tier alias. Non-claude runtimes return full
IDs verbatim (unchanged). resolveModelForTier is intentionally unchanged.

The warn-dedupe cache lives in model-resolver.cts; core.cts composes the
exported _resetRuntimeWarningCacheForTests to clear both that cache and the
config-loader warning cache (config-loader cannot import model-resolver --
circular dependency).

Ports the 6 #1133 tests (rewriting the old claude-no-op test that asserted
the bug) plus one added test covering the MODEL_ALIAS_MAP reverse-map path.

Forward-port of #1133

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 19:56:54 -04:00
Tom Boucher
827011b865 fix(#1098): guard generate-claude-md against clobbering hand-crafted files; redirect default to .claude/CLAUDE.md (#1118)
/gsd-new-project wrote a repo-root CLAUDE.md full of broad project docs,
overwriting/diluting a hand-crafted instruction file. --force was parsed but
silently dropped, and nothing guarded an existing non-GSD file.

- Guard: an existing instruction file with no `<!-- GSD:<section>-start` markers
  (hand-crafted) is left untouched; report action:"skipped". --force (now wired
  through CmdGenerateClaudeMdOptions) overwrites intentionally. The marker check
  uses /<!-- GSD:[a-z]+-start/ so a file merely documenting GSD syntax is safe.
- Redirect: the Claude-family default output is now ./.claude/CLAUDE.md (a valid
  auto-loaded project-memory location) instead of repo-root ./CLAUDE.md. Aligned
  across the handler default, config-defaults.manifest.json, buildNewProjectConfig,
  the config template, new-project.md, and cmdGenerateClaudeProfile; advisory
  read-CLAUDE.md hints in plan-phase/quick/profile-user updated. Codex still
  writes AGENTS.md.

Closes #1098

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 16:03:51 -04:00
Tom Boucher
93c5ecd645 feat(#792): add devin-desktop runtime alias for windsurf (#1086)
windsurf now also answers to devin-desktop (CLI --devin-desktop) for the Windsurf→Devin Desktop rebrand; all paths unchanged. The .devin/skills/ workspace migration is split to #1085. Closes #792.
2026-06-11 22:32:49 -04:00
Tom Boucher
9e3b056b15 fix(#779): correct stale model-catalog model IDs verified against live providers (#1047)
Verify-first audit: gemini opus gemini-3-pro→gemini-3.1-pro-preview (undefined in gemini-cli source), codex sonnet gpt-5.3-codex→gpt-5.4 (deprecated per OpenAI); qwen3-coder-next verified valid, unchanged. Adds a regression guard + sourcing note. Closes #779.
2026-06-11 13:36:23 -04:00
Tom Boucher
e4dfa6b9ea fix(#1012): invoke fallow with its real CLI and wire the report normalizer (#1044)
* fix(#1012): invoke fallow with its real CLI and wire the report normalizer

The /gsd-code-review structural pre-pass invoked fallow with flags no published
fallow version accepts (--json, --profile, --stdin-files), so it failed on every
run and degraded silently per REQ-FALLOW-02 — the feature never delivered on any
fallow version. Three compounding defects:

1. Invalid flags. Real fallow audit uses --format json (not --json), -q/--quiet,
   --changed-since/--base for changed-files scoping (no file-list input), and
   --max-crap for thresholds. There is no --profile or --stdin-files.
2. Exit-code handling. fallow audit exits 1 when it FINDS issues (verdict=fail),
   0 when clean. The pre-pass treated any non-zero exit as a crash and discarded
   the output — i.e. it threw away exactly the findings it exists to surface.
   Success is now decided by whether a valid fallow JSON report was produced,
   not by the exit code.
3. Schema mismatch. normalizeFallowReport parsed a fictional top-level schema
   (unusedExports/duplicates/circularDependencies) fallow never shipped, and was
   dead code (the workflow embedded raw JSON; its tests asserted the fictional
   schema, one even calling a non-existent runFallowAudit and passing vacuously).

Fixes: align the invocation to fallow's documented agent-facing pattern; map the
profile preset (minimal/standard/strict) to --max-crap (50/30/15); scope phase
runs via --changed-since with a repo-scope fallback; rewrite the normalizer to
fallow's real schema (dead_code.unused_exports/unused_files/circular_dependencies
+ duplication.clone_groups) and wire it into the workflow so the reviewer
receives normalized findings; replace the fictional-schema fixtures and tests
with real-schema ones and delete the vacuous runFallowAudit test.

Closes #1012

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1012): backfill changeset PR number to 1044

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:36:50 -04:00
Colin Johnson
76f42ddb4b feat(#1014): add Claude Fable 5 model config (#1015) 2026-06-10 20:32:26 -04:00
Jeremy McSpadden
f61b97276e fix(#724): block convergence on actionable review findings (#728)
* fix(#724): block convergence on actionable review findings

* merge: integrate clean next (#936 inline) onto author tip + re-apply cursor fixes (Mode field, REVIEWS.md extraction) and review hardening (#724)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 14:54:15 -04:00
Tom Boucher
29c0a2f5a1 docs(#849): capture 1.4.0 release features across the docs base (#850)
Diataxis review of the 1.4.0 content (52 changesets, multi-runtime maturation
plus native packaging and new flags) against the existing docs base found most
per-feature docs already landed with their PRs. Fill the four remaining gaps,
each in its Diataxis quadrant:

- Reference: FEATURES.md Feature #36 (Multi-Runtime Support) updated in place
  with 1.4.0 additions — native skills emission (Cline/Kilo/OpenCode), new
  slash-command surfaces (CodeBuddy/Augment/Cursor), cross-runtime lifecycle
  hooks for context-headroom tracking, and the Gemini CLI extension package.
- Reference: CONFIGURATION.md gains a dedicated worktree.baseRef entry (values,
  .claude/settings.local.json location, auto-set-on-install behaviour).
- How-to: plan-a-phase.md gains an 'override planning granularity for one phase'
  section for the --granularity flag.
- Explanation: context-engineering.md gains a 'Lifecycle hooks and context
  headroom' section (the why of lifecycle hooks + forked context), cross-linked
  from multi-agent-orchestration.md.

Docs-only; documents already-shipped features, so no changeset required
(docs/ is not in the changeset-lint user-facing prefixes).

Closes #849

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 23:22:15 -04:00
Tom Boucher
f7e902f1cf feat(#52): add agent_skills_security.trusted_global_roots allowlist for global skills (#754)
* feat(#52): add agent_skills_security.trusted_global_roots allowlist

Opt-in allowlist so a global: agent skill whose SKILL.md realpath resolves
outside the default global skills base (e.g. ~/.claude/skills) is accepted
when its real target lies under a user-declared trusted root. Default [] is
byte-identical to prior behavior; the symlink-escape guard is preserved and
simply re-applied against each declared root.

- src/security.cts: loadTrustedGlobalRoots — tilde-expand (~ and ~/), reject
  project-relative and dangerously broad roots (filesystem/UNC root, homedir),
  realpath-canonicalize each root every run and drop non-existent ones.
- src/init.cts: on base-check failure the guard consults the trusted roots
  (hoisted out of the loop); emits a stderr NOTE when a skill is accepted via
  a trusted root so the widened boundary is visible.
- src/core.cts: thread agent_skills_security through loadConfig.
- config-schema.manifest.json: allow the new key path.
- docs/CONFIGURATION.md: document the option and its security model.
- tests/agent-skills.test.cjs: unit + end-to-end CLI coverage (regression,
  feature, negative, broad-root hardening, stderr NOTE).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#52): add changeset fragment for trusted_global_roots (#754)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 01:06:41 -04:00
Tom Boucher
cf8bd3cd5e fix(#683): auto-degrade phase execution to sequential on worktree base mismatch (#749)
* fix(#683): auto-degrade phase execution to sequential on worktree base mismatch

Claude Code forks worktree-isolated executors off the repository default
branch (origin/HEAD), not the orchestrator's HEAD. Running /gsd-execute-phase
on a branch diverged from the default (unmerged milestone/feature branch) left
every executor without the phase's plan files and tripped the
worktree-branch-check guard with `exit 42` — 100% reproducible, all OSes.

- New module src/worktree-base-ref.cts: HEAD-vs-fork-base drift detection
  (origin/HEAD with symbolic-ref fallback) and no-clobber worktree.baseRef
  management, exposed as `worktree base-check` / `worktree set-baseref`.
- execute-phase.md: pre-dispatch, for Claude Code with worktrees enabled,
  auto-degrades the run to sequential on the main tree when a base mismatch
  is detected, recommending worktree.baseRef:"head". The exit-42 guard stays
  as a backstop.
- Installer: fresh local Claude installs set worktree.baseRef:"head" in
  .claude/settings.local.json (no-clobber, respecting an explicit shared
  settings.json value); upgrades print an opt-in notice pointing at
  `gsd-tools worktree set-baseref`.
- Docs: how-to guide, CLI/config reference, planning-config cross-ref.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#683): auto-apply worktree.baseRef on upgrade; gate fresh+upgrade on use_worktrees

Per maintainer direction: on a local Claude Code UPGRADE, set
worktree.baseRef:"head" automatically (no opt-in notice) when the project's
workflow.use_worktrees is enabled, instead of merely printing a remediation
notice. For consistency the FRESH path is now gated the same way: both paths
compute worktrees-enabled once (bounded walk-up read of .planning/config.json,
default enabled unless workflow.use_worktrees === false) and apply the
no-clobber baseRef only when enabled — never overwriting an explicit value in
settings.local.json or a shared settings.json. gsd-tools worktree set-baseref
remains for manual use. Docs + changeset updated; tests hardened (file-exists
assertions, fresh+disabled case, upgrade idempotency).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#683): measure workflow byte-budget on LF, fixing Windows-only CI failure

The workflow-size-budget test failed only on Windows: git checks out the .md
files as CRLF (no eol=lf in .gitattributes) and byteCount used
fs.statSync().size (raw on-disk bytes), counting an extra \r per line. That
inflated execute-phase.md — the XL high-water-mark file pinned near its ceiling
by the tighten-only ratchet — from 88492 LF bytes to ~90245 on Windows, over
the 90000 XL ceiling, while passing on the LF-checkout Mac/Linux runners.

The ceilings are explicitly "calibrated against raw `wc -c`" on an LF checkout,
so the measurement should be LF-based on every platform. byteCount now reads the
file and counts Buffer.byteLength after stripping CR, making the budget
platform-independent (a no-op on LF checkouts; verified statSync === normalized
for all 88 workflow files). No ceilings changed. Added a regression test
asserting CRLF and LF content of the same file count identically.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#683): make worktree-base-ref test path mocks Windows-safe (path.join)

tests/worktree-base-ref.test.cjs keyed its injected readFile/writeFile mocks
(and a few expected `file` values) with forward-slash template literals like
`${claudeDir}/settings.local.json`. The module composes those paths with
path.join(), which emits backslashes on Windows, so the mock keys never matched
the module's lookup → readFile returned null → resolveEffectiveBaseRef /
cmdWorktreeBaseCheck / cmdWorktreeSetBaseRef (and the JSONC variants) failed on
the Windows full-test runner only (they passed on Mac/Linux, and the install
tests passed because they use the real filesystem). The module is correct;
only the test fixtures hardcoded '/'.

All mock keys and path assertions now use path.join(base, ...) mirroring the
module, so they match on every platform (no-op on POSIX). 19 path references
across 16 lines.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 23:40:24 -04:00
Tom Boucher
463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00
Tom Boucher
3bb2f8f1c5 docs: rebrand to GSD Core and restructure docs with Diataxis (#605)
* chore: wire docs/agents config into AGENTS.md Agent skills section

Add the `## Agent skills` discovery block pointing the engineering
skills at the existing docs/agents/{issue-tracker,triage-labels,domain}.md
files (issue tracker, triage label mapping, single-context domain docs).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: rebrand to GSD Core and restructure docs with Diataxis

Reorganise the root README and docs/ around the Diataxis framework
(tutorials, how-to guides, reference, explanation), add new how-to
guides and schema references (STATE.md / CONTEXT.md / PLAN.md /
planning artifacts), and cross-link the whole set. Update the lone
legacy gsd-build reference to open-gsd; keep internal get-shit-done/
filesystem paths unchanged (directory rename tracked separately in
open-gsd/gsd-core#604). Regenerate the ja-JP, ko-KR, pt-BR and zh-CN
localised trees to mirror the new structure.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: backfill changeset PR number (#605)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 08:13:09 -04:00
Tom Boucher
a11ba2dfcb feat(#68): per-phase granularity overrides (granularities.<phaseType>) (#595)
Closes #68. Per-phase-type granularity overrides via granularities.<phaseType>, mirroring models.<phaseType>. Includes maintainer-authorized sdk-seam reference cleanup.
2026-06-01 20:53:46 -04:00
Tom Boucher
9ffe45a7c3 feat(#163): tighten gsd-roadmapper granularity defaults to reduce thin-phase fragmentation (#591)
* feat(#163): tighten gsd-roadmapper granularity defaults to reduce thin-phase fragmentation

Tighten the Granularity Calibration buckets in gsd-roadmapper (Coarse 3-5->2-4,
Standard 5-8->4-6, Fine 8-12->6-10) and append inline Key guidance naming the
thin-phase failure pattern (single requirement / internal-quality goal /
task-shaped success criteria) with instruction to fold into a neighbor rather
than create a standalone phase. Implements the maintainer-approved proposal
verbatim.

Update the canonical English docs that hardcoded the old phase-count numbers:
docs/CONFIGURATION.md and docs/FEATURES.md. Translated docs are
community-maintained and are not updated per-PR (CONTRIBUTING.md language
policy).

Prompt/doc text only; no code, format, or downstream-consumer changes. Agent
size-budget and skills-awareness tests pass; full suite green.

Closes #163

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#163): add Changed changeset for roadmapper granularity tightening

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#163): lock tightened gsd-roadmapper granularity buckets

source-text-is-the-product test asserting the Granularity Calibration table
holds the tightened ranges (Coarse 2-4, Standard 4-6, Fine 6-10), that no row
maps to an old bucket, and that the Key paragraph carries the thin-phase
folding guidance. Would fail if the values regress.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 20:37:49 -04:00
Tom Boucher
2ba6b69d53 feat(#49): provider-neutral model policy presets
* feat(#49): provider-neutral model policy presets

Adds model_policy config surface with known-provider presets (openai/anthropic/google/qwen) and generic provider escape hatch. model_policy.runtime_tiers resolves before legacy model_profile_overrides. reasoning_effort is stripped for unsupported runtimes.

Closes #49

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#49): replace unregistered /gsd-settings-advanced token in docs

docs-parity-live-registry enforces every /token in docs/*.md maps to
a live command. /gsd-settings-advanced is a workflow filename, not a
registered command — use /gsd:settings instead.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#49): update INVENTORY.md count and manifest for config-types.cjs

inventory-counts and inventory-manifest-sync tests require the headline
count and INVENTORY-MANIFEST.json to reflect every file in bin/lib/.
config-types.cjs (new module added by feat(#49)) was missing from both.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-01 08:46:55 -04:00
Tom Boucher
0a12b06381 feat(#39): milestone-prefixed phase IDs (M-NN convention) + migration tool + validation (#565)
* feat(#39): milestone-prefixed phase IDs (M-NN convention) + migration tool + validation

- Add getMilestoneFromPhaseId() / getPhaseDirFromPhaseId() helpers to core.cjs
- Fix isDirInMilestone to match M-NN-style dirs (02-01-setup) against M-NN ROADMAP headings
- Extend heading regex to tolerate [bracket-token] scope prefix on phase headings
- Add W021 validation rule for milestone prefix mismatch
- Add gsd-tools roadmap validate + roadmap upgrade --convention milestone-prefixed
- Add phase_id_convention config field (null default, backwards-compatible)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#39): address 4 Codex review findings in milestone-prefixed phase ID implementation

- getMilestoneFromPhaseId: tighten regex to require a digit after the hyphen (rejects '1-' and '1-abc')
- isDirInMilestone: use convention-aware regex — only capture M-NN segments when ROADMAP itself uses hyphenated phase IDs, preventing legacy dirs like '01-02-setup' from being misread as phase '1-02'
- checkW021: add UNPREFIXED_PHASE_RE path so unprefixed headings (### Phase 1:) also fire W021 when convention is milestone-prefixed
- roadmap-upgrade: remove isMigratedDirName dir-name check (false-positive for legacy dirs); config + ROADMAP heading checks at lines 194 and 212 are sufficient

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: update changeset pr reference to #565

* fix(#39): restore phaseDirNameRe 2-digit minimum; add roadmap-upgrade to inventory

- validate.cjs: \d{1,} → \d{2,} to keep single-digit prefix rejection per W005 contract
- docs/INVENTORY.md: 79 → 80, add roadmap-upgrade.cjs row
- docs/INVENTORY-MANIFEST.json: regenerated (roadmap-upgrade.cjs entry)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-31 23:15:55 -04:00
Tom Boucher
0fbe1d899e chore(#191): retire the gsd-sdk shim — route everything at gsd-tools (#522)
* chore(#191): migrate gsd-sdk query call sites to gsd-tools query

Retiring the gsd-sdk shim. gsd-tools.cjs already accepts `query` as a
meta-prefix (gsd-tools query <command>), so this is a behavior-preserving 1:1
swap across the runtime reference prompts, the graphify hook's commit-detection
gate, and two bin/lib comment/message references.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#191): remove vestigial gsd-sdk shim code from installer + projection

The gsd-sdk shim was already not wired up (no gsd-sdk bin in package.json;
buildWindowsShimTriple had zero call sites). Remove the dead code:
- shell-command-projection.cjs: buildWindowsShimTriple + formatSdkPathDiagnostic
  (+ their now-unused PACKAGE_NAME import) and exports
- install.js: the re-export wrappers + imports, the #3406 stale-standalone-sdk
  detection (detectStaleStandaloneSdk/formatStaleStandaloneSdkWarning + its
  global-install call site), and the exports

Preserved (retained, not gsd-sdk): buildCodexHookWindowsShimIR (#3426) — only
its comments referenced the gsd-sdk pattern; reworded. Also kept the
homePathCoveredByRc 'reopen your shell' branch in maybeSuggestPathExport — its
logic is bin-dir-agnostic, only the message mentioned gsd-sdk; reworded to use
the actual bin dir.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#191): update tests for retired gsd-sdk shim

- bug-3441/bug-3442: drop the formatSdkPathDiagnostic / buildWindowsShimTriple
  assertions (functions removed); retained PATH-action + drift-guard tests stay
- bug-505: remove the 'still exported' assertions for detectStaleStandaloneSdk /
  formatStaleStandaloneSdkWarning / the shim contract surface (#505 kept them;
  #191 removes them)
- graphify-auto-update: migrate the hook-dispatch inputs gsd-sdk query commit ->
  gsd-tools query commit to match the migrated commit hook

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#191): point active docs at gsd-tools query (gsd-sdk shim retired)

Update the user/agent-facing docs (AGENTS, COMMANDS, CONFIGURATION, USER-GUIDE,
ship-pr-body-sections) that presented gsd-sdk query as a current command to
gsd-tools query. Historical docs (ADRs, PRDs, release notes) left untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#191): correct state.load vs state.json description for gsd-tools query

Adversarial-review (codex) finding: the migrated USER-GUIDE line claimed both
'gsd-tools query state.json' and 'state.load' resolve to the frontmatter-rebuild
handler. Verified they don't — state.load returns the CJS load shape
(config + state_raw + flags), state.json returns the frontmatter shape. Both are
available via gsd-tools query; corrected the text to say so.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#191): add changeset for gsd-sdk shim retirement

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-05-30 19:19:14 -04:00
Tom Boucher
79002a00cb chore(#518): rename npm package + bin to @opengsd/gsd-core (#519)
* chore: rename npm package + bin to @opengsd/gsd-core (functional)

- package.json: name @opengsd/get-shit-done-redux → @opengsd/gsd-core,
  bin key get-shit-done-redux → gsd-core, repository/homepage/bugs URLs
- package-lock.json: regenerated (npm install --package-lock-only)
- tests/**, scripts/**, bin/**, .github/**, agents/**, commands/**,
  get-shit-done/bin/**, get-shit-done/workflows/**:
  applied the 4-rule replacement (scoped npm ref, GitHub repo path,
  bin/clone invocations) per #505 single-source refactor

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: sweep live references to @opengsd/gsd-core

Update all live documentation (README.md + translations, docs/**,
CONTRIBUTING.md, VERSIONING.md, SECURITY.md, CONTEXT.md,
docs/CANARY.md) to reflect the renamed package and repository.

Rules applied:
- @opengsd/get-shit-done-redux → @opengsd/gsd-core (scoped npm name)
- open-gsd/get-shit-done-redux → open-gsd/gsd-core (GitHub repo)
- GSD-redux/get-shit-done-redux → open-gsd/gsd-core (stale badge org)
- bare bin/clone refs → gsd-core

CHANGELOG.md, docs/adr/**, docs/RELEASE-*.md, docs/research/**,
and .changeset/** are preserved byte-identical.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: add negative lookbehind to slash-command regex in bug-2954 test

The extractSlashReferences regex matched /gsd-core inside npm package
URLs (@opengsd/gsd-core), producing a false /gsd:core command reference.
Adding a negative lookbehind (?<![a-z]) excludes matches preceded by a
letter, so only standalone /gsd-<cmd> and /gsd:<cmd> tokens are found.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#518): add changeset for package rename

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#518): update package-identity expectations to the renamed coordinates

The rebase regenerated the seam to @opengsd/gsd-core (bin gsd-core, repo
open-gsd/gsd-core). The #498 seam tests assert deriveIdentity against the REAL
package.json, so their expected literals must follow the rename. The drift-lint
unit test is left as-is — its SEAM is a self-consistent fixture and its
stale-literal detection cases would shift if altered; the live-repo scan in it
already passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-05-30 17:25:02 -04:00
Tom Boucher
05cdec5f47 feat(#22): plan-vs-codebase drift guard (source-grounded reviewer + intel surface) (#487)
* feat(#22): add plan_review.source_grounding + _authority config keys

Two additive opt-out keys for the drift guard: source_grounding (bool,
default true) gates the source-grounded reviewer pass; _authority (enum
grep|intel|treesitter|lsp|scip, default grep) selects the resolver rung.
No existing default changed.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#22): add intel api-surface renderer + CLI subcommand

Renders .planning/intel/api-map.json into a human-readable API-SURFACE.md
for planner injection. Empty/missing map still writes a surface that
announces itself incomplete (absence = unknown, not 'does not exist').
Gated on intel.enabled like all intel functions.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#22): add source-grounding pass to plan-review-convergence

Default-on reviewer pass (plan_review.source_grounding) that enumerates
every symbol a plan cites, excludes declared new artifacts, resolves each
against source via the configured authority adapter, and records
three-valued verdicts. rung-0/1 MISSING is needs-acknowledgement, not a
hard block; UNCHECKABLE is logged in a REVIEWS.md coverage section.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#22): inject API-SURFACE.md into planner + require Artifacts section

When intel.enabled, plan-phase regenerates API-SURFACE.md and injects it
as a HINT (prefer, may be incomplete, absence = unknown), never a hard
rule. Every plan must now emit an 'Artifacts this phase produces' section
so the source-grounding reviewer can separate new symbols from references
to existing code.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#22): surface drift-guard in setup + settings, add docs

/gsd:new-project asks to enable plan_review.source_grounding (default Y);
/gsd:settings exposes the toggle and authority knob. Documents both config
keys in CONFIGURATION.md, the intel api-surface command in COMMANDS.md,
the drift guard in USER-GUIDE.md, and links ADR 22 from ARCHITECTURE.md.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#22): respect AskUserQuestion 4-option cap and plan-phase XL line budget

settings drift-guard toggle moved to its own 2-option question; #22
plan-phase additions condensed to bring the file back under the 1810-line
XL budget without dropping the intel gate, the incomplete-surface hint, or
the Artifacts-section requirement.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#22): use live slash-command forms in drift-guard docs

Doc-parity gate requires every slash-command token in docs/*.md to resolve
to a registered command. Corrected the command form(s) referenced in the
#22 drift-guard / api-surface documentation.

The unresolved token was /gsd-core, matched from the GitHub repo reference
"open-gsd/gsd-core#22" in docs/adr/22-plan-drift-guard.md. This is the
same pattern as the existing 'test-runner' exemption (open-gsd/gsd-test-runner).
Added 'core' to INTERNAL_COMPONENT_SLUGS with a matching explanatory comment.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#22): add changeset fragment for drift guard (PR #487)

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 17:08:11 -04:00
Tom Boucher
c7e5a88353 enh(#466): refresh opus-tier model IDs to current GA (Opus 4.8 / codex gpt-5.5) (#467)
* enh: bump opus-tier model IDs to current GA (Opus 4.8 / codex gpt-5.5)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#466): changeset for opus-tier model-ID refresh

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-29 11:52:47 -04:00
Tom Boucher
5ca646f015 feat(#443): unified cross-provider effort controls + fast-mode-aware routing (#463)
* test(#443): RED unified effort + fast_mode + resolve-execution

All 68 tests failing as expected — no implementation yet.
Covers: effort cascade (tier defaults, overrides, invalid fallthrough),
fast_mode cascade (boolean-only, tier defaults), resolveEffortForTier
escalation, renderEffortForRuntime clamping, resolve-execution CLI,
config schema new keys, QA hostile-input matrix.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(#443): unified cross-provider effort + fast_mode knobs and resolve-execution query

Adds config-driven effort control (universal ladder: minimal<low<medium<high<xhigh<max)
and fast_mode propagation knobs, with per-runtime rendering that clamps the unique
tail values (max=Anthropic-only clamps to xhigh on Codex; minimal=Codex-only clamps
to low on Claude).

Key changes:
- config-schema.manifest.json: add effort.default, fast_mode.enabled as validKeys;
  add 4 dynamicKeyPatterns for effort.routing_tier_defaults, effort.agent_overrides,
  fast_mode.routing_tier_defaults, fast_mode.agent_overrides; fix stale _comment
- config-defaults.manifest.json: add effort and fast_mode blocks with tier defaults
- model-catalog.cjs: add EFFORT_RENDERING map, renderEffortForRuntime(), RUNTIMES_WITH_FAST_MODE
- model-profiles.cjs: re-export new catalog exports
- core.cjs: add resolveEffortInternal, resolveFastModeInternal, resolveEffortForTier,
  VALID_EFFORTS, EFFORT_SET, nextEffort; pass effort/fast_mode through loadConfig
- commands.cjs: replace reasoning_effort in cmdResolveModel with unified effort;
  add cmdResolveExecution (superset command with effort_rendered, effort_param,
  effort_propagation, fast_mode, fast_mode_supported)
- gsd-tools.cjs: add resolve-execution case with --effort/--fast-mode/--attempt flags
- tests/feat-443: 69 tests covering cascade, rendering, escalation, CLI, schema, QA matrix
- tests/commands.test.cjs: convert 3 reasoning_effort assertions to unified effort
- docs/CONFIGURATION.md: document effort + fast_mode + resolve-execution sections
- settings-advanced.md: list new effort/fast_mode keys in confirmation table

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#443): remove dead catalog effort lane; unify codex effort through renderEffortForRuntime

- Remove resolveReasoningEffortInternal (catalog-driven effort function) from
  core.cjs and its export; remove from commands.cjs destructure import
- Convert tests/issue-2517-runtime-aware-profiles.test.cjs: all 11 effort
  assertions now use resolveEffortInternal + renderEffortForRuntime; Claude
  effort is first-class (output_config.effort); unknown runtimes assert param===null
- Convert tests/feat-3023-model-phase-types.test.cjs: replace the entire
  resolveReasoningEffortInternal describe with unified effort assertions;
  effort derives from AGENT_DEFAULT_TIERS routing tier, not phase-type tier

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#443): ADR for unified cross-provider effort + fast-mode routing

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* test(#443): architecture-level QA invariants + test-strategy doc

Add 48-test integration suite (feat-443-effort-fast-mode.integration.test.cjs)
covering 8 architectural invariants: cross-provider validity (never emit a value
the real API would 400 on), param/channel contract stability, resolve-execution
JSON contract (all 8 keys + correct types), totality across the full 33-agent
registry, fast-mode honesty (claude always fast_mode_supported=false), precedence
first-valid-wins matrix for both effort and fast_mode cascades, dynamic-routing
composition (effort escalation independent of model tier), and config-set round-trip
for all new effort/* and fast_mode/* key namespaces. Append test-strategy section
with invariant rationale and E2E gap documentation to docs/TESTING-SUITES.md.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#443): add failing install-wiring tests for effort per-runtime injection (RED)

TDD RED: 10 failing tests covering:
- Claude .md gets effort: injected per tier (planner=xhigh, mapper=low, executor=high)
- Gemini .md does NOT get effort: (already passing — Gemini-safe)
- Codex .toml gets model_reasoning_effort via unified resolver
- Config-driven: effort.agent_overrides drives both Claude .md and Codex .toml
- Source purity: agents/*.md have no effort: key (already passing)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(#443): wire effort per-runtime at install (Claude .md frontmatter + Codex .toml unified)

- Import AGENT_DEFAULT_TIERS and renderEffortForRuntime from model-catalog.cjs
- Add readGsdEffectiveEffortConfig(targetDir): reads merged effort config from
  .planning/config.json (per-project wins) + ~/.gsd/defaults.json (global fallback),
  same probe pattern as readGsdRuntimeProfileResolver
- Add resolveInstallTimeEffort(effortCfg, agentName): pure function matching
  resolveEffortInternal() precedence (agent_overrides > routing_tier_defaults > default > 'high')
  without loadConfig side-effects (no sub-repo detection, no migration writes)
- Claude agent copy loop: inject `effort: <value>` into frontmatter ONLY for
  runtime === 'claude'; all other .md runtimes (Gemini, Qwen, Hermes, etc.) stay
  effort-free (Gemini-safe source contract preserved in agents/*.md)
- generateCodexAgentToml: add effortCfg param; emit model_reasoning_effort from
  unified resolver (replaces old catalog entry.reasoning_effort); Codex clamps
  max → xhigh via renderEffortForRuntime('codex', ...)
- installCodexConfig: pass readGsdEffectiveEffortConfig(targetDir) to
  generateCodexAgentToml so per-project config wins for Codex .toml too
- Update failing tests to GREEN: 12/12 pass; all 17 install tests pass;
  2847/2848 unit tests pass (1 pre-existing failure: policy-shell-pinning)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#443): source install effort defaults from manifest (kill drift) + guard test

Replace hardcoded _GSD_EFFORT_MANIFEST_TIER_DEFAULTS and the 'high' fallback in
resolveInstallTimeEffort with values read from config-defaults.manifest.json at
module init, using the same __dirname-relative path install.js already uses for
all shared manifests. Add feat-443-effort-defaults-drift.test.cjs to assert
equality between install.js's runtime constants and the manifest on every CI run.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): reconcile Codex TOML tests with unified effort design

The #443 unified effort resolver makes generateCodexAgentToml always emit
model_reasoning_effort (driven by resolveInstallTimeEffort, not model_profile_overrides).
The test 'generated TOML omits reasoning_effort when runtime has none' had an
obsolete premise — model_profile_overrides.reasoning_effort:'' no longer suppresses
unified effort. Convert it to assert the new invariant: Codex TOML always carries a
valid model_reasoning_effort from the agent's routing tier (xhigh for gsd-planner,
a heavy-tier agent), while model_profile_overrides model override is still respected.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): make install.js effort resolution lazy (no load-time side effects breaking launcher-parity)

Replace module-load-time IIFE + hard throw (config-defaults.manifest.json read)
and top-level require of model-catalog.cjs with a lazy _getGsdEffortCatalog()
getter that initialises on first call from resolveInstallTimeEffort /
generateCodexAgentToml / Claude .md effort injection.  Requiring install.js in
unrelated test contexts (e.g. runtime-launcher-parity) no longer triggers
manifest IO or throws, eliminating the load-time side effect that changed
subprocess exit codes / stderr on the bench.

Drift-guard exports (_GSD_EFFORT_MANIFEST_TIER_DEFAULTS / _GSD_EFFORT_MANIFEST_DEFAULT)
preserved as lazy getter properties on module.exports so feat-443-effort-defaults-drift
still validates them without forcing eager load.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): isolate install-wiring test HOME to stop \$HOME/.claude pollution breaking launcher-parity

runGlobalInstall() now redirects HOME to a per-call isolated tmpdir in addition
to the existing runtime-specific env-var redirects (CLAUDE_CONFIG_DIR,
GEMINI_CONFIG_DIR, CODEX_HOME). This ensures install.js code that uses
os.homedir() directly — including the ~/.cache/gsd update-check deletion,
~/.gsd/defaults.json reads, and any HOME-relative npm subprocess writes —
never touches the real \$HOME during the test.

Without the HOME isolation the install test (which is new to this branch and
is now picked up by Docker's raw \`tests/*.test.cjs\` glob) could write or
delete files under the real \$HOME, causing runtime-launcher-parity test (D)
to fail: (D) asserts a loud non-zero exit when \$RUNTIME_DIR/gsd-tools.cjs is
absent and gsd-tools is not on PATH, but the launcher's \$HOME/.claude fallback
arm succeeds if \$HOME/.claude/get-shit-done/bin/gsd-tools.cjs exists.

Also sets GSD_SKIP_STALE_SDK_CHECK=1 to suppress the \`npm ls -g\` subprocess
that the global installer spawns — irrelevant to effort-wiring assertions,
slow, and potentially writes to ~/.npm cache.

All 12 feat-443 install-wiring assertions preserved. Drift-guard 5/5. Unit
suite 2848/2850 (pre-existing policy-shell-pinning.test.cjs failure on next).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#443): add changeset fragment for effort + fast-mode routing

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(#443): set GSD_TEST_MODE before requiring install.js in drift-guard test to prevent HOME leak

Without GSD_TEST_MODE=1, require('bin/install.js') runs the module's main
install block (guarded by !GSD_TEST_MODE), performing a real global Claude
install into $HOME/.claude/. On CI ubuntu where node is on standard PATH,
the launcher's $HOME/.claude fallback arm then finds gsd-tools.cjs, causing
runtime-launcher-parity test (D) to exit zero when it must exit non-zero.

Root cause: feat-443-effort-defaults-drift.test.cjs (unit suite) runs
alphabetically before runtime-launcher-parity.test.cjs in the same node
--test invocation. Each runs in a separate worker process but shares the
same HOME. The drift test's install leaks gsd-tools.cjs into that HOME,
then the launcher test's bash subprocess finds it via the $HOME/.claude arm.

Fix: add process.env.GSD_TEST_MODE = '1' at the top of the drift-guard
test, before the require(installPath) call. This matches the pattern used
by feat-443-effort-fast-mode.test.cjs and feat-443-effort-install-wiring
.install.test.cjs.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): deterministic resolve-execution arg parsing + validate install-time effort (Codex adversarial findings)

Finding 1: resolve-execution --effort low gsd-planner misrouted 'low' as the agent.
Replace find(non-dash) with a proper flag-consuming loop that collects a single
positional; validate missing/extra positionals and malformed --attempt values.

Finding 2: resolveInstallTimeEffort returned unvalidated effort strings (e.g. "ultra")
verbatim. Each precedence layer now checks GSD_EFFORT_SET (imported once from
core.cjs) before accepting a value, mirroring resolveEffortInternal exactly.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#443): newline-agnostic effort frontmatter injection (Windows CRLF) + CRLF-safe assertions

Extracts injectEffortFrontmatter(content, effortValue) pure helper that detects
EOL (LF vs CRLF) from the opening '---' line and inserts 'effort: <value>'
before the closing '---' delimiter using the same EOL as the surrounding
frontmatter. Regex now uses /^---\r?\n([\s\S]*?)^---\r?$/m instead of the
LF-only /^(---\n[\s\S]*?)(---)(\n|$)/ that silently skipped CRLF files on
Windows (git core.autocrlf=true checkout).

Also adds 7 unit tests covering LF, CRLF, idempotency, no-frontmatter, and
complex frontmatter cases. Exports injectEffortFrontmatter from module.exports.

Fixes 6 CI failures in tests/feat-443-effort-install-wiring.install.test.cjs
on windows-latest runners (lines 138, 145, 152, 261, 345, 356).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-29 10:32:58 -04:00
Tom Boucher
5f3eb42864 feat(observability): propagate parentTraceId on DispatchEvent — ADR-0174 SDK retirement Phase 1.4 (#178) (#225)
* test(#178): update DispatchEvent factory tests to propagate parentTraceId

P1.3 test 'parentTraceId is always undefined' replaced with four P1.4
contracts: absent → undefined, string → propagated, null → undefined,
non-string → undefined (defensive normalization policy).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(#178): propagate parentTraceId through DispatchEvent factory

Stop ignoring the parentTraceId parameter added as a forward-compat hook
in P1.3. Defensive normalization: only non-null strings are propagated;
null, non-string values, and absent callers all yield undefined, keeping
P1.3 behavior intact for all existing dispatch call sites.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#178): add Hub-level parentTraceId propagation tests

Four new assertions: req.parentTraceId propagates to event, absent →
undefined (P1.3 regression), shared parentTraceId across multiple
dispatches, and unique traceId invariant despite shared parentTraceId.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(#178): plumb parentTraceId through Hub dispatch and _notifyLogger

dispatch() now reads req.parentTraceId and passes it to _notifyLogger,
which forwards it to makeDispatchEvent. Backward-compatible: callers
that omit parentTraceId emit events with parentTraceId: undefined,
identical to P1.3 behavior.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#178): add trace correlation end-to-end test

Dispatches a root command then 3 children with parentTraceId=rootTraceId.
Reads the real .gsd-trace.jsonl audit file and verifies: 4 events total,
root has no parentTraceId, all children carry rootTraceId, all traceIds
unique, JS filter returns exactly the 3 children given the root's traceId.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#178): document traceId/parentTraceId in audit file

Update Observability section to note that audit events now carry both
traceId and parentTraceId, and explain the correlation filter pattern.
Note that leaf dispatches emit parentTraceId: undefined until the Phase 2
composer wires it automatically.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#178): add changeset for trace correlation seam

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#178): cover invalid parentTraceId values in DispatchEvent factory

Adds 9 new test cases for UUID v4 validation of parentTraceId:
empty string, whitespace, non-UUID, oversized, UUID v1, missing-hyphen,
extra-char (all dropped to undefined), plus UPPERCASE and lowercase v4
(both propagated). Tests are intentionally red until the implementation
commit that follows.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(#178): validate parentTraceId against UUID v4 before propagation

Adds UUID_V4_REGEX constant and isValidParentTraceId() helper to
event.cjs. makeDispatchEvent now silently coerces any parentTraceId that
fails the UUID v4 check (wrong version nibble, wrong variant, missing
hyphens, oversized, empty, etc.) to undefined. No stderr warn is emitted
— the factory remains pure and side-effect-free. Closes the correlation-
poisoning vector identified in the Codex adversarial review of PR #225.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#178): assert Hub silently drops invalid parentTraceId at the seam

Adds two tests to hub-logger-integration.test.cjs:
1. dispatch with 'junk' parentTraceId emits event with parentTraceId===undefined.
2. The logger-failure warn path is NOT triggered — the factory coerces the bad
   value before onEvent is called, confirmed by zero stderr output even when a
   logger that would throw on non-undefined parentTraceId is installed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#178): assert invalid parentTraceId does not poison correlation siblings

Adds one test to trace-correlation.test.cjs: dispatches a root, a valid
child (parentTraceId = rootTraceId), and an invalid child (parentTraceId =
'junk'). Asserts: valid child carries correct parentTraceId, invalid child
has parentTraceId dropped to undefined, filtering by rootTraceId yields
exactly 1 event (the valid child only), and all 3 events have unique
traceIds. Uses an isolated Hub + tmpdir to avoid shared fixture interference.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#178): document UUID v4 contract for parentTraceId

Appends one sentence to the Observability audit-trail paragraph in
CONFIGURATION.md: parentTraceId must be canonical UUID v4 (RFC 4122);
values that don't match are silently dropped from audit output. No section
restructuring — single sentence addition only.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-24 16:32:28 -04:00
Tom Boucher
2d14eb8873 feat(observability): add DispatchLogger seam — ADR-0174 SDK retirement Phase 1.3 (#177) (#223)
* test(#177): add DispatchEvent factory failing tests

Red tests for makeDispatchEvent shape, traceId UUID v4, uniqueness,
parentTraceId-always-undefined (P1.3), args redaction toggle, ISO 8601
timestamp, and all result variant passthrough.

* feat(#177): introduce DispatchEvent factory

makeDispatchEvent produces an immutable event record per dispatch:
- traceId: crypto.randomUUID() (UUID v4)
- parentTraceId: always undefined (P1.4 wires composer)
- command, result, timestamp (ISO 8601)
- args only included when includeArgs === true (default: omitted)

* test(#177): add arg redaction policy failing tests

Red tests for shouldIncludeArgs (GSD_AUDIT_ARGS env gating) and
redactEvent (strips args from frozen events, preserves all other
fields, returns a new object, never mutates the source).

* feat(#177): introduce arg redaction policy

shouldIncludeArgs(): only GSD_AUDIT_ARGS==='1' opts in; all other
values (unset, '', '0', 'true') default to omitting args.

redactEvent(event): returns a shallow copy of the event, dropping the
args field unless opted in. Never mutates the (frozen) source event.

* test(#177): add DispatchLogger interface failing tests

Red tests covering:
- no-op logger: silent on all events, never throws
- default logger: silent on ok, one flattened JSON line to stderr on error
- default logger: audit file creation + append-only + redaction + config gate
- GSD_AUDIT env var and config.audit.enabled config gate
- GSD_AUDIT_ARGS opt-in for args inclusion
All tests use real fs under os.tmpdir() — no mocked appendFileSync.

* feat(#177): introduce DispatchLogger with default and no-op implementations

createNoOpLogger(): silent on all events — Hub default when no logger injected.
createDefaultLogger({ cwd, config }):
  - Silent on ok result
  - Flattened JSON line to stderr on error: { kind, traceId, ...typedPayload }
  - Append-only audit at .planning/.gsd-trace.jsonl when GSD_AUDIT=1 or config.audit.enabled
  - Args redacted by default; GSD_AUDIT_ARGS=1 opts in
  - Logger errors caught internally; never break dispatch callers

* test(#177): add Hub+logger integration failing tests

Red tests verifying:
- onEvent called exactly once per dispatch (ok, error, handler-throw, unknown)
- DispatchEvent shape: traceId uniqueness, command, result.kind, parentTraceId
- Logger errors contained (dispatch still returns Result, warn line to stderr)
- Hub defaults to no-op when no logger injected
- End-to-end with createDefaultLogger: silent on success, stderr on error, audit file

* feat(#177): wire DispatchLogger into CommandRoutingHub

Add optional logger param to createHub({ ..., logger }).
Defaults to createNoOpLogger() — silent, no behaviour change for callers
that don't inject a logger.

After every dispatch (success and error):
- Normalises HubResult { ok } to DispatchEvent { kind: 'ok'|error-kind }
- Calls makeDispatchEvent({ command, args, result }) to mint the event
- Calls logger.onEvent(event) exactly once
- Wraps in try/catch: logger errors emit { level:'warn', source:'DispatchLogger' }
  to stderr but never propagate to dispatch callers

* chore(#177): gitignore .planning/.gsd-trace.jsonl audit file

The audit trail is local-only, append-only, and must never be committed.
Slotted under the existing "Local scratch + Claude-test artifacts" block.

* docs(#177): document GSD_AUDIT, GSD_AUDIT_ARGS, config.audit.enabled

New ## Observability section at end of CONFIGURATION.md covering:
- Default silent/stderr behaviour overview
- Stderr error JSON format
- Audit file opt-in (env var and config key)
- Args redaction policy and GSD_AUDIT_ARGS opt-in

Also slots GSD_AUDIT and GSD_AUDIT_ARGS into the existing
## Environment Variables table (alphabetical order).

* chore(#177): add changeset for observability seam

type: Added — new DispatchLogger seam with default silent/stderr/audit behaviour.
2026-05-24 15:22:36 -04:00
Tom Boucher
2a915c1b82 chore: migrate references from gsd-build to open-gsd/get-shit-done-redux (#120) (#121)
Security-motivated migration of all stale repository and npm-scope references.

Three categories of changes (58 files, 174 substitutions):

1. gsd-build → open-gsd (security-critical):
   - .github/workflows/release-sdk.yml — npm token comment, tarball filename pattern
   - .github/workflows/hotfix.yml — same
   - .changeset/fix-3406-detect-stale-sdk-shadow.md — @gsd-build/sdk → @open-gsd/sdk
   - .changeset/sharp-quails-leap.md — same
   - get-shit-done/workflows/update.md — CHANGELOG raw GitHub URL

2. GSD-redux org slug → open-gsd (canonical rename):
   - package.json + sdk/package.json — repository/homepage/bugs metadata
   - All README.*.md — live badge and link sections
   - CONTRIBUTING.md, CONTEXT.md, QUICK-WINS-CONFIRMED-BUGS.md
   - .coderabbit.yaml, .release-monitor.sh, scripts/sync-rulesets.sh
   - docs/** — all live agent/ADR/user-facing documentation
   - tests/** — repo slug assertions and test fixtures
   - scripts/changeset/cli.cjs + github-release-notes.cjs
   - .github/ISSUE_TEMPLATE/*, .github/pull_request_template.md
   - bin/install.js, get-shit-done/bin/lib/model-catalog.cjs
   - sdk/HANDOVER-*.md, sdk/src/*.test.ts

3. CLAUDE.md (gitignored local file — not in this commit):
   Updated separately outside git: --repo gsd-build/get-shit-done →
   --repo open-gsd/get-shit-done-redux with security warning.

Intentionally unchanged: CHANGELOG.md, docs/RELEASE-*.md,
.changeset/README.md, .changeset/build-hooks-atomic-write.md,
README.md migration table (historical fork record),
tests/changeset-serialize.test.cjs line 78 (serialization fixture).

The gsd-build/get-shit-done repo is compromised (rug-pull documented in
README.md). Do not push to or interact with that repo.

Closes #120
2026-05-22 12:28:16 -04:00
Tom Boucher
74cb493373 fix(3784): expose adaptive in model_profile settings flow (#91)
* fix(3784): expose adaptive in model_profile settings flow

Split the single 4-option model-profile AskUserQuestion into a two-question
flow: Q1 (Adaptive / Standard tier / Inherit) routes top-level intent; Q2
(Quality / Balanced / Budget) appears only when Standard tier is chosen.
Updates the confirm table and success_criteria to include adaptive.

Adds regression test asserting all five valid profiles are reachable
interactively via the settings UI.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* changeset: add Fixed entry for #3784 / PR #3795

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(3784): correct Q2-skip comment and remove duplicate brace in settings.md

Codex review followup:
- Replaced vague "preserve existing config" comment with accurate description:
  Q1 still writes model_profile on Adaptive/Inherit branches; only Q2 is skipped.
- Removed stray duplicate `{` line before the Spawn Plan Researcher question block
  (pseudocode had two consecutive `{` openers, one spurious).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(3784): address review — gate Q2 structurally, define cancel rule, harden tests

Addresses gsd-code-reviewer (M1/M2/m1/m2/m3/m4) and codex adversarial
(Q2 gating, save-mapping, Claude-only wording, step-of-2 wording).

- F1: Replace //comment-only Q2 gating with Conditional visibility block
  (mirrors code_review_depth / graphify.auto_update structural pattern)
- F2: Define model_profile cancel rule in update_config step (leave
  existing value unchanged when Q1="Standard tier…" but Q2 cancelled)
- F3: Fix Adaptive description — remove "Claude only" tail; describe
  heavy/light role tiers across all supported runtimes
- F4: Remove "step 1 of 2 for standard profiles" from Q1 question text
  (2-step nature now structurally documented by Conditional visibility)
- F5: Fix vacuously-true test disjunct (|| content.includes('Adaptive')
  always true — 6+ occurrences); assertion now requires role-based cost
  optimization + heavy roles wording
- F6: Add 4-option cap enforcement test (ASK_USER_QUESTION_OPTION_CAP=4
  named constant, counts per question object not per AskUserQuestion call)
  and brace-balance regression test (guards against bd53925f recurrence)

* docs(3784): list adaptive in model_profile reference docs

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-22 11:39:57 -04:00
Tom Boucher
dff176bfd2 chore: rebrand to GSD-redux/get-shit-done-redux
Mirror of code, issues, and PRs from the upstream gsd-build/get-shit-done,
which appears compromised or abandoned (maintainer unreachable since
2026-04-01; $GSD token linked to rug-pull).

- Adds rebrand notice block at top of English README
- Removes $GSD token badge and @gsd_foundation X badge (keeps Discord)
- Renames npm packages: get-shit-done-cc -> get-shit-done-redux,
  @gsd-build/sdk -> @gsd-redux/sdk
- Updates all repo URLs across docs, workflows, package.json, bin/
- Updates ci@gsd-build -> ci@gsd-redux in workflow git identities
- Leaves CHANGELOG and .changeset/* alone (historical, time-stamped)
2026-05-22 08:27:07 -04:00
Tom Boucher
6a5fa59129 feat(3081): auto-trim review prompts for small-context model reviewers (#3708)
* feat(3081): auto-trim review prompts for small-context model reviewers

Adds review.max_prompt_tokens and review.max_prompt_tokens_per_reviewer
config keys. When configured, the /gsd-review workflow deterministically
trims the assembled prompt before sending to each reviewer (drop CONTEXT
→ RESEARCH → REQUIREMENTS; head-shrink PROJECT.md; tail-truncate PLANs
proportionally; reserve disclosure-note tokens upfront). Trim metadata
is recorded in REVIEWS.md frontmatter. Reviewer is skipped with a
warning if even the minimum review set exceeds the budget.

Closes #3081

* fix(3081): register prompt-budget in SDK query registry and update inventory manifest

review.md references `gsd-sdk query prompt-budget` at three call sites, but the
command had no handler in the SDK registry — failing the registry-integration
drift-guard test on all 6 CI matrix legs. Added a native TypeScript SDK handler
(sdk/src/query/prompt-budget.ts) that ports the applyBudget logic from the CJS
module, registered it in DOMAIN_STATIC_CATALOG, and regenerated
docs/INVENTORY-MANIFEST.json to include the new cli_modules/prompt-budget.cjs entry.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(3081): bump ws to 8.20.1 and allowlist prompt-budget sibling pair

Two additional CI failures after the registry fix:

1. ws moderate CVE (GHSA-58qx-3vcg-4xpx, uninitialized memory disclosure):
   The advisory covers ws >=8.0.0 <8.20.1. Both root and sdk/package.json
   pinned ^8.20.0 which resolved to 8.20.0. Bumped both to 8.20.1 to clear
   the npm audit drift-guard test (bug-3588-npm-audit-clean.test.cjs).

2. lint-shared-module-handsync detected the new prompt-budget.ts / prompt-budget.cjs
   sibling pair without an allowlist entry. Added a cooperatingSiblings entry
   to scripts/shared-module-handsync-allowlist.json with classification and
   justification matching the established pattern.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(3081): align prompt-budget skip semantics across CJS and SDK dispatch paths

Replace brittle `[ $EXIT -eq 2 ]` guards with `[ $EXIT -ne 0 ]` in all three
local-reviewer blocks (Ollama, LM Studio, llama.cpp) in workflows/review.md.
Any non-zero exit from prompt-budget now triggers a skip with a descriptive
warning — exit 2/11 prints "budget too small", any other non-zero prints
"unexpected exit code". This ensures the SDK bridge dispatch path (exit 11
via GSDError(Blocked)) triggers the same skip as the CJS path (exit 2).

The SDK handler (sdk/src/query/prompt-budget.ts) already writes both metadata
and prompt files before throwing, so no change needed there.

The Ollama block also gains the missing OLLAMA_SKIP guard so the reviewer
invocation is actually skipped (previously the block only suppressed the
OLLAMA_PROMPT_FILE update but still ran the curl invocation).

SDK integration path (hardFailed via GSDError(Blocked) → exit 11) is covered
by handler unit tests in tests/prompt-budget.test.cjs; no gsd-sdk-*.test.cjs
exercising the full bridge dispatch for this command exists yet — that gap
remains and is documented here.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix prompt-budget trim ordering and review guard follow-ups

* perf: optimize prompt-budget and dedup reviewer trim workflow

* fix(3708): drop source-grep theater tests to satisfy lint-no-source-grep

All four test files added in commit 2df566ed were pure source-grep theater:
they read .cjs / .ts / .md source files and asserted that specific string
literals were present or absent. None exercised runtime behaviour.

Deleted:
- tests/gsd-tools-memory-optimizer.test.cjs   — 7 includes() on gsd-tools.cjs
- tests/prompt-budget-hotpath-optimizer.test.cjs — includes() on prompt-budget.cjs + .ts
- tests/prompt-budget-io-optimizer.test.cjs   — includes() on prompt-budget.ts + gsd-tools.cjs
- tests/review-workflow-budget-dedup.test.cjs — includes() on review.md

Behavioural coverage for the prompt-budget feature already exists in
tests/prompt-budget.test.cjs and tests/prompt-budget-cli.test.cjs (also
added by this PR). No replacement tests needed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(3708): correct budget-pressure threshold and minSet accounting

Two bugs in applyBudget caused premature trimming and false hard-fails:

1. UNNEEDED_TRIM: budgetUnderPressure compared baseTokens against
   effectiveBudget - NOTE_RESERVE_TOKENS, triggering trim pressure 80
   tokens before the budget was actually exceeded. Fix: compare against
   effectiveBudget directly; NOTE_RESERVE_TOKENS are still reserved in
   contentBudget once real pressure is confirmed.

2. FALSE_HARDFAIL: minSet included NOTE_RESERVE_TOKENS unconditionally,
   treating the note as mandatory even when no trim would occur and no
   note would be injected. Fix: exclude NOTE_RESERVE_TOKENS from minSet;
   a prompt that fits untrimmed needs no note and must not hard-fail.

Both fixes applied in CJS and TypeScript implementations. Two regression
tests added (cycles 11 and 12) that reproduce each case behaviorally.

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-18 23:13:09 -04:00
Tom Boucher
08848df839 docs(3562): pin minimum Codex CLI version (0.130.0) and explain the seam
Rationale for the version pin (the timeline that produced the oscillation):

  2026-05-08  Codex CLI 0.130.0 ships, dropping extra-skills-roots
              discovery via openai/codex#21485 (scans only ~/.codex/skills,
              cwd .codex/skills, and registered plugin roots).
  2026-05-14  GSD PR #3512 lands, removing ~/.codex/skills/gsd-* under the
              assumption Codex would auto-discover from extra roots.
              That assumption was already obsolete in shipped Codex.
  2026-05-15  #3562 filed — Codex CLI 0.130.0 users have zero $gsd-*
              commands after install.

The previous fix (#3427) was for Codex Desktop's official-skills surface,
which is a different product; that surface still exists on Desktop and
remains harmless duplication when both root scans see the gsd-* dirs.

Documents the supported version inline at the Codex sections of both
USER-GUIDE.md and CONFIGURATION.md, plus a one-line note in README's
Troubleshooting block. No runtime version-detection added — out of scope
and brittle against future Codex changes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-15 13:46:12 -04:00
Tom Boucher
088fc204ee fix(3347): clear CI gates raised by initial commit
- docs/CONFIGURATION.md: document graphify.auto_update key (config-schema-docs-parity)
- hooks/gsd-check-update-worker.js: add gsd-graphify-update.sh to MANAGED_HOOKS (managed-hooks)
- agents/gsd-planner.md + gsd-phase-researcher.md: slim auto-update awareness block to
  a one-line @-reference; extract full instructions to a new reference file
- get-shit-done/references/planner-graphify-auto-update.md: new reference with the
  status-file schema, the four annotation cases (running/failed/ok-current/ok-stale),
  and interaction with the existing stale-mtime annotation
- docs/INVENTORY.md: References (60 → 61 shipped) + row for new reference; regenerate
  docs/INVENTORY-MANIFEST.json via gen-inventory-manifest.cjs

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-15 11:59:43 -04:00
Tom Boucher
5c202372ed docs: update v1.42.1 release documentation 2026-05-14 20:42:35 -04:00
Tom Boucher
d4d4178603 Merge pull request #3483 from radioflyer28/feat/agent-launch-reasoning-transport-3474
feat: transport resolved reasoning effort to agent launches
2026-05-14 20:01:32 -04:00
Tom Boucher
e0adba7e08 feat(statusline): add opt-in context_position config for narrow terminals (#2937)
Extract composeStatusline() helper from duplicated inline template logic in
runStatusline() and renderStatusline(). Both call sites now route through the
helper, which accepts a position param ('end' | 'front', default 'end').

- 'end' (default) preserves byte-identical output to v1.38.x and earlier
- 'front' renders ctx immediately after model name, before the first │
- Invalid values silently coerce to 'end' at runtime (belt-and-suspenders;
  config-set rejects invalid values upfront via enum validator)

Adds statusline.context_position to VALID_CONFIG_KEYS in both CJS and TS
schemas, enum validator in config.cjs, docs row in CONFIGURATION.md,
and a changeset. Closes #2937.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-14 15:45:41 -04:00
Tom Boucher
da21edfb59 feat(workflow): add git.create_tag config to disable milestone tagging
Adds boolean config key `git.create_tag` (default: true, fully backcompat)
so projects with their own release flow can disable GSD's automatic
`git tag -a v[X.Y]` on milestone completion. Also adds tag-collision
pre-check to prevent silent failure on re-run. Closes #3086

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-14 12:06:36 -04:00
radioflyer28
810778e137 fix(sdk): gate reasoning effort by runtime allowlist 2026-05-13 22:27:15 -04:00
Tom Boucher
75d5ca5875 feat(code-review): integrate fallow structural pre-pass for /gsd-code-review (#3424)
* feat(code-review): add optional fallow structural pre-pass

* fix(ci): sync lockfile for fallow optional binaries

* fix(test): make fallow integration tests cross-platform

* fix(review): require executable fallow binary paths

* docs(review): clarify structural findings usage and size guard

* fix(fallow): preserve line:0, prefer node_modules/.bin, sync SDK twin (H1, M2, N1 from #3424 review)

* fix(workflow): harden fallow pre-pass — exit check, timeout, atomic write, size-guard order (B1, H2-H4, M1, M3 from #3424 review)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(deps): pin fallow floor to ^2.70.0 matching lockfile (H7 from #3424 review)

* fix(config): enum-validate fallow.scope/profile + group code_quality.* contiguously (H5, N3 from #3424 review)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(fallow): label mcp gate reserved, version-pin install, expand context schema (B3, H8, M4, M8, L1 from #3424 review)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(fallow): replace source-grep with behavioral tests, expand fixtures, fail-loud tmpdir (B4, H6, L2, L3, M5, M6, N2 from #3424 review)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(workflow): escape closing structural_findings tag in JSON payload (CR #3424 inline finding)

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-13 21:19:47 -04:00
radioflyer28
c20f714d68 docs: document reasoning effort transport 2026-05-13 19:43:42 -04:00
Tom Boucher
245d5f66a1 feat: add review.default_reviewers config for /gsd-review defaults (#3464)
* feat(review): add review.default_reviewers selection policy

* docs(review): explain default reviewer config and precedence

* chore(changeset): add feature entry for review.default_reviewers

* chore(changeset): set pr field for #3464

* test(review): cover unavailable default-reviewer failure path

* fix(review): sync sdk and inventory parity for default reviewers

* fix(review): use canonical /gsd:review namespace in source
2026-05-13 14:10:15 -04:00
Tom Boucher
e79c472d7b Feat(ship): add configurable PR body sections (#3391)
* feat: add configurable ship PR body sections

* chore: add changeset for ship PR sections

* docs: avoid prompt scanner trigger in PR body guide

* docs: escape pr body source separator
2026-05-11 15:27:40 -04:00
Tom Boucher
25fb81d01e feat(3309): workflow.human_verify_mode = end-of-phase (new default; mid-flight opt-back-in) (#3325)
* test(3309): red — workflow.human_verify_mode contract

New behavioral test file covers:
- workflow.human_verify_mode is a recognized config key (VALID_CONFIG_KEYS)
- defaults to 'mid-flight' (preserves current behavior)
- config-set / config-get round-trips for both values
- persists in config.json as string
- planner agent file references the flag with canonical wording, couples
  end-of-phase mode with the rule that checkpoint:human-verify is not
  emitted, and documents the <verify><human-check> deferred-item shape
- verifier agent file references harvesting <verify><human-check> blocks
- references/checkpoints.md documents the cost-control alternative

Source-text assertions on agent .md files are exempted via
allow-test-rule: source-text-is-the-product — those files ARE the
runtime contract loaded by AI runtimes, so asserting their wording is
the only way to verify the agents will respect the flag.

Fails 10/11 against current source. Will pass after the fix.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(3309): add workflow.human_verify_mode = end-of-phase opt-out

Each mid-flight checkpoint:human-verify halt costs a full executor
cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on every
respawn) because subagent context is discarded across the pause. A plan
with N human-verify checkpoints pays the cold-start cost N+1 times. The
reporter (rentanything-nb) measured this at "tens of thousands of tokens"
per round-trip and "hundreds of thousands per week."

This adds workflow.human_verify_mode (default 'mid-flight') with an
'end-of-phase' value that:
- instructs gsd-planner to NOT emit <task type="checkpoint:human-verify">
  tasks; verification details go into a <verify><human-check> sub-block
  on the relevant auto task instead
- instructs gsd-verifier (Step 8) to harvest those <verify><human-check>
  blocks at end-of-phase and merge them into its own human-verification
  list
- the existing human_needed → HUMAN-UAT.md flow in execute-phase.md is
  the single sink — no new file/writer is created

checkpoint:decision and checkpoint:human-action are unaffected — those
gate the work itself, not post-hoc verification.

Surfaces touched:
- bin/lib/config-schema.cjs, bin/lib/config.cjs — register key + default
- sdk/src/config.ts, sdk/src/query/config-schema.ts — SDK parity
- agents/gsd-planner.md — slim Detection section + reference link
- agents/gsd-verifier.md — Step 8 harvest instruction
- get-shit-done/references/planner-human-verify-mode.md — full rules,
  loaded conditionally to keep planner.md under its size budget
- get-shit-done/references/checkpoints.md — surface the alternative
- docs/CONFIGURATION.md — config table row
- docs/INVENTORY.md, docs/INVENTORY-MANIFEST.json — track new reference

Tag name <human-check> chosen instead of <human> to avoid the
prompt-injection scan pattern that flags <system|assistant|human> tags.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(3309): align changeset pr: to actual PR number

The pr: field was authored as 3319 (a guess at the next number) before
the PR was opened. Actual PR is #3325.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(3309): flip workflow.human_verify_mode default to end-of-phase

Per maintainer direction on PR #3325, end-of-phase is the new project
default. Mid-flight checkpoint:human-verify halts cost a full executor
cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) per
round-trip — reported at "tens of thousands of tokens" per round-trip,
"hundreds of thousands per week" on real projects. The cost-control
mode is what new projects should get out of the box.

mid-flight remains a one-line opt-back-in via:

    gsd config-set workflow.human_verify_mode mid-flight

Behavior change for existing projects: the new default takes effect
when .planning/config.json is rewritten (config-set, fresh project).
Existing in-flight PLAN.md files with checkpoint:human-verify tasks
continue to work in either mode — the flag only changes what the
planner emits next time it runs.

Surfaces updated:
- bin/lib/config.cjs, sdk/src/config.ts — default flipped
- sdk/src/config.ts docstring — describes new default + opt-back-in
- agents/gsd-planner.md — Detection section explains new default
- references/planner-human-verify-mode.md — reordered modes; added
  guidance on when to opt back into mid-flight
- references/checkpoints.md — surface the default flip and the why
- docs/CONFIGURATION.md — table row reflects new default + reason
- tests/feat-3309-human-verify-mode.test.cjs — default test asserts
  end-of-phase
- .changeset/fierce-geese-march.md — describes the default flip and
  the migration semantics

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: address human verify mode review

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 23:29:11 -04:00
Tom Boucher
e7942c21b3 fix: add executor stall recovery contract (#3329)
* fix: add executor stall recovery contract

* chore: add changeset for executor recovery

* chore: keep execute phase within size budget
2026-05-09 23:28:52 -04:00