Commit Graph

22 Commits

Author SHA1 Message Date
Tom Boucher
07603df8f2 fix(#2647): code-fixer worktree under .claude/worktrees/, not a hardcoded /tmp path (#2942)
* test(#2647): failing-first — fixer worktree path must be repo-relative not /tmp

* fix(#2647): place code-fixer worktree under .claude/worktrees/, not /tmp

The gsd-code-fixer agent hand-rolled its worktree at a hardcoded
`/tmp/sv-${padded_phase}-reviewfix-XXXXXX` mktemp path. On Windows/Git Bash
that landed OUTSIDE the project tree — outside the agent session's permission
allowlist, so every Read inside the worktree prompted (~25/run) — and mktemp's
MAX_PATH-avoidance substitute produced an un-removable `C:/mvwtNN` path.

Place the worktree repo-relative under `.claude/worktrees/` (the same dir the
harness-managed executor worktrees use: gitignored via `.claude/`, inside the
session's permission scope), with a $$-PID + epoch suffix for concurrency
uniqueness (replacing mktemp's XXXXXX). $main_repo is resolved the same way
the cleanup tail already resolves it.

Three sites updated: setup_worktree bash, concrete-steps prose, critical_rules.
The #2990 `-b "$reviewfix_branch"` invariant is preserved (the folded test
asserts it). Failing-first regression added to the #2990 suite in
tests/agent-frontmatter.test.cjs.

* test(#2647): update #2686 path assertion to expect .claude/worktrees/, not /tmp

The #2686 regression test encoded the worktree location as a hardcoded
`/tmp/sv-` path (matching sibling GSD agents at the time). #2647 showed that
breaks Windows/Git Bash (worktree outside the project tree → permission prompts;
mktemp MAX_PATH substitute un-removable). Update the #2686 path assertion to
require the repo-relative `.claude/worktrees/` location and forbid `/tmp/sv-`.
The #2686 isolation + cleanup assertions are unchanged.

* fix(#2647): word-boundary wt= parse + ack the fixer growth vs next

Two follow-ups to the #2647 GREEN run:
- parseWtAssignments matched `prior_wt=` (no word boundary), polluting the
  set and tripping the repo-relative + concurrency-unique assertions. Anchor
  on (?:^|\s)wt= so only the real worktree-path assignment is captured.
- emitted-attribution: gsd-code-fixer.md grew 1875 bytes vs origin/next. Update
  the emitted-drift-ack entry to attribute the #2647 worktree-path change
  (supersedes the prior #2825 attribution, whose growth is already in next).

* fix(#2647): address review — validate padded_phase at the sink + tighten test

Code-review + security-review both APPROVED with one actionable minor:
padded_phase is interpolated into a worktree PATH and a git BRANCH NAME, but
was only validated by the orchestrator (code-review-fix.md), not at the agent
sink. The agent prompt is a literal bash contract any caller can spawn, so add
a `[[ =~ ^[0-9]+(\.[0-9]+)?$ ]]` self-defense check rejecting traversal/shell
metachars (defense-in-depth; not a present vuln — the only caller validates).

Also tighten the concurrency-uniqueness test to require BOTH $$ AND $(date +%s)
(either-alone was too lax per review). Update the emitted-drift-ack reason to
cover the added validation growth.

* changeset(#2647): code-fixer worktree under .claude/worktrees not /tmp

* changeset(#2647): backfill PR number 2942

* chore(#2938): regenerate stale docs/CONTEXT-INDEX.json on next

#2938 (#2928) updated the CONTEXT.md RULESET prose for the new per-PR
emitted-drift-ack fragment mechanism (#2914) but shipped a CONTEXT-INDEX.json
generated from the OLD prose. lint:generated-sync fails on every PR that
rebases onto next after #2938 (the regen produces a 3-line diff bringing three
RULESET entries — AGENT_SIZE_BUDGET, EMITTED_ATTRIBUTION, WORKFLOW_SIZE_BUDGET
— in sync with the prose already on next). Mechanical regen via
`node scripts/gen-context-index.cjs --write`; idempotent; surfaced by the
#2647 rebase. No behavioral change.

---------

Co-authored-by: sim <sim@users.noreply.github.com>
2026-07-31 14:45:51 -04:00
Tom Boucher
799230873f test(#1976): scope GSD_TEST_MODE in folded bug-2990 block (adversarial review)
Codex review: the folded bug-2990 block set process.env.GSD_TEST_MODE='1' at
collection time, which persisted into sibling folded suites in agent-frontmatter.test.cjs
(process-isolated when standalone). Scope it to before/after so it no longer leaks.
Assertions unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 09:51:29 -04:00
Tom Boucher
cb7ca4f3dc test(#1976): consolidate 17 agent regression tests into agent-contract suites
Fold 17 issue-named agent-markdown regression files into their canonical agent
suites: agent-frontmatter (shared agent-body-contract catch-all: read-loop guards,
retired-slash-ref bans, sycophancy hardening, doc-writer/code-fixer/review-fix/mapper
contracts, write-truncation), planner-language-regression (planner directive/grep/phase
contract), executor-mvp-tdd-section, verifier-behavior-unverified, intel, and
research-agent-profiles. Verbatim block-scoped describe wrappers; 98 subtests conserved
1:1. All unit (markdown-content) tests — no CLI/env surface. No new files.

Regenerates regression-name allowlist (222->207), ratchets file-count allowlist (intel
entry removed, phase 4->3), makes 13 relocated allow-test-rule exemptions issue-ref-
compliant (ADR-456; prunes stale ids). No CONTEXT.md/ADR refs named the moved files.
lint:ci green.

Part of epic #1969. Closes #1976.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 09:51:29 -04:00
Tom Boucher
f08b177215 feat(#1726): G1-G6 portability AST rules; fix all offenders; delete the ratchet (Phase 4) (#1731)
Phase 4 of epic #1702. Closes #1726.
2026-06-25 18:24:39 -04:00
Behruz Nassre Esfahani
fc5ca178a2 test(#1178): consolidate duplicated agent-roster helper into tests/helpers (#1420)
* test(#1178): consolidate duplicated agent-roster helper into tests/helpers

The "list gsd-*.md agent files, strip .md, sort" derivation was hand-duplicated
across the suite (two listAgentFiles(), an identical agentFilesOnDisk(), and
inline readdir blocks). Add tests/helpers/agent-roster.cjs exporting
listAgentFiles(agentsDir?) and route the genuinely-identical source-roster sites
through it. Semantically-different sites (installed-dest dirs, absolute-path
returns, .toml-inclusive Codex rosters, full-.md-filename readers, the uniform
multi-family inventory table) are left intact, each with a one-line comment.

Test-only; no production code touched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1178): note AGENTS_DIR export is for future call sites

Review nit: clarify that the currently-unused AGENTS_DIR export is intentional
— available for future tests needing the canonical source agents path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-23 19:34:08 -04:00
Tom Boucher
3a3b2135c2 chore(#1073): purge phantom pre-migration issue refs from source, tests, docs (#1471)
#2551/#3182/#2361 are pre-migration get-shit-done-redux issue numbers with no
equivalent in open-gsd/gsd-core; they mislead triage and manufacture phantom
blockers. Repoint to real successors (#717 byte-budget rework, #720) or rewrite
as prose referencing the discuss-phase/modes progressive-disclosure split.
Correct co-located 'line budget'/'<500 lines' framing to the byte-based reality
(#717). Add a CI guard (tests/no-phantom-issue-refs.test.cjs) that fails if a
phantom ref is reintroduced. SSH-key patterns (id_ed25519) left untouched.
No user-facing runtime behavior change.

Closes #1073
2026-06-20 13:36:57 -04:00
Tom Boucher
2981983bae fix(#973): add Edit to gsd-planner tools and forbid whole-file Write of ROADMAP.md (#989)
* fix(#973): add Edit to gsd-planner tools and forbid whole-file Write of ROADMAP.md

gsd-planner shipped Write but not Edit — the same writer-agent gap fixed for six
agents in #571/#581. Without Edit, an in-place ROADMAP update fell back to a
whole-file Write that truncated committed milestone history (292→16 lines in a
real incident).

Changes:
- agents/gsd-planner.md: add Edit to tools: frontmatter (adjacent to Write)
- agents/gsd-planner.md: update_roadmap step now directs Edit (scoped), with an
  explicit blocking prohibition on whole-file Write of ROADMAP.md or any existing
  curated .planning/ file
- agents/gsd-planner.md: Write contract section clarifies Write is authorized only
  for net-new PLAN.md creation; existing files must use Edit
- tests/agent-frontmatter.test.cjs: extend SECTION_WRITER_AGENTS list (#581 test)
  to cover gsd-planner — fails before fix, passes after
- .changeset/973-gsd-planner-edit-tool.md: Fixed changeset, pr:0

Closes #973

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#973): backfill changeset pr number (989)

* fix(#973): trim gsd-planner.md prose under agent size cap (keep Edit + scoped-Edit-for-ROADMAP rule)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 10:44:20 -04:00
Tom Boucher
3aed02822d chore(#771): convert agent color: hex/magenta values to documented named colors (#823)
* chore(#771): convert agent color: hex/magenta values to documented named colors

Claude Code's sub-agent `color:` field documents only 8 named colors
(red, blue, green, yellow, purple, orange, pink, cyan). Twelve agent
files used hex values and two used the undocumented `magenta`; convert
each to the nearest documented named color so the intended per-agent
TUI color differentiation is spec-compliant.

- agents/*.md: 14 color values hex/magenta -> nearest named color
- scripts/research-profiles.cjs: update the 3 generated research-agent
  profiles (source of truth) so gen-research-agents stays in sync
- docs/AGENTS.md: update documented colors; add missing Color rows for
  gsd-nyquist-auditor, gsd-project-researcher, gsd-phase-researcher
- tests/agent-frontmatter.test.cjs: add regression guard asserting every
  agent color: is in the documented named-color set

Closes #771

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#771): add changeset

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 18:50:50 -04:00
Tom Boucher
463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00
Tom Boucher
692343f8cc fix(#581): add Edit to six writer agents' tools so Edit-only discipline is enforceable (#582)
* fix(#581): add Edit to six writer agents' tools so Edit-only discipline is enforceable

Six writer agents (gsd-eval-planner, gsd-ai-researcher, gsd-domain-researcher,
gsd-phase-researcher, gsd-ui-researcher, gsd-debug-session-manager) shipped with
Write but no Edit in their tools: frontmatter. Their spawn prompts instruct
surgical in-place section edits on existing/shared files (notably the AI-SPEC.md
trio writing disjoint sections of the same file), but with no Edit tool they
fall back to whole-file Write — silently clobbering sibling sections
(last-writer-wins) while still reporting success.

Same bug class as #571, fixed for gsd-doc-writer in #575. This adds Edit
alongside the existing Write for all six (Edit placed adjacent to Write, mirroring
the gsd-doc-writer fix). Write is retained; no prompt-body changes; no other agents
touched.

Adds a regression test (tests/agent-frontmatter.test.cjs) asserting each of the
six section-writer agents carries both Write and Edit.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#581): add changeset fragment for writer-agent Edit fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#581): regenerate changeset via npm run changeset

Replace hand-authored fragment with one generated by the official
scripts/changeset/new.cjs script (correct <adjective>-<noun>-<noun>
filename convention).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-01 14:25:12 -04:00
Tom Boucher
7c6f8005f3 test: destroy 9 config-schema.cjs/core.cjs source-grep tests, replace with behavioral config-set (#2696)
* test: destroy 9 config-schema.cjs/core.cjs source-grep tests, add behavioral config-set tests (#2691, #2693)

Replace source-grep theater with config-set behavioral tests:
- execute-phase-wave: config-set workflow.use_worktrees replaces VALID_CONFIG_KEYS grep
- inline-plan-threshold: delete redundant source-grep (behavioral test at L36 already covered it)
- plan-bounce: config-set for plan_bounce / plan_bounce_script / plan_bounce_passes replaces 3 key-presence greps
- code-review: config-set for code_review / code_review_depth replaces 2 greps; removes CONFIG_PATH constant
- thinking-partner: config-set features.thinking_partner replaces two greps (config-schema.cjs AND core.cjs)

Behavioral tests survive refactors (no path constants, no file reads). The config-schema.cjs →
core.cjs migration commit 990c3e64 happened because these tests groped source paths.

Add allow-test-rule: source-text-is-the-product annotations to legitimate product-content tests:
autonomous-allowed-tools, agent-frontmatter, agent-skills-awareness, bug-2334, bug-2346,
execute-phase-wave (MD reads), plan-bounce (workflow reads). Annotations explain WHY text
inspection is the right level of testing for AI instruction files.

Closes #2691
Closes #2693

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test: address CodeRabbit findings on #2696

- agent-frontmatter.test.cjs: move allow-test-rule annotation from block comment
  to standalone // line comment so rule scanners can detect it
- thinking-partner.test.cjs: strengthen config-set test with config-get read-back
  assertion to verify the value was persisted, not just accepted (exit 0)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test: tighten thinking_partner config assertion per CodeRabbit (#2696)

Replace config-get output substring check (includes('true') false-positive
risk) with a direct JSON read of .planning/config.json, asserting the
exact persisted value via strictEqual. This also validates the config file
was created, catching silent key-acceptance without persistence.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-25 10:50:54 -04:00
Tom Boucher
41dc475c46 refactor(workflows): extract discuss-phase modes/templates/advisor for progressive disclosure (closes #2551) (#2607)
* refactor(workflows): extract discuss-phase modes/templates/advisor for progressive disclosure (closes #2551)

Splits 1,347-line workflows/discuss-phase.md into a 495-line dispatcher plus
per-mode files in workflows/discuss-phase/modes/ and templates in
workflows/discuss-phase/templates/. Mirrors the progressive-disclosure
pattern that #2361 enforced for agents.

- Per-mode files: power, all, auto, chain, text, batch, analyze, default, advisor
- Templates lazy-loaded at the step that produces the artifact (CONTEXT.md
  template at write_context, DISCUSSION-LOG.md template at git_commit,
  checkpoint.json schema when checkpointing)
- Advisor mode gated behind `[ -f $HOME/.claude/get-shit-done/USER-PROFILE.md ]`
  — inverse of #2174's --advisor flag (don't pay the cost when unused)
- scout_codebase phase-type→map selection table extracted to
  references/scout-codebase.md
- New tests/workflow-size-budget.test.cjs enforces tiered budgets across
  all workflows/*.md (XL=1700 / LARGE=1500 / DEFAULT=1000) plus the
  explicit <500 ceiling for discuss-phase.md per #2551
- Existing tests updated to read from the new file locations after the
  split (functional equivalence preserved — content moved, not removed)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(#2607): align modes/auto.md check_existing with parent (Update it, not Skip)

CodeRabbit flagged drift between the parent step (which auto-selects "Update
it") and modes/auto.md (which documented "Skip"). The pre-refactor file had
both — line 182 said "Skip" in the overview, line 250 said "Update it" in the
actual step. The step is authoritative. Fix the new mode file to match.

Refs: PR #2607 review comment 3127783430

* test(#2607): harden discuss-phase regression tests after #2551 split

CodeRabbit identified four test smells where the split weakened coverage:

- workflow-size-budget: assertion was unreachable (entered if-block on match,
  then asserted occurrences === 0 — always failed). Now unconditional.
- bug-2549-2550-2552: bounded-read assertion checked concatenated source, so
  src.includes('3') was satisfied by unrelated content in scout-codebase.md
  (e.g., "3-5 most relevant files"). Now reads parent only with a stricter
  regex. Also asserts SCOUT_REF exists.
- chain-flag-plan-phase: filter(existsSync) silently skipped a missing
  modes/chain.md. Now fails loudly via explicit asserts.
- discuss-checkpoint: same silent-filter pattern across three sources. Now
  asserts each required path before reading.

Refs: PR #2607 review comments 3127783457, 3127783452, plus nitpicks for
chain-flag-plan-phase.test.cjs:21-24 and discuss-checkpoint.test.cjs:22-27

* docs(#2607): fix INVENTORY count, context.md placeholders, scout grep portability

- INVENTORY.md: subdirectory note said "50 top-level references" but the
  section header now says 51. Updated to 51.
- templates/context.md: footer hardcoded XX-name instead of declared
  placeholders [X]/[Name], which would leak sample text into generated
  CONTEXT.md files. Now uses the declared placeholders.
- references/scout-codebase.md: no-maps fallback used grep -rl with
  "\\|" alternation (GNU grep only — silent on BSD/macOS grep). Switched
  to grep -rlE with extended regex for portability.

Refs: PR #2607 review comments 3127783404, 3127783448, plus nitpick for
scout-codebase.md:32-40

* docs(#2607): label fenced examples + clarify overlay/advisor precedence

- analyze.md / text.md / default.md: add language tags (markdown/text) to
  fenced example blocks to silence markdownlint MD040 warnings flagged by
  CodeRabbit (one fence in analyze.md, two in text.md, five in default.md).
- discuss-phase.md: document overlay stacking rules in discuss_areas — fixed
  outer→inner order --analyze → --batch → --text, with a pointer to each
  overlay file for mode-specific precedence.
- advisor.md: add tie-breaker rules for NON_TECHNICAL_OWNER signals — explicit
  technical_background overrides inferred signals; otherwise OR-aggregate;
  contradictory explanation_depth values resolve by most-recent-wins.

Refs: PR #2607 review comments 3127783415, 3127783437, plus nitpicks for
default.md:24, discuss-phase.md:345-365, and advisor.md:51-56

* fix(#2607): extract codebase_drift_gate body to keep execute-phase under XL budget

PR #2605 added 80 lines to execute-phase.md (1622 -> 1702), pushing it over
the XL_BUDGET=1700 line cap enforced by tests/workflow-size-budget.test.cjs
(introduced by this PR). Per the test's own remediation hint and #2551's
progressive-disclosure pattern, extract the codebase_drift_gate step body to
get-shit-done/workflows/execute-phase/steps/codebase-drift-gate.md and leave
a brief pointer in the workflow. execute-phase.md is now 1633 lines.

Budget is NOT relaxed; the offending workflow is tightened.

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 21:57:24 -04:00
Tom Boucher
2703422be8 refactor(tests): standardize to node:assert/strict and t.after() per CONTRIBUTING.md (#1675)
* refactor(tests): standardize to node:assert/strict and t.after() per CONTRIBUTING.md

- Replace require('node:assert') with require('node:assert/strict') across
  all 73 test files to enforce strict equality (no type coercion)
- Replace try/finally cleanup blocks with t.after() hooks in core.test.cjs
  and hooks-opt-in.test.cjs per the test lifecycle standards
- Utility functions in codex-config and security-scan retain try/finally
  as that is appropriate for per-function resource guards, not lifecycle hooks

Closes #1674

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(tests): add --test-concurrency=4 to test runner for parallel file execution

Node.js --test-concurrency controls how many test files run as parallel child
processes. Set to 4 by default, configurable via TEST_CONCURRENCY env var.
Fixes tests at a known level rather than inheriting os.availableParallelism()
which varies across CI environments.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(security): allowlist verify.test.cjs in prompt-injection scanner

tests/verify.test.cjs uses <human>...</human> as GSD phase task-type
XML (meaning "a human should verify this step"), which matches the
scanner's fake-message-boundary pattern for LLM APIs. This is a
false positive — add it to the allowlist alongside the other test files
that legitimately contain injection-adjacent patterns.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-04 14:29:03 -04:00
Tibsfox
9ddf004368 fix(agents): remove permissionMode that breaks Gemini CLI agent loading (#1522)
permissionMode: acceptEdits in gsd-executor and gsd-debugger frontmatter
is Claude Code-specific and causes Gemini CLI to hard-fail on agent load
with "Unrecognized key(s) in object: 'permissionMode'". The field also
has no effect in Claude Code (subagent Write permissions are controlled
at runtime level regardless). Remove it from both agents and update
tests to enforce cross-runtime compatibility.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 14:37:41 -07:00
Tom Boucher
0ce31ae882 fix: add <available_agent_types> to all workflows spawning named agents
PR #1139 added <available_agent_types> sections to execute-phase.md and
plan-phase.md to prevent /clear from causing silent fallback to
general-purpose. However, 14 other workflows and 2 commands that also
spawn named GSD agents were missed, leaving them vulnerable to the same
regression after /clear.

Added <available_agent_types> listing to: research-phase, quick,
audit-milestone, diagnose-issues, discuss-phase-assumptions,
execute-plan, map-codebase, new-milestone, new-project, ui-phase,
ui-review, validate-phase, verify-work (workflows) and debug,
research-phase (commands).

Added regression test that enforces every workflow/command spawning
named subagent_type must have a matching <available_agent_types>
section listing all spawned types.

Fixes #1357

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 13:17:18 -04:00
Tom Boucher
a6939f135f fix: add permissionMode: acceptEdits to worktree agents (#1334)
Worktree agents (gsd-executor, gsd-debugger) prompt for edit permissions
on every new directory they touch, even when the user has "accept edits"
enabled. This is caused by Claude Code's directory-scoped permission
model not propagating to worktree paths.

Setting permissionMode: acceptEdits in the agent frontmatter tells Claude
Code to auto-approve file edits for these agents, bypassing the per-
directory prompts. This is safe because these agents are already granted
Write/Edit in their tools list and are spawned in isolated worktrees.

- Add permissionMode: acceptEdits to gsd-executor.md frontmatter
- Add permissionMode: acceptEdits to gsd-debugger.md frontmatter
- Add regression tests verifying worktree agents have the field
- Add test ensuring all isolation="worktree" spawns are covered

Upstream: anthropics/claude-code#29110, anthropics/claude-code#28041

Fixes #1334

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 14:05:33 -04:00
Tom Boucher
e98b41aa15 feat: add data-flow tracing, environment audit, and behavioral spot-checks
Verification checked structure but not whether data actually flows
end-to-end or whether external dependencies are available. Adds:

- Step 4b (Data-Flow Trace): Level 4 verification traces upstream from
  wired artifacts to verify data sources produce real data, catching
  hollow components that render empty/hardcoded values
- Step 7b (Behavioral Spot-Checks): lightweight smoke tests that verify
  runnable code produces expected output, not just that it exists
- Step 2.6 (Environment Audit): researcher probes target machine for
  external tools/services/runtimes before planning, so plans include
  fallback strategies for missing dependencies

Closes #1245

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-20 22:26:22 -04:00
Tom Boucher
d032322bcb feat: add CLAUDE.md compliance as plan-checker Dimension 10
Add CLAUDE.md enforcement across the three core agents:
- gsd-plan-checker: new Dimension 10 verifies plans respect project
  conventions, forbidden patterns, and required tools from CLAUDE.md
- gsd-phase-researcher: outputs Project Constraints section from
  CLAUDE.md so planner can verify compliance
- gsd-executor: treats CLAUDE.md directives as hard constraints,
  with precedence over plan instructions

Includes 4 regression tests validating the new dimension and
enforcement directives across all three agents.

Closes #1260

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-20 15:41:28 -04:00
Tom Boucher
61d82dc3c0 feat(discuss-phase): auto-generate DISCUSSION-LOG.md for audit trail (#1209)
- Generate {phase_num}-DISCUSSION-LOG.md alongside CONTEXT.md during
  discuss-phase sessions
- Log captures all options presented per gray area (not just the
  selected one), user's choice, notes, Claude's discretion items,
  and deferred ideas
- File is explicitly marked as audit-only — not for agent consumption
- Add discussion-log.md template with format specification
- Track Q&A data accumulation instruction in discuss_areas step
- Commit discussion log alongside CONTEXT.md in same git commit
- Add regression tests for workflow reference and template existence
2026-03-19 12:25:59 -04:00
Tom Boucher
8a6cdf5f25 fix(execute-phase): add Copilot sequential fallback and spot-check completion detection (#1128)
- Default Copilot runtime to sequential inline execution instead of
  unreliable subagent spawning
- Add spot-check fallback in step 3 (SUMMARY.md exists + git commits
  found) so orchestrator never blocks indefinitely on missing completion
  signals
- Add runtime detection guidance in init step to force sequential mode
  under Copilot
- Add regression test verifying execute-phase contains Copilot fallback
  and spot-check documentation
2026-03-19 09:15:06 -04:00
TÂCHES
1d3d1f3f5e fix: strip skills: from agent frontmatter for Gemini compatibility (#1045)
* fix: remove dangling skills: from agent frontmatter and strip in Gemini converter (closes #1023, closes #953, closes #930)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: invert skills frontmatter test to assert absence (fixes CI)

The PR deliberately removed skills: from agent frontmatter (breaks
Gemini CLI), but the test still asserted its presence. Inverted the
assertion to ensure skills: stays removed.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 21:25:02 -06:00
Tibsfox
3d926496d9 test(agents): add 47 agent frontmatter and spawn consistency tests
New test suite covering:
- HDOC: anti-heredoc instruction present in all 9 file-writing agents
- SKILL: skills: frontmatter present in all 11 agents
- HOOK: commented hooks pattern in file-writing agents
- SPAWN: no stale workaround patterns, valid agent type references
- AGENT: required frontmatter fields (name, description, tools, color)

509 total tests (462 existing + 47 new), 0 failures.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-03 08:01:54 -06:00