a33cbe72f569e75e72d94de85d0930296d1ca1df
1015 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6dce1de4a7 |
fix: gap-analysis parses mixed requirement prefixes and skips table headers (#2902)
* fix: parse non-REQ IDs in gap-analysis and ignore table headers * fix: parse requirement IDs from first traceability column only --------- Co-authored-by: Tom Boucher <thomas.boucher@sas.com> |
||
|
|
abb2cb63f6 |
refactor: extract planning-workspace seam from core.cjs (#2901)
* refactor: extract planning workspace seam from core * docs: document planning-workspace module and inventory updates * fix: harden planning lock timeout and preserve workstream set contract --------- Co-authored-by: Tom Boucher <thomas.boucher@sas.com> |
||
|
|
951d5bf7c0 |
fix(#2893): surface non-canonical plan filenames instead of silently returning zero plans (#2896)
* fix(#2893): surface non-canonical plan filenames instead of silently returning zero plans Reporter saw `plan_count: 0` from `/gsd:execute-phase` even though five plan files existed on disk. Investigation showed the planner had written files like `01-PLAN-01-foundation.md`, while `phase-plan-index`'s strict filter (`f.endsWith('-PLAN.md') || f === 'PLAN.md'`) rejected them silently — collapsing two distinct states into the same `plans: []` return: - directory truly has no plans (legit empty) - directory has plans but the filter rejected them (user/agent error) The canonical contract is documented in three places: - `agents/gsd-planner.md` write_phase_prompt step (lines 1063-1080) - `commands/gsd/plan-phase.md` - `references/universal-anti-patterns.md` (rule 26) It mandates `{padded_phase}-{NN}-PLAN.md` and explicitly forbids `PLAN-NN.md` / `01-PLAN-01.md` / `plan-NN.md` etc. The strict filter is correct per that contract. The bug is that the executor never tells the user when the contract was violated — they just see `plan_count: 0` with no signal. Fix: add a diagnostic helper `describeNonCanonicalPlans()` that scans the phase directory for files matching `*PLAN*.md` (the diagnostic net) that the canonical filter rejected, excluding legit derivatives like `*-PLAN-OUTLINE.md` and `*-PLAN.pre-bounce.md`. When offenders exist, return a `warning` field naming each one and citing the canonical pattern so the user knows what to rename to. Wired into the three filter sites: - `phase-plan-index` (the executor's main entry point) - `phases list --type plans` - `find-phase` The strict filter itself is unchanged — existing canonical plans behave identically. This is purely a diagnostic that converts silent-empty into loud-with-actionable-error. Tests: - `phase-plan-index returns warning for reporter's exact filename pattern (`01-PLAN-01-foundation.md`)` - `truly empty dir does not emit a warning` - `canonical plans + outline + pre-bounce files do not emit a warning` Closes #2893 * test(#2893): add parity tests for find-phase and phases list --type plans warnings CodeRabbit's only finding on the prior commit: I wired the warning into three filter sites (`phase-plan-index`, `find-phase`, `phases list --type plans`) but only `phase-plan-index` had test coverage for the warning shape. The other two paths could silently diverge during future refactors — exactly the silent-drift class of bug this fix exists to prevent. Add four parity tests mirroring the existing two: - find-phase: non-canonical filenames produce a warning naming each offender + citing the canonical pattern. - find-phase: canonical plan + derivative files (PLAN-OUTLINE, pre-bounce) produce no warning. - phases list --type plans: same non-canonical case, but assert the warning is prefixed with `${dir}: ` (this path aggregates across phase directories so each offender is tagged with its dir). - phases list --type plans: canonical case, no warning. `node --test tests/phase.test.cjs`: 98/98 pass (was 94, +4 new). |
||
|
|
5fdc950eb7 |
feat(#2792): namespace meta-skills + keyword-tag descriptions + context utilization guard (#2825)
* feat(#2792): namespace meta-skills retargeted at the post-#2790 surface This branch is now based on #2790's HEAD (the consolidation PR) instead of main, and every routing table targets the consolidated surface so a user routed by a namespace meta-skill never lands at a deleted / folded sub-skill. Cross-PR inconsistencies the original PR #2825 carried (vs #2790): - ns-ideate routed to gsd-note / gsd-add-todo / gsd-add-backlog / gsd-plant-seed → all folded into gsd-capture by #2790. Now routes to gsd-capture (the parent picks the mode from the user's intent). - ns-context routed to gsd-scan and gsd-intel → folded into gsd-map-codebase --fast / --query by #2790. Now routes to those flag forms. - ns-manage routed all workspace intent to gsd-list-workspaces (a list-only entry) → CR also flagged the over-narrow target. #2790 folds into gsd-workspace; routing now points there. - ns-workflow routed to gsd-research-phase → deleted outright by #2790. Removed. - ns-project routed to gsd-plan-milestone-gaps → deleted outright by #2790. Removed. - None of the namespaces previously surfaced #2790's new consolidated skills (gsd-capture, gsd-phase, gsd-config, gsd-workspace, gsd-progress). All five are now reachable through the routers. - extract_learnings → extract-learnings (canonicalized by #2858). Defect fixes within the namespace skills: - Hyphen-form `name:` (gsd-workflow, …) per the canonical naming contract — the colon-form addressed CR's drift complaint. - `Skill` added to allowed-tools on every router. The body instructs "Invoke the matched skill directly using the Skill tool" — without Skill in the permission list the meta-skill cannot route at all. New regression guard in tests/enh-2792-namespace-skills.test.cjs: every gsd-* token in any namespace router's table column resolves to a surviving commands/gsd/*.md file (or to a known consolidated parent for flag-form targets like gsd-map-codebase --fast). This single test would have caught every dead-end route the original PR shipped with. Skill-count cap in tests/enh-2790-skill-consolidation.test.cjs now filters out ns-*.md from its <= 63 cap. Namespace routers are descriptor-only entries, not part of the consolidation surface that cap is policing — they have their own contract in tests/enh-2792-namespace-skills.test.cjs. INVENTORY.md gains a "Namespace Meta-Skills" section with the 6 router rows; INVENTORY-MANIFEST.json gains 6 entries; the headline count moves 59 → 65 to match. Out of scope for this rebase: the gsd-health --context flag (PR #2825 advertised the contract but didn't implement it). That's a separate feature concern and is left untouched here. 5908/5908 on `npm test`. * feat(#2792): implement gsd-health --context utilization guard The original PR #2825 advertised a `--context` flag on gsd-health with a 60%/70% utilization threshold table but never implemented the workflow logic — CR caught it as a contract leak, the rebase deferred it. This commit closes the gap with TDD red/green/refactor. Math layer (pure): - get-shit-done/bin/lib/context-utilization.cjs classifyContextUtilization(tokensUsed, contextWindow) → { percent, state } State boundaries use the exact ratio: < 60% healthy / 60–70% warning / ≥ 70% critical (fracture point) Display percent rounded for humans. Throws TypeError on non-integer or out-of-range inputs. - STATES = Object.freeze({ HEALTHY, WARNING, CRITICAL }) exported so callers reference the names by symbol, not by literal string. SDK CLI integration: - get-shit-done/bin/gsd-tools.cjs `validate context --tokens-used N --context-window M [--json]` routes to the classifier, owns the recommendation copy (the classifier intentionally does not — keeps the renderer free to evolve without touching the math layer or its tests), and uses core.output's rawValue path for the sync-flush guarantee. - sdk/src/query/validate.ts + sdk/src/query/index.ts TypeScript validateContext handler registered at 'validate.context' and 'validate context'. Mirrors the CJS classifier inline (15 lines of arithmetic; not worth a shared cross-language module). User-facing wiring: - commands/gsd/health.md frontmatter advertises --context, body documents the three-state threshold table. - get-shit-done/workflows/health.md adds a `context_check` step that's reached only when --context is set. Step calls `gsd-sdk query validate.context` with self-reported tokensUsed and contextWindow, prints the SDK output verbatim, and ends. Includes a TEXT_MODE plain-text fallback for non-Claude runtimes per #2012. Tests: - tests/context-utilization.test.cjs (17 tests) — pure-function contract: state thresholds at every boundary, percent rounding, input validation, return-shape (no recommendation field — that's the renderer's job). - tests/validate-context.test.cjs (9 tests) — SDK CLI plumbing: arg parsing errors, JSON vs human rendering, recommendation copy pinned per state. - tests/enh-2792-namespace-skills.test.cjs (4 new tests) — markdown contract: --context advertised in argument-hint, threshold table in command body, context_check step exists in workflow, step invokes gsd-sdk query validate.context with both flags. Inventory bookkeeping: - docs/INVENTORY.md "CLI Modules" 31 → 32; new row for context-utilization.cjs. - docs/INVENTORY-MANIFEST.json mirror. 5939/5939 on `npm test`. |
||
|
|
87917131f2 |
refactor(#2790): consolidate 86 gsd-* skills to 59 — fold flags, delete dead skills (#2824)
* feat(#2790): consolidate 86 gsd-* skills to 59 — zero functional loss Closes #2790 - `capture.md` — absorbs add-todo (default), note (--note), add-backlog (--backlog), plant-seed (--seed), check-todos (--list) - `phase.md` — absorbs add-phase (default), insert-phase (--insert), remove-phase (--remove), edit-phase (--edit) - `config.md` — absorbs settings-advanced (--advanced), settings-integrations (--integrations), set-profile (--profile); settings.md retained as-is - `workspace.md` — absorbs new-workspace (--new), list-workspaces (--list), remove-workspace (--remove) - `update.md` — adds --sync (absorbs sync-skills) and --reapply (absorbs reapply-patches) - `sketch.md` — adds --wrap-up (absorbs sketch-wrap-up) - `spike.md` — adds --wrap-up (absorbs spike-wrap-up) - `map-codebase.md` — adds --fast (absorbs scan) and --query (absorbs intel) - `code-review.md` — adds --fix (absorbs code-review-fix) - `progress.md` — adds --next (absorbs next) and --do (absorbs do) join-discord, research-phase, session-report, from-gsd2, analyze-dependencies, list-phase-assumptions, plan-milestone-gaps autonomous.md: updated Skill(skill="gsd:code-review-fix") → Skill(skill="gsd:code-review", args="--fix --auto") to match the consolidated skill name - New: tests/enh-2790-skill-consolidation.test.cjs (48 tests) - Updated: 14 existing test files redirected from deleted command paths to their consolidated equivalents - docs/INVENTORY.md: Commands count 86→59, ghost rows removed, new consolidated rows added - docs/INVENTORY-MANIFEST.json: regenerated to match filesystem Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * docs(#2790): add CHANGELOG entry for skill consolidation * docs(#2790): update COMMANDS.md for 86→59 skill consolidation Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#2790): address CodeRabbit review findings - CHANGELOG.md: add --next alongside --do in progress flag list - config.md: remove trailing space from --profile code span (MD038) - COMMANDS.md: add required descriptions to /gsd-phase examples; /gsd-phase without args errors, not interactive - COMMANDS.md: add --next and --do to /gsd-progress flags table + examples - test: convert content.includes('--reapply') to structural frontmatter parse; add allow-test-rule comment for workflow content assertions - test: replace redundant existsSync duplicate with assertion that verifies the full consolidated flag surface (--sync | --reapply) in argument-hint Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#2790): restore reapply-patches workflow and strengthen test assertions - Create get-shit-done/workflows/reapply-patches.md: the #2790 consolidation deleted the 14K combined command+workflow file (reapply-patches.md) but update.md already referenced the workflow via execution_context_extended. Restoring it fixes a silent behavioral gap where --reapply had no workflow to load. Includes full three-way merge logic, hunk verification table (Step 4), and the Hunk Verification Gate (Step 5) that blocks cleanup until all user-added hunks are confirmed present in the merged output. - Fix update.md: /gsd-reapply-patches → /gsd-update --reapply (stale ref) - Fix reapply-verify-hunks.test.cjs: was checking existsSync(update.md) 8×; now points to the workflow file and asserts real behavioral content (Post-merge verification, Hunk presence check, Line-count check, backup reference, per-file tracking, structural ordering) - Fix reapply-patches.test.cjs: replace content.includes() stubs with frontmatter-parsed argument-hint assertions; replace 4 existsSync(update.md) no-ops with real assertions against the workflow content - Fix edit-phase.test.cjs: /gsd-edit-phase → /gsd-phase (COMMANDS.md now documents the consolidated command with --edit flag) - Fix next-safety-gates.test.cjs: split OR predicates into independent assertions — --next in progress.md and --force in next.md workflow - Fix workspace.test.cjs: add allow-test-rule comment for routing content checks (command routing text IS the deployed behavioral contract) - Fix bug-2439 test: strengthen pre-flight assertion to verify gsd-sdk is referenced (not just --profile) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: address CodeRabbit review findings (CR round 2) - INVENTORY.md: update sync-skills.md row to reference /gsd-update --sync instead of stale /gsd-sync-skills (absorbed in #2790) - enh-2380-sync-skills.test.cjs: align INVENTORY.md assertion with the corrected reference; was asserting the old /gsd-sync-skills name while the manifest test correctly asserted /gsd-update, creating conflicting expectations in the same suite - reapply-verify-hunks.test.cjs: add explicit notEqual(-1) assertions for all three anchors before the ordering check so a missing anchor produces a clear failure instead of a false positive (writeIdx=-1 < verifyIdx=5 is true) - bug-2439-set-profile-gsd-sdk-preflight.test.cjs: defer fs.readFileSync until after the existence assertion; eager describe-level read caused the suite to crash before the existence test could run, making it effectively dead code Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#2790): address CR — INVENTORY routing + reapply test contract wording Two unresolved CodeRabbit findings (Major): - docs/INVENTORY.md: workflow-file table still pointed at obsolete /gsd-do, /gsd-next, /gsd-note, /gsd-add-todo, /gsd-add-backlog, /gsd-check-todos, /gsd-plant-seed slash commands. Re-route to the consolidated /gsd-progress (--next, --do) and /gsd-capture (--note, --backlog, --seed, --list) so the inventory is internally consistent. - tests/reapply-verify-hunks.test.cjs: 'verification tracks per-file status' asserted on phrasing that doesn't appear in reapply-patches.md (the 'per-file' substring only matched accidentally via 'sequential integer per file'). Switch to the actual contract text — Hunk Verification Table, one row per hunk per file, verified column. * test(#2790): update CR-INTEGRATION tests for consolidated --fix invocation After the merge of main (which carries #2843's hyphen-form fix), the consolidation in this branch absorbs gsd-code-review-fix into gsd-code-review as the --fix flag. Update the two CR-INTEGRATION tests that previously asserted on the standalone gsd-code-review-fix skill name to instead assert on a gsd-code-review invocation carrying --fix in its arg tokens. Tests still parse Skill() invocations structurally; only the asserted skill-name + arg-token shape changed. * test(#2790): scope success_criteria check to the <success_criteria> block CodeRabbit nitpick: 'success criteria includes verification' did a whole-file substring check, which can false-pass if the phrase appears elsewhere in the document. Extract the <success_criteria>...</success_criteria> block first via extractTagBlock() and assert against that scope only. * fix(#2790): post-rebase reconciliation with main - INVENTORY.md/JSON: add reapply-patches workflow row + bump count to 85 - autonomous.md: switch consolidated --fix invocation to hyphen Skill name - analyze-dependencies test: assert COMMANDS.md does NOT document the consolidated-away /gsd-analyze-dependencies entry (was: bare .includes()) * fix(#2790): address remaining CR findings — strengthen contract tests Doc-fixes: - INVENTORY.md: route transition.md & edit-phase.md rows to consolidated /gsd-progress --next and /gsd-phase --edit (was: deleted /gsd-next, /gsd-edit-phase) - config.md --profile branch: document #2439 pre-flight `command -v gsd-sdk` guard + install hint BEFORE the gsd-sdk invocation (closes opaque "command not found: gsd-sdk" regression path) Test discipline (no-source-grep contract): - bug-2439: replace bare `content.includes('gsd-sdk')` with structured parse of <context> block + --profile branch; assert pre-flight token, install hint, #2439 citation, and ordering vs gsd-sdk invocation - edit-phase: parse INVENTORY.md edit-phase.md row's "Invoked by" column and assert `/gsd-phase --edit` (not the deleted /gsd-edit-phase) - next-safety-gates: tighten `--next` documentation contract — require --next AND --force AND completeness routing (was OR-based, passed when only --next present) - reapply-patches: parse argument-hint flag list structurally; scan ALL <execution_context*> blocks for the @-include of reapply-patches.md; parse Hunk Verification Table header columns directly; locate Step 5 via heading parsing then assert (i) table reference, (ii) verified=no gate, (iii) STOP/halt directive, (iv) explicit absent-table halt path - workspace: parse frontmatter, tokenize argument-hint across multiple bracketed segments, parse @-include targets from <execution_context> rather than substring-matching the file body --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
4d394a249d |
fix(commands): normalize gsd slash namespace drift (#2858)
* fix(commands): normalize gsd slash namespace drift * fix(#2855): address CodeRabbit findings on namespace drift PR Three CR findings, all valid: 1. autonomous.md line 783 still had `gsd:discuss-phase` (the PR's own normalization missed this line). Switched to `gsd-discuss-phase` and updated the matching test in autonomous-interactive.test.cjs that was asserting the now-retired colon form. 2. tests/bug-2543-gsd-slash-namespace.test.cjs source-grepped the fix-slash-commands.cjs script with .includes() rather than driving its transform behaviour. Refactored fix-slash-commands.cjs to export a pure transformContent(src, cmdNames) function, kept the CLI behaviour unchanged via require.main, and replaced the source-grep block with five behavioural cases: rewrite, multi-occurrence, idempotence on canonical input, no-op on gsd-sdk/gsd-tools, and word-boundary safety. 3. tests/bug-2808-skill-hyphen-name.test.cjs matched `name:` anywhere in SKILL.md; a stray name: in the body could satisfy the assertion. Scoped the lookup to the YAML frontmatter block via the suggested diff (parse the leading --- ... --- region first, then find name: inside it). Full suite: 5854/5854 passing. * fix(#2855): address remaining CodeRabbit findings on PR #2858 Three structural concerns flagged on the namespace-drift fix PR: 1. scripts/fix-slash-commands.cjs:24 — `buildPattern([])` compiled `/gsd:()(?=[^a-zA-Z0-9_-]|$)/g`. The empty capture group still matches any `/gsd:` token followed by a non-word boundary (whitespace, EOL, punctuation), rewriting it to a stray `/gsd-`. Verified live: `transformContent("/gsd:", [])` → `"/gsd-"`. Added a guard returning null from `buildPattern` on empty input and updated `transformContent` and `processDir` to no-op when the pattern is null. 2. tests/autonomous-interactive.test.cjs:44-47 — assertion was `content.includes('gsd-discuss-phase') && content.includes('INTERACTIVE')`, which would false-pass on any unrelated co-occurrence (e.g. `INTERACTIVE=""` initialization plus a stray `gsd-discuss-phase` prose mention). Replaced with a structural extraction: locate the `**If \`INTERACTIVE\` is set:**` branch, bound it by the next `**If` / `<step>` boundary, and assert the `Skill(skill="gsd-discuss-phase", ...)` invocation lives inside that region. Tolerates whitespace around `(`, `skill`, and `=`. 3. tests/bug-2808-skill-hyphen-name.test.cjs:104 — colon-call regex was `Skill\(skill=...` and missed valid formatting like `Skill(skill = "gsd:cmd")` or `Skill( skill = ...)`. Loosened to `Skill\(\s*skill\s*=\s*...` so reformatting drift can't slip past the namespace guard. Verification: 5854/5854 pass on `npm test` from the rebased branch. * fix(#2855): drop pre-validation filter that hid namespace drift CR finding on tests/bug-2808-skill-hyphen-name.test.cjs:128: the test collected generated skill directories with `.filter(entry => entry.isDirectory() && entry.name.startsWith('gsd-'))`, then validated namespace invariants over that filtered list. Anything that violated the prefix invariant — `gsd:extract-learnings` (colon form), `extract_learnings` without prefix, `Gsd-foo` mis-cased — would silently disappear from the iteration and the test would falsely pass. Drop the `startsWith('gsd-')` filter so every generated directory shows up. Add explicit assertions before the existing per-skill loop: - directory list is non-empty (catches a broken converter that produces nothing) - every directory begins with `gsd-` - every directory contains no `:` - every directory contains no `_` Re-audited the full PR diff for the same anti-pattern: only this one site filtered before validating the namespace; bug-2643 and commands-doc-parity also use `readdirSync().filter()` but only by file extension, which is correct. 5854/5854 on `npm test`. * fix(#2855): address remaining CR findings (1 active + 2 nitpicks) Three findings on PR #2858, all the same root cause: input narrowing before validation lets drift slip past the guards. 1. tests/bug-2808-...:104 (active) — `colonCallRe` captured local names with `[a-z0-9-]+`, which excluded the underscore. A drift like `Skill(skill="gsd:extract_learnings")` (deprecated colon syntax with the old underscore filename) silently slid through. Broadened the capture to `[^'"\s)]+` so any malformed local name is surfaced; surrounding pattern (whitespace tolerance, escape support, flags) unchanged. 2. tests/bug-2643-...:43 (nitpick) — `extractSkillNamesHyphen` and `extractSkillNamesColon` had the same over-strict capture plus relied on a single regex over raw bytes, which the project test- rigor memory bans (`feedback_no_source_grep_tests.md`). Replaced with `extractSkillCalls(content)` — a small structural extractor that walks `Skill(` openers, locates each call's matching `)`, parses the body's `skill = "..."` keyword argument with permissive whitespace + quoting + escape handling, and returns `{ name, raw }` records. The two namespace-form helpers become thin filters over the structured output. Tightened the body class to `[^'"\\]+` so a trailing escape `\` before the closing quote (as in `Skill(skill=\"gsd-foo\", …)` written inside another string context) doesn't get included in the captured name. 3. tests/bug-2543-...:44 (nitpick) — `DOC_SEARCH_FILES` was a hand- curated 7-entry array. Every doc added in the future would silently weaken drift detection until someone remembered to extend the list. Replaced with `discoverDocSearchFiles(ROOT)`: globs every `.md` under `docs/` and adds `README.md` if present. New docs are picked up automatically. Re-audited the diff surface for similar narrowings; no other sites filter or constrain before validating namespace invariants. 5854/5854 on `npm test`. * fix(#2855): recurse docs/ tree so localized translations are scanned too CR finding: discoverDocSearchFiles() stopped at docs/*.md, leaving localized translation trees (docs/ja-JP/, docs/zh-CN/, docs/ko-KR/, docs/pt-BR/) and other nested doc collections (docs/skills/, docs/superpowers/) invisible to the namespace-drift invariant. Verified the gap: docs/ has 6 nested directories with ~30 .md files that the previous top-level-only scan was skipping. None contain /gsd: references today, but a future translation update or new doc subdir could leak drift. Switch to an iterative stack walk so every .md under docs/ is scanned regardless of depth. Stack form (rather than recursion) avoids the risk of running into the call-stack limit on deep doc trees. 5854/5854 on `npm test`. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
99af76b3ba |
fix(#2851): replace bare gsd-tools invocations with absolute path (#2869)
* fix(#2851): replace bare gsd-tools invocations with absolute path `gsd-tools` is not a published bin entry — package.json declares only get-shit-done-cc and gsd-sdk. The shipped invocation pattern is `node "$HOME/.claude/get-shit-done/bin/gsd-tools.cjs" <subcommand>`, used by every other workflow file. Two leaked bare invocations: - get-shit-done/workflows/plan-phase.md §13e (gap-analysis) — reported in #2851; gap-analysis silently skipped on every plan-phase run - get-shit-done/workflows/ingest-docs.md §finalize (commit) — caught by the new structural test; ingest-docs commit step was broken Both updated to canonical absolute-path form. Adds tests/bug-2851-workflow-bare-gsd-tools.test.cjs which parses every markdown file under get-shit-done/workflows/, extracts shell-fenced code blocks, tokenizes each line, and asserts no token in command position is the bare string `gsd-tools` (the trailing `.cjs` is a different token). The test also asserts plan-phase.md's gap-analysis call uses the canonical `node …/gsd-tools.cjs` form. Closes #2851 * fix(#2851): catch third bare gsd-tools call in ingest-docs.md init After the first commit, the structural test was strengthened to detect bare `gsd-tools` inside `$(...)` and backtick command-substitution forms. The improved test surfaced a third leak: ingest-docs.md:55: INIT=$(gsd-tools init ingest-docs) Fixed to canonical form INIT=$(node "$HOME/.claude/get-shit-done/bin/gsd-tools.cjs" init ingest-docs) plus the standard `@file:` handoff line that every other workflow uses when capturing INIT (required by tests/windows-robustness.test.cjs). Updated tests/bug-2801-ingest-docs-handler.test.cjs to match either the bare `gsd-tools init ingest-docs` or canonical `gsd-tools.cjs" init ingest-docs` form — the test's intent is to verify the dispatch handler is wired, not to lock the bare-bin form that #2851 removes. Closes #2851 * test(#2851): tighten ingest-docs and gap-analysis assertions to canonical form CodeRabbit caught two soft assertions in the regression tests: 1. tests/bug-2801: the init-ingest-docs assertion accepted both the legacy bare `gsd-tools` form and the canonical node-path form. Since #2851 is the fix that removes the bare form, the test should only accept the canonical absolute-path invocation. Switched to parsed-bash-block extraction with an anchored regex on the full `node "$HOME/.claude/get-shit-done/bin/gsd-tools.cjs"` path. 2. tests/bug-2851: the gap-analysis assertion used two loose .includes()/word-boundary checks. Replaced with a single assert.match() against the full canonical path so non-canonical forms fail. * test(#2851): env-assignment skip accepts lowercase identifiers too CodeRabbit caught: the cmdIdx-skip regex /^[A-Z_][A-Z0-9_]*=/ only matched uppercase variable names, so a line like `tmp=1 gsd-tools init` would tokenize to ['tmp=1','gsd-tools','init'], the regex would fail on 'tmp=1', cmdIdx would stay at 0, and the command-position check would compare 'tmp=1' against 'gsd-tools' — false negative. POSIX shell variable names are [A-Za-z_][A-Za-z0-9_]*. Widen the regex to match the actual lexical rule. Existing uppercase forms still work (FOO=bar gsd-tools); now lowercase forms (tmp=1 gsd-tools) and mixed case forms are also detected. |
||
|
|
4815b3c972 |
fix(#2838): SUMMARY rescue handles gitignored .planning (#2850)
* fix(#2838): SUMMARY rescue handles gitignored .planning explicitly The pre-fix rescue used `git ls-files --modified --others --exclude-standard` to detect uncommitted SUMMARY.md before worktree removal. When projects gitignore .planning/, --exclude-standard filters out the very files the rescue is meant to save, the rescue branch is skipped, and `git worktree remove --force` permanently deletes the SUMMARY. Replace both rescue blocks (quick.md, execute-phase.md) with a filesystem-level find + cp rescue that bypasses gitignore entirely and avoids the worktree↔main commit/merge cascade. cmp -s makes it idempotent. Adds tests/bug-2838-summary-rescue-gitignored-planning.test.cjs which extracts each rescue block, runs it against a real temp repo with a gitignored .planning/, and asserts the SUMMARY survives worktree removal. * test(#2838): assert rescue block exits 0 in idempotency test CodeRabbit (Minor): the idempotency test pre-creates the destination SUMMARY.md, so even a syntax/runtime error in the rescue block would silently false-pass. Add an explicit r.status === 0 assertion. |
||
|
|
74b81379cf |
fix(#2836): audit-open quick SUMMARY filename + UAT terminal-status drift (#2847)
* fix(#2836): audit-open quick SUMMARY filename + UAT terminal-status drift Fixes two convention drifts in bin/lib/audit.cjs that produced false-positive "open" items at every milestone close: 1. scanQuickTasks: looked for bare `SUMMARY.md`, but workflows/quick.md mandates `${quick_id}-SUMMARY.md`. Now matches either filename so quick tasks created via the documented workflow are recognized. 2. scanUatGaps: only treated `status: complete` as terminal, but workflows/execute-phase.md uses `status: resolved` post-gap-closure. Now treats both `complete` and `resolved` as terminal, with `result: all_pass` as a fallback when status is absent. Also reconciles workflows/help.md one-liner that referenced bare `SUMMARY.md` so docs match the authoritative quick.md workflow. Adds tests/bug-2836-audit-open-summary-uat-drift.test.cjs with 6 structural regression tests covering both fixes plus no-regression cases. * refactor(#2836): hoist TERMINAL_UAT_STATUSES outside scanUatGaps loop Address CodeRabbit nitpick: the Set was being recreated on each UAT file iteration. Hoist to module scope so it is constructed once. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
c3a42d66f9 | Revert "feat(install): add Hermes Agent runtime support" (#2849) | ||
|
|
5a636bc90a |
feat(install): add Hermes Agent runtime support (#2841)
Adds Hermes Agent as a supported installation target. Users can run
\`npx get-shit-done-cc --hermes\` to install all 86 GSD commands as
skills under \`~/.hermes/skills/gsd-*/SKILL.md\`, following the same
open skill standard as Claude Code 2.1.88+, Qwen Code, Antigravity,
Trae, Augment, and Codebuddy.
Hermes Agent is an open-source AI agent framework by Nous Research
(NousResearch/hermes-agent, MIT). Its skill loader accepts the Claude
skill format as-is: frontmatter parsed with PyYAML SafeLoader (unknown
keys like \`allowed-tools\` / \`argument-hint\` ignored), body XML tags
(\`<objective>\`, \`<execution_context>\`, \`<process>\`) passed directly
to the model. Compatibility proven end-to-end with all 86 GSD skills
loading cleanly, \`skill_view()\` returning full bodies, and
\`build_skills_system_prompt()\` emitting them into the agent system
prompt — zero Hermes code changes required.
Changes:
- \`bin/install.js\`: --hermes flag, getDirName/getGlobalDir/getConfigDirFromHome
support, HERMES_HOME env var (native to Hermes — used for profile
mode / Docker deploys), install/uninstall pipelines, interactive
picker option 10 (alphabetical: between Gemini and Kilo), .hermes
path replacements in copyCommandsAsClaudeSkills and
copyWithPathReplacement, legacy commands/gsd cleanup, CLAUDE.md ->
HERMES.md and "Claude Code" -> "Hermes Agent" content rewrites in
skills/agents/hooks, runtime-appropriate finish message.
- \`get-shit-done/bin/lib/core.cjs\`: add hermes to KNOWN_RUNTIMES;
add RUNTIME_PROFILE_MAP.hermes with OpenRouter-slug defaults
(Hermes is provider-agnostic; these defaults resolve across
OpenRouter, native Anthropic, and Copilot via Hermes' aggregator-
aware resolver, and are overridable per-tier via
model_profile_overrides.hermes.{opus,sonnet,haiku}).
- \`README.md\`: Hermes Agent in tagline, runtime list, verification
command, install/uninstall examples, \`--hermes\` flag reference.
- \`tests/hermes-install.test.cjs\`: new, 14 tests covering directory
mapping, HERMES_HOME env var precedence, install/uninstall
lifecycle, user-skill preservation, engine cleanup.
- \`tests/hermes-skills-migration.test.cjs\`: new, 11 tests covering
frontmatter conversion, path replacement (~/.claude/ ->
\$HERMES_HOME/skills/), CLAUDE.md -> HERMES.md, "Claude Code" ->
"Hermes Agent", stale skill cleanup, SKILL.md format validation.
- \`tests/multi-runtime-select.test.cjs\`: updated for new option
numbering (hermes=10, kilo=11, opencode=12, qwen=13, trae=14,
windsurf=15, all=16).
- \`tests/kilo-install.test.cjs\`: updated assertions for Kilo having
moved from option 10 to option 11.
Closes #2841
Implementation notes:
- Zero custom code paths: Hermes reuses copyCommandsAsClaudeSkills()
identical to Qwen Code / Antigravity pattern.
- Path replacement: ~/.claude/, \$HOME/.claude/, ./.claude/ ->
.hermes equivalents in skill/agent/hook content.
- Config precedence: --config-dir > HERMES_HOME > ~/.hermes (matches
how Hermes itself resolves its home directory).
- Legacy cleanup: removes commands/gsd/ if present from a prior
install, preserving dev-preferences.md (same as Qwen).
- No external dependencies added.
Testing: 5841 / 5841 tests pass (0 failures, 0 regressions)
- 14 new tests in hermes-install.test.cjs
- 11 new tests in hermes-skills-migration.test.cjs
- multi-runtime-select.test.cjs renumbered + 1 new test (single choice for hermes)
|
||
|
|
eeaf9c556f |
fix(#2787): track fenced code blocks in extractCurrentMilestone (#2812)
* fix(#2787): track fenced code blocks in extractCurrentMilestone The milestone-end search used a multiline regex against the raw restContent string. Lines inside fenced code blocks (``` or ~~~) that matched the milestone-heading pattern (e.g. `# note v1.0`) prematurely set sectionEnd, hiding all phases after the block from roadmap analyze, roadmap get-phase, and every downstream command. Replace the regex match with a line-by-line scan that tracks fence state. Lines inside an open fence are skipped regardless of content. Adds three regression tests covering backtick fences, tilde fences, and the roadmap get-phase code path. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#2787): track fence delimiter instead of toggling bare boolean Replace the inFence boolean with fenceChar/fenceLen tracking so that indented fences (up to 3 leading spaces) and mixed-delimiter content (~~~ inside a backtick fence) are parsed correctly. A closing fence is only recognised when it uses the same character as the opening delimiter and has at least the same run length, matching the CommonMark spec. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#2787): require fence-only closing line — reject info-string lines as closers A closing fence delimiter must contain only optional trailing whitespace. A line like \`\`\`js inside an open fence has an info string and must not close it. The previous regex /^\s{0,3}([`~]{3,})/ matched the opening of any such line, so the closing check could toggle fenceChar off on an info-string line and expose subsequent heading-like content to the milestone-end detector. Fix: capture the trailing portion of every fence-candidate line and only clear fenceChar when trailing matches /^\s*$/ (per CommonMark §4.5). Adds a regression test covering the ```text / ```js nesting scenario. --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
eddb2a205b |
fix(#2801): add ingest-docs handler to gsd-tools init dispatch (#2820)
* fix(#2801): add ingest-docs handler to gsd-tools init dispatch The `/gsd-ingest-docs` workflow was broken because `workflows/ingest-docs.md` called `gsd-sdk query init.ingest-docs` but the installed binary is `gsd-tools`, and `gsd-tools init` had no `ingest-docs` case in its dispatch switch. - Added `cmdInitIngestDocs` function to `init.cjs` and exported it; returns `project_exists`, `planning_exists`, `has_git`, `project_path`, `commit_docs` - Added `case 'ingest-docs'` to the `init` switch in `gsd-tools.cjs` - Updated `workflows/ingest-docs.md` to call `gsd-tools init ingest-docs` (line 55) and `gsd-tools commit` (line 292) instead of `gsd-sdk query ...` - Regression test: `tests/bug-2801-ingest-docs-handler.test.cjs` Closes #2801 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#2801): address CodeRabbit — commit_docs assertion, broader gsd-sdk detection, bash fence --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
5fe1f00a0d |
fix(#2808): SKILL.md files use hyphen name form (gsd-cmd not gsd:cmd) (#2819)
* fix(#2808): SKILL.md name uses hyphen form for Claude Code autocomplete skillFrontmatterName() was converting gsd-<cmd> to gsd:<cmd> (colon) so installed SKILL.md files had name: gsd:add-phase etc. Claude Code surfaces this name in autocomplete, showing the deprecated colon form to users even though the hyphen form is canonical everywhere else. Root cause: the colon form was needed because workflows called Skill(skill="gsd:<cmd>"). All 4 remaining colon-form Skill() calls in autonomous.md and execute-phase.md are updated to hyphen form. skillFrontmatterName() now returns the hyphen dir name unchanged. Updated 4 existing tests that asserted colon form. Regression test: tests/bug-2808-skill-hyphen-name.test.cjs * fix(#2808): address CodeRabbit — bash/text fences, structured test assertions, fail-loud on errors |
||
|
|
d46efb4790 |
fix(#2784): clear shared ~/.cache/gsd/ update-check cache in update workflow (#2813)
* fix(#2784): clear shared ~/.cache/gsd/ cache in update workflow The SessionStart hook (hooks/gsd-check-update.js) writes update-check results to $HOME/.cache/gsd/gsd-update-check.json (shared, tool-agnostic). The update.md run_update step only cleared per-runtime paths like ~/.claude/cache/gsd-update-check.json, so the statusline kept showing the stale upgrade indicator after a successful update. Fix: add rm -f "$HOME/.cache/gsd/gsd-update-check.json" to the cache-clear block in the run_update step. Regression test: tests/bug-2784-update-cache-clear-path.test.cjs * fix(#2784): address CodeRabbit review — four edge-cases count, bash fence, structured test assertions |
||
|
|
936cf26706 |
fix(#2769): tolerate Requirements header with colon inside bold delimiters (#2782)
* fix(#2769): tolerate Requirements header with colon inside bold delimiters extractReqIds in sdk/src/query/init.ts and the legacy init.cjs port only matched `**Requirements**:` (colon outside bold), so phases declared with the equally-valid markdown form `**Requirements:**` (colon inside bold, which is what the project's own templates emit) returned phase_req_ids: null for both `init plan-phase` and `init execute-phase`. The mirror-image bug in `phase complete`'s REQUIREMENTS.md traceability sweep at get-shit-done/bin/lib/phase.cjs:871 only matched the inside-bold form, silently skipping the REQ-ID checkbox flips for any roadmap that used the outside-bold form. Both parsers now share the same canonical regex that accepts all three rendered-identical variants: **Requirements:** (colon inside bold) **Requirements**: (colon outside bold) **Requirements** : (space before outside colon) Tests: - tests/init.test.cjs — parameterized over the three header variants for both init plan-phase and init execute-phase (6 new behavioral cases). - sdk/src/query/init.test.ts — describe.each over the same variants exercising initPlanPhase through the SDK. - tests/bug-2769-requirements-header-variants.test.cjs — phase complete flips REQ-001 in REQUIREMENTS.md across all three header variants. Closes #2769 * refactor(#2769): centralize REQUIREMENTS_HEADER_RE constant per CodeRabbit |
||
|
|
54e6da3126 |
fix(#2767): pass paths via --files to gsd-sdk query commit + lint guard (#2781)
* fix(#2767): pass paths via --files to gsd-sdk query commit + lint guard Workflows, agents, commands, and references passed file paths positionally to `gsd-sdk query commit`, which silently appended them to the commit subject and triggered the `.planning/` wholesale-stage fallback in sdk/src/query/commit.ts:136. Regression of #733/#798. Inserted `--files` before the path list at every site (81 invocations across 50 files). Added tests/bug-2767-gsd-sdk-commit-files-flag.test.cjs as a permanent lint that scans every shipped .md file and asserts each `gsd-sdk query commit[-to-subrepo]` invocation either uses `--files` or carries no path arguments. Closes #2767 * test(#2767): replace source-grep with behavioral SDK test The original test walked every shipped .md file and regex-tokenized `gsd-sdk query commit` invocations to assert `--files` was present. CONTRIBUTING.md prohibits this source-grep pattern. Rewrite as behavioral SDK tests against `sdk/dist/cli.js` over a real tmp git project (createTempGitProject helper). Cover both the well-formed (`--files <paths>`) form — clean subject, exactly-staged files, .planning/ left untouched — and the buggy positional form, asserting the documented misbehavior (paths leak into subject + the `.planning/` wholesale-stage fallback at commit.ts:136). Also asserts `commit-to-subrepo` rejects when `--files` is omitted (commit.ts:258). The doc-lint is retained as a supplementary defense-in-depth guard since agent-prompt markdown invocations cannot be exercised end-to-end — but it is no longer the primary contract. * docs(#2767): correct contradictory --files guidance in zh-CN/en docs + fix test docstring |
||
|
|
3ac3a2ae70 |
fix(#2770): coerce non-string depends_on YAML values to preserve dependencies (#2780)
* fix(#2770): coerce non-string truths to preserve cross-cutting constraints `cmdRoadmapAnnotateDependencies` skipped non-string truth entries via `if (typeof t !== 'string') continue`. That avoided the TypeError reported in #2770 but silently dropped legitimate constraints — numeric YAML scalars (`- 3`) and kv-shaped truths from parseMustHavesBlock's continuation-kv path (#2757) — from the cross-cutting analysis, leaving ROADMAP.md under-annotated. Replace the skip-guard with a `coerceTruthToString` helper that: * passes strings through * `String()`-coerces numbers, booleans, bigints * extracts a string field (title, text, name, rule, path, provides) from object-shaped items Composes cleanly with #2757 (objects from kv continuation lines now contribute their title rather than being dropped) and the existing `splitInlineArray` quote-aware parser. Tests: tests/bug-2770-annotate-deps-int-coerce.test.cjs - numeric scalar truth shared across plans surfaces as constraint - kv-shaped truth surfaces via title field - bare-int depends_on regression guards on extractFrontmatter Full suite: 5678 pass, 0 fail. Closes #2770 * test(#2770): use array join() for multi-line fixtures per CONTRIBUTING * refactor(#2770): cache trim() and avoid no-op truthCounts.set in aggregation |
||
|
|
8b6c44433f |
fix(#2772): only disable worktree isolation when planned paths touch submodules (#2779)
* fix(#2772): only disable worktree isolation when planned paths touch submodules The previous guard in execute-phase.md and quick.md unconditionally set USE_WORKTREES=false whenever .gitmodules existed, penalising every plan in a submodule project even when no plan touched a submodule path. Replace with submodule-path parsing + per-plan path intersection: - Parse SUBMODULE_PATHS once from .gitmodules via `git config --file .gitmodules --get-regexp '^submodule\..*\.path$'`. - In execute-phase.md, intersect SUBMODULE_PATHS with each plan's files_modified frontmatter; disable worktree isolation only for plans with non-empty intersection. Fall back to safe-disable for that plan when files_modified is missing/unparseable, with a log line explaining why. - In quick.md (no pre-declared paths), keep submodule-path parsing and document a fail-loud commit-time guard so the executor aborts only when it actually stages a submodule path. Add tests/bug-2772-gitmodules-path-intersection.test.cjs covering both files: no unconditional disable, submodule paths are parsed, intersection logic exists in execute-phase, fallback path is documented. Full suite: 5680 / 5680 pass. Closes #2772 * test(#2772): replace source-grep with behavioral test of submodule path intersection * fix(#2772): wire USE_WORKTREES_FOR_PLAN into dispatch + fix glob matcher + add quick.md commit guard Address CodeRabbit review on PR #2779 — the original fix computed USE_WORKTREES_FOR_PLAN but never read it, so the per-plan submodule intersection was dead code. Dispatch sites still branched on the project-level USE_WORKTREES. Changes: 1. execute-phase.md (CRITICAL — dispatch wiring): Move per-plan computation into execute_waves as sub-step 2.5, run it for each plan before its dispatch, and gate all four dispatch sites on USE_WORKTREES_FOR_PLAN: worktree-mode header, sequential-mode header, "worktrees disabled" sequential rule, and post-wave cleanup. Document PLAN_FILES extraction via jq from the phase-plan-index JSON. Track WAVE_WORKTREE_PLANS so post-wave cleanup only runs when at least one plan in the wave actually used worktrees. 2. Per-plan gate matcher (MAJOR — glob safety): Strip leading "./" and trailing "/" from both submodule and planned paths. Match bidirectionally (pf inside sm AND sm inside pf). Handle globby planned paths like "vendor/**/*.c" by extracting the literal prefix before the first glob metachar and re-checking. Wrap the iteration in set -f / set +f so glob expansion does not corrupt patterns. Extracted the gate (~92 lines) into workflows/execute-phase/steps/per-plan-worktree-gate.md to keep execute-phase.md under the 1700-line XL budget. 3. quick.md (CRITICAL — fail-loud guard): Inject SUBMODULE_PATHS into the executor Task prompt and add a <submodule_commit_guard> bash block the executor must run before every git commit. The guard inspects staged paths via `git diff --cached --name-only`, normalizes paths, and aborts with a clear ABORT message + recovery instruction ("re-run with workflow.use_worktrees=false") when any staged path falls inside a submodule. 4. tests/bug-2772-gitmodules-path-intersection.test.cjs: 25 tests total. Updated GATE_SNIPPET to match the new bash matcher. Added normalization tests (./ prefix, trailing /, glob "vendor/**/*.c", parent directory, ./ in .gitmodules). Added workflow-markdown wiring assertions for all 4 dispatch sites + per-plan gate file extraction. Added quick.md guard tests: prompt injection assertion + behavioral fixture-repo tests that stage a submodule path and assert the guard exits non-zero with the ABORT message. Test count: 5701 pass / 0 fail (was 5698/1 before). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
77d929429f |
fix(#2774): inclusion-based worktree cleanup to protect workspace .git (#2778)
* fix(#2774): inclusion-based worktree cleanup to protect workspace .git The cleanup blocks in execute-phase.md and quick.md used an exclusion filter (`grep -v "$(pwd)$"`) to skip the current worktree before calling `git worktree remove --force` on everything else. The exclusion fails whenever the current workspace is itself a worktree of an upstream repo: - multi-workspace setups where `git worktree list` reports the registry path as a different absolute path than `$(pwd)` - the cross-drive Windows case where the registry reports `E:/...` while `$(pwd)` resolves to `C:/...` — the equality test never holds, every other worktree (including the workspace itself) is removed, and the workspace's `.git` pointer file is destroyed. Switches both cleanup blocks to an inclusion-based filter that targets only agent-spawned worktrees under `.claude/worktrees/agent-`, the namespace Claude Code's `isolation="worktree"` always uses for executor worktrees. The workspace path can never collide with that prefix. Adds tests/bug-2774-worktree-cleanup-workspace-safety.test.cjs covering: - both workflow files use the inclusion filter - neither falls back to the broken `grep -v "$(pwd)$"` guard - end-to-end simulation of porcelain output with workspace + agent worktrees yields only the agent worktree Closes #2774 * test(#2774): replace source-grep with behavioral test of cleanup pipeline * fix(#2774): whitespace-safe worktree iteration with while/read CodeRabbit review on PR #2778 flagged that `for WT in $WORKTREES` splits on whitespace. Any agent worktree path containing a space (e.g. a workspace under '/Users/dev/My Workspace/') would be torn into broken half-paths, `git -C` would fail on each fragment, and the executor branch would never be deleted. Switch both cleanup blocks (quick.md and execute-phase.md) to: while IFS= read -r WT; do [ -z "$WT" ] && continue ... done < <(git worktree list --porcelain | grep ... | sed ...) Process substitution feeds the pipeline output line-by-line — IFS= and -r preserve every byte of the path including embedded spaces. Also rename the misleading `makeBareTempGitRepo` helper to `makeTempUpstreamRepo` (it does not pass --bare; it inits a normal repo with an initial commit so worktree-add works). Add two new behavioral tests: - discovery pipeline yields whitespace paths intact on a single line - the actual while/read loop iterates each whitespace-bearing path exactly once (would fail with the previous `for WT in` form) Tests: 5681 pass, 0 fail. |
||
|
|
dc9b712967 |
refactor(state): drop unused args + lift currentPhase in cmdStateCompletePhase (#2761)
* refactor(state): drop unused args param and lift currentPhase in cmdStateCompletePhase Two cleanup items surfaced by CodeRabbit review of PR #2759: 1. cmdStateCompletePhase(cwd, args, raw) — args is never read inside the function. All sibling state subcommands use the leaner (cwd, raw) shape. Remove the unused parameter and update the dispatch call in gsd-tools.cjs. 2. output() at line 1754 called fs.readFileSync(statePath) after readModifyWriteStateMd had already released the lock, re-extracting Current Phase via an extra fs read. The closure already computed currentPhase at line 1704; lifting resolvedPhase into outer scope and capturing it in the callback eliminates the post-lock read and closes the small race window. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#2761): apply CodeRabbit nitpicks with regression tests Two CodeRabbit nitpicks from PR #2761 review, each landed with a regression test so a future refactor can't unwind them. 1. tests/dispatcher.test.cjs — pin the enumerated subcommand list: the 'state unknown subcommand errors' test now also asserts that the dispatcher's error string includes 'complete-phase'. Without this, a future reformat of the available-subcommands enumeration could silently drop entries and the existing 'Unknown state subcommand' substring check would still pass. 2. get-shit-done/bin/lib/state.cjs — tighten the Phase fallback in cmdStateCompletePhase: when STATE.md is missing the canonical '**Current Phase:**' field and the only phase signal is the decorated body line under '## Current Position' (e.g. 'Phase: 01 (Foo) — EXECUTING'), the previous fallback returned the entire decorated string, producing messy downstream output: Status: Phase 01 (Foo) — EXECUTING complete Phase: 01 (Foo) — EXECUTING — COMPLETE The fallback now strips everything past the leading numeric/decimal token via /^\\s*([\\w.-]+)/ so degraded inputs produce clean output identical to the canonical path. 3. tests/state.test.cjs — two new tests in a dedicated describe block: - decorated Phase line writes clean Phase identifier - canonical Current Phase wins over Current Position decoration Both run real `gsd state complete-phase` against synthetic STATE.md fixtures and assert on the rendered Status field. --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
9472f343db |
feat(#2762): --minimal install profile (≥94% cold-start token reduction) (#2764)
* feat(#2762): add --minimal install profile to cut cold-start token cost Eager system-prompt load from 86 gsd-* skill descriptions plus 33 subagent descriptions costs ~12k tokens per turn even in directories with no .planning/. Frontier models (Sonnet 4.6 / Opus 4.7) with 200K-1M context don't feel it; local LLMs with 32K-128K do. --minimal (alias --core-only) installs only the main GSD loop: new-project, discuss-phase, plan-phase, execute-phase, plus help/update. Zero gsd-* subagents are written. Re-running gsd update without --minimal expands to the full surface. Default install behavior is unchanged. DRY: a single stageSkillsForMode() helper filters the source dir; all 13 runtime-specific copy fns are unchanged because they recurse the staged dir. Allowlist + helpers live in get-shit-done/bin/lib/install- profiles.cjs as the single source of truth. Manifest now records mode: 'minimal' | 'full' so future commands can detect install profile. Tested end-to-end: --minimal yields 6 skill folders + 0 agents; default yields 86 + 33 (unchanged). * docs(#2762): document --minimal install in README Adds a collapsible 'Minimal Install' section under Getting Started covering: who it's for (local LLMs, token-billed APIs), what you get (6 skills, 0 subagents, ~700 token floor vs ~12k), and the critical caveat that re-installing without --minimal restores the full surface and erases the savings. Includes a comparison table, the manifest inspection one-liner, and the use-case decision matrix. * fix(#2762): address CodeRabbit review + CI failures CodeRabbit findings: 1. Temp dir leak (Minor): stageSkillsForMode created tmp dirs that were never cleaned up. Added a module-level Set tracking every staged dir plus a process.on('exit') handler that rm -rf's them. Also wrap the copy loop in try/catch to remove a partially-populated tmp dir on mid-flight failure. Verified end-to-end: 0 leaked dirs in /tmp after a real install. 2. Codex full -> minimal stale state (Major): a previous full Codex install left agents/gsd-*.toml files plus [agents.gsd-*] sections in config.toml. The original cleanup only removed .md files, so a switch to --minimal would leave Codex still advertising the full agent surface. Cleanup now also handles .toml under isCodex, and minimal mode strips GSD sections from config.toml via the existing stripGsdFromCodexConfig helper (same path used by --uninstall). 3. Nitpick — Codex downgrade regression test: added a spawnSync-based end-to-end test that fakes a previous full install (stale gsd-*.md + gsd-*.toml + GSD-marked config.toml + a user-owned agent/setting), runs install.js --codex --minimal, and asserts stale GSD files + sections are gone while user content is preserved. CI failures (inventory parity): - docs/INVENTORY.md CLI Modules table now lists install-profiles.cjs with the correct headline count (30 -> 31). - docs/INVENTORY-MANIFEST.json regenerated via gen-inventory-manifest.cjs. Test count: 149 pass (was 116 in last commit; +14 new install-minimal + all previously-failing inventory tests now green). * test(#2762): expand install-minimal test coverage for future-proofing Each new test pins a specific guarantee that closes off a future regression class — turning every CodeRabbit finding (including the nitpicky one) into a permanent guard. cleanupStagedSkills suite (+3 tests): - 'full mode does not register a staged dir' — catches a future regression where someone forgets the early-return in stageSkillsForMode and starts polluting STAGED_DIRS in default installs. - 'exit handler registers exactly once across many calls' — catches removal of the exitHandlerRegistered guard. install.js has 13 dispatch sites, so a missing guard would attach 13 listeners. - 'mid-copy failure removes partial staged dir and re-throws' — intercepts fs.copyFileSync to throw mid-loop and asserts the staged dir count in /tmp is unchanged after the throw. Pins the exact CodeRabbit-flagged leak. Claude full -> minimal downgrade (+1 test): - Mirrors the Codex downgrade test for the .md-only path that the other 12 runtimes share. Asserts user-owned agents are preserved. Manifest mode round-trip (+3 tests): - Default install -> mode: 'full' with >6 skills and >0 agents - --minimal -> mode: 'minimal' with exactly 6 skills and 0 agents - --core-only alias produces identical manifest to --minimal Allowlist scope guards (+3 tests): - Every main-loop command IS in allowlist (positive) - Off-loop commands (autonomous, ship, do, progress, next, fast, quick, debug, code-review, verify-work) are NOT (guards against silent scope creep — future contributor adds 'autonomous' to core and the floor erodes) - Unknown mode strings fall through to full behavior — pre-emptive guard for future 'compact'/'tier2' modes that might forget to update the predicate. Total: 25 tests in this file (was 15), 159/159 passing across the install + inventory suites. * fix(#2762): clean up staged tmp dirs on SIGINT/SIGTERM/SIGHUP CodeRabbit follow-up review on c727bf5f flagged that process.on('exit') does not fire on signal-driven termination. An installer is exactly the kind of process users abort mid-run with Ctrl+C, so without explicit signal handlers the staged tmp dirs in STAGED_DIRS would be left behind until the OS reaps tmpdir. Fix: ensureExitCleanup now also registers process.once handlers for SIGINT, SIGTERM, SIGHUP. Each handler runs cleanupStagedSkills then re-raises the same signal via process.kill(pid, sig) so the OS-default handler takes over and the parent shell sees the correct exit code (130 for SIGINT, etc.) — CI scripts and interactive users see the abort the way they expect. Test: spawns a child that stages a tmp dir then blocks; parent captures the staged path from stdout, sends SIGINT, asserts (a) the staged dir is gone after child exit, (b) child exits via the signal not via code 0. Skipped on Windows (signal semantics differ; the natural-exit cleanup test covers the Windows CI matrix). Total: 26 tests in install-minimal.test.cjs (was 25). |
||
|
|
ab5ad6c8bc |
fix(#2757): unquoted truths with colons crash annotate-dependencies (#2759)
parseMustHavesBlock dispatched on `includes(':')` to detect key-value pairs,
but unquoted YAML strings like `GET /foo/:id resolves...` and
`Class::Method is idempotent` also contain colons. When the KV regex failed
to match, `current` was left as `{}` (the empty object initialized before
the branch), which then caused `t.trim()` in roadmap.cjs to throw
`TypeError: t.trim is not a function`.
Two fixes:
- frontmatter.cjs: tighten the KV regex to require at least one space after
the colon (`\s+` instead of `\s*`), matching YAML convention. When the
regex still fails to match, fall back to treating the item as a plain
string instead of leaving `current` as `{}`.
- roadmap.cjs: add `typeof t !== 'string'` guard before `.trim()` as a
cheap safety net against any future parser anomaly.
Closes #2757
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
||
|
|
1405728292 |
fix: parseMustHavesBlock quoted strings + gsd state complete-phase (#2744)
* fix: parseMustHavesBlock quoted strings + gsd state complete-phase Bug #2734: parseMustHavesBlock dropped quoted truths containing ':' because fully-quoted strings like `"App-side UUIDv4: generated locally"` fell into the kv-parse branch, the regex failed (value starts with '"'), and current stayed as empty {}. Fix: detect fully-quoted strings before the ':' check and extract them directly. Two regression tests added to frontmatter.test.cjs. Bug #2735: `gsd state complete-phase` subcommand was missing — unknown subcommands fell through to cmdStateLoad. Added cmdStateCompletePhase to state.cjs (updates Status, Last Activity, and Current Position to COMPLETE), exported it, and wired it into the case 'state': dispatch in gsd-tools.cjs. Closes #2734 Closes #2735 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(state): unknown subcommand returns explicit error instead of silent fallthrough Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
3246810876 |
feat: extend RUNTIME_PROFILE_MAP to gemini/qwen/opencode/copilot + settings-advanced UI (#2754)
Closes #2612 - Add gemini, qwen, opencode, and copilot entries to RUNTIME_PROFILE_MAP in core.cjs - Group B runtimes (kilo, cline, cursor, windsurf, augment, trae, codebuddy, antigravity) intentionally have no built-in map and fall through to the existing unknown-runtime fallback - Add 40 new tests to tests/issue-2517-runtime-aware-profiles.test.cjs covering each new runtime's three tiers, Group B fall-through, and partial override merge semantics - Add Section 7 "Runtime Model Tiers" to settings-advanced.md with interactive UI to view and override built-in tier defaults per runtime - Update docs/CONFIGURATION.md built-in tier table to include all four new runtimes Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
e0b4561fa9 |
feat: add /gsd-edit-phase command to modify roadmap phases in place (#2753)
Adds a new slash command that lets developers modify any field of an existing phase in ROADMAP.md without affecting phase number or position. - commands/gsd/edit-phase.md: command file with --force flag support - get-shit-done/workflows/edit-phase.md: full workflow with status guard, depends_on validation, diff+confirmation, and STATE.md update - tests/edit-phase.test.cjs: 32 tests covering all acceptance criteria - docs/INVENTORY.md, INVENTORY-MANIFEST.json, COMMANDS.md: registered Closes #2617 Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
8788ab2381 |
feat: post-merge build & test gate — Build step, iOS/Xcode, serial mode (#2751)
* feat: post-merge build & test gate — Build step, iOS/Xcode, serial mode Step 5.6 of execute-phase is extended per #2720: - Renamed from "Post-merge test gate (parallel mode only)" to "Post-merge build & test gate" - Gate now runs in both parallel mode (after worktree merge) and serial mode (after last plan) - Added Step A: Build gate resolving BUILD_CMD from workflow.build_command config key, then auto-detecting via priority: config override → Xcode (.xcodeproj) → Makefile build: → Justfile → Cargo/Go/Python/npm. Xcode uses xcodebuild -list -json to get first scheme, then xcodebuild build -scheme ... -destination 'platform=iOS Simulator,name=iPhone 16'. Build failure increments WAVE_FAILURE_COUNT. - Added Xcode/iOS detection to Step B (Test gate): when *.xcodeproj present and no workflow.test_command configured, uses xcodebuild test instead of the previous "no test runner detected" skip. Scheme reused from Step A when available. - Documented workflow.build_command and workflow.test_command in docs/CONFIGURATION.md (table + JSON schema) Closes #2720 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(execute-phase): extract Step 5.6 body to post-merge-gate.md sub-file Moves the build-detection logic and xcodebuild commands from the inline Step 5.6 body into execute-phase/steps/post-merge-gate.md, replacing it with a single Read() reference. Reduces execute-phase.md from 1755 to 1647 lines, satisfying the ≤1700 XL-tier budget enforced by tests/workflow-size-budget.test.cjs. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
8f2ec0e8f7 |
fix: add explicit wait-for-subagent instructions in orchestration workflows (#2755)
Adds ORCHESTRATOR RULE blockquotes immediately after every Task() spawn in 26 GSD workflow files, instructing the parent orchestrator to stop working on the task while the subagent is active. This prevents the parallel-work anti-pattern on Codex runtime where the parent continues reading files and producing duplicate/conflicting output after spawning. Rules are placed inline at each spawn point (not as generic headers) so they are adjacent to and unambiguously associated with each Task() call. Background Task() spawns get a variant noting not to return to the spawning context until the subagent reports back. Closes #2729 Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
7255539ff9 |
fix: validate LM Studio model identity in review workflow (#2746)
* fix: validate LM Studio model identity in review workflow Capture the full API response before extracting content, then compare the top-level `.model` field against the configured LM_STUDIO_MODEL. Emits a warning to stderr if LM Studio served a different model than requested, while still proceeding with the review response. Closes #2721 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(review): skip LM Studio review file when content is empty instead of writing error text Also applies the same fix to llama.cpp which had the identical pattern of writing a literal error string into the review temp file when content was empty/null. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
2b95ccbddd |
Merge pull request #2743 from gsd-build/fix/2714-workstream-config-inheritance
feat(config): workstream config.json inherits from root .planning/config.json |
||
|
|
0eef943f0a |
Merge pull request #2739 from gsd-build/fix/2732-graphify-cli-invocation
fix(graphify): update CLI invocation from legacy flag form to subcommand |
||
|
|
cb149383c1 |
fix(config): bump MODEL_ALIAS_MAP to claude-opus-4-7 (#2733)
* fix(config): bump MODEL_ALIAS_MAP and RUNTIME_PROFILE_MAP to claude-opus-4-7 Opus 4.7 shipped Q1 2026 but MODEL_ALIAS_MAP and RUNTIME_PROFILE_MAP.claude.opus were still pinned to claude-opus-4-6. Users with resolve_model_ids: true received stale model IDs in logs and agent-tool calls. Also adds a resolve_model_ids: true test suite — this path had zero coverage, which is why the stale ID survived undetected. Closes #2712 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(config): derive RUNTIME_PROFILE_MAP.claude from MODEL_ALIAS_MAP (coderabbit) RUNTIME_PROFILE_MAP.claude was duplicating model IDs that MODEL_ALIAS_MAP already owns. Future model bumps now only require updating MODEL_ALIAS_MAP. Also fixes stale test assertion (claude-opus-4-6 → claude-opus-4-7). * fix(tests): update stale claude-opus-4-6 refs to claude-opus-4-7; DRY: derive RUNTIME_PROFILE_MAP.claude from MODEL_ALIAS_MAP - Update 3 hardcoded `claude-opus-4-6` assertions in tests/issue-2517-runtime-aware-profiles.test.cjs to `claude-opus-4-7` - Update comment on line 128 that referenced the old model ID - Replace manual per-tier expansion of RUNTIME_PROFILE_MAP.claude with Object.fromEntries so future alias bumps only require updating MODEL_ALIAS_MAP Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
1cb4bebcf5 |
feat(config): workstream config.json inherits from root .planning/config.json
- Add _deepMergeConfig() with correct null-override semantics - loadConfig() reads root config.json when GSD_WORKSTREAM is set, then deep-merges with workstream config (workstream wins on conflict) - Workstream without config.json falls back to root config entirely - Migrations and disk writes operate on fileData (on-disk content) only, never on the merged result, to prevent workstream pollution - Fixes null-override bug from PR #2717: explicit null in workstream now correctly overrides root value instead of falling back to root - Tests: inherit root model_overrides, workstream override, nested workflow.* deep merge, explicit null override, missing workstream config Closes #2714 |
||
|
|
f89a56eb55 |
fix(graphify): update CLI invocation from legacy flag form to subcommand
graphify . --update was removed in favor of graphify update . in v0.4.x. Also improves version detection to try `graphify --version` before falling back to python3 importlib query. Closes #2732 |
||
|
|
470c1a0bff |
fix(#2722): forensics gh commands pin --repo gsd-build/get-shit-done (#2723)
* fix(#2722): forensics gh commands pin --repo gsd-build/get-shit-done gh issue create and gh label list both defaulted to the repo inferred from $PWD, causing issues to be submitted to the user's current project instead of this repo. Added --repo gsd-build/get-shit-done to both commands. Added two regression tests covering both gh calls. Closes #2722 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#2723): scope forensics tests to specific gh commands, not whole file CodeRabbit found that the gh issue create test searched the whole workflow file, so it would pass even if gh issue create lacked --repo (because gh label list already contains the repo string elsewhere). Also replaced the brittle 200-char slice in the label-list test with a regex. Both tests now use assert.match() with command-scoped regexes. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
b40110111d |
feat(#2306): plan-review-convergence v2 — CYCLE_SUMMARY contract, config gate, local model reviewers (#2718)
* feat(#2306): plan-review-convergence v2 — CYCLE_SUMMARY contract, config gate, local model reviewers Fixes the false-stall detection bug in the plan→review→replan convergence loop. REVIEWS.md accumulates history across cycles so raw grep inflated HIGH counts; HIGH count now comes from a per-cycle CYCLE_SUMMARY contract emitted in the review agent's return message. Key changes: - workflow.plan_review_convergence config gate (disabled by default, same pattern as workflow.code_review / workflow.nyquist_validation) - Review agent prompt defines CYCLE_SUMMARY: current_high=<N> contract with PARTIALLY RESOLVED / FULLY RESOLVED counting rules - Orchestrator aborts on absent/malformed CYCLE_SUMMARY (distinguishes both) - Warns when HIGH_COUNT > 0 but ## Current HIGH Concerns section is missing - Stall detection and --ws forwarding preserved and tested - Local model reviewers: --ollama, --lm-studio, --llama-cpp flags added to convergence workflow and review workflow; all three use OpenAI-compatible /v1/chat/completions endpoint with jq --rawfile for safe JSON encoding - review.ollama_host / review.lm_studio_host / review.llama_cpp_host config keys registered and documented (default to localhost:11434/1234/8080) - review.models.ollama / .lm_studio / .llama_cpp model-name config support - 58 tests (up from 29 in PR #2339), all passing Closes #2306 Closes #2339 Co-authored-by: Tom Boucher <trekkie@nomorestars.com> Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(ci): sync sdk/src/query/config-schema.ts with CJS schema (#2306) Add workflow.plan_review_convergence, review.ollama_host, review.lm_studio_host, and review.llama_cpp_host to the SDK-side TypeScript mirror — required by the CJS↔SDK parity test (#2653). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#2306): resolve CodeRabbit review findings - Anchor HIGH_COUNT extraction with head -1 to prevent multi-match when agent return message contains multiple CYCLE_SUMMARY lines (e.g. quoted back from prompt context) - Replace hardcoded reviewers list in REVIEWS.md frontmatter template with runtime-derived placeholder — the static list did not reflect which reviewers were actually invoked - Broaden workflow.plan_review_convergence docs to include local reviewers (Ollama, LM Studio, llama.cpp) alongside cloud reviewers Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(ci): restore reviewers frontmatter list with runtime note The cursor-reviewer.test.cjs (and equivalent per-reviewer tests) assert that each supported reviewer appears on the reviewers: line — these are wiring tests that catch when a new reviewer is added to invocation but not to the REVIEWS.md template. Replacing the list with a placeholder broke those tests. Restore the full static list and add an inline comment clarifying that the actual committed frontmatter should be filtered to only the reviewers invoked that run — satisfying both the per-reviewer tests and the CodeRabbit correctness note. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
8393f4b355 |
fix(#2601): config-set-model-profile now accepts 'inherit' value (#2707)
VALID_PROFILES was derived solely from Object.keys(MODEL_PROFILES['gsd-planner']), which only contained the named tiers (quality/balanced/budget/adaptive). The cmdConfigSetModelProfile validator rejected 'inherit' even though the runtime has supported it since #1829. Fix: append 'inherit' to VALID_PROFILES and handle it in getAgentToModelMapForProfile so the agent→model table shows 'inherit' instead of undefined. Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
b7ff14fe51 |
fix(#2687): derive KNOWN_TOP_LEVEL from DYNAMIC_KEY_PATTERNS to eliminate read-side drift (#2706)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
b1a670e662 |
fix(#2697): replace retired /gsd: prefix with /gsd- in all user-facing text (#2699)
All workflow, command, reference, template, and tool-output files that surfaced /gsd:<cmd> as a user-typed slash command have been updated to use /gsd-<cmd>, matching the Claude Code skill directory name. Closes #2697 Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
cd05725576 |
fix(#2661): unconditional plan-checkbox sync in execute-plan (#2682)
* fix(#2661): unconditional plan-checkbox sync in execute-plan
Checkpoint A in execute-plan.md was wrapped in a "Skip in parallel mode"
guard that also short-circuited the parallelization-without-worktrees
case. With `parallelization: true, use_worktrees: false`, only
Checkpoint C (phase.complete) then remained, and any interruption
between the final SUMMARY write and phase complete left ROADMAP.md
plan checkboxes stale.
Remove the guard: `roadmap update-plan-progress` is idempotent and
atomically serialized via readModifyWriteRoadmapMd's lockfile, so
concurrent invocations from parallel plans converge safely.
Checkpoint B (worktree-merge post-step) and Checkpoint C
(phase.complete) become redundant after A is unconditional; their
removal is deferred to a follow-up per the RCA.
Closes #2661
* fix(#2661): gate ROADMAP sync on use_worktrees=false to preserve single-writer contract
Adversarial review of PR #2682 found that unconditionally removing the
IS_WORKTREE guard violates the single-writer contract for shared
ROADMAP.md established by commit
|
||
|
|
c811792967 |
fix(#2660): capture prose after labeled bold in extractOneLinerFromBody (#2679)
* fix(#2660): capture prose after label in extractOneLinerFromBody The regex `\*\*([^*]+)\*\*` matched the first bold span, so for the new SUMMARY template `**One-liner:** Real prose here.` it captured the label `One-liner:` instead of the prose. MILESTONES.md then wrote bullets like `- One-liner:` with no content. Handle both template forms: - Labeled: `**One-liner:** prose` → prose - Bare: `**prose**` → prose (legacy) Empty prose after a label returns null so no bogus bullets are emitted. Note: existing MILESTONES.md entries generated under the bug are not regenerated here — that is a follow-up. Closes #2660 * fix(#2660): normalize CRLF before one-liner extraction Windows-authored SUMMARY files use CRLF line endings; the LF-only regex in extractOneLinerFromBody would fail to match. Normalize \r\n and \r to \n before stripping frontmatter and matching the one-liner pattern. Adds test case (h) covering CRLF input. |
||
|
|
303fd26b45 |
fix(#2662): add state.add-roadmap-evolution SDK handler; insert-phase uses it (#2683)
/gsd-insert-phase step 4 instructed the agent to directly Edit/Write
.planning/STATE.md to append a Roadmap Evolution entry. Projects that
ship a protect-files.sh PreToolUse hook (a recommended hardening
pattern) blocked the raw write, silently leaving STATE.md out of sync
with ROADMAP.md.
Adds a dedicated SDK handler state.add-roadmap-evolution (plus space
alias) that:
- Reads STATE.md through the shared readModifyWriteStateMd lockfile
path (matches sibling mutation handlers — atomic against
concurrent writers).
- Locates ### Roadmap Evolution under ## Accumulated Context, or
creates both sections as needed.
- Dedupes on exact-line match so idempotent retries are no-ops
({ added: false, reason: "duplicate" }).
- Validates --phase / --action presence and action membership,
throwing GSDError(Validation) for bad input (no silent
{ ok: false } swallow).
Workflow change (insert-phase.md step 4):
- Replaces the raw Edit/Write instructions for STATE.md with
gsd-sdk query state.patch (for the next-phase pointer) and
gsd-sdk query state.add-roadmap-evolution (for the evolution
log).
- Updates success criteria to check handler responses.
- Drops "Write" from commands/gsd/insert-phase.md allowed-tools
(no step in the workflow needs it any more).
Tests (vitest, sdk/src/query/state-mutation.test.ts): subsection
creation when missing; append-preserving-order when present;
duplicate -> reason=duplicate; idempotence over two calls; three
validation cases covering missing --phase, missing --action, and
invalid action.
This is the first SDK handler dedicated to STATE.md Roadmap
Evolution mutations. Other workflows with similar raw STATE.md
edits (/gsd-pause-work, /gsd-resume-work, /gsd-new-project,
/gsd-complete-milestone, /gsd-add-phase) remain on raw Edit/Write
and will need follow-up issues to migrate — out of scope for this
fix.
Closes #2662
|
||
|
|
7b470f2625 |
fix(#2633): ROADMAP.md is the authority for current-milestone phase counts (#2665)
* fix(#2633): use ROADMAP.md as authority for current-milestone phase counts initMilestoneOp (SDK + CJS) derives phase_count and completed_phases from the current milestone section of ROADMAP.md instead of counting on-disk `.planning/phases/` directories. After `phases clear` at the start of a new milestone the on-disk set is a subset of the roadmap, causing premature `all_phases_complete: true`. validateHealth W002 now unions ROADMAP.md phase declarations (all milestones — current, shipped, backlog) with on-disk dirs when checking STATE.md phase refs. Eliminates false positives for future-phase refs in the current milestone and history-phase refs from shipped milestones. Falls back to legacy on-disk counting when ROADMAP.md is missing or unparseable so no-roadmap fixtures still work. Adds vitest regressions for both handlers; all 66 SDK + 118 CJS tests pass. * fix(#2633): preserve full phase tokens in W002 + completion lookup CodeRabbit flagged that the parseInt-based normalization collapses distinct phase IDs (3, 3A, 3.1) into the same integer bucket, masking real STATE/ROADMAP mismatches and miscounting completions in milestones with inserted/sub-phases. Index disk dirs and validate STATE.md refs by canonical full phase token — strip leading zeros from the integer head only, preserve [A-Z] suffix and dotted segments, and accept just the leading-zero variant of the integer prefix as a tolerated alias. 3A and 3 never share a bucket. Also widens the disk and STATE.md regexes to accept [A-Z]? suffix tokens. |
||
|
|
c8ae6b3b4f |
fix(#2636): surface gsd-sdk query failures and add workflow↔handler parity check (#2656)
* fix(#2636): surface gsd-sdk query failures and add workflow↔handler parity check Root cause: workflows invoked `gsd-sdk query agent-skills <slug>` with a trailing `2>/dev/null`, swallowing stderr and exit code. When the installed `@gsd-build/sdk` npm was stale (pre-query), the call resolved to an empty string and `agent_skills.<slug>` config was never injected into spawn prompts — silently. The handler exists on main (sdk/src/query/skills.ts), so this is a publish-drift + silent-fallback bug, not a missing handler. Fix: - Remove bare `2>/dev/null` from every `gsd-sdk query agent-skills …` invocation in workflows so SDK failures surface to stderr. - Apply the same rule to other no-fallback calls (audit-open, write-profile, generate-* profile handlers, frontmatter.get in commands). Best-effort cleanup calls (config-set workflow._auto_chain_active false) keep exit-code forgiveness via `|| true` but no longer suppress stderr. Parity tests: - New: tests/bug-2636-gsd-sdk-query-silent-swallow.test.cjs — fails if any `gsd-sdk query agent-skills … 2>/dev/null` is reintroduced. - Existing: tests/gsd-sdk-query-registry-integration.test.cjs already asserts every workflow noun resolves to a registered handler; confirmed passing post-change. Note: npm republish of @gsd-build/sdk is a separate release concern and is not included in this PR. * fix(#2636): address review — restore broken markdown fences and shell syntax The previous commit's mass removal of '2>/dev/null' suffixes also collapsed adjacent closing code fences and 'fi' tokens onto the command line, producing malformed markdown blocks and 'truefi' / 'true fi' shell syntax errors in the workflows. Repaired sites: - commands/gsd/quick.md, thread.md (frontmatter.get fences) - workflows/complete-milestone.md (audit-open fence) - workflows/profile-user.md (write-profile + generate-* fences) - workflows/verify-work.md (audit-open --json fence) - workflows/execute-phase.md (truefi -> true / fi) - workflows/plan-phase.md, discuss-phase-assumptions.md, discuss-phase/modes/chain.md (true fi -> true / fi) All 5450 tests pass. |
||
|
|
a6e692f789 |
fix(#2646): honor ROADMAP [x] checkboxes when no phases/ directory exists (#2669)
initProgress (and its CJS twin) hardcoded `not_started` for ROADMAP-only phases, so `completed_count` stayed at 0 even when the ROADMAP showed `- [x] Phase N`. Extract ROADMAP checkbox states into a shared helper and use `- [x]` as the completion signal when no phase directory is present. Disk status continues to win when both exist. Adds a regression test that reproduces the bug with no phases/ dir and one `[x]` / one `[ ]` phase, asserting completed_count===1. Fixes #2646 |
||
|
|
06463860e4 |
fix(#2638): write sub_repos to canonical planning.sub_repos (#2668)
loadConfig's multiRepo migration and filesystem-sync writers targeted the top-level parsed.sub_repos, but KNOWN_TOP_LEVEL (the unknown-key validator's allowlist) only recognizes planning.sub_repos (canonical per #2561). Each migration/sync therefore persisted a key the next loadConfig call warned was unknown. Redirect both writers to parsed.planning.sub_repos, ensuring parsed.planning is initialized first. Also self-heal legacy/buggy installs by stripping any stale top-level sub_repos on load, preserving its value as the planning.sub_repos seed if that slot is empty. Tests cover: (a) canonical planning.sub_repos emits no warning, (b) multiRepo migration writes to planning.sub_repos with no top-level residue, (c) filesystem sync relocates to planning.sub_repos, (d) stale top-level sub_repos from older buggy installs is stripped on load. Closes #2638 |
||
|
|
387c8a1f9c |
fix(#2653): eliminate SDK↔CJS config-schema drift (#2670)
The SDK's config-set kept its own hand-maintained allowlist (28-key drift vs. get-shit-done/bin/lib/config-schema.cjs), so documented keys accepted by the CJS config-set — planning.sub_repos, workflow.code_review_command, workflow.security_*, review.models.*, model_profile_overrides.*, etc. — were rejected with "Unknown config key" when routed through the SDK. Changes: - New sdk/src/query/config-schema.ts mirrors the CJS schema exactly (exact-match keys + dynamic regex sources). - config-mutation.ts imports VALID_CONFIG_KEYS / DYNAMIC_KEY_PATTERNS from the shared module instead of rolling its own set and regex branches. - Drop hand-coded agent_skills.* / features.* regex branches — now schema-driven so claude_md_assembly.blocks.*, review.models.*, and model_profile_overrides.<runtime>.<tier> are also accepted. - Add tests/config-schema-sdk-parity.test.cjs (node:test) as the CI drift guard: asserts CJS VALID_CONFIG_KEYS set-equals the literal set parsed from config-schema.ts, and that every CJS dynamic pattern source has an identical counterpart in the SDK. Parallel to the CJS↔docs parity added in #2479. - Vitest #2653 specs iterate every CJS key through the SDK validator, spot-check each dynamic pattern, and lock in planning.sub_repos. - While here: add workflow.context_coverage_gate to the CJS schema (already in docs and SDK; CJS previously rejected it) and sync the missing curated typo-suggestions (review.model, sub_repos, plan_checker, workflow.review_command) into the SDK. Fixes #2653. |
||
|
|
e973ff4cb6 |
fix(#2630): reset STATE.md frontmatter atomically on milestone switch (#2666)
The /gsd:new-milestone workflow Step 5 rewrote STATE.md's Current Position body but never touched the YAML frontmatter, so every downstream reader (state.json, getMilestoneInfo, progress bars) kept reporting the stale milestone until the first phase advance forced a resync. Asymmetric with milestone.complete, which uses readModifyWriteStateMdFull. Add a new `state milestone-switch` handler (both SDK and CJS) that atomically: - Stomps frontmatter milestone/milestone_name with caller-supplied values - Resets status to 'planning' and progress counters to zero - Rewrites the ## Current Position section to the new-milestone template - Preserves Accumulated Context (decisions, blockers, todos) Wire the workflow Step 5 to invoke `state.milestone-switch` instead of the manual body rewrite. Note the flag is `--milestone` not `--version`: gsd-tools reserves `--version` as a globally-invalid help flag. Red vitest in sdk/src/query/state-mutation.test.ts asserts the frontmatter reset. Regression guard via node:test in tests/bug-2630-*.test.cjs runs through gsd-tools end-to-end. Fixes #2630 Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
a72bebb379 |
fix(workflows): agent-skills query keys must match subagent_type (follow-up to #2555) (#2616)
* fix(workflows): agent-skills query keys must match subagent_type Eight workflow files called `gsd-sdk query agent-skills <KEY>` with a key that did not match any `subagent_type` Task() spawns in the same workflow (or any existing `agents/<KEY>.md`): - research-phase.md:45 — gsd-researcher → gsd-phase-researcher - plan-phase.md:36 — gsd-researcher → gsd-phase-researcher - plan-phase.md:38 — gsd-checker → gsd-plan-checker - quick.md:145 — gsd-checker → gsd-plan-checker - verify-work.md:36 — gsd-checker → gsd-plan-checker - new-milestone.md:207 — gsd-synthesizer → gsd-research-synthesizer - new-project.md:63 — gsd-synthesizer → gsd-research-synthesizer - ui-review.md:21 — gsd-ui-reviewer → gsd-ui-auditor - discuss-phase.md:114 — gsd-advisor → gsd-advisor-researcher Effect before this fix: users configuring `agent_skills.<correct-type>` in .planning/config.json got no injection on these paths because the workflow asked the SDK for a different (non-existent) key. The SDK correctly returned "" for the unknown key, which then interpolated as an empty string into the Task() prompt. Silent no-op. The discuss-phase advisor case is a subtle variant — the spawn site uses `subagent_type="general-purpose"` and loads the agent role via `Read(~/.claude/agents/gsd-advisor-researcher.md)`. The injection key must follow the agent identity (gsd-advisor-researcher), not the technical spawn type. This is a follow-up to #2555 — the SDK-side fix in that PR (#2587) only becomes fully effective once the call sites use the right keys. Adds `sdk/src/workflow-agent-skills-consistency.test.ts` as a contract test: every `agent-skills <slug>` invocation in `get-shit-done/workflows/**/*.md` must reference an existing `agents/<slug>.md`. Fails loudly on future key typos. Closes #2615 * test: harden workflow agent-skills regex per review feedback Review (#2616): CodeRabbit flagged the `agent-skills <slug>` pattern as too permissive (can match prose mentions of the string) and the per-line scan as brittle (misses commands wrapped across lines). - Require full `gsd-sdk query agent-skills` prefix before capture + `\b` around the pattern so prose references no longer match. - Scan each file's full content (not line-by-line) so `\s+` can span newlines; resolve 1-based line number from match index. - Add JSDoc on helpers and on QUERY_KEY_PATTERN. Verified: RED against base (`f30da83`) produces the same 9 violations as before; GREEN on fixed tree. --------- Co-authored-by: forfrossen <forfrossensvart@gmail.com> |
||
|
|
df0ab0c0c9 |
fix(#2410): emit wave + plan checkpoint heartbeats to prevent stream idle timeout (#2626)
/gsd:manager's background execute-phase Task fails with
"Stream idle timeout - partial response received" on multi-plan phases
(Claude Code + Opus 4.7 at ~200K+ cache_read) because the long subagent
never emits tokens fast enough between large tool_results — the SSE layer
times out mid-assistant-turn and the harness retries hit the same TTFT
wall after prompt cache TTL expires.
Root cause: no orchestrator-level activity at wave/plan boundaries.
Fix (maintainer-approved A+B):
- A (wave boundary): execute-phase.md now emits a `[checkpoint]`
heartbeat before each wave spawns and after each wave completes.
- B (plan boundary): also emit `[checkpoint]` before each Task()
dispatch and after each executor returns (complete/failed/checkpoint).
Heartbeats are literal assistant-text lines (no tool call) with a
monotonic `{P}/{Q} plans done` counter so partial-transcript recovery
tools can grep progress even when a run dies mid-phase.
Docs: COMMANDS.md /gsd-manager section documents the marker format.
Tests: tests/bug-2410-stream-checkpoint-heartbeats.test.cjs (12 cases)
asserts the heartbeats exist at every boundary and in the right workflow
step. Full suite: 5422 node:test cases pass. Pre-existing vitest
failures on main are unrelated to this change.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|