* test(3309): red — workflow.human_verify_mode contract
New behavioral test file covers:
- workflow.human_verify_mode is a recognized config key (VALID_CONFIG_KEYS)
- defaults to 'mid-flight' (preserves current behavior)
- config-set / config-get round-trips for both values
- persists in config.json as string
- planner agent file references the flag with canonical wording, couples
end-of-phase mode with the rule that checkpoint:human-verify is not
emitted, and documents the <verify><human-check> deferred-item shape
- verifier agent file references harvesting <verify><human-check> blocks
- references/checkpoints.md documents the cost-control alternative
Source-text assertions on agent .md files are exempted via
allow-test-rule: source-text-is-the-product — those files ARE the
runtime contract loaded by AI runtimes, so asserting their wording is
the only way to verify the agents will respect the flag.
Fails 10/11 against current source. Will pass after the fix.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(3309): add workflow.human_verify_mode = end-of-phase opt-out
Each mid-flight checkpoint:human-verify halt costs a full executor
cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on every
respawn) because subagent context is discarded across the pause. A plan
with N human-verify checkpoints pays the cold-start cost N+1 times. The
reporter (rentanything-nb) measured this at "tens of thousands of tokens"
per round-trip and "hundreds of thousands per week."
This adds workflow.human_verify_mode (default 'mid-flight') with an
'end-of-phase' value that:
- instructs gsd-planner to NOT emit <task type="checkpoint:human-verify">
tasks; verification details go into a <verify><human-check> sub-block
on the relevant auto task instead
- instructs gsd-verifier (Step 8) to harvest those <verify><human-check>
blocks at end-of-phase and merge them into its own human-verification
list
- the existing human_needed → HUMAN-UAT.md flow in execute-phase.md is
the single sink — no new file/writer is created
checkpoint:decision and checkpoint:human-action are unaffected — those
gate the work itself, not post-hoc verification.
Surfaces touched:
- bin/lib/config-schema.cjs, bin/lib/config.cjs — register key + default
- sdk/src/config.ts, sdk/src/query/config-schema.ts — SDK parity
- agents/gsd-planner.md — slim Detection section + reference link
- agents/gsd-verifier.md — Step 8 harvest instruction
- get-shit-done/references/planner-human-verify-mode.md — full rules,
loaded conditionally to keep planner.md under its size budget
- get-shit-done/references/checkpoints.md — surface the alternative
- docs/CONFIGURATION.md — config table row
- docs/INVENTORY.md, docs/INVENTORY-MANIFEST.json — track new reference
Tag name <human-check> chosen instead of <human> to avoid the
prompt-injection scan pattern that flags <system|assistant|human> tags.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore(3309): align changeset pr: to actual PR number
The pr: field was authored as 3319 (a guess at the next number) before
the PR was opened. Actual PR is #3325.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(3309): flip workflow.human_verify_mode default to end-of-phase
Per maintainer direction on PR #3325, end-of-phase is the new project
default. Mid-flight checkpoint:human-verify halts cost a full executor
cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) per
round-trip — reported at "tens of thousands of tokens" per round-trip,
"hundreds of thousands per week" on real projects. The cost-control
mode is what new projects should get out of the box.
mid-flight remains a one-line opt-back-in via:
gsd config-set workflow.human_verify_mode mid-flight
Behavior change for existing projects: the new default takes effect
when .planning/config.json is rewritten (config-set, fresh project).
Existing in-flight PLAN.md files with checkpoint:human-verify tasks
continue to work in either mode — the flag only changes what the
planner emits next time it runs.
Surfaces updated:
- bin/lib/config.cjs, sdk/src/config.ts — default flipped
- sdk/src/config.ts docstring — describes new default + opt-back-in
- agents/gsd-planner.md — Detection section explains new default
- references/planner-human-verify-mode.md — reordered modes; added
guidance on when to opt back into mid-flight
- references/checkpoints.md — surface the default flip and the why
- docs/CONFIGURATION.md — table row reflects new default + reason
- tests/feat-3309-human-verify-mode.test.cjs — default test asserts
end-of-phase
- .changeset/fierce-geese-march.md — describes the default flip and
the migration semantics
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: address human verify mode review
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(3317): red — SDK detect-custom-files must scan skills/
Mirrors tests/bug-2942-detect-custom-skills.test.cjs on the SDK side.
The SDK's GSD_MANAGED_DIRS array omits 'skills', so user-added skills
under <config-dir>/skills/<name>/ are never returned and get destroyed
on /gsd-update. New vitest covers:
- detects custom skill at skills/<name>/SKILL.md
- does not flag manifest-tracked skill as custom
- still detects custom files under get-shit-done/workflows/ (regression)
- custom_count matches custom_files.length across multiple skills
Fails 2/4 against current SDK source. Will pass after fix.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(3317): SDK detect-custom-files now scans skills/ (parity with CJS)
The SDK port of detect-custom-files declared GSD_MANAGED_DIRS without
'skills', while the canonical bin/gsd-tools.cjs port (which the SDK
docstring explicitly cites as its source) had the entry. Because
update.md prefers gsd-sdk over the CJS shim, real-world users with the
SDK installed never had their custom skills detected — the installer's
'Installed N skills to skills/' step then wiped any non-manifest skill
without backing it up to gsd-user-files-backup/. Real-world incident:
skills/gsd-roadmap/SKILL.md (fully user-owned) destroyed during the
1.40.0 → 1.41.0 update, recoverable only via Time Machine.
One-line fix adds 'skills' to the SDK's GSD_MANAGED_DIRS, matching the
CJS source the port was supposed to mirror. The 4 vitest cases added in
the prior commit now all pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore: correct changeset pr number
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: reconcile planner action contract
The deep_work_rules workflow block was stronger than the planner agent contract: it required verbatim context copies and self-sufficient action text, so planners were incentivized to inline implementation code. Bound action content to directive prose with concrete identifiers, allow behavior/test acceptance criteria, and pin the cross-file contract with a regression test. Closes#3320.
* chore: add changeset for planner contract fix
* refactor: tighten sdk-first architecture seams
Refs #3312
* refactor: finish state document seam cleanup
Refs #3312
* test: harden minimal install cleanup assertion
* ci: support sdk-scoped package lock
* fix(3316): restore root package-lock.json and align changeset pr ref
Reverts dec57a83 ("ci: support sdk-scoped package lock") and restores
the root package-lock.json that c249d34d deleted. The deletion was the
wrong direction:
- The root package.json declares its own runtime and dev deps
(@anthropic-ai/claude-agent-sdk, ws, c8). Without a root lockfile,
`npm install --no-package-lock` resolves whatever satisfies semver at
install time — CI today and CI in six months can install different
transitive trees, defeating reproducibility.
- The lockfile has been part of every release on this repo (long
history on main); removing it loses the npm audit / Dependabot
target without compensating benefit.
- The CI workaround pattern (cache-dependency-path: sdk/package-lock.json
+ `npm install --no-package-lock`) papered over the symptom rather
than fix the cause.
Also fix the changeset pr: from 3312 (issue) to 3316 (PR). CONTEXT.md
flags this exact failure mode as a recurring CodeRabbit finding.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: address coderabbit review findings
* fix: close remaining coderabbit threads
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes#3310
Wires `ERROR_REASON.SDK_UNKNOWN_COMMAND` and `ERROR_REASON.USAGE` into the remaining untyped error paths in `gsd-tools.cjs` (template, frontmatter, requirements, milestone, uat, todo, workstream, graphify, learnings subcommand routers and `--cwd`/missing-required-arg paths). Tests assert via `JSON.parse(stderr).reason` rather than substring matching on prose. Closure regression guard locks the canonical `{ok, reason, message}` shape for every newly-typed path.
Also moves `--json-errors` activation up to the top of `main()` so `--cwd` validation paths emit JSON rather than plain text when `GSD_JSON_ERRORS=1` is set.
* chore: revert redundant CHANGELOG.md row from #3308
PR #3308 added a Shared scanPhasePlans helper row to CHANGELOG.md while
also dropping the canonical .changeset/3262-extract-scan-phase-plans.md
fragment. CONTRIBUTING.md (line 110) prohibits hand-editing CHANGELOG.md
since the release workflow folds .changeset/*.md fragments into the
file at release time — the manual row would duplicate at next release.
Removes only the 4-line Enhancement block added by #3308. The fragment
remains unchanged and is the single source of truth for this entry.
Refs #3262
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore(changeset): add Removed fragment for redundant CHANGELOG row cleanup
Documents the deletion in PR #3313 of the 4-line ### Enhancement block
that #3308 hand-wrote into CHANGELOG.md alongside its canonical
.changeset/3262-extract-scan-phase-plans.md fragment.
Refs #3262
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(phase-dir): add red test for k015 prefix-drift in plan-milestone-gaps and import workflows (#3298)
Asserts that:
- plan-milestone-gaps.md step 8 does not use bare {NN}-{name} mkdir pattern
- plan-milestone-gaps.md step 8 uses phase.add or expected_phase_dir
- import.md plan_convert does not use bare {NN}-{slug} mkdir pattern
- import.md plan_convert uses expected_phase_dir from init.phase-op
These 5 tests are RED until the fix lands.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(phase-dir): add projectCode prefix to phase-dir construction in plan-milestone-gaps and import workflows (#3298)
Both plan-milestone-gaps.md step 8 and import.md plan_convert step were
constructing phase directories using raw {NN}-{name}/{NN}-{slug} template
patterns, bypassing the project_code prefix from .planning/config.json.
Fix: both steps now call `gsd-sdk query init.phase-op <N>` and consume the
`expected_phase_dir` field (which includes the `<CODE>-<NN>-<slug>` prefix
when project_code is set), matching the pattern already established by
PR #3292 for /gsd-discuss-phase and /gsd-plan-phase.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(phase-dir): apply project_code prefix to backlog phase dir in add-backlog workflow (k015 sibling, #3298)
Sibling k015 audit found a third drift site: add-backlog.md step 4 was
constructing the 999.x backlog phase directory using raw ${NEXT}-${SLUG}
without applying the project_code prefix from .planning/config.json.
Fix: read project_code via `gsd-sdk query config-get project_code --raw`
and prepend `${CODE}-` when set, matching the pattern used by phase.insert
(which already applies project_code to decimal phases in phase.cjs line 736).
Also extends bug-3298 regression test to cover this third site.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(changeset): add changeset for #3298 phase-dir prefix drift fix
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(changeset): set pr: 3306 in changeset for #3298; resolve stash conflict in live-command-registry.cjs
The conflict was cosmetic (string concat → template literals) introduced by
accidental git stash during test verification. Taking the newer template-literal
form throughout.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(3251): red — assert 14 missing commands in command-aliases.generated.cjs
Parametrized test covering all 14 commands from issue #3251. Requires the
CJS manifest and asserts structurally (never greps source). Currently fails
because NON_FAMILY_COMMAND_ALIASES is not exported and all 14 are missing.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(3251): add 14 missing commands to command-aliases.generated.cjs
Adds NON_FAMILY_COMMAND_ALIASES export to command-aliases.generated.cjs and
extends sdk/src/query/command-manifest.non-family.ts with the 10 commands that
were registered in static catalogs but absent from the manifest source-of-truth:
- check.decision-coverage-plan / check.decision-coverage-verify
- frontmatter.get
- phase.mvp-mode
- progress.bar
- stats.json
- task.is-behavior-adding
- todo.match-phase
- uat.render-checkpoint
- workstream.list
The other 4 (frontmatter.set, learnings.copy, milestone.complete,
requirements.mark-complete) were already in non-family.ts but unexported.
Generator (sdk/scripts/gen-command-aliases.ts) now produces both the TS and CJS
artifacts including the non-family section, sorted by canonical for determinism.
Freshness check (sdk/scripts/check-command-aliases-fresh.mjs) verifies TS and CJS
non-family parity against the manifest source-of-truth.
Fixes#3251
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(3251): add changeset and CHANGELOG entry for PR #3305
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(gen-command-aliases): preserve TS type interfaces and annotations on regen (#3251)
- Add FamilyCommandAlias and NonFamilyCommandAlias interface declarations to tsBody so regen never strips them
- Replace JSON.stringify with single-line compact serialisers for TS output (matches committed file format)
- Apply typed annotations (readonly FamilyCommandAlias[] / readonly NonFamilyCommandAlias[]) to all exported TS constants
- CJS path unchanged: remains untyped pure-JS with JSON.stringify multi-line format
- Regenerate sdk/src/query/command-aliases.generated.ts to sync with updated generator
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: drop redundant CHANGELOG.md edit (use .changeset/ fragment per CONTRIBUTING.md)
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(3255): add red/green tests for --json-errors structured error mode
Ten tests covering the --json-errors mode contract:
- Unknown command → sdk_unknown_command
- Dotted unknown command → sdk_unknown_command
- Missing --pick value → usage
- Config key not found → config_key_not_found
- Unknown subcommand → sdk_unknown_command
- GSD_JSON_ERRORS=1 env var activation
- Successful command unaffected
- Stable error shape ({ok, reason, message})
- Single error line per invocation
- Unknown flag → usage
All assertions use JSON.parse on stderr captures, never .includes() on
text (#2974 / CONTRIBUTING.md "Prohibited: Raw Text Matching" rule).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(3255): add typed ERROR_REASON codes and GSD_JSON_ERRORS env var support
- Destructure ERROR_REASON from core in gsd-tools.cjs
- Add GSD_JSON_ERRORS=1 env var as alternative to --json-errors CLI flag
- Pass ERROR_REASON.SDK_UNKNOWN_COMMAND to unknown top-level command default path
- Pass ERROR_REASON.SDK_UNKNOWN_COMMAND to unknown intel subcommand path
- Pass ERROR_REASON.USAGE to --pick missing value error path
- Pass ERROR_REASON.USAGE to --version flag rejection path
All ten tests in feat-3255-json-errors-mode.test.cjs pass.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(3255): add json-errors taxonomy doc, changeset, and CHANGELOG entry
- docs/json-errors.md: full error code taxonomy, wire format spec, and
test-authoring guidelines for the --json-errors mode
- .changeset/gentle-tigers-roar.md: changeset fragment (pr will be updated
after PR is opened)
- CHANGELOG.md: Unreleased → Added entry for the new structured error mode
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: update changeset PR number to 3304
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(gsd-tools): document --json-errors in usage/help text (#3255)
Add [--json-errors] to the TOP_LEVEL_USAGE synopsis line and introduce a
"Global flags:" section describing all four global flags (--raw, --pick,
--cwd, --ws) plus --json-errors with its GSD_JSON_ERRORS=1 env-var
alternative, so operators can discover the flag via `gsd-tools --help`.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: drop redundant CHANGELOG.md edit (use .changeset/ fragment per CONTRIBUTING.md)
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(adr): add docs/adr/README.md index and structural ADR test (#3271)
- Add docs/adr/README.md as an indexed entry point linking all 7 ADRs
- Add tests/enh-3271-sdk-adr-structure.test.cjs: structural assertions that
ADR 0005 and 0006 exist, have required headings and Status/Date metadata,
and that README links every ADR file by filename
- Update CHANGELOG.md with Enhancement entry
- Add .changeset/3271-sdk-adr-structure.md
ADRs 0005 (SDK architecture seam-map) and 0006 (planning-path projection
module) already landed on main. This PR completes issue #3271 by adding the
README index and the structural test gate that enforces ADR completeness
going forward.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: set changeset pr: 3302
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(test): exclude self-reference from ADR 0005 cross-ref count (#3271)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: drop redundant CHANGELOG.md edit (use .changeset/ fragment per CONTRIBUTING.md)
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(phase-plans): red — shared scanPhasePlans contract + parity across call sites (#3262)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(phase-plans): extract shared scanPhasePlans helper (k014) (#3262)
Eliminates four divergent copies of the plan-scan algorithm:
- roadmap.cjs:countPhasePlansAndSummaries (root call site)
- state.cjs:buildStateFrontmatter (1 of 3)
- state.cjs:cmdStateValidate (2 of 3)
- state.cjs:cmdStateSync (3 of 3)
- init.cjs:listPhasePlanFiles / listPhaseSummaryFiles
New bin/lib/plan-scan.cjs exports scanPhasePlans(phaseDir) → {
planCount, summaryCount, completed, hasNestedPlans,
planFiles, summaryFiles
}
Divergences resolved:
- roadmap.cjs used a broad isPlanFile (any .md containing PLAN in name,
matching the extended layout 5-PLAN-01-setup.md); canonical helper
adopts this wider pattern as the reference implementation.
- state.cjs used a strict endsWith(-PLAN.md) filter, missing extended-
layout root files; now unified with roadmap.cjs semantics.
- init.cjs listPhasePlanFiles used ^PLAN-\d+ for nested, missing
the -PLAN-\d+ variant state.cjs also matched; helper includes both.
- pre-bounce exclusion broadened to /.pre-bounce.md$/i (any pre-bounce
file), not just -PLAN.*\.pre-bounce\.md (roadmap form) or flat
.pre-bounce.md (state form).
- OUTLINE exclusion broadened to /-OUTLINE\.md$/i to catch both
flat (-PLAN-OUTLINE.md) and nested (PLAN-01-OUTLINE.md) forms.
Sibling audit: no 5th call site found. phase.cjs:looksLikePlanFile is
a diagnostic probe for non-canonical naming (not a counter) — left
in place per its distinct purpose.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(changelog): add entry for #3262 scanPhasePlans extraction
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(changeset): add changeset fragment for #3262
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(inventory): add plan-scan.cjs row to INVENTORY.md CLI Modules table (#3262)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(3262): update bug-3128 test + INVENTORY counts for plan-scan.cjs
- Update tests/bug-3128-roadmap-plan-count-slug-layout.test.cjs to verify
that roadmap.cjs delegates to plan-scan.cjs (require check) and that the
extended filter lives in plan-scan.cjs as isRootPlanFile with /PLAN/i
- Bump docs/INVENTORY.md CLI Modules headline from 46 to 47 (plan-scan.cjs)
- Regenerate docs/INVENTORY-MANIFEST.json to include cli_modules/plan-scan.cjs
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(3262): migrate two missed call sites to scanPhasePlans (k014)
- init.cjs cmdInitExecutePhase: replace inline /-PLAN\.md$/i filter
with listPhasePlanFiles(path) to honour nested, extended-layout, OUTLINE
and pre-bounce exclusions (CR finding)
- state.cjs cmdStateUpdateProgress: replace dual /-PLAN\.md$/i and
/-SUMMARY\.md$/i filters with scanPhasePlans() so the progress-bar
body field uses the same counts as buildStateFrontmatter frontmatter
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(3262): correct INVENTORY-MANIFEST.json to tracked files only
Remove 3 untracked local entries from cli_modules so the manifest matches
what CI sees (47 tracked .cjs files, not 50 local). Previous regeneration
ran against the local filesystem which included cjs-command-router-adapter.cjs,
state-document.cjs, and workstream-inventory.cjs (all untracked on this branch).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(3232): add contributor-standards.md (CONTEXT.md + ADR + AI-agent pillars)
Codifies contributor expectations around the three pillars called out in
issue #3232: CONTEXT.md format and governance, ADR naming/status/amendment
conventions, and AI-agent-assisted work requirements (worktree isolation,
TDD discipline, adversarial review, CR-loop). Updates CONTRIBUTING.md to
link the new doc and adds an AI-agent bullet to the architecture-standards
summary. Structural test asserts all three pillars and the CONTRIBUTING.md
cross-link exist.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: add changeset for PR #3301 (contributor-standards)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: fix worktree example to use generic branch name placeholder
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(docs): add lang tags to fenced code blocks (MD040) (#3232)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: gsd-intel-updater layout-detection block must be gated or removed (#3290 RED)
Group A asserts the bare `ls -d .kilo ... || echo unknown` detection invocation
is absent or wrapped in a framework-repo gate (fails RED: currently unconditional).
Group B confirms zero downstream consumers of the verdict (passes GREEN: none exist).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(intel): gate layout-detection block on framework-repo check (#3290)
The "Runtime layout detection" bash block in gsd-intel-updater ran
unconditionally on every project analysed, emitting a noisy:
Layout detection returned "unknown" — this project is not a GSD-system
installation (no `.claude/get-shit-done/` or `.kilo/` runtime root).
for every ordinary (non-GSD-framework) user project. Group B audit confirmed
zero downstream consumers of the verdict outside the file itself.
Fix (option A): wrap the detection bash block in a positive framework-repo
gate — `jq -r '.name' package.json == "get-shit-done-cc"` — so it runs
only when analysing the GSD framework's own repo. The layout table (.kilo/*
paths) is retained for kilo-layout coverage (required by bug #2351 regression
test).
Dead-code vintage: byte-identical from v1.21.0 through v1.41.1 per reporter.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* changeset: pr=3299 for #3290
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: codex hooks.state.<key> tables must validate as regular tables (#3285 RED)
Drive validateCodexConfigSchema with a fixture containing both [hooks.state]
and [[hooks.SessionStart]] entries. Expect the state tables to pass as regular
tables. Currently fails — validator over-classifies every hooks.* path as AoT.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(install): treat codex hooks.state as regular table not AoT (#3285)
validateCodexConfigSchema was over-classifying every hooks.* section header
as an event-handler array-of-tables, rejecting the hooks.state.* namespace
that Codex CLI 0.130.0+ uses for per-hook trust persistence.
Fix:
1. Section-header check: carve out `hooks.state` and `hooks.state.*` from
the AoT-required rule — only paths that are neither of those still
require double-bracket form.
2. Parsed-object check: skip the `state` key when iterating Object.entries
(parsed.hooks) so the "must be array" guard does not fire for the trust
namespace object.
All other hooks.<EVENT> validation (SessionStart AoT, handler-field placement)
is unchanged.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* changeset: pr=3289 for #3285
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(install): codex hooks.state must be regular-table, reject AoT/scalar (CR finding 1)
- migrateCodexHooksMapFormat: exclude hooks.state and hooks.state.* from
legacy-map detection so [hooks.state] is never promoted to [[hooks.state]] AoT
- validateCodexConfigSchema: reject [[hooks.state]] / [[hooks.state.*]] AoT at
section level; reject Array/scalar values at parsed-object level
- Accept only plain-object shape for hooks.state and hooks.state.* (Codex
CLI 0.130.0+ trust-persistence namespace)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: assert codex hooks.state trust entry preserved with original values (CR finding 2)
Strengthen post-install preservation assertion to verify the actual trust
entry key and its enabled/trusted_hash values survive — not just that
hooks.state is an object. Add two validator-reject tests for [[hooks.state]]
and [[hooks.state.foo]] AoT forms.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: state record-metric/add-decision: auto-create + workstream routing (#3286 RED)
Three failing test groups:
- Bug A: silent no-op contract (exit 0 with recorded/added:false)
- Bug B: auto-create ## Performance Metrics / ## Decisions when absent
- Bug C: --ws routing writes to workstream STATE.md, not root
* fix(state): workstream-route + auto-create sections in record-metric/add-decision (#3286)
Bug B: cmdStateRecordMetric now auto-creates ## Performance Metrics (with
canonical table header) when absent, instead of silently returning
{ recorded: false }. cmdStateAddDecision does the same for ## Decisions.
Both functions return created: true in the JSON when the section was
newly scaffolded — matching the DWIM behavior of state begin-phase and
advance-plan.
Bug A: silent no-op disappears — sections are auto-created, so recorded/added
is always true on valid input.
Bug C: workstream routing via planningPaths(cwd) already reads GSD_WORKSTREAM
which gsd-tools.cjs sets from --ws before calling state handlers. Confirmed
by the passing --ws routing tests.
Fixes#3286
* refactor(state): extend DWIM auto-create to cmdStateAddBlocker (k302/k014)
Applies the same section auto-create pattern from record-metric/add-decision
to add-blocker: when ## Blockers / ### Blockers is absent, auto-scaffold it
instead of silently returning added:false. Returns created:true in the JSON
when newly scaffolded, matching the uniform shape across all three write verbs.
Per k014 (duplicate algorithm drift), the pattern is now applied uniformly
across all cmdState* functions that mutate a named section.
* changeset: pr=3291 for #3286
* fix(state): remove dead else-branches flagged by CodeRabbit (#3291)
After the auto-create fallback, recorded/added is always true — the else
blocks emitting { recorded: false } / { added: false } were unreachable.
Removed all three dead branches (cmdStateRecordMetric, cmdStateAddDecision,
cmdStateAddBlocker) and replaced with an explanatory comment per CR nitpick.
* test: reproduce model-catalog MODULE_NOT_FOUND in install layout (#3288)
Tests A/B/C exercise the install-layout regression introduced by #3230:
- A: confirms the old 3-level __dirname path fails when sdk/shared/ is absent
- B: confirms the new co-located bin/shared/ path resolves correctly (RED)
- C: confirms install() copies model-catalog.json to co-located path (RED)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(install): copy sdk/shared/model-catalog.json + resolve chain in CJS (#3288)
Two-part fix for the CRITICAL regression introduced by #3230:
Install-side: bin/install.js now copies sdk/shared/model-catalog.json into
get-shit-done/bin/shared/model-catalog.json immediately after the main
get-shit-done/ copy step. Every runtime install (Claude Code, Codex, OpenCode,
Gemini, etc.) now includes this file in the payload.
CJS-side: model-catalog.cjs replaces the brittle single-path require with a
resolve-chain that checks candidates in order:
1. get-shit-done/bin/shared/model-catalog.json (co-located, preferred post-install)
2. sdk/shared/model-catalog.json (source-repo dev path, legacy fallback)
3. GSD_MODEL_CATALOG env override (custom deployments / test harnesses)
When no candidate resolves, throws with a diagnostic listing all tried paths
(PRED.k301 — throw must include candidate paths for debuggability).
REFACTOR audit: three other __dirname traversals in bin/lib/ were inspected:
- core.cjs:1242 (3 levels up → agents/) — safe; agents/ IS copied to targetDir
- profile-output.cjs:547,740 (2 levels up → templates/) — safe; templates/ is
inside get-shit-done/ and IS copied
Only model-catalog.cjs traversed outside the installed payload.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* changeset: pr=3293 for #3288
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(model-catalog): narrow catch to missing-file errors; clear env in test (#3288)
Two CR findings from PR #3293 review:
1. model-catalog.cjs catch block swallowed ALL errors — malformed JSON,
permission errors, and other real failures were silently absorbed into
the fallback chain. Now only MODULE_NOT_FOUND (with matching path in
message) and ENOENT are treated as recoverable; any other error is
rethrown immediately.
2. test beforeEach saved GSD_EXPLICIT_CONFIG_DIR but didn't clear it —
a CI-set value could leak into install() and redirect the install to
an unexpected directory, making test C non-deterministic. Added
`delete process.env.GSD_EXPLICIT_CONFIG_DIR` to beforeEach.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: phase-dir prefix parity across creation paths (#3287 RED)
Add failing tests asserting that init.phase-op and init.plan-phase
expose expected_phase_dir with the project_code prefix when the phase
directory does not yet exist — matching the prefix applied by phase.add.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(phase): unify phase-dir naming via shared getPhaseDirName helper (#3287)
Both init.phase-op (discuss-phase workflow) and init.plan-phase
(plan-phase workflow) now compute expected_phase_dir — the canonical
directory name including the project_code prefix when set — and expose
it in their JSON bundle.
Workflow fallback mkdir calls are updated to use ${expected_phase_dir}
instead of constructing the path from padded_phase + phase_slug, which
was missing the project_code prefix.
This eliminates the two-headed naming convention where phase.add/insert
produced XR-01-foundation/ while the first-touch paths produced
01-foundation/ for the same project.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(phase): apply project_code prefix in scaffold phase-dir (#3287)
Audit finding: phase.scaffold (CJS commands.cjs + SDK phase-lifecycle.ts)
also constructed phase dir names without project_code prefix.
Both implementations now read config.project_code and apply the same
prefix logic as phase.add/phase.insert.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* changeset: pr=3292 for #3287
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(changelog): add entry for #3287 phase-dir prefix parity fix
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: reproduce Windows SDK not found after fresh npx install (#3211)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(install): Windows persistent Path probe + npx-PATH filter on Windows (#3211)
Add getUserShellWindowsPersistentPath() — the Windows counterpart to
getUserShellPath(). Probes the user-level 'Path' registry key via
powershell.exe so the installer can verify gsd-sdk is reachable from
PowerShell/cmd.exe/Git Bash post-install, not just in the transient
npx subprocess PATH.
Wire it into installSdkIfNeeded: on Windows, use the registry-derived
persistent Path (with npx dirs stripped) as the cross-shell reachability
gate, instead of skipping the check entirely. This is the Windows sibling
of the Linux fix in #3249/#3231.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* changeset: pr=3282 for #3211
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(install): include Machine+User Path in Windows persistent probe (#3211)
getUserShellWindowsPersistentPath now merges Machine-level and User-level
registry Path entries (matching the effective PATH that PowerShell, cmd.exe,
and Git Bash inherit), instead of reading only User-level. Reading User-only
would produce a false warning when gsd-sdk is installed in a machine-level
bin dir (e.g. C:\Program Files\nodejs).
Addresses CodeRabbit finding on PR #3282.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: red — bounded git subprocess + structured worktree warnings (#3281)
Regression tests for #3281: worktree-related git subprocess calls have no
timeout bound, and timeout/error outcomes are not surfaced as structured signals.
Failing assertions:
- planWorktreePrune / listLinkedWorktreePaths / snapshotWorktreeInventory must
return reason=git_timed_out (not generic git_list_failed) when execGit returns
timedOut:true — enables callers to distinguish timeout from auth failure
- executeWorktreePrunePlan must include timedOut:true in result when the git
prune call itself times out
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(worktree): bounded git subprocess + structured warning surfacing (#3281)
Root cause (PRED.k014): execGit / execGitDefault called spawnSync with no
timeout, so `git worktree list --porcelain` against a hung/locked repo
blocked the parent process indefinitely. Downstream callers in core.cjs
and verify.cjs then swallowed any resulting failure silently via
catch { /* intentionally empty */ } (PRED.k302).
Fix:
- worktree-safety.cjs: execGitDefault now passes timeout:10000 to spawnSync.
Detects SIGTERM+ETIMEDOUT and returns { timedOut:true } in the result shape.
readWorktreeList maps timedOut:true -> reason:'git_timed_out' (distinct from
generic git_list_failed) so callers can emit a structured warning.
executeWorktreePrunePlan propagates timedOut:true as a first-class result field.
- core.cjs: execGit receives the same timeout+timedOut treatment (PRED.k014
uniform-fix discipline). pruneOrphanedWorktrees now emits a [gsd-tools]
WARNING to stderr when the git prune call times out instead of silent-catch.
- verify.cjs: Check 11 branches on worktreeHealth.ok to surface W018 warning
when the worktree list times out, instead of silent-catch on ok:false.
Backward-compatible: exitCode/stdout/stderr continue to work for all existing
callers; timedOut and error are additive new fields.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* changeset: pr=3283 for #3281
* fix(verify): rename W020 for worktree-timeout warning to avoid W018 collision
W018 is already used for milestone archive drift (Check 12). The new
worktree-health-degraded timeout warning was assigned W018, causing
warning-code ambiguity in triage. Rename to W020 (next available code).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(3266): preserve wave 0 and bucket plans by depends_on DAG in phase-plan-index
Fixes two cooperating bugs in the phase-plan-index builder:
1. Wave 0 collapse: `parseInt(...) || 1` coerced parsed value `0` to `1` due to
JS falsy default. Fixed with `Number.isNaN` guard.
2. depends_on ignored: wave-bucketing used only the `wave:` frontmatter field.
Now replaced with Kahn's topological-level algorithm over `depends_on`:
source nodes (no in-phase deps) → lowest level; each plan's level = max(deps'
levels) + 1. Declared `wave:` that disagrees with computed level emits a
non-fatal warning on the result. Cycle detection throws GSDError.
`PlanInfo` gains `depends_on: string[]`. `PhasePlanIndex` gains `warnings?: string[]`.
Both TS (`sdk/src/query/phase.ts`) and CJS twin (`get-shit-done/bin/lib/phase.cjs`)
fixed identically.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: add changeset for #3276
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(phase): resolve depends_on against canonical plan id (#3276 CR)
Build a secondary `canonicalToId` index alongside `planMap` so that a
dependency declared as '03-01' resolves to a descriptive plan stored
under '03-01-auth-hardening', preventing silent wave-ordering failures.
Applied at both DAG construction sites in phase.cjs and the SDK's
phase.ts (k014 parity). Regression test added.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(workstream): normalize migrate-name to valid slug
* docs(context): record workstream migrate-name slug invariant
* fix(catalog-cjs): balanced fallback for unknown profile (CR finding A)
profiles[profile] could return undefined for any profile key absent from
the catalog entry, causing downstream callers like formatAgentToModelMapAsTable
to crash on .length. Add ?? profiles.balanced fallback to match the SDK adapter.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(sdk): anchor path resolution on import.meta.url not cwd (CR finding B)
resolve(process.cwd(), '..') breaks when Vitest is invoked from the repo root
because cwd is already the repo root and '..' goes one level above. Replace
with a file-relative path using fileURLToPath(new URL('../../../', import.meta.url))
anchored at the test file's location (sdk/src/query/).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: derive Group B runtime list from catalog (CR finding C)
Hardcoded ['kilo', 'cline', ...] throws TypeError if a runtime name is
removed from the catalog. Derive group B dynamically via
Object.keys(catalog.runtimeTierDefaults).filter(r => !r.opus) so the
test never goes stale and auto-covers future Group B additions.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(workflow): add hermes to Step B runtime options (CR finding D)
hermes appears in the Group A built-in defaults table but was missing from
the AskUserQuestion options in Step B, forcing users to manually type it via
'Other (Group B or custom)'. Add explicit hermes entry for UI consistency.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(config): refresh dynamic_routing tier table; fix stale L671 (findings E+F)
Finding E: tier table was missing 6 heavy-tier agents and 15 standard/light
agents added by this PR. Updated all three rows to match catalog routingTier
assignments (33 agents total).
Finding F: removed stale '18 of 31' claim and agent enumeration; replaced
with accurate note that all 33 agents have explicit catalog entries. Updated
authoritative source pointers to model-catalog.cjs / model-catalog.ts.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(core): add profile-fallback unit tests for quality and budget (CR nitpick G)
The PR introduced quality→opus and budget→haiku unknown-agent fallbacks but
only balanced→sonnet and inherit→inherit were tested. Add two tests covering
the remaining two branches to complete coverage.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* adr: define planning workspace and worktree seam
* refactor(worktree): extract worktree safety policy module
* refactor(workstream): extract active workstream pointer store seam
* test(worktree): cover policy branch paths and persist seam guardrails
* refactor(worktree): centralize health inventory seam for W017
* fix(workspace): align SDK project path policy with CJS planningDir
* refactor(query): unify SDK planning path projection seam
* refactor(init): route workspace projection through planningPaths seam
* docs(adr): add SDK architecture and planning path ADRs
* refactor(worktree): deepen name, pointer, inventory, and config seams
* docs(config): harmonize claude-opus-4-6 to 4-7 in resolve_model_ids example (CR finding 2)
* fix(sdk): return undefined for model_profile='inherit' sentinel (CR finding 3)
* docs(adr): renumber conflicting 0003-sdk-package-seam-module to 0007, update seam-map reference (CR finding 4)
* fix(workstream): align CJS and SDK name validation to accept dots, guard path traversal via includes('..') (CR finding 5)
* fix(sdk): guard writeActiveWorkstream against non-existent workstream directory, k014/k031 parity (CR finding 6)
* chore(changeset): add #3269 changeset (CR finding 1 — proper changeset for this PR)
* docs(inventory): register 3 new CLI modules in INVENTORY.md/MANIFEST (active-workstream-store, workstream-name-policy, worktree-safety)
* fix(sdk): use relPlanningPath(workstream) in planningPaths, fix setActiveWorkstream/getActiveWorkstream name errors in workstream.ts
* fix(sdk): validate GSD_WORKSTREAM in planningPaths before use (#3269 regression)
planningPaths() called resolveWorkspaceContext() which returned GSD_WORKSTREAM
raw (no validation). An invalid value like '../evil' was used as effectiveWorkstream,
constructing a bad path; roadmapAnalyze() caught the ENOENT and returned a
no-phase_count error object instead of the root ROADMAP result.
Fix: validate envCtx.workstream with validateWorkstreamName() in planningPaths()
before accepting it as effectiveWorkstream. Invalid env → null → root .planning/
fallback, preserving the bug-2791 contract: invalid GSD_WORKSTREAM is silently
ignored and falls back to the root context (phase_count: 0 for empty root ROADMAP).
The bug-2791 regression test now passes. No other call sites read GSD_WORKSTREAM
without validation: query-runtime-context.ts already validates; cli.ts already
validates; context-engine.ts takes a caller-validated workstream parameter.
Closes#3268 (regression introduced by #3269 workstream-name-policy work).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(3263): harden code-review SUMMARY parser; accept BL-/blocker as Critical-tier across pipeline
Bug 1: compute_file_scope Node script used ^\s*\w+: boundary regex, which excluded
hyphens and left inSection sticky after key-decisions:/patterns-established:/
requirements-completed: blocks. Prose bullets were captured as file paths. Fixed
to [\w-]+ boundary and added em-dash/parenthetical stripping with a path validity
guard so only path-shaped strings are emitted.
Bug 2: present_results grep matched only critical: in frontmatter. When reviewer
emitted blocker:, CRITICAL was silently empty. Fixed grep to accept both keys via
-E "^\s*(critical |blocker):". Top-issues preview also missed BL-* headings; fixed
to include ### BL-\ in the grep pattern.
Bug 3: gsd-code-fixer finding_parser documented CR-\d+ only. BL-* findings from
a drifted reviewer were silently dropped from critical_warning scope. Updated ID
alphabet, severity description, filter sets, and sort order to treat BL-* as
Critical-tier-equivalent to CR-*.
Reviewer contract: gsd-code-reviewer write_review step now declares blocker:/BL-
as accepted tier-equivalent alternatives to critical:/CR-, so the contract
acknowledges the reality the workflow defenses accept.
Regression tests: tests/code-review-pipeline-regression.test.cjs (18 tests)
covers all three bugs behaviourally (pure-function parsers) plus docs-parity
assertions on the workflow and agent .md files.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* changeset: add fragment for PR 3274 (fix(3263) code-review parser)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(workflow): use POSIX [[:space:]] instead of \s in grep -E (CR finding 1)
BSD grep on macOS does not support \s in ERE; replace with the POSIX
[[:space:]] character class so the critical/blocker grep works on both
GNU and BSD grep. Also update the corresponding docs-parity test assertion.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: tighten em-dash and grep docs-parity assertions (CR finding 2)
- Replace `includes('split(/\\s+')` with `includes('split(/\\s+—\\s')`
so the assertion actually enforces the em-dash narrative strip and
cannot be satisfied by a bare whitespace split.
- Update the present_results grep assertion to expect [[:space:]] after
the workflow portability fix.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(3259): non-mutating --help guard for native query handlers; reject --help as milestone version
Adds a dispatcher-level guard in query-dispatch.ts that short-circuits
to a non-mutating help stub whenever --help/-h appears in args destined
for a native mutating handler (fail-closed by default). Adds defense-
in-depth in milestoneComplete to reject --help/-h as a version value
before any disk write. Regression tests cover: per-handler --help guard,
registry-driven invariant across all mutating commands, handler-level
GSDError for both flags, and preservation of the #3019 CJS fallback
contract.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: add changeset fragment for #3272
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(3265): prefer YAML frontmatter for state-snapshot canonical fields
stateSnapshot in both sdk/src/query/state.ts and the CJS twin
(get-shit-done/bin/lib/state.cjs cmdStateSnapshot) passed the whole
STATE.md blob to stateExtractField, whose bold pattern (**Field:**)
has no line anchor. A body table cell such as
"**Status:** to ✅ COMPLETE" therefore silenced the correct YAML
frontmatter value.
Fix: extractFrontmatter(content) first; stripFrontmatter(content) for
the body passed to stateExtractField; for each canonical scalar field
prefer the non-empty frontmatter value, falling back to body extraction
when the key is absent or the file has no frontmatter block at all.
Regression tests added in sdk/src/query/state.test.ts (vitest) and
tests/state.test.cjs (node:test) covering:
- frontmatter status beats **Status:** inside a table cell
- frontmatter current_plan beats bold body value
- no-frontmatter files continue to extract from body
- field absent from frontmatter falls through to body extractor
Fixes#3265
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: add changeset for #3275
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: reproduce fmStr drops non-string YAML scalars (#3275 CR finding)
Add tests/bug-3275-fmstr-non-string-scalars.test.cjs with 5 cases covering
CJS state-snapshot with numeric frontmatter scalars (current_phase: 19,
total_phases: 7, total_plans_in_phase: 5), string regression, and
no-frontmatter body fallback regression.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(state): fmStr accepts numeric/boolean YAML scalars (CR finding)
Rename `fmStr` to `fmScalar` in both state.cjs and sdk/src/query/state.ts
and broaden the type guard so that non-null number/boolean frontmatter values
are coerced to String(v) instead of being discarded.
The previous `typeof v === 'string'` check was a latent bug: if the YAML
parser ever returns typed scalars (e.g. `current_phase: 19` as the number 19),
the frontmatter value would be silently dropped and the stale body value used
instead. Both files are updated identically (k014 parity).
Also adds three SDK vitest regression cases (numeric current_phase,
total_phases, total_plans_in_phase) in sdk/src/query/state.test.ts.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(adr): add ADR-0003 model catalog module
* fix(#3229): add shared model catalog as source of truth for agent profiles and runtime tier defaults
Research / design (ADR-0003):
- Existing drift came from 4 independent model truths:
1. CJS model-profiles.cjs
2. SDK config-query.ts stale copy (18 agents)
3. settings-advanced.md runtime tier table
4. session-runner Claude-only profile map
- New design: one machine-readable Model Catalog Module in sdk/shared/
that both packages ship and consume.
Implementation:
- sdk/shared/model-catalog.json — canonical source of truth for:
- full 33-agent registry
- per-agent golden (quality) alias + balanced/budget aliases
- adaptive derivation from routingTier
- agent→phaseType map
- agent→dynamic-routing default tier map
- runtime tier defaults for all supported runtimes
- get-shit-done/bin/lib/model-catalog.cjs — CJS adapter over the catalog
- sdk/src/model-catalog.ts — SDK adapter over the same catalog
- CJS model-profiles.cjs now re-exports derived data from model-catalog.cjs
- SDK config-query.ts now re-exports MODEL_PROFILES/VALID_PROFILES from
model-catalog.ts instead of maintaining its own list
- sdk/src/query/helpers.ts runtime list now comes from the catalog (fixes hermes drift)
- sdk/src/session-runner.ts Claude profile→model-id mapping now resolves via catalog
- docs/CONFIGURATION.md + settings-advanced.md runtime tables updated to match catalog
Behavior changes:
- resolve-model now covers every shipped agent file on disk (33 agents)
- unknown-agent fallback is profile-semantic, not hardcoded sonnet:
quality→opus, budget→haiku, balanced/adaptive→sonnet, inherit→inherit
- Group B runtimes remain known runtimes but do not get built-in tier defaults
Tests (RED→GREEN):
- root tests: shipped agent files must equal MODEL_PROFILES keys
- sdk tests: shipped agent files must equal MODEL_PROFILES keys
- direct fix assertion: gsd-code-reviewer resolves to opus under quality with no unknown_agent
- runtime defaults parity test: settings-advanced.md + CONFIGURATION.md tables must match catalog
- helper tests: hermes included in SUPPORTED_RUNTIMES and getRuntimeConfigDir()
Closes#3229
* chore(changeset): update #3229 changeset pr field to 3230
* fix(ci): update inherit fallback expectations and inventory parity for model catalog
* test: reproduce extractFrontmatter LAST-block bug (#3240)
* test: reproduce state.update progress trampling and percent formula (#3242)
Two failing regression tests:
- Bug A: state.update "Last Activity" tramples curated progress.* frontmatter via readModifyWriteStateMd → syncStateFrontmatter
- Bug B: 12 declared ROADMAP phases / 6 realized / 6/6 plans done → percent: 100 instead of 50 (phase-fraction ignored)
* test: reproduce TOML float rejection and partial rollback (#3245)
Two failing regression tests:
1. parseTomlToObject rejects valid Codex TOML floats (tool_timeout_sec = 20.0)
2. Post-install validation failure leaves skills/, agents/, VERSION on disk
despite restoring config.toml — hybrid state after abort
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(install): accept TOML floats; idempotent codex rollback (#3245)
Two fixes for the Codex install failure introduced by #2760 CR4 finding 3:
1. parseTomlValue now accepts TOML 1.0 float literals (decimals,
exponents, underscore separators, signed). Codex CLI's serde schema
requires f64 for tool_timeout_sec / startup_timeout_sec — the prior
strict-integer-only check was the inverse of what Codex requires,
causing every config with a float to trigger a fatal schema validation
failure. Date/time separators (-/:T/Z) are still rejected.
2. restoreCodexSnapshot is extended into a unified idempotent rollback
that reverts ALL Codex-specific mutations on failure:
- config.toml (existing behavior)
- skills/gsd-* directories (new)
- agents/gsd-*.{md,toml} files (new)
- get-shit-done/VERSION (new)
- orphaned atomic-write temp files (new)
Pre-install state is captured before the first Codex write so the
rollback reflects the true pre-GSD state. Non-gsd-* user content is
untouched. The rollback is safe to call multiple times and before any
snapshots are captured.
Fixes#3245
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* changeset: pr=3254 for #3245
* test: fix source-grep lint violation in bug-3242 test (#3242)
Replace content.includes() check with line-by-line parse of STATE.md body.
The lint enforces structural assertions over raw text matching.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: mark #3242 RED tests as todo pending fix (#3242)
The three failing tests are intentional regression tests for bugs in
state.cjs that will be fixed in a separate PR. Mark them { todo: true }
so they don't block CI on this branch.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(install): tighten TOML underscore placement validation (CR finding 1)
The float regex used [\d_]* which accepts invalid forms like 1__0, 1_.0,
and 1._0. TOML 1.0 §2 requires underscores only between digits. Switch
both the integer pre-check and the full float pattern to (?:_?\d)* so
consecutive underscores, leading underscores on a segment, and trailing
underscores on a segment are all rejected before replace(/_/g,'') can
silently normalize them into valid JS numbers.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(install): restore pre-existing gsd-* content on rollback (CR finding 2)
The snapshot only recorded names of pre-existing skills/gsd-* dirs and
agents/gsd-* files. On a failed reinstall the rollback could delete
newly-created dirs but could not restore the bytes of dirs/files that
were overwritten, leaving the user in a hybrid state (old config.toml,
new skill files).
Now snapshot the full file tree of every pre-existing gsd-* skill dir
into codexPreInstallSkillContents (Map<name, Map<relPath, Buffer>>) and
every pre-existing agent file into codexPreInstallAgentContents
(Map<filename, Buffer>). restoreCodexSnapshot() uses these maps to
wipe-and-restore overwritten entries and only removes entries that had
no pre-install state, giving a true atomic rollback guarantee.
Reads are best-effort so a partial snapshot is still better than none.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(install): scope temp-file cleanup to installer-owned writes (CR finding 3)
_cleanTmpFiles() was deleting any *.tmp-<pid>-<n> file found under
targetDir. This is too broad: other tools in the user's Codex/home
directory may create temp files matching the same suffix pattern, and a
GSD install rollback would silently delete them.
Add __atomicWrittenTmps (a module-level Set<string>) populated by
atomicWriteFileSync for every temp path it creates. _cleanTmpFiles()
now checks __atomicWrittenTmps.has(full) before unlinking, so only temp
files this installer process actually wrote are eligible for cleanup.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(test): remove no-op doesNotThrow wrapping try/catch (CR finding 4)
assert.doesNotThrow(() => { try { f(); } catch(_){} }) always passes
because the catch block swallows every exception before the outer
assertion can see it. This meant the rollback-idempotency guarantee was
never actually verified.
Replace with an explicit threw flag around runCodexInstall, assert that
the install did throw (validation failure is expected), and add a
post-rollback state assertion that skills/ was not created. This gives
a loud failure surface if runCodexInstall starts crashing from inside
the rollback path, matching the intent described in the test comment.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(test): correct describe title for float-acceptance tests (CR nitpick 1)
The describe block title said 'rejects malformed input that previously
slipped through', but the test inside now asserts that TOML floats are
accepted (the #3245 inversion). This misled readers expecting every
sub-test to assert rejection. Update the title to reflect the mixed
behaviour: floats are accepted; dates and trailing-garbage are rejected.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(test): rename test to match what the assertion actually checks (CR nitpick 2)
The test name 'post-install config retains float literal form (20.0 not
truncated to 20)' promised a string-form invariant, but the assertion
uses numeric equality (assert.strictEqual(parsed.tool_timeout_sec, 20))
which cannot distinguish 20 from 20.0 in JS. Rename to 'post-install
config round-trips tool_timeout_sec as numeric 20' so the description
matches what the test actually verifies.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(test): replace raw text scan with state json assertion (CR nitpick 3)
The 'Last Activity updates the body field' test was reading STATE.md as
raw text, splitting on newlines, and using lines.find/startsWith to
locate the 'Last Activity:' line — the exact pattern-match-on-source
approach prohibited by the no-source-grep testing standard.
Replace with runGsdTools('state json', tmpDir) which surfaces the body-
extracted Last Activity value as fm.last_activity in its JSON output,
and assert against that structured field instead.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(test): correct post-rollback state assertion for early-failure case
The previous assertion checked that skills/ didn't exist, but the
installer writes skills/ before the schema validator fires. Rollback
removes gsd-* dirs inside skills/, not skills/ itself. Update the
assertion to verify that no gsd-* skill dirs survive rollback, which
is the actual invariant the test name describes.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* changeset: document full rollback scope (CR finding 1)
Adds config.toml restoration and orphaned atomic-write temp-file
cleanup to the changeset description — the previous text only listed
skills/, agents/, and VERSION.
* fix(install): wrap post-snapshot scope in rollback handler (CR finding 2)
Any throw between the pre-install snapshot capture and the Codex config
block (skills copy, agents copy, VERSION write, manifest write, leaked-
path scan, etc.) now triggers _codexPreConfigRollback() so the caller
is never left in a partially-installed state. Previously only the later
config.toml mutation paths had rollback wired in.
Introduces _codexPreConfigRollback (defined right after snapshot capture)
and wraps the intervening operations in a try/catch that invokes it on
error for Codex installs; non-Codex paths are unaffected.
* test: assert threw=true to prevent vacuous pass (CR finding 4)
Two tests used bare try/catch without asserting threw === true, so they
would silently pass even if runCodexInstall never threw (k060 pattern).
Each bare catch block is replaced with a threw flag and a
strictEqual(threw, true, ...) assertion.
CR findings 2+3 are both addressed in the preceding install commit:
finding 3 (restore from snapshot manifest, not current FS state) lands
alongside the rollback-wrapper change as part of the restoreCodexSnapshot
refactor.
* fix(install): reject leading zeros in TOML float integer part per TOML 1.0 (CR finding round 4)
TOML 1.0 §2 disallows leading zeros in the integer part of numeric
literals — `01`, `00`, `01.5`, `00e2`, `+01.0`, `-01.0` are all invalid.
The pre-check and float regexes in parseTomlValue used `\d(?:_?\d)*` which
accepted any digit as the leading digit.
Both regexes are tightened to `(0|[1-9](?:_?\d)*)` for the integer part:
- `0` alone is valid
- a non-zero leading digit followed by optional underscored digits is valid
- `01`, `00`, and any variant with a leading zero and further digits is rejected
The "still rejects bare time (07:32:00)" test assertion is broadened from
`/unsupported TOML value/` to `/unsupported TOML value|trailing bytes/`
because the parser now stops at `0` and the remainder `7:32:00` is rejected
as trailing bytes — the invariant (time literals are not accepted) is unchanged.
25 new regression tests cover all rejection cases and valid TOML forms.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: reproduce false GSD SDK ready signals on Linux (#3231)
* fix(install): require persistent SDK reachability before reporting ready (#3231)
* changeset: pr=3249 for #3231
* fix(install): filter _npx from login-shell PATH probe (CR finding 1)
Apply filterNpxFromPath() to the getUserShellPath() result before passing
it to isGsdSdkOnPath(), mirroring the same filtering already applied to
process.env.PATH. Without this, a transient _npx entry in the login-shell
PATH can falsely satisfy the cross-shell reachability check and reintroduce
the false-ready condition this PR fixes.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(test): unconditional legacy-shim replacement assertion (CR finding 2)
Replace readFileSync+includes source-grep check with isLegacyGsdSdkShim()
and add an else branch asserting that when sdkReady is false, a warning/error
was emitted. Previously the sdkReady===false path had no assertion at all,
allowing the test to pass without verifying any postcondition.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: replace text-grep assertions with structured ones (CR finding 2 + nitpick)
Finding 2: restructure the legacy-shim replacement assertion to branch on
isLegacyGsdSdkShim() state (a behavioral fact) rather than console output,
and add an unconditional postcondition for both branches.
Nitpick 3 (4 locations):
- lines 149-153: replace /GSD SDK ready/.test(combined) with
isGsdSdkOnPath(filterNpxFromPath(PATH)) === false
- lines 167-169, 185-189: split filterNpxFromPath result into segments array
and use array.includes() instead of string.includes() on the raw PATH string
- lines 375-377: replace /GSD SDK ready/.test(combined) with
fs.existsSync(shimPath) + isGsdSdkOnPath(filterNpxFromPath(localBin))
All 8 tests pass. lint-no-source-grep: 0 violations.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(build-hooks): per-PID staging dir eliminates concurrent-cleanup TOCTOU race
When multiple test before() hooks spawned build-hooks.js concurrently
(--test-concurrency=4), a race existed: Process A would finish all copies,
call rmdirSync('.dist-staging/') in cleanup, then Process B — still in its
copy loop — would call copyFileSync(src, '.dist-staging/hook.pid.ts') and
get ENOENT because the staging directory was gone.
On macOS/Linux, copyFileSync reports the SOURCE path in ENOENT errors when
the destination directory is missing, making the failure appear to be a
missing source file (hooks/gsd-statusline.js) rather than a missing
destination directory. This misled the diagnosis.
Fix: make STAGE_DIR per-PID ('.dist-staging-<pid>/') so each builder owns
its own staging directory. No other process touches it, eliminating all
contention on staging-dir creation and cleanup. Update .gitignore to match
the new 'hooks/.dist-staging-*/' glob.
Reproduces as: CI test matrix (macos-24, ubuntu-22, ubuntu-24) all failing
with ENOENT on hooks/gsd-statusline.js in bug-2136 before() hook. The new
test file added in this PR (bug-3231) shifts the concurrency schedule just
enough to expose the race on every CI run.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: assert on captured console output, not tautological PATH state (CR finding)
The two discarded `captureConsole()` return values in the bug-3231 test
were flagged by CodeRabbit as tautological assertions. Fix:
- Test 1 (transient _npx PATH): capture stdout/stderr and assert the
installer does NOT emit "GSD SDK ready" (the false-positive the PR
fixes), and that it does emit some diagnostic output instead.
- Test 3 (clean install): capture stdout/stderr and assert the installer
DOES emit "GSD SDK ready" after successfully self-linking into a
persistent PATH dir — confirming the positive path works correctly.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: reproduce nested plans/ undercount in buildStateFrontmatter (#3257)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(state): count nested plans/<N>-PLAN-<NN>-<slug>.md in buildStateFrontmatter (#3257)
`buildStateFrontmatter` did a flat `readdirSync` on each phase directory and
missed plan files inside the nested `plans/` subdirectory written by
gsd-plan-phase (post-#3139 / #3115). Every state mutation flowing through
`syncStateFrontmatter` overwrote the curated `progress.*` frontmatter block
with the under-counted disk scan.
The fix adds a `plans/` descent using the same regex shapes as
`roadmap.cjs:countPhasePlansAndSummaries` and `phase.cjs:looksLikePlanFile`
(#2893/#3128). Both the `{N}-PLAN-{NN}-{slug}.md` (agent-emitted) and
`PLAN-{NN}-{slug}.md` (bare-prefix) forms are now matched. Outline files
(`-PLAN-OUTLINE.md`) and pre-bounce files are excluded. Flat-layout repos
are unaffected.
Note: the same algorithm now lives in 4 places (state.cjs, roadmap.cjs,
init.cjs, phase.cjs). Shared-helper extraction per CONTEXT.md k014 is
tracked in the follow-on issue filed with this PR.
Sibling fix to #3115 / #3139 / #3191 — state.cjs was missed in the
post-#3139 migration that updated the other three files.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* changeset: pr=3261 for #3257
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(changelog): add entry for #3257 nested plans/ fix (#3261)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(state): broaden PLAN_PRE_BOUNCE_RE to match bare PLAN- prefix (CR)
PLAN_PRE_BOUNCE_RE was /-PLAN.*\.pre-bounce\.md$/i, which missed bare-prefix
files like PLAN-01-foo.pre-bounce.md in the nested plans/ scan — those would
incorrectly count as real plans. Broadened to /\.pre-bounce\.md$/i to exclude
any .pre-bounce.md file regardless of prefix shape.
Adds regression test for this exclusion.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(state): extend nested plans scan to cmdStateValidate and cmdStateSync (CR finding)
`buildStateFrontmatter` already received the nested-aware scan in this PR, but
`cmdStateValidate` and `cmdStateSync` still did flat-only `readdirSync` on the
phase root, producing false plan-count drift warnings and under-counted totals
on `phases/<N>/plans/` repos. Extend the identical scan pattern to both sites
(regex byte-identical to the `buildStateFrontmatter` site, k014). Regression
tests added for all three commands.
Closes#3257
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(bug-3257): replace readFileSync+.includes() with structural dry-run idempotency check
The lint-no-source-grep rule flags readFileSync-bound variables used with
text-match methods (.includes, .match, etc.). Replace the afterContent.includes()
check with a structural idempotency assertion: run state sync --verify twice and
confirm the second run still reports a pending change, proving the first dry-run
did not mutate STATE.md.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(bug-3257): fix progress assertion to use min(plan,phase) formula (#3242)
After rebasing onto main, computeProgressPercent now applies
min(plan_fraction, phase_fraction) per #3242 Bug B. Update the
multi-phase sync test to assert 50% (min(3/5, 1/2)) instead of 60%.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>