The descriptor-driven install path (resolveRuntimeArtifactLayout) installed no
agents for copilot/antigravity/cursor/windsurf/augment/trae/codebuddy/cline —
their per-runtime agent conversion ran only via the legacy bin/install.js loop.
This wires each runtime's agent converter into the descriptor so the new path
applies per-runtime conversion (follow-up to #1099; ADR-1235 cutover).
- capabilities/<rt>/capability.json: declare an `agents` kind with the runtime's
converter (global+local; cline global-only). Regenerated capability-registry.cjs.
- convertedAgentsKind threads install scope -> isGlobal so the scope-aware
copilot/antigravity converters choose global vs workspace-relative paths; the
six single-arg converters ignore the extra arg. stageAgentsForRuntimeWithConverter
passes isGlobal to the converter.
- Tests: feat-1173 gains a real-registry block asserting each runtime's descriptor
applies the correct converter (== conv(src, isGlobal), != raw copy) with scope
threading (fails-first on pristine next). The ADR-857 equivalence golden +
per-runtime kind-count assertions now include the agents kind, each annotated as
an intentional #1173 change.
Scope: this wires the per-runtime CONVERTER. The remaining byte-parity behaviors of
the legacy loop (copilot `.agent.md` rename, cross-cutting path/attribution rewrites,
config-reading) stay with the legacy loop -- which runs after installRuntimeArtifacts
and is authoritative for the real install -- and are tracked by ADR-1235's later
cutover steps. No user-facing change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs
Generalizes the Codex-only #1515/#1519 fix to every non-Claude runtime, and
wires it into the real install path (where it was previously dead-on-arrival).
Root causes:
1. The runtime-default stamping lived only in `_applyRuntimeRewrites`, but the
installer emits `gsd-core/workflows/*.md` via `copyWithPathReplacement`, which
never calls it — so a real `--codex`/`--cursor`/etc. install emitted
`--default claude` and worktrees-on. RUNTIME mis-resolved to claude and the
workflow ran executors unisolated against the main checkout. (#1515/#1519 were
also dead-on-arrival in real installs; this repairs them.)
2. Only `case 'codex'` was stamped; every other non-Claude runtime kept the
Claude default.
Fix:
- New `_stampNonClaudeRuntimeDefaults(content, runtime)` (single shared helper)
stamps `--default <runtime>` + `use_worktrees=false` for every `runtime !=
claude`; called from both `_applyRuntimeRewrites` and, crucially,
`copyWithPathReplacement` in bin/install.js (the real workflow emit path).
- Generalize the fail-closed worktree guard `= codex` -> `!= claude` in
execute-phase/quick/diagnose-issues (worktree isolation is Claude-Code-only).
- Flip manager/autonomous inline-vs-background gating to `codex -> background,
everything-else -> inline` (research: only Codex can background-nest the
pipeline's subagents; all others run inline, which they support).
Worktree-capability determination is research-backed (official docs for all 14
non-Claude runtimes: none honor GSD's isolation="worktree" mechanism, only Codex
background-nests). New end-to-end real-install test asserts the EMITTED workflow
is stamped — the regression guard that would have caught the dead-on-arrival bug.
Closes#1521
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r
* chore(#1521): backfill changeset PR number (#1537)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Root cause: in acquireStateLock, once openSync(O_CREAT|O_EXCL) created the
lock file, the subsequent writeSync(pid)/closeSync were unguarded. A
recoverable errno (EAGAIN etc., in ACQUIRE_LOCK_RETRY_ERRNOS) made the catch
do checkBudgetAndSleep + continue WITHOUT closing the fd or unlinking the
just-created empty lock — leaking a descriptor every occurrence and stranding
a content-less lock (the #500/#905/#1230 STATE.md write-corruption family).
Fix: wrap writeSync/closeSync in an inner try that guardedly closeSync(fd) +
unlinkSync(lockPath) then re-throws to the existing outer catch (DRY errno
classification). Recoverable errno retries from a clean slate; a FATAL errno
(e.g. ENOSPC, not recoverable) still propagates after cleanup — not masked.
Mirrors the already-shipped capability-lock.cts:415-425 pattern.
Extends the M8 test seam with a one-shot simulateWriteError errno + an
onLoopIteration snapshot hook so the orphan-before-retry is deterministically
observable. New tests prove RED (orphan stranded / fatal leaves orphan)
before the cleanup and GREEN after.
Source of truth src/state.cts (ADR-457); bin/lib/state.cjs is generated.
Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz
Root cause: writeStateMd computed its frontmatter disk scan
(syncStateFrontmatter — the READ half of a read-modify-write) BEFORE
acquireStateLock, leaving a TOCTOU window. A concurrent writer that
committed a new PLAN/SUMMARY between our scan and our lock made writeStateMd
stamp stale progress counts (lost update — the #500/#905/#1230 family). The
atomic sibling readModifyWriteStateMd already scans inside its lock.
Fix: move _diskScanCache.delete + syncStateFrontmatter inside the
acquireStateLock-held try, before platformWriteSync. Byte-for-behaviour
identical for single-threaded callers — only the concurrent-writer window
closes. Adds an afterAcquire test seam (mirrors the M1 _setLockProbes seam)
to make the window deterministic; new test proves RED (stale count) before
the reorder and GREEN after.
Source of truth src/state.cts (ADR-457); bin/lib/state.cjs is generated.
Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz
The M1/M2 lock fix (4903ee04) was asymmetric: acquireStateLock got a 60s deadman
ceiling (recovers a lock once its age crosses an absolute bound ABOVE the wait
budget) but withPlanningLock did not. The .lock body carries no startTime, so
_planningHolderVerifiedLive can only check pid liveness — it cannot detect pid
reuse. A false-alive holder (original holder crashed, pid recycled by an unrelated
live process) would therefore make withPlanningLock throw on every call with no
self-heal until the reused pid happens to die.
Mirror acquireStateLock: in the EEXIST branch, steal a verified-live holder anyway
once the lock ages past deadmanCeilingMs (60000 > lockTimeout 10000). mtime age is
measured from lock creation, so a stuck lock self-heals on a subsequent call.
Adds a clock+probe-seam regression test (false-alive holder past the ceiling IS
stolen). planning-workspace 13/13, clock-seam 34/34, locking suites green; lint clean.
Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz
The STATE.md write lock (acquireStateLock) and the .planning/ workspace lock
(withPlanningLock) stole contended locks on mtime age alone with no
process.kill(pid,0) liveness check, and mis-ordered stale-vs-wait so a
live-but-slow holder could be robbed mid-critical-section.
M1 (lost update / STATE.md corruption): a live writer whose critical section
ran past the stale threshold aged out and a waiter unlinked its lock and
acquired -> two writers in STATE.md's read-modify-write window. mtime is a
leaky proxy for "holder is alive"; it leaks under exactly the slow-holder
condition the lock guards against.
M2 (uncaught EEXIST): withPlanningLock's timeout fallback unconditionally
unlinked whatever lock existed (even a live holder's) and re-acquired OUTSIDE
any try -- a concurrent re-create raced a raw EEXIST out of the helper.
Fix backports capability-lock.cts's liveness gate (process.kill(pid,0) via a
_setLockProbes/_resetLockProbes test seam):
- acquireStateLock: steal when holder pid is DEAD (any age) OR age exceeds a
deadman ceiling (60000ms, ABOVE maxWaitMs=30000) so a verified-live holder is
never stolen within budget; garbage/legacy bodies stay recoverable.
- withPlanningLock: same gate in the EEXIST path (dead stolen promptly, live
waited on); removed the unconditional force-steal -> clear timeout throw,
which also closes M2 (no re-acquire outside try).
Uncontended path unchanged byte-for-behaviour; realClock + real process.kill
remain the defaults. Pid-reuse residual fails safe (waits/times out, never
corrupts) and recovers at the deadman ceiling.
Tests: TDD red->green via the clock + new pid-liveness probe seams (no
wall-clock; #453 deleted the race tests). 8 new behavioural tests across
tests/clock-seam.test.cjs and tests/planning-workspace.test.cjs; the prior
withPlanningLock timeout test rewritten to pin the no-force-steal contract.
Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz
* fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees
A Codex install with a runtime-neutral .planning/config.json resolved
RUNTIME=claude and enabled git worktree isolation, which Codex's
spawn_agent cannot honor. Two root causes:
1. Workflows read `config-get runtime` / `config-get workflow.use_worktrees`
without `--raw`, so config-get's JSON-quoted output ("codex") was captured
verbatim into the bash var and broke every `[ "$RUNTIME" = ... ]` check —
the Codex fail-closed guard was dead even when runtime:codex was explicit,
and Claude's own worktree degrade-check was dead too. Add `--raw` to those
reads across execute-phase, autonomous, manager, diagnose-issues, quick.
2. The conversion engine emitted `--default claude` for every runtime. Stamp
the codex-emitted workflows to `--default codex` (runtime) and
`--default false` (use_worktrees) in _applyRuntimeRewrites case 'codex', so
a neutral config on a Codex install resolves runtime=codex / worktrees off.
Also extend the Codex fail-closed worktree guard to quick.md and
diagnose-issues.md (they spawned isolation="worktree" with no runtime guard).
Regression test asserts source<->engine parity across all five workflows
(DEFECT.GENERATIVE-FIX) plus fast-check property coverage of the stamping.
Closes#1515
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r
* chore(#1515): backfill changeset PR number (#1519)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The #1279 node-test machine-proof confirmed a known-bad subject drives the
negative test RED, but could not distinguish a genuine content-violation from a
deceptive test that reds merely because GSD_PROHIB_SUBJECT is set. Add an
optional fifth flat scalar `check_clean_fixture` (-> CheckDescriptor.cleanFixture)
threading a KNOWN-CLEAN control subject through projectProhibitions +
descriptorFromProjection. When present, the prover also runs the check against
the clean subject and requires GREEN, so fail-first is proven only when the check
is RED on the violation AND GREEN on the clean subject (content-dependent).
Opt-in and additive: absent a clean fixture the prover behaves exactly as
post-#1314 (no control, documented residual), preserving the zero-authoring
compose path; the lint-rule kind needs no analog (its subject IS the linted
file, no env indirection). Coverage: RED-first deceptive case, positive,
missing-clean fail-closed, round-trip read-back/emit, fast-check property
extended to the 5th scalar, and an end-to-end COMPOSE capstone (honest vs
deceptive). Docs: ADR-550 dated addendum, prohibition-probe reference,
spec-phase + verify-phase workflows.
Closes#1346
Claude-Session: https://claude.ai/code/session_01GsPRb8zvpcT7Eat6vZw8PX
* refactor(#1511): move content-rewrite engine to conversion module, delete the install.js relay
Phase 2 of epic #1507 (ADR-1508). Behavior-preserving: makes the Runtime
Artifact Conversion Module the single owner of per-runtime content rewriting and
removes the last upward .cts -> bin/install.js dependency.
- src/runtime-artifact-conversion.cts now owns the engine (_applyRuntimeRewrites,
5-arg with INJECTED attribution), the staged-content walkers
(applyRuntimeContentRewritesInPlace / ...ForCommandsInPlace), computePathPrefix
(private, exported as _computePathPrefix for tests), and the deep public seam
rewriteStagedSkillBodies / rewriteStagedCommandBodies({runtime, configDir,
scope, homedir?, platform?, resolveAttribution?}).
- src/surface.cts:applySurface calls rewriteStagedSkillBodies directly (no
resolveAttribution -> undefined). Co-Authored-By is absent from ALL rewritten
content, so processAttribution is vacuous there and undefined is provably
behavior-identical. surface no longer imports getInstallExports.
- src/runtime-artifact-layout.cts: deleted getInstallExports / loadInstallExports
/ InstallExports + the GSD_TEST_MODE require('bin/install.js') relay.
- bin/install.js: binds computePathPrefix / the two walkers / _applyRuntimeRewrites
from the conversion module (single implementation, exports preserved for Hyrum);
install callsites pass getCommitAttribution(runtime) as the injected attribution.
getCommitAttribution stays here (impure install-time config I/O).
- DEFECT.GENERATIVE-FIX guard: tests assert install.X === conversion.X reference
identity for computePathPrefix + both walkers (no drift).
New tests/enh-1511-*.test.cjs (engine, attribution injection, deep seam, prefix,
layout-no-relay guard, reference-identity). 316 affected-suite tests green; lint clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0187qgypdy1wkWRpdaf2hRuD
* test(#1511): make rewrite-engine path assertions Windows-robust
The deep seam normalizes paths as path.resolve(configDir).replace(/\\/g,'/')
and compares homedir().replace(/\\/g,'/'). Three assertions in the new test
rebuilt expected paths without that normalization, so they passed on Mac/Linux
but failed on Windows CI (PR #1513):
- two absolute-branch asserts rebuilt resolvedTarget via path.resolve(configDir)
without the backslash→slash replace → mismatch on Windows.
- the $HOME-branch test fed a POSIX-literal /home/testuser, which Windows
path.resolve re-roots onto the cwd drive (D:/home/...), so the
resolvedTarget.startsWith(homeDir) check failed and the $HOME shorthand was
never produced.
Fix is test-only (engine unchanged, still behavior-preserving): mirror the
engine's .replace(/\\/g,'/') in the two absolute-branch asserts, and use a real
absolute path (path.resolve(os.tmpdir(), ...)) + platform: process.platform for
the $HOME-branch test so the comparison holds on all platforms.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0187qgypdy1wkWRpdaf2hRuD
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phase 1 of epic #1507 (ADR-1508): behavior-preserving relocation of the pure
rewrite-engine helpers out of the hand-authored installer so the conversion
module can own them without importing bin/install.js.
- getDirName -> src/runtime-name-policy.cts (pure runtime->dir-name switch).
- processAttribution -> src/runtime-artifact-conversion.cts (pure Co-Authored-By
content transform).
- bin/install.js imports both back via destructure/binding and re-exports
getDirName unchanged (Hyrum: install.test.cjs + runtime install tests import
getDirName from bin/install.js).
Two refinements to ADR-1508's Phase 1 (verified against the source):
- getCommitAttribution STAYS in bin/install.js: it is impure install-time
config I/O (reads runtime settings.json, uses install-time config-dir state +
attributionCache), not a content transform. Phase 2 will inject the resolved
attribution into the engine rather than move this function.
- The convertClaudeToAugmentMarkdown dedup is deferred to Phase 2's cleanup: the
local copy is entangled with a converter cluster (convertSlashCommandsTo
AugmentSkillMentions is used only by it; the family is partly dead-local via
the ...runtimeArtifactConversion export spread), deduping only augment would be
arbitrary, and it is not required to unblock Phase 2 (the engine will call the
conversion module's own copy when it moves).
New tests/enh-1510-*.test.cjs: getDirName at its new home (all runtimes +
fallback + install.js re-export identity) and processAttribution
(null/undefined/string/$-escape/CRLF/global). 487 affected-suite tests green.
Claude-Session: https://claude.ai/code/session_0187qgypdy1wkWRpdaf2hRuD
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phase 0 of epic #1507. Records the decision to make the Runtime Artifact
Conversion Module the single owner of per-runtime content rewriting, flip the
dependency direction to installer/layout -> conversion, and close the
surface.cts -> bin/install.js getInstallExports relay. Doc-only.
Resolves ADR-3660 Initial-Scope deferral; distinct from epic #1258.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0187qgypdy1wkWRpdaf2hRuD
cmdAgentSkills' plain (non-JSON) path wrote the <agent_skills> block with
process.stdout.write then immediately called process.exit(0). stdout.write is
async on pipes/files, so on Windows the process tore down before Node flushed
the buffer — the workflows' `$(gsd_run query agent-skills <type>)` capture
received 0 bytes and every ${AGENT_SKILLS_*} substitution expanded empty,
silently dropping configured per-agent skills.
Route the plain path through the existing synchronous output() helper
(writeAllSync, src/io.cts) + return — the same flush-safe mechanism the --json
branch already uses — instead of write + process.exit(0).
Adds a #1400 regression block to tests/agent-skills.test.cjs that captures
stdout via a real file descriptor and asserts the block is non-empty and
byte-identical to the --json .block.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Proactive checkpoint guard fires at each wave boundary before spawning agents.
Self-assesses context pressure against context-budget.md degradation tiers and
warns (warn, default) or auto-invokes /gsd:pause-work (auto) when POOR tier
(70%+) is detected. Config key validated; defaults to \"warn\".
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
loadConfig() returns a flattened object with no nested `workflow` key, so
config?.workflow was always undefined, making drift_action permanently 'warn'
and drift_threshold permanently 3 regardless of .planning/config.json. Fixes
by reading the raw config.json directly (matching the pattern in
check-command-router.cts:readWorkflowConfig). Adds two behavioral regression
tests that fail under the old code and pass under the fix.
Closes#1493
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
product-name-purity (#1777) rejects 'Gemini (…)' parentheticals that render
verbatim into CHANGELOG.md. Reword to 'Gemini and Gemini-backed Antigravity';
no behavior change. (Reword was left uncommitted in the prior rebase push.)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Addresses review: the regression unit test required the tsc build artifact
(gsd-core/bin/lib/runtime-artifact-conversion.cjs) instead of the live install
path. Destructure convertGeminiToolName from the existing ../bin/install.js
import so a stale build can't false-green while the live copy is broken, and
assert the full AskUserQuestion/ask_user exclusion group.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
convertGeminiToolName lowercased any unmapped Claude tool, so Skill and
SlashCommand became an invalid `skill`/`slashcommand` tool name. Gemini CLI
has no such built-in tool, so frontmatter validation failed (tools.N: Invalid
tool name) and aborted the entire agent load — 22 of 34 GSD agents were dead
on Gemini.
Add Skill and SlashCommand to the same `return null` exclusion branch that
already handles AskUserQuestion, in both the canonical src converter and the
hand-maintained bin/install.js copy that runs on the live --gemini install
path. Antigravity reuses this converter (it runs on the Gemini backend) and is
intentionally covered by the same exclusion — it surfaces GSD skills via the
skill surface (SKILL.md), not the agent tools: allowlist — locked by an added
Antigravity regression test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1498): regenerate capability-registry.cjs in npm version lifecycle hook
The version npm lifecycle script now runs gen-capability-registry.cjs --write
after sync-manifest-versions.cjs stamps all capability/*.json version fields.
Without this, every npm version X.Y.Z call left capability-registry.cjs stale
(capability JSONs got the new version string, but the registry still embedded
the old ones), causing gen-capability-registry.cjs --check to fail in the RC
test suite.
- package.json version script: add gen-capability-registry --write + git add
- tests/issue-844-manifest-version-sync.test.cjs: regression guard (describe F)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: backfill changeset pr: 1499
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Address second maintainer review on #1487:
- Blocker 1: getInstallExports now caches per runtimeConfigDir (Map) so a
no-arg warm-up via the legacy walk-up path cannot poison a later
getInstallExports(configDir) call. Add a regression test (verified to fail
on the singleton cache) proving the no-arg warm-up does not poison the
configDir-keyed resolution.
- Blocker 3: document the .gsd-source two-party provisioning contract in the
CONTEXT.md Runtime Artifact Layout Module entry (writer, reader, content,
guard, fall-through, package-root invariant, per-configDir cache).
Claude-Session: https://claude.ai/code/session_01XNT3SWgzjmEycNuweURDme
Address review on #1487:
- Scope the marker write to runtime === 'claude' && isGlobal, matching
issue #1477. The Claude global skills layout is the only install path
that ships the skills layout without a commands/gsd source tree; every
other runtime/scope deploys commands/gsd, so walk-up already resolves.
- Add the trailing (#1487) reference to the changeset body per PRED.k329.body.
Claude-Session: https://claude.ai/code/session_01XNT3SWgzjmEycNuweURDme
* fix(#1369): refresh wave manifest and re-check base before each wave in execute-phase
Two compounding issues caused wave N+1 worktrees to fork from the stale
pre-wave-N commit, immediately tripping the worktree_branch_check FATAL
guard in every executor:
1. worktree.base-check auto-degrade only ran once at initialize time.
After wave N merges advanced orchestrator HEAD past origin/HEAD, new
worktrees were still forked from origin/HEAD (Claude Code "fresh" base).
2. WAVE_WORKTREE_MANIFEST was never unset between waves, so wave N+1
reused the consumed wave-N manifest file, which would have blocked the
step 5.5 manifest guard (#3384) on subsequent waves.
Fix: add two safeguards in execute-phase.md —
- Step 0.5 (start of each wave): re-runs worktree.base-check; auto-degrades
USE_WORKTREES=false for that wave when HEAD has diverged from origin/HEAD.
- Step 7c (end of each wave): unsets WAVE_WORKTREE_MANIFEST so wave N+1
creates a fresh per-wave manifest; re-asserts worktree.set-baseref
(idempotent) and re-evaluates base degradation after wave merges land.
17 regression tests added in tests/bug-1369-wave-stale-base.test.cjs.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1369): rename test to fix-NNN convention; update workflow size baseline
Rename tests/bug-1369-wave-stale-base.test.cjs → tests/fix-1369-wave-stale-base.test.cjs
to satisfy the lint-regression-test-names gate (new files cannot use bug-NNN prefix).
Update tests/workflow-size-baseline.json for execute-phase.md: 93157 → 97393
(LF-normalized byte count after adding step 0.5 inter-wave base re-check and
step 7c between-wave manifest reset).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1369): add issue reference to allow-test-rule comment
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1369): extract new execute-phase steps to references; satisfy ADR-857 cap
Step 0.5 (inter-wave worktree base re-check) and steps 7b–7c (pre-wave
dependency check + between-wave manifest reset/base refresh) added by this
PR grew execute-phase.md to 97393 bytes, violating the ADR-857 phase-6
architectural mandate that host-loop bodies remain strictly below the
pre-phase-6 baseline of 93166 bytes.
Extract both new blocks into dedicated reference files:
- gsd-core/references/execute-phase-wave-guard.md (step 0.5)
- gsd-core/references/execute-phase-between-wave-reset.md (steps 7b + 7c)
Replace inline prose with @-reference pointers. File now measures 92851
bytes (LF-normalized), satisfying the ADR-857 capstone conformance gate.
Also update tests/workflow-size-baseline.json to 92851 and add both new
reference files to docs/INVENTORY-MANIFEST.json.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1369): update regression tests to read from extracted reference files
Steps 0.5 and 7b+7c were moved to reference files to satisfy the ADR-857
size cap on execute-phase.md. Tests now check @-reference pointers in the
workflow for ordering and read content assertions from the reference files.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#1464): add behavioral manifest-validation test for capability tutorial docs
Extracts JSON capability manifests from tutorial/reference docs and validates
them through the real validateCapability — closing the test gap that let
issue #1464's broken tutorial manifest (step missing ref) pass undetected.
Adds fix-1464-docs-manifest-validation.test.cjs with:
- Suite 1: build/install/reference docs' complete manifests pass validateCapability
- Suite 2: adversarial fixtures prove the original #1464 bug shapes are caught
(step without ref → "steps[0].ref must be an object…"; id/folderId mismatch)
- Suite 3: extractManifests helper unit tests (complete vs. partial block filtering)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(#1496): allowlist docs module for 3-file test cluster
docs-parity-live-registry, docs-update, and the new
fix-1464-docs-manifest-validation sit on the same docs production
module; add the docs entry to lint-test-file-count.allowlist.json
so the novel-offender CI gate passes.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1453): clean up stale get-shit-done paths in Codex skill mirror on upgrade
Extends planLegacyCleanup in gsd-core/bin/lib/legacy-cleanup.cjs to scan
skills/gsd-* subdirectories for .md files that still embed the pre-rename
get-shit-done/ path (e.g. ~/.agents/skills/gsd-docs-update/SKILL.md).
These stale copies are removed by cleanupLegacyGsdCc during the next
install/upgrade so Codex can no longer discover and select them.
Adds 5 regression tests to tests/issue-607-legacy-cleanup.test.cjs covering
the stale path detection, non-flagging of fresh skills and user-owned dirs,
and the end-to-end ~/.agents/skills scenario from the issue.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1453): add gsd-allow-legacy-name markers to intentional legacy name references
Comments and test descriptions in legacy-cleanup.cjs and its test file
legitimately cite the old 'get-shit-done' directory name to explain what
the cleanup logic removes. Add the lint-exemption marker to each line so
lint-legacy-dir-name passes.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1445,#1446): exclude 999.x backlog phases from milestone totals; allow total_phases downward correction
#1445: deriveProgressFromRoadmap (phase-lifecycle.cts), the roadmapPhaseCount
loop (state.cts), and getMilestonePhaseFilter (roadmap-parser.cts) all now
skip phase tokens matching /^999\b/ — consistent with the existing init.cts
filter. 999.x backlog dirs are consequently excluded from phaseDirs too.
#1446: shouldPreserveExistingProgress (state-document.cts) no longer includes
total_phases in its ratchet check. total_phases always takes the freshly
derived value; only completed_phases, total_plans, and completed_plans retain
ratchet behaviour.
Regression tests added for both bugs.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: add changeset for #1445/#1446 progress-backlog-exclusion-and-ratchet
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1445,#1446): rename test files to fix-NNN convention; fix changeset pr: null
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1437): add phase.list-plans to gsd-tools
Register phase.list-plans in PHASE_COMMAND_ALIASES, implement
cmdPhaseListPlans in src/phase.cts (uses findPhaseInternal + scanPhasePlans
to return plan_count/has_plans/plans/phase_dir), and wire the handler in
phase-command-router. Previously every call produced "Unknown phase
subcommand".
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1437): register new test file in lint-test-file-count allowlist
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1437): rename test to fix-NNN convention; update file-count allowlist
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1422): sub_repos config takes precedence over .git in findProjectRoot
When heuristic-3 (.git + parent .planning/) fires, do a lookahead walk
over ancestors above the matching parent to check if any further ancestor
has a sub_repos entry that explicitly claims the starting directory. If
found, return that ancestor instead — explicit config wins over the
implicit .git signal.
Regression tests updated: the heuristic-3 precedence test now asserts the
correct new behavior (sub_repos wins), and two additional coverage cases
added (startDir directly in sub_repo child, and startDir nested 2+ levels
inside the child).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1447): guard against uncommitted changes before deleting phase dirs in new-milestone
cmdPhasesClear now runs `git status --porcelain <phasesDir>` before
executing any rmSync. If uncommitted changes are detected it calls
error() and aborts, preventing silent data loss when new-milestone's
§6 "phases.clear --confirm" fires before the operator has archived
or committed outgoing phase work.
A new --force flag is added to bypass the guard for callers that have
already verified archival is done (or explicitly accept the loss).
When git is unavailable or the directory is not inside a git repo the
guard silently skips, preserving the existing behaviour for non-git
projects.
Five new regression tests cover: untracked files abort, staged-but-
uncommitted abort, --force bypasses, committed files pass, non-git
project passes without guard.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: add changeset for #1422 and #1447
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: correct changeset format
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1472,#1454): validate health workstream-aware paths; exclude active worktree from W017
#1472: cmdValidateHealth now uses planningRoot(cwd) for shared-root files
(PROJECT.md, config.json, MILESTONES.md) and planningDir(cwd) for
workstream-scoped files (ROADMAP.md, STATE.md, phases/). Previously a
single planningDir() call was used for all paths, causing false
E002/E003/E004/W003 when GSD_WORKSTREAM is set.
#1454: W017 no longer fires for a stale worktree whose path equals or is
an ancestor of process.cwd(), preventing advice to remove the active
session's own worktree.
Regression tests added for both bugs; all 40 existing health tests pass.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: correct changeset format
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1478,#1479,#1480): prohibit ungrounded baselines, error-suppressing fallbacks, and stale-artifact authority in verify blocks
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: add changeset for #1478/#1479/#1480 planner verify gate fix
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: correct changeset format
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1478,#1479,#1480): fix test contract violations from new planner dimensions
gsd-planner.md exceeded both the planner-decomposition 48K char limit and
the reachability-check 50K char limit after the new HARD RULE blocks were
added inline. The full rule details already exist in planner-antipatterns.md
(added in the same PR); replace the verbose inline blocks with a single
@-reference pointer to the antipatterns file, reducing the file from 50981
to 49130 chars (under both limits).
Also regenerate tests/agent-size-baseline.json to reflect the new sizes of
gsd-planner.md and gsd-plan-checker.md.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
#2551/#3182/#2361 are pre-migration get-shit-done-redux issue numbers with no
equivalent in open-gsd/gsd-core; they mislead triage and manufacture phantom
blockers. Repoint to real successors (#717 byte-budget rework, #720) or rewrite
as prose referencing the discuss-phase/modes progressive-disclosure split.
Correct co-located 'line budget'/'<500 lines' framing to the byte-based reality
(#717). Add a CI guard (tests/no-phantom-issue-refs.test.cjs) that fails if a
phantom ref is reintroduced. SSH-key patterns (id_ed25519) left untouched.
No user-facing runtime behavior change.
Closes#1073