- restructure _loadFlatCommandsGsdManifest try/catch to wrap read+parse+set
together (mirrors loadSkillsManifest exactly), so a thrown parser degrades
both keys to [] — closes the latent catch-scope parity drift [Nit-1]
- add boundary test: gsd-.md (empty stem) skipped, gsd-x.md single-char stem
kept (slice(4,-3) boundary) [Low-1]
- add unreadable-file test (POSIX-gated): both keys degrade to [] [Low-2]
- strengthen parity test to also compare requires + _calls_agents_ VALUES,
not just the stem set [Nit-2]
Both orthogonal reviews returned APPROVE with no Critical/High findings.
Security review confirmed no new trust-boundary crossing, prototype-pollution
immune (Map/Set throughout), and symlink/path-traversal surface identical to
the pre-existing nested loader (not a regression).
Mirrors the #1160 installed-layout tests for the flat source layout
(<repo>/commands/gsd-<stem>.md, no commands/gsd/ subdir). Covers stem
extraction (strip gsd- prefix), requires parsing via shared parseRequires,
companion _calls_agents_ key parity, _resolveManifest flat-branch detection,
precedence (flat-empty falls through to installed), and a generative-parity
assertion that the flat loader and nested loader produce identical stem sets
for the real command tree.
Expected RED against unfixed capability-state.cts (_resolveManifest has no
flat branch; _loadFlatCommandsGsdManifest not exported).
- add typeof guard so a non-string override passes through verbatim instead of
crashing on .startsWith (preserves pre-fix no-crash behaviour) [LOW-1]
- use Object.hasOwn() for the alias lookup so __proto__/constructor cannot
return a truthy non-string from the plain object literal [LOW-D3]
- cap the unmappable-override stderr warning at 64 chars so an oversized or
secret-shaped value cannot leak in full to stderr/logs [LOW-D4]
- remove the unused mapClaudeOverrideForRuntime export (helpers are covered
behaviourally via resolveModelInternal/resolveModelForTier) [NIT]
- add resolveModelForTier unmappable-override fall-through test (closes the
mutation-score gap) [MEDIUM-1]
- add case-sensitivity contract test (Claude-Sonnet-5 passes through verbatim) [LOW-2]
Both orthogonal reviews returned APPROVE with no Critical/High findings.
model_overrides values that are full Claude model IDs (claude-sonnet-5,
claude-opus-4-8, claude-haiku-4-5, claude-fable-5) were returned verbatim on
the claude runtime and handed to the Claude Agent tool, whose typed model
parameter documents only tier aliases (opus/sonnet/haiku/fable). The
model_policy path already mapped full IDs -> aliases via
CLAUDE_POLICY_ID_TO_ALIAS (#1144); model_overrides skipped that mapping, so
the two resolver paths produced different shapes for the same underlying
Claude model. The fix mirrors #1144 on the override path via a shared
mapClaudeOverrideForRuntime helper used by both resolveModelInternal and
resolveModelForTier. Bare aliases pass through verbatim; non-Claude runtimes
and non-Claude custom/vendor values keep full IDs verbatim (parity). An
unmappable Claude ID (e.g. claude-opus-4-5) warns once to stderr and falls
through to tier resolution, exactly as the model_policy path already does.
Alias mapping is also the documented best practice (prevents staleness when
new model versions ship).
Mirrors the #1133 model_policy alias-mapping tests for the model_overrides
path. Covers AC1-AC6: mappable Claude full IDs (claude-sonnet-5/opus-4-8/
haiku-4-5/fable-5) resolve to aliases on runtime:claude; bare aliases pass
through; non-claude runtimes keep full IDs verbatim; unmappable Claude IDs
warn-once + fall through; resolveModelForTier escalation path also maps;
non-Claude custom/vendor values pass through verbatim (regression guards).
Expected RED against unfixed model-resolver.cts (override short-circuit at
lines 162-167 / 288-290 returns override verbatim with no alias mapping).
- compareSemver: implement full SemVer 2.0.0 §11 pre-release identifier
comparison (two pre-releases of the same triple now order correctly; was 0).
- capability description + fragment: scope the plan-checker/verifier claim
(this capability delivers the parallel-execution backend; those gates remain
inline until separately wired). Correct the 'each wave is one barrier' prose
(a wave splits into multiple sequential parallel() barriers on files_modified
overlap). Frame detect-backend CLI as a simulation harness; the pure function
with the live host descriptor is the real detection seam.
- partitionStages docstring: 'near-minimal via greedy first-fit' (not 'fewest');
document empty-files_modified behavior.
AGENTS.md is legitimately created by gsd install copilot (issue #786) when
the installer runs inside a repo checkout. The file may exist on disk — that
is expected. The test must verify it is not git-tracked, not that it is absent
from the working tree. Uses git ls-files --error-unmatch instead of fs.existsSync.
- Teach applySurface to build agentCtx (pathPrefix + attribution) and pass it
to kind.stage() for agents kind, mirroring createRuntimeArtifactInstallPlan
(ADR-1235 §1). Surface-path agents now receive path-rewrite + attribution +
converter + normalize, matching install output byte-for-byte.
- Pass skills:'*' sentinel for agents staging when no surface state modifications
exist, so ALL agents are staged (not just those referenced by _calls_agents_).
- Declare converted agents kind in copilot and antigravity capability.json;
add to _DESCRIPTOR_AGENTS_RUNTIMES in bin/install.js.
- Handle copilot .agent.md filename rename in both _copyStaged (install path)
and _syncGsdDir (surface path).
- Ship golden-parity harness (ADR-1235 §0): tests/issue-1575-agent-descriptor-
parity.test.cjs asserts applySurface output is byte-identical to
installRuntimeArtifacts for all 7 descriptor-driven runtimes, plus stale-
cleanup convergence and prune data-loss coverage.
- Update ADR-1235 with cutover progress.
Cline remains deferred (rules-only local branch + local/global complication).
The descriptor-driven gate must treat a runtime as layout-driven when its scoped
artifactLayout is non-empty (any kind), not only when it has a skills kind —
windsurf's global layout is agents-only and is a legitimate layout-driven runtime.
Preserve the three legacy special-cased paths (opencode/kilo combined path,
claude-local copyWithPathReplacement).
zcode falls into the default package.json+hooks path (not in the CommonJS-mode
exclusion roster), so its install contract is packageJson: true — matching
observed install output. No runtime === 'zcode' branch is added (AC#3): zcode
gets the default by not being excluded.
The legacy _isSkillsRuntime gate was a hardcoded isCodex || isCopilot || ...
roster; zcode was absent so the layout-driven skills-install path was skipped
(commands + agents installed, but no skills). Adding || isZcode would violate
AC#3 (no runtime === 'zcode' branches), so the gate is now derived from the
descriptor: a runtime takes the skills path iff its scoped artifactLayout
declares a skills kind. Behavior-preserving for all 15 existing runtimes
(verified by their install contracts); includes zcode (and any future skills
runtime) with zero per-runtime branches. opencode/kilo keep their specialized
combined path.
Regenerate the golden-install-parity fixtures: the non-zcode runtimes shift only
the shared model-catalog.json hash (zcode was added to the catalog); zcode now
includes its skills tree.
Resolve the zcode test cascade exposed by gsd-test:
- model-catalog.json: add zcode runtimeTierDefaults (null tiers, matching the
other profile-marker-only runtimes) so KNOWN_RUNTIMES stays parity with allRuntimes.
- runtime-name-policy: add zcode to GLOBAL_CONFIG_HOME_FRAGMENTS (~/.zcode) so
getGlobalConfigHomeFragment stops falling through to the .claude default.
- installer-migration-install: add the zcode fresh-install contract (flat-skills,
no settings, no package.json — same shape as trae).
- golden-install-parity: capture the zcode fixture via a standalone generator
script (not node --test — the gate stays gsd-test).
- capability-matrix.md: regenerate so zcode appears as a first-party row.
Fluidify the remaining count-pinned guards so adding a runtime no longer trips a
hand-pinned snapshot: gemini-runtime-removed (flag count derives from the
registry), non-claude-runtimes-registry-derivation (golden list derived from the
registry).
Add ZCode as a first-party runtime via a declarative capability descriptor
(capabilities/zcode/capability.json) with zero hardcoded runtime === 'zcode'
branches — exercising the de-hardcoded, data-driven runtime path that 1.7.0
(ADR-1016 / ADR-1239) enables.
Descriptor (all axes sourced verbatim from zcode.z.ai docs):
- configHome ~/.zcode; nested skills + flat commands/agents; profile-marker install
- Claude-shaped skill format reuses convertClaudeCommandToClaudeSkill (no new converter)
- hostIntegration: declarative / slash-file / mcp / electron; dispatch background=false
(foreground-only per docs); nested+maxDepth undocumented; passive model mode
Installer registration (data, not branches): --zcode flag, allRuntimes, runtimeMap,
interactive menu, --all list. getGlobalConfigDir + resolveRuntimeArtifactLayout +
ALLOWED_CONFIG_RUNTIMES + resolveInstallPlan all derive from the descriptor.
Revamped the brittle per-runtime golden-master tests to be count-agnostic,
descriptor-derived property tests (1.7.0 makes runtimes pluggable data, so pinning
frozen '15 runtime' snapshots is the wrong invariant): getdirname, label-policy,
config-adapter-registry (intent + install-plan golden master), capability-registry,
host-integration-descriptors (counts derive from curated maps). Adding a runtime
descriptor now extends coverage with zero edits to those suites.
Welcome banner, --zcode help, supported-runtimes how-to, and the host-integration
capability matrix (every axis cited) updated.
The compiled gsd-core/bin/lib/*.cjs modules are gitignored build artifacts
(ADR-457) shipped prebuilt in the npm tarball. A Claude Code plugin-marketplace
/ git-clone install materializes the repo tree directly and never runs
`npm run build:lib`, so those ~148 files are absent and every CLI command dies
at load with `Cannot find module './lib/cli-exit.cjs'`.
gsd-tools.cjs now calls ensureRuntimeBuild() before its ./lib requires: a new
committed bin helper that compiles the tree once on demand (lock-guarded,
incremental-cache-cleared, portable `node <tsc>` invocation) when the sentinel
cli-exit.cjs is missing, and surfaces an actionable error when TypeScript is
unavailable. The already-built npm path is a single fs.existsSync no-op.
ADR-457 (gitignored, build-at-publish) is preserved; no release-pipeline change.
Verified against a simulated marketplace checkout (148 build outputs removed):
the CLI self-heals and the result is byte-identical to `npm run build:lib`.
Closes#2002
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The #2018 regression test used raw fs.rmSync() in afterEach, tripping the
local/no-raw-rmsync-in-tests eslint rule (Windows-EBUSY retry budget). CI's
lint-tests job (no eslint cache) caught it on next; local --cache runs had
false-greened it. Switch to helpers.cleanup(). Also drop two unused
destructured imports (evaluateCommandExitZero, interpolate) in the
gate-predicate-evaluator test that emitted no-unused-vars warnings.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2012): scope Progress-row regex to ## Progress section (was binding to earlier table)
The Progress-row writer used a non-global regex that matched ANY table row
starting with the phase number. When an earlier table (e.g. Requirements
coverage | Phase | Requirements | Count |) preceded ## Progress, the regex
bound to the wrong row (3-column), no-op'd, and never reached the real
Progress row. roadmap_updated stayed true (it's existsSync), masking the
failure.
- src/phase.cts: scope the tableRowPattern regex to the ## Progress section
(indexOf + slice) so it only matches Progress-table rows.
- tests/phase.test.cjs: regression test — ROADMAP with a phase-numbered
Requirements table before ## Progress → Progress row updated, Requirements
row untouched.
Closes#2012
* docs(#2012): backfill changeset pr 2032
* fix(#2018): applySurface empty manifest no longer deletes gsd-* agents
The agent-prune loop in _syncGsdDir deleted any gsd-*.md not in the staged
set. When the manifest resolved empty (null/non-object, no entries, no files
key, or unresolvable install source root), the staged set was empty → every
gsd-* agent was deleted. Skills were guarded by pruneSkillDirs' manifest-
membership check; agents had no equivalent.
- src/surface.cts: skip the agent-prune loop when the manifest is empty/absent
(copy still runs — genuinely new agents are added).
- tests/surface-empty-manifest-agents.test.cjs: boundary matrix — empty
manifest preserves agents (Map() + undefined); non-empty manifest + empty
staged still prunes (boundary); empty manifest + new staged copies new +
preserves existing.
Closes#2018
* docs(#2018): backfill changeset pr 2031
* fix(#2017: grant mcp__plugin_context7_context7__* for plugin-marketplace context7
The 8 context7-using agents granted only mcp__context7__* (standalone server
form). Claude Code's plugin-marketplace context7 install names tools
mcp__plugin_context7_context7__*, so the grant never matched and every
researcher/planner/executor silently lost doc lookup (fell back to WebSearch).
- 8 agents: add mcp__plugin_context7_context7__* alongside mcp__context7__*.
- scripts/research-profiles.cjs: update the researcher profile tools to match.
- tests/context7-plugin-grant-parity.test.cjs: regression guard — no agent
grants the standalone form without the plugin form.
Closes#2017
* docs(#2017): backfill changeset pr 2029
* fix(#2022): gate roadmap update-plan-progress checkbox on verification passed
cmdRoadmapUpdatePlanProgress stamped the phase checkbox + completion date
the moment all summaries landed, with NO verification gate — unlike
cmdPhaseComplete (phase.cts:1436) which checks readVerificationStatus. Since
update-plan-progress is called after every wave and every plan, the checkbox
fired before gsd-verifier confirmed the phase.
- src/roadmap.cts: isComplete now requires summaryCount >= planCount AND
readVerificationStatus(phaseDir).status === 'passed'.
- tests/roadmap.test.cjs: 2 regression tests (no VERIFICATION.md → not
complete; gaps_found → not complete) + updated 3 existing complete tests
to include a passed VERIFICATION.md.
Closes#2022
* docs(#2022): backfill changeset pr 2030
* fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows
gsd-executor.md referenced sdk/src/query/QUERY-HANDLERS.md and reapply-patches.md
referenced sdk/dist/cli.js — both retired with the SDK (ADR-0174). AI runtimes
that locate doc refs via filesystem search ran find /, which on Git Bash for
Windows traverses the whole drive (14h+, orphaned find.exe, 4M+ handles).
- agents/gsd-executor.md: drop dead QUERY-HANDLERS.md ref.
- workflows/reapply-patches.md: drop dead sdk/dist/cli.js clause.
- tests/no-dead-sdk-refs.test.cjs: regression guard — no sdk/src|dist|handlers
file refs in agents/workflows/references markdown.
Closes#2020
* docs(#2020): backfill changeset pr 2027
* fix(#2019): planning-config.md global-learnings path ~/.gsd/learnings → ~/.gsd/knowledge
The features.global_learnings row pointed at ~/.gsd/learnings/ but the
implementation (src/learnings.cts, execute-phase.md) uses ~/.gsd/knowledge/.
Docs now match the code.
Closes#2019
* docs(#2019): backfill changeset pr 2026
* fix(#1921): preserve verify-work state across gap-closure + defer follow-ups
Resuming /gsd:verify-work after /gsd:execute-phase --gaps-only lost the
verification state: UAT ## Gaps still read status: failed even after their
fix plans executed, so verify-work re-diagnosed them as fresh blockers,
spawned a new gap plan, and reported only the new plan verified. A new
state contract links each gap to its fix plan so fixed gaps are recognized.
- verify-work.md gap YAML: stable gap_id (G-{phase}-{N}) per gap.
- plan_gap_closure planner prompt: each *-PLAN.md tags the gap_ids it
addresses in its frontmatter (gap_closure: true, gap_ids: [...]).
- new reconcile_gaps step (run at resume_from_file entry): marks a gap
status: resolved when its plan has a matching *-SUMMARY.md — fixed gaps
are not re-diagnosed and do not spawn new gap plans; a re-reported break
is treated as a fresh regression with a new gap_id.
- deferred-follow-up branch: a future-work idea (signals: 'later',
'next version', 'out of scope', ...) is captured to UAT ## Deferred
Follow-Ups instead of becoming a blocking gap/plan.
- workflow-size baseline recaptured (verify-work LARGE tier, 38221/61440).
Closes#1921
* docs(#1921): backfill changeset pr 2025
* fix(#1921): balance <step> tags in verify-work.md via 4-backtick outer fence
The plan_gap_closure step nested a ```yaml example inside a ``` block;
same-length fences made stripFencedCode close the outer block early,
swallowing the step's </step> (14 opens / 13 closes). Widen the outer
fence to 4 backticks so the nested yaml is contained. Regenerate golden
fixtures + size baselines.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Wire the two forward-declared mempalace.memory_mode modes so they actually
route recall/capture instead of silently behaving as `augment`:
- kg_backend: the palace temporal KG is the primary knowledge-graph source;
native .planning/graphs/ is the fallback. Non-KG drawer recall stays additive.
- replace: recall resolves through the palace as the source of truth; native
artifacts are the fallback.
Every mode stays onError:skip and default-resilient — an unreachable palace
degrades to native memory and GSD keeps writing .planning/graphs/, so no memory
is lost. Cross-mode .planning/graphs/ migration remains a documented open
question (PRD/ADR §17), out of scope here.
Surfaces updated (instruction-only contract): recall/capture commands (+ generated
skills), discuss/wave fragments, curator agent, capability.json schema. Docs:
how-to Step 3, CONFIGURATION, FEATURES, CONTEXT glossary. Regenerated
capability-registry, golden install-parity fixtures (mempalace hashes only),
agent-size-baseline. Added a routing-contract + cross-surface parity test.
Incidental (folded per no-defer rule): removed pre-existing unused imports
(spawnSync in capability-registry.test.cjs; fs in issue-498-package-identity.test.cjs)
that eslint flagged in/alongside the touched files.
Closes#2007
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1825): configurable graphify graph location (graphify.graph_path)
Add a graphify.graph_path config key (.planning/config.json) that overrides
where /gsd-graphify query|status|diff read the knowledge graph, so one curated
umbrella-level cross-repo graph can serve multiple sibling projects without N
drifting ~5 MB mirror copies. Previously the graph location was hardcoded to
<cwd>/.planning/graphs/.
- src/graphify.cts: resolveGraphLocation(cwd, planningDir) honors the key
(resolved relative to project root; absolute paths honored via path.resolve);
falls back to the historical .planning/graphs/graph.json when unset/blank/
non-string (byte-identical). Wired into graphifyQuery, graphifyStatus,
graphifyDiff (snapshot travels with the configured graph via dirname), and
writeSnapshot. Configured-but-missing -> actionable error naming the path.
Build stays project-scoped (skill hardcodes the cp dest); umbrella graph is
built in the umbrella project, sub-projects only READ it.
- config-schema.manifest.json: register graphify.graph_path in validKeys.
- tests/graphify-graph-path.test.cjs: boundary matrix (unset byte-identical,
set+present reads configured graph not default, set+missing actionable error,
relative resolved vs project root, blank treated as unset, snapshot alongside
configured graph, diff from configured dir, build project-scoped) +
VALID_CONFIG_KEYS registration.
- docs: CONFIGURATION.md row, FEATURES.md REQ-GRAPH-06, CONTEXT.md module note,
.changeset (Added).
Closes#1825
* docs(#1825): backfill changeset pr number 2013
* fix(#1864): balance <step> tags in settings-advanced.md §8 Model Policy
§8 Model Policy ended with </step> but had no matching opening tag (5 opens /
6 closes), leaving it as loose inter-step content. Add the missing
<step name="model_policy"> opener so the section is a proper step.
- gsd-core/workflows/settings-advanced.md: add <step name="model_policy">
- tests/workflow-step-tag-balance.test.cjs: regression guard — every top-level
workflow must have balanced <step>/</step> (fenced code stripped), plus a
focused assertion that §8 is wrapped in model_policy.
- goldens + workflow-size baseline recaptured.
Closes#1864
* docs(#1864): backfill changeset pr 2014
* fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR
The gsd_run preamble resolved the Claude global install only at
$HOME/.claude/gsd-core/bin/, but the installer honors CLAUDE_CONFIG_DIR —
so a global install redirected via CLAUDE_CONFIG_DIR was invisible to every
gsd_run call (every command failed with 'gsd-tools.cjs not found').
The Claude resolver arm now uses ${CLAUDE_CONFIG_DIR:-$HOME/.claude},
matching the installer + the other runtimes' ${VAR:-default} pattern.
Default $HOME/.claude behavior is unchanged.
- _runtime-launcher.snippet.sh: Claude arm honors CLAUDE_CONFIG_DIR.
- sync-runtime-launcher.cjs re-run: 95 workflows/agents re-synced.
- review.md / discuss-phase.md: trimmed to stay under their byte budgets.
- runtime-launcher-parity.test.cjs: (A) substring updated for the new form
+ explicit #1865 assertion that the snippet honors CLAUDE_CONFIG_DIR.
- goldens + size baselines recaptured.
Closes#1865
* docs(#1865): backfill changeset pr 2024
* fix(#1936): reconstruct OpenCode review from JSON events; diagnosable empty-output stub
On a large review prompt, OpenCode's default `build` agent runs a few read
tool calls then ends its turn with zero output tokens (reason:"stop",
output:0), so `opencode run --format default` emits empty stdout. The reviewer
block redirected stderr to /dev/null and wrote a generic "failed or returned
empty output" stub — so the phase silently lost its second independent reviewer
with no diagnostic and no timeout.
Rewrite the OpenCode reviewer block to invoke `--format json` as the primary
call and reconstruct the review from the assistant `text` parts (jq). Capture
stderr to a `.err` sidecar (mirrors the Codex block). When the agent emits no
text, surface the stop reason, output-token count, and stderr so the failure is
diagnosable. Gate the stub on the extracted CONTENT, not the output file size —
an empty jq extraction still prints a lone newline that a `[ -s file ]` check
would treat as populated. Document the wall-clock timeout as a Bash-tool param
(macOS lacks GNU timeout; opencode has no native timeout flag).
review.md was already at the DEFAULT size-tier ceiling (40956/40960), so the
fix cannot fit without reclassifying it into the LARGE tier (it is a
multi-reviewer orchestration file that outgrew "focused single-purpose"; 43.4 KB
sits well under the LARGE high-water mark). Recapture the 16 golden-install
fixtures — the diff is exactly one review.md hash per runtime. Regression block
folded into review-default-reviewers-workflow.test.cjs (new bug-NNNN test files
are not accepted).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1936): add changeset
* test(#1936): property-test the OpenCode review jq reconstruction
Address the re-review's one actionable finding: the jq JSON-event → text
reconstruction had no fast-check property test.
Add tests/opencode-review-reconstruction.property.test.cjs. It extracts the two
shipped jq programs (OPENCODE_REVIEW, OPENCODE_DIAG) verbatim from
gsd-core/workflows/review.md and runs the real jq — not a reimplementation — so
the shipped logic is what gets tested. Properties: the reconstructed review
equals the newline-join of every assistant text part (order preserved); a stream
with no text part reconstructs to empty (drives the #1936 stub); null/absent text
parts are dropped, never rendered as "null". Plus example-based coverage of the
diagnostic edges the reviewer cited: missing .tokens.output and no step_finish
degrade to "?"; non-JSON stdout makes jq fail rather than masquerade as a review.
Verified the invariant has teeth (a comma-join jq fails the property).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1936): skip jq reconstruction property test when jq is absent
The property test shells out to `jq`, which GitHub's windows-latest runners do
not ship (macOS/Linux runners do). `execFileSync('jq')` therefore ENOENT-failed
the whole file on `test (windows-latest, *)`. Probe `jq --version` at load and
skip the suite when jq is not on PATH — the reconstruction logic is
platform-independent, so the assertions still run in full on every jq-present
runner (mirrors how golden-install-parity skips on win32).
Verified: jq present → 7 pass; jq removed from PATH → 7 skipped, 0 fail.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1936): skip jq reconstruction property test on Windows, not just when jq is absent
The prior guard skipped only when `jq` was absent from PATH — but the
windows-latest runners DO ship jq, so the suite still ran there and failed with
`jq: parse error: Invalid numeric literal` (confirmed from the CI job log). Root
cause is Node's child_process argument quoting mangling the jq program (it embeds
double quotes) on Windows, not the shipped review.md logic — the macOS/Linux legs
pass. Gate the suite on `process.platform === 'win32'` (still also skipping when
jq is absent), mirroring golden-install-parity's win32 skip. Logic is
platform-independent and fully asserted on every macOS/Linux CI leg.
Verified: macOS → 7 pass; simulated win32 → skips.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Refinements A-D to the external-job capability (PR #1998 follow-up):
A. Document why the contribution registers at execute:wave:post: #1164 asks
for wave:pre, but execute-phase.md only dispatches wave:post today (wave:pre
is declared in the loop host contract but not rendered). Wiring wave:pre is
a core-loop change #1164 puts out of scope; the executor honors the
runtime_budget classification guidance before running any tagged task.
B. external_job.artifact_dir is now consumed (was declared but unused): the
adapter resolves it via the canonical capability-config seam and surfaces
the resolved root in submit output.
C. external_job.submit_timeout_ms / poll_timeout_ms are now read from config
(were shadowed by env-only reads). Precedence: env > config > registry
default; non-numeric config values fall back (no guessing, no NaN).
D. CLI surface gains unit coverage: parseFlags, findPlanningDir,
resolveExternalJobSettings, formatShowReport.
Regenerates capability-registry.cjs from the updated capability.json.
* fix(#1993): milestone --ws requirements archive header points at workstream path
The requirements archive header hardcoded the root .planning/REQUIREMENTS.md
path, so a workstream (--ws) archive pointed readers at the wrong file even
though #1917 fixed the archive LOCATIONS to land inside the workstream.
Derive the display path from the same workstream-aware reqPath the writer
already uses (path.relative(cwd, reqPath)). Root behavior is byte-identical
('.planning/REQUIREMENTS.md'); the --ws case now correctly reads
'.planning/workstreams/<ws>/REQUIREMENTS.md'.
- src/milestone.cts: reqDisplay interpolation in the archive header.
- tests/milestone.test.cjs: #1993 regression in the #1911 --ws block — header
references the workstream path, not the root literal.
Closes#1993
* docs(#1993): backfill changeset pr 2015
* fix(#1993): use posix separators in archive header (Windows CI) + CRLF-safe test split
- src/milestone.cts: normalize path.relative output to POSIX separators so
the workstream archive header renders forward slashes on Windows too
(path.relative yields backslashes there; the original literal was posix).
- tests/milestone.test.cjs: .split(/\r?\n/) for the CRLF-fragile lint rule.
* fix(#1988): exclude stray non-plan *-SUMMARY.md from phase completion count
Stray remediation/gap-closure summaries (30-FIX-CR02-SUMMARY.md,
30-GAPCLOSURE-SUMMARY.md, …) inflated summary_count, and once
summary_count >= plan_count the phase silently flipped to Complete even
though several plans had no summary. A summary now counts toward completion
only if it pairs with a real plan file.
- core-utils.cts: new countMatchedSummaries(planFiles, summaryFiles) —
layout-agnostic pairing via the PLAN→SUMMARY marker swap (root/nested/bare)
plus the <stem>-SUMMARY.md form (bare PLAN.md↔PLAN-SUMMARY.md); the swap is
applied to the basename only so a 'plans/' dir prefix isn't corrupted.
- plan-scan.cts: scanPhasePlans.summaryCount/.completed use the matched count
(summaryFiles array still holds every summary on disk for listing/reading).
Fixes roadmap listing, state sync, verification, workstream inventory.
- roadmap.cts: cmdRoadmapUpdatePlanProgress uses the matched count.
- tests/roadmap.test.cjs: countMatchedSummaries unit tests (root/nested/bare/
stray) + E2E reproducing the exact #1988 report (4 plans, 1 plan summary,
3 strays → 1/4 In Progress, NOT Complete).
Closes#1988
* docs(#1988): backfill changeset pr 2016
* test(#1988): strengthen countMatchedSummaries unit tests for mutation coverage
Add direct unit tests for the extended (N-PLAN-MM-slug↔N-MM-SUMMARY), bare
(PLAN↔SUMMARY, PLAN↔PLAN-SUMMARY), legacy (N-PLAN-NN↔N-PLAN-NN-SUMMARY), and
stray-exclusion pairings so every branch of countMatchedSummaries is exercised
(Stryker mutation-score coverage).
* test(#1988): move countMatchedSummaries unit tests into core-utils.test.cjs
The Stryker core-utils shard runs ONLY tests/core-utils.test.cjs (per
scripts/mutation-matrix.cjs), so the unit tests for countMatchedSummaries
must live there to be mutation-covered (previously in roadmap.test.cjs, the
shard never ran them → mutants survived → Stryker gate failed). The E2E
#1988 reproduction stays in roadmap.test.cjs. Added an absolute-path case to
guard the lastIndexOf('/') >= 0 boundary.
* fix(#1581): config-set no longer silently coerces Infinity/project_code
The value parser used !isNaN(Number(val)), which admits Infinity/-Infinity;
JSON.stringify then renders those as null on disk while the CLI echoed the
non-finite value (output ≠ disk). Leading-zero strings like project_code
'007' were also silently number-coerced to 7.
- config.cts parser: Number.isFinite instead of !isNaN, so Infinity falls
through to the JSON branch (rejected) and stays a string.
- project_code: always persisted as a string (identifier; leading zeros
matter), bypassing number coercion.
- context_window: new per-key validator — must be a finite positive integer
(rejects Infinity/0/negatives/non-integers with a non-zero exit).
- tests/config.test.cjs: #1581 regression (Infinity rejected, 0 rejected,
200000 accepted finite, project_code '007' string-preserved, granularity
numeric coercion unchanged).
Closes#1581
* docs(#1581): backfill changeset pr 2023
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity
Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap).
--gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1928): backfill changeset PR number (#1996)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1928): drop Gemini CLI from issue templates (review nit)
Removes the sunset Gemini CLI runtime from the two GitHub issue-template
runtime lists that the removal PR missed, per @davesienkowski's review nit:
- feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise
request a feature for a runtime GSD no longer supports)
- bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json
retrieval-help line
Leaves the post-removal templates fully consistent with the Antigravity redirect.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1906): require node-test clean-fixture causation control
The node-test fail-first proof accepted a deceptive content-independent
negative test — one that reds merely because GSD_PROHIB_SUBJECT is set,
ignoring the subject's content — whenever no cleanFixture was supplied,
because #1346's causation control was opt-in. The proof's observed signal
(RED) thus diverged from its target (RED caused by content) by default.
Make the causation control mandatory for the node-test kind: a descriptor
that omits cleanFixture is un-provable (fail-closed), never accepted under
the weaker violation-only proof. When a clean fixture is present, fail-first
is proven exactly as before (RED on violation AND non-vacuous GREEN on clean).
The lint-rule kind is unchanged (its subject IS the linted file; no
GSD_PROHIB_SUBJECT indirection).
Breaking (Hyrum): a previously-green node-test prohibition with no clean
fixture now hard-gates — blast radius is zero in-tree (no node-test
prohibition ships today; only the lint-rule local/no-source-grep dogfood).
Supersedes ADR-1606 Decision 4 / ADR-550 #1346 addendum's opt-in.
Closes#1906
Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ
* docs(#1906): supersede the #1346 opt-in causation control (mandatory for node-test)
Record the node-test mandatory-causation-control supersede across the
governing surfaces:
- ADR-1606 (the enforcement decision-of-record): addendum + Decision 4
annotated + the "Mandatory causation control — REJECTED" alternative
flipped to accepted (premise no longer holds: zero in-tree node-test
consumers).
- ADR-550: the 2026-06-21 #1346 "Why opt-in, not required" paragraph
marked SUPERSEDED, pointing at ADR-1606.
- spec-phase.md: check_clean_fixture is now REQUIRED for node-test
(was "optional").
- CONTEXT.md: PROHIB.enforce.causation predicate updated.
Regenerated the shipped-artifact cascade from the spec-phase.md edit
(+149 B, well under the 40960 cap): 16 golden-install-parity fixtures
and the workflow size baseline.
Refs #1906
Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ
* feat(#1105): add external-job capability (SLURM scheduler-adapter producer half)
The async external-job consumer half (#1165) shipped long ago: the core loop
reads .planning/async-jobs/<job>.json manifests and treats a non-terminal one
as the legal external_job_waiting half-state. The PRODUCER half (#1164) was
the remaining unimplemented piece of #1105.
This adds the producer as a default-off capability:
- capabilities/external-job/ — capability.json (execute:wave:post -> executor,
plan:post -> planner contributions, external_job.* config keys, default-off)
+ fragments teaching runtime-budget classification and externalization.
- src/external-job.cts -> gsd-core/bin/lib/external-job.cjs — pure producer
module: SLURM state -> manifest-status map (no guessing), manifest
build/validate (versioned stability contract), sbatch/squeue/sacct parsers,
and a fail-closed manifest writer (refuses a second non-terminal job for a
plan_id already in flight; refuses to clobber a malformed manifest). fs/clock
seams for deterministic tests.
- scripts/slurm-adapter.cjs — operator CLI (submit/poll/show) wrapping bounded
sbatch/squeue/sacct subprocesses; surfaces manifest commands for confirmation
and never auto-runs them (trust boundary).
- tests/external-job.test.cjs — 23 behavioral + fast-check property tests.
- docs/reference/long-running-operations.md + docs/how-to/async-external-jobs.md.
- CONTEXT.md glossary entry for the External-job Capability.
- Regenerated capability-registry.cjs; pruned the now-stale test-file-count
allowlist entry (external-job is at the 2-file cap).
* chore(#1105): backfill PR number in changeset
* fix(#1105): sync capability artifacts + update registry shape-pin tests
gsd-test caught that adding the external-job capability requires its
dependent artifacts regenerated and its registry-shape drift absorbed:
- sync-manifest-versions: stamp 1.7.0-rc.2 into capability.json (was 1.0.0).
- gen-capability-matrix --write: regenerate docs/reference/capability-matrix.md.
- gen-inventory-manifest --write: regenerate docs/INVENTORY-MANIFEST.json.
- check-gap-analysis-plan-post-e2e: plan:post now has 1 contribution
(external-job planner fragment) instead of 0.
- execute-wave-post-gate-pipeline-e2e: execute:wave:post now has 2
contributions (mempalace + external-job) instead of 1.
* fix(#1105): regenerate capability-registry after version stamp
sync-manifest-versions re-stamped external-job/capability.json from
1.0.0 to 1.7.0-rc.2 after the last registry regeneration, leaving the
committed capability-registry.cjs stale (CI gen-capability-registry
--check failed). gsd-test masked this because its setup runs the full
'npm run build' (which regenerates the registry); CI's 'npm test'
pretest only runs build:lib.
gsd-test's build leg runs the full 'npm run build' (which regenerates
capability-registry.cjs, loop-host-contract.cjs, package-identity.cjs, etc.),
so committed-freshness guards that lived in the unit suite were masked there:
gsd-test passed a stale-commit that CI's shard-1/3 test then red-flagged
(caught live on PR #1998). The mandated pre-push gate was green on a commit
CI correctly flagged.
Move the committed-state --check guards into a new 'lint:generated-sync'
script wired into lint:ci (the single orchestrated entry point the lint-tests
CI job already runs on a build:lib-only tree, so the committed artifacts are
checked without regeneration). gsd-test no longer contains these guards, so
it can no longer mask them.
- package.json: add lint:generated-sync (7 generators --check); wire into lint:ci.
- generate-package-identity.cjs: add --check mode (was the only generator
without it); no-arg behaviour unchanged (still writes, as build expects).
- Remove the committed-freshness guards from the unit suite, keeping all
behavioral/structural tests:
- capability-registry.test.cjs: drop the --check describe.
- loop-host-contract.test.cjs: drop the committed-file staleness test
(keep the normalizeLineEndings unit test).
- capability-matrix-sync.test.cjs: drop --check + byte-for-byte (keep the
architectural content invariants: every cap appears, security ship:pre).
- issue-844-manifest-version-sync.test.cjs: drop describe D (--check).
- issue-498-package-identity.test.cjs: drop the drift-check test (keep
behavioral module-export tests); drop the now-unused render import and
its allow-test-rule exemption (allowlist ratcheted 175 -> 174).
* docs: design spec for /gsd smart-entry command
Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.
* feat: add /gsd-start smart-entry command
State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.
- src/smart-entry.cts: deterministic situation classifier (no-project,
paused, blocked, verify-failed, needs-first-phase, planning, executing,
verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
dispatcher presenting an AskUserQuestion menu (with --text fallback for
non-Claude runtimes) and dispatching to existing commands. Falls back
to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
(markdown-layer invariants + every emitted command resolves to a real
slash command).
Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.
* refactor: rename smart-entry command to /gsd:next
Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.
All affected tests (188) pass; lint:ci clean.
* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)
Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.
- detectSignals now reads total_phases + percent from nested progress{}
first, then scalar fm, then body; current_phase falls back to the
body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
body Phase field) covering verify-pending + executing situations.
Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.
* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)
Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.
Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
layer is down — read .planning/STATE.md directly with the Read tool
and synthesize a minimal situation + actions menu so /gsd:next stays
useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
/gsd:progress as before.
Matches the direct-read resilience the live agent already did by hand.
* docs: add gsd-next skill surface
* chore: trigger no-mistakes validation
* no-mistakes(review): Fix smart-entry phase ordering
* no-mistakes(review): Fix decimal smart-entry phase ordering
* no-mistakes(test): Fix smart-entry next test contracts
* no-mistakes(document): Docs synced for smart entry
* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: shorten next.md description and update golden install parity fixtures
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: trigger no-mistakes validation
* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files
Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.
* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: regenerate golden install parity fixtures for /gsd:next
Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.
* Fix smart-entry verify-failed phase scoping and empty resolve shim step
Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.
* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: recapture all 16 golden fixtures with updated smart-entry.md hash
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: regenerate fixtures + inventory manifest after rebase onto next
Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next
Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).
Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)
The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs
Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo
Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
recommended action, 1-4 unique-id /gsd:* actions (previously the
one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout
Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:
- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
consolidation file doing dozens of real installs — 41s even on a fast Mac
(much worse on the slow Windows I/O path), plus an install-heavy cluster.
Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.
Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This consolidation PR's breadth (28 changed test files) exposed a scoped-test-lane
capacity limit: ci-test-scope pulls the 3–6 min release-tarball-smoke.install.test.cjs
(npm pack + npm install -g, 10MB/1499 files) into the targeted+windows lane whenever
install files change AND when it is itself a changed file — bundling it with the other
27 files overran the 600s per-chunk timeout on windows-latest-24 (deterministic).
release-tarball-smoke has its OWN dedicated workflow (.github/workflows/install-smoke.yml,
triggered on the production install paths), so its scoped-lane run is redundant. Add a
SCOPED_LANE_EXCLUDE guard that drops it from both targeted_tests and windows_tests however
it entered (matched rule OR changed-file), and remove it from the install rule's tests list.
Update the ci-test-scope.test.cjs assertion accordingly. No coverage lost.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
gsd-test surfaced 21 failures:
- 19: folded real-install suites (bug-1834 .sh hooks, enh-2380 --skills-root, fix-1521
install stamping, bug-2136 .sh hook version) spawn install.js and assert side effects,
but their host suites (install-minimal-hooks/install.test/managed-hooks) set
GSD_TEST_MODE=1 at collection time — the install child inherited it and suppressed
the writes. Clear GSD_TEST_MODE in each of those blocks (before/after; standalone had
it unset), so the child performs a real install.
- 2: ci-test-scope A1 used deleted tests/bug-1974-context-exhaustion-record.test.cjs as a
fixture; scopeFor filters nonexistent paths, so it fell back to ['unit']. Repointed to
its consolidation destination tests/perf-317-context-monitor-fs.test.cjs (an existing test).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Fold 51 issue-named CLI black-box + scripts-tooling regression files into their
canonical module suites (runtime-launcher-parity, worktree-safety, install-*, managed-hooks,
read-guard, capability-registry, etc.), plus a NEW slash-command-namespace.test.cjs grouping
the 4 slash/colon-namespace-leak invariant suites that had no canonical owner. Verbatim
block-scoped describe wrappers; 427 subtests conserved 1:1.
Host-env pre-check (per B2): no CLI-receiving host sets a redirecting GSD_WORKSTREAM/GSD_PROJECT
value. One folded suite (bug-3668 runtime resolver) creates an extension-less PATH gsd-tools
stub + bash -c; co-locating it with the host's chmodSync tripped local/no-unguarded-nonportable-exec,
so it's now Windows-guarded (skip on win32) matching the host suite's own bash -c guard.
Regenerates regression-name allowlist (222->182), ratchets file-count allowlist (graphify 7->6,
docs entry removed), makes 26 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456;
prunes stale ids). Repoints 13 tests/ references across CONTEXT.md, COMMANDS.md/FEATURES.md
(EN + ja/ko/pt/zh) and ADR-0002. lint:ci green.
Part of epic #1969. Closes#1975.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Final epic-#1969 batch. Fold 22 issue-named files: the 4 genuine repo-wide invariant
scans (551-eslint-bin-lib-coverage, bug-3054 stale /gsd-next, bug-3810 no-gsd-sdk-runtime-refs,
feat-3593 cli-negative-universal) into a NEW shared repo-invariants.test.cjs; the other 18 as
singletons into their nearest module suite (model-resolver, codex-config, runtime-converters,
security, state-transition, worktree-safety, roadmap-parser, etc.). Verbatim block-scoped
describe wrappers; 334 subtests conserved 1:1.
Host-env pre-check (B2+B6): the 6 CLI folds into GSD_TEST_MODE-setting hosts (model-resolver/
codex-config/runtime-converters) are benign — each origin independently sets GSD_TEST_MODE=1
itself (idempotent), unlike the B6 real-install case.
Regenerates regression-name allowlist (222->213), ratchets file-count allowlist (state 17->16),
makes 7 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes stale ids).
Repoints 2 tests/ refs in docs/TESTING-SUITES.md. lint:ci green.
Part of epic #1969. Closes#1977.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Codex review: the folded bug-2990 block set process.env.GSD_TEST_MODE='1' at
collection time, which persisted into sibling folded suites in agent-frontmatter.test.cjs
(process-isolated when standalone). Scope it to before/after so it no longer leaks.
Assertions unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Fold 27 issue-named installer/hooks/statusline/migration/reapply regression files
into their canonical module suites (installer-migrations, installer-migration-report,
gsd-statusline, reapply-verify-hunks, install-*, gsd-check-update-worker-platform-gate,
etc.). Verbatim block-scoped describe wrappers; 276 subtests conserved 1:1. No new files.
The one subdir origin (tests/installer-migrations/001-legacy-orphan-files) moved up one
level into installer-migrations.test.cjs; its single ../../ module require corrected to
../ so it resolves from tests/ root (verified). Host-env pre-check: no CLI-receiving host
sets a redirecting GSD_WORKSTREAM/GSD_PROJECT value.
Regenerates regression-name allowlist (222->205), ratchets file-count allowlist (verify
11->8, validate entry removed), makes 16 relocated allow-test-rule exemptions issue-ref-
compliant (ADR-456; prunes stale ids). Repoints 15 tests/ references across state-md.md
(EN + ja/ko/pt/zh). lint:ci green.
Part of epic #1969. Closes#1974.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>