Commit Graph

1962 Commits

Author SHA1 Message Date
Jeremy McSpadden
64d9f84cf2 no-mistakes(review): Forward onboard text flag 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
c9d964243a no-mistakes(review): Fix onboard fast-map and skip routing 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
cd5971fd3f no-mistakes(review): Fix onboard skip handoffs 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
2a38ea5331 no-mistakes(review): Harden onboarding projection routing 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
a5298c1fc0 refactor: project onboard routing in init 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
afc3309b61 no-mistakes(review): Forward onboard fast init flag 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
1c992bf57d no-mistakes(review): Stop fast onboard dead-end 2026-07-05 19:16:07 +00:00
Jeremy McSpadden
41015129e1 no-mistakes(review): Derive onboard map status before summary prompt 2026-07-05 19:16:07 +00:00
Jeremy McSpadden
3e0dbf6cf4 no-mistakes(review): Label fast onboard partial maps 2026-07-05 19:16:07 +00:00
Jeremy McSpadden
6f2c2c2fd6 no-mistakes(review): Anchor onboard summary writes 2026-07-05 19:16:07 +00:00
Jeremy McSpadden
b3555b103d no-mistakes(review): fix onboard fast root anchoring 2026-07-05 19:16:07 +00:00
jeremymcs
1797207280 fix(onboard): route onboard under ns-project and drop from core profile
Integrate the brownfield /gsd:onboard skill into the skill subsystems so the
full CI suite passes:

- Route onboard under commands/gsd/ns-project.md (requires + routing row) so it
  nests as gsd-ns-project/skills/onboard on nested-layout runtimes instead of
  leaking as a 7th top-level skill dir (fixes install-nested-layout + issue-69).
- Remove onboard from PROFILES.core (src/install-profiles.cts) so the frozen
  main-loop core stays at 8 skills; onboard remains in standard/full.
- Add the TEXT_MODE plain-text fallback note to gsd-core/workflows/onboard.md
  for non-Claude runtimes (#2012).
- Allowlist onboard.md as a user-invocable skill (enh-2790 ratchet).
- Regenerate docs/INVENTORY-MANIFEST.json, golden-install-parity fixtures, and
  the workflow size baseline to match.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-05 19:16:07 +00:00
jeremymcs
20accf67ac test(onboard): track allow-test-rule ref and recapture golden install fixtures
Fixes the two Tests-workflow failures on the onboard PR.

lint-tests (lint-allow-test-rule-refs): the new
tests/onboard-command.test.cjs carried a `source-text-is-the-product`
allow-test-rule exemption with no tracking ref, tripping the novel-offender
gate. Add `(see #1990)` per ADR-456 so the exemption is traceable.

golden-install-parity: the /gsd:onboard feature adds
gsd-core/workflows/onboard.md and skills/gsd-onboard, and updates
templates/project.md, workflows/do.md, the help modes, and map-codebase.
Recapture the golden fixtures (UPDATE_GOLDEN=1) for all 16 runtimes so the
installed-output manifest matches the intended source changes.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-05 19:16:07 +00:00
Cursor Agent
7b3bf9be3f fix: align new-project map gate and guard onboarding summary overwrite
init new-project now uses the same seven-file codebase map completeness
check as init onboard, so partial .planning/codebase/ directories no longer
skip the brownfield mapping offer after onboarding warns about an incomplete
map.

The onboard workflow now branches on onboarding_summary_exists and asks for
confirmation before regenerating SUMMARY.md on repeat runs.
2026-07-05 19:16:06 +00:00
Jeremy McSpadden
b006b1a23e no-mistakes(review): Fix doc-only onboarding route 2026-07-05 19:16:06 +00:00
Jeremy McSpadden
66ff521c71 no-mistakes(review): Fix onboard runtime and doc detection 2026-07-05 19:16:06 +00:00
Jeremy McSpadden
f29f981486 no-mistakes(review): Fix onboarding docs gate detection 2026-07-05 19:16:06 +00:00
Jeremy McSpadden
6b0b5f1b2c no-mistakes(review): Fix onboard planning detection 2026-07-05 19:16:06 +00:00
Jeremy McSpadden
896c2740d3 feat(#1990): add onboard command for brownfield setup 2026-07-05 19:16:06 +00:00
Tom Boucher
ed79902509 feat(#2007): implement mempalace memory_mode kg_backend and replace routing (#2010)
Wire the two forward-declared mempalace.memory_mode modes so they actually
route recall/capture instead of silently behaving as `augment`:

- kg_backend: the palace temporal KG is the primary knowledge-graph source;
  native .planning/graphs/ is the fallback. Non-KG drawer recall stays additive.
- replace: recall resolves through the palace as the source of truth; native
  artifacts are the fallback.

Every mode stays onError:skip and default-resilient — an unreachable palace
degrades to native memory and GSD keeps writing .planning/graphs/, so no memory
is lost. Cross-mode .planning/graphs/ migration remains a documented open
question (PRD/ADR §17), out of scope here.

Surfaces updated (instruction-only contract): recall/capture commands (+ generated
skills), discuss/wave fragments, curator agent, capability.json schema. Docs:
how-to Step 3, CONFIGURATION, FEATURES, CONTEXT glossary. Regenerated
capability-registry, golden install-parity fixtures (mempalace hashes only),
agent-size-baseline. Added a routing-contract + cross-surface parity test.

Incidental (folded per no-defer rule): removed pre-existing unused imports
(spawnSync in capability-registry.test.cjs; fs in issue-498-package-identity.test.cjs)
that eslint flagged in/alongside the touched files.

Closes #2007

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 15:01:22 -04:00
Tom Boucher
bd77b40107 feat(#1825): configurable graphify graph location (graphify.graph_path) (#2013)
* feat(#1825): configurable graphify graph location (graphify.graph_path)

Add a graphify.graph_path config key (.planning/config.json) that overrides
where /gsd-graphify query|status|diff read the knowledge graph, so one curated
umbrella-level cross-repo graph can serve multiple sibling projects without N
drifting ~5 MB mirror copies. Previously the graph location was hardcoded to
<cwd>/.planning/graphs/.

- src/graphify.cts: resolveGraphLocation(cwd, planningDir) honors the key
  (resolved relative to project root; absolute paths honored via path.resolve);
  falls back to the historical .planning/graphs/graph.json when unset/blank/
  non-string (byte-identical). Wired into graphifyQuery, graphifyStatus,
  graphifyDiff (snapshot travels with the configured graph via dirname), and
  writeSnapshot. Configured-but-missing -> actionable error naming the path.
  Build stays project-scoped (skill hardcodes the cp dest); umbrella graph is
  built in the umbrella project, sub-projects only READ it.
- config-schema.manifest.json: register graphify.graph_path in validKeys.
- tests/graphify-graph-path.test.cjs: boundary matrix (unset byte-identical,
  set+present reads configured graph not default, set+missing actionable error,
  relative resolved vs project root, blank treated as unset, snapshot alongside
  configured graph, diff from configured dir, build project-scoped) +
  VALID_CONFIG_KEYS registration.
- docs: CONFIGURATION.md row, FEATURES.md REQ-GRAPH-06, CONTEXT.md module note,
  .changeset (Added).

Closes #1825

* docs(#1825): backfill changeset pr number 2013
2026-07-05 14:50:01 -04:00
Tom Boucher
ef8a3e27d4 fix(#1864): balance <step> tags in settings-advanced.md §8 Model Policy (#2014)
* fix(#1864): balance <step> tags in settings-advanced.md §8 Model Policy

§8 Model Policy ended with </step> but had no matching opening tag (5 opens /
6 closes), leaving it as loose inter-step content. Add the missing
<step name="model_policy"> opener so the section is a proper step.

- gsd-core/workflows/settings-advanced.md: add <step name="model_policy">
- tests/workflow-step-tag-balance.test.cjs: regression guard — every top-level
  workflow must have balanced <step>/</step> (fenced code stripped), plus a
  focused assertion that §8 is wrapped in model_policy.
- goldens + workflow-size baseline recaptured.

Closes #1864

* docs(#1864): backfill changeset pr 2014
2026-07-05 14:38:24 -04:00
Tom Boucher
a62079b2da fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR (#2024)
* fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR

The gsd_run preamble resolved the Claude global install only at
$HOME/.claude/gsd-core/bin/, but the installer honors CLAUDE_CONFIG_DIR —
so a global install redirected via CLAUDE_CONFIG_DIR was invisible to every
gsd_run call (every command failed with 'gsd-tools.cjs not found').

The Claude resolver arm now uses ${CLAUDE_CONFIG_DIR:-$HOME/.claude},
matching the installer + the other runtimes' ${VAR:-default} pattern.
Default $HOME/.claude behavior is unchanged.

- _runtime-launcher.snippet.sh: Claude arm honors CLAUDE_CONFIG_DIR.
- sync-runtime-launcher.cjs re-run: 95 workflows/agents re-synced.
- review.md / discuss-phase.md: trimmed to stay under their byte budgets.
- runtime-launcher-parity.test.cjs: (A) substring updated for the new form
  + explicit #1865 assertion that the snippet honors CLAUDE_CONFIG_DIR.
- goldens + size baselines recaptured.

Closes #1865

* docs(#1865): backfill changeset pr 2024
2026-07-05 14:20:52 -04:00
Tom Boucher
0656f2e831 fix(#1477): write .gsd-source marker at install for Claude-global (#1487)
fix(#1477): write .gsd-source marker at install for the Claude-global layout
2026-07-05 14:09:17 -04:00
Behruz Nassre Esfahani
23254ca5a7 fix(#1936): reconstruct OpenCode review from JSON events (#1992)
* fix(#1936): reconstruct OpenCode review from JSON events; diagnosable empty-output stub

On a large review prompt, OpenCode's default `build` agent runs a few read
tool calls then ends its turn with zero output tokens (reason:"stop",
output:0), so `opencode run --format default` emits empty stdout. The reviewer
block redirected stderr to /dev/null and wrote a generic "failed or returned
empty output" stub — so the phase silently lost its second independent reviewer
with no diagnostic and no timeout.

Rewrite the OpenCode reviewer block to invoke `--format json` as the primary
call and reconstruct the review from the assistant `text` parts (jq). Capture
stderr to a `.err` sidecar (mirrors the Codex block). When the agent emits no
text, surface the stop reason, output-token count, and stderr so the failure is
diagnosable. Gate the stub on the extracted CONTENT, not the output file size —
an empty jq extraction still prints a lone newline that a `[ -s file ]` check
would treat as populated. Document the wall-clock timeout as a Bash-tool param
(macOS lacks GNU timeout; opencode has no native timeout flag).

review.md was already at the DEFAULT size-tier ceiling (40956/40960), so the
fix cannot fit without reclassifying it into the LARGE tier (it is a
multi-reviewer orchestration file that outgrew "focused single-purpose"; 43.4 KB
sits well under the LARGE high-water mark). Recapture the 16 golden-install
fixtures — the diff is exactly one review.md hash per runtime. Regression block
folded into review-default-reviewers-workflow.test.cjs (new bug-NNNN test files
are not accepted).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1936): add changeset

* test(#1936): property-test the OpenCode review jq reconstruction

Address the re-review's one actionable finding: the jq JSON-event → text
reconstruction had no fast-check property test.

Add tests/opencode-review-reconstruction.property.test.cjs. It extracts the two
shipped jq programs (OPENCODE_REVIEW, OPENCODE_DIAG) verbatim from
gsd-core/workflows/review.md and runs the real jq — not a reimplementation — so
the shipped logic is what gets tested. Properties: the reconstructed review
equals the newline-join of every assistant text part (order preserved); a stream
with no text part reconstructs to empty (drives the #1936 stub); null/absent text
parts are dropped, never rendered as "null". Plus example-based coverage of the
diagnostic edges the reviewer cited: missing .tokens.output and no step_finish
degrade to "?"; non-JSON stdout makes jq fail rather than masquerade as a review.

Verified the invariant has teeth (a comma-join jq fails the property).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1936): skip jq reconstruction property test when jq is absent

The property test shells out to `jq`, which GitHub's windows-latest runners do
not ship (macOS/Linux runners do). `execFileSync('jq')` therefore ENOENT-failed
the whole file on `test (windows-latest, *)`. Probe `jq --version` at load and
skip the suite when jq is not on PATH — the reconstruction logic is
platform-independent, so the assertions still run in full on every jq-present
runner (mirrors how golden-install-parity skips on win32).

Verified: jq present → 7 pass; jq removed from PATH → 7 skipped, 0 fail.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1936): skip jq reconstruction property test on Windows, not just when jq is absent

The prior guard skipped only when `jq` was absent from PATH — but the
windows-latest runners DO ship jq, so the suite still ran there and failed with
`jq: parse error: Invalid numeric literal` (confirmed from the CI job log). Root
cause is Node's child_process argument quoting mangling the jq program (it embeds
double quotes) on Windows, not the shipped review.md logic — the macOS/Linux legs
pass. Gate the suite on `process.platform === 'win32'` (still also skipping when
jq is absent), mirroring golden-install-parity's win32 skip. Logic is
platform-independent and fully asserted on every macOS/Linux CI leg.

Verified: macOS → 7 pass; simulated win32 → skips.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 14:07:07 -04:00
Tom Boucher
8de2ff9121 feat(#2008): generic command-exit-zero gate-predicate evaluator (#2011)
* feat(#2008): add generic command-exit-zero gate-predicate evaluator

Third-party capability gates declared via check.predicate were rendered for
display but never evaluated (only built-in check.query gates fired; the
security capability's gate worked solely via a hard-coded ship.md branch).

Add a generic, deps-injected gate-predicate evaluator (src/gate-predicate-evaluator.cts)
that dispatches by predicate.kind. Built-in kind: command-exit-zero — runs a
bounded sh -c command at the project root (via shell-command-projection.execTool),
inherits env, exit 0 => pass, non-zero => block, timeout => block, fail-closed.

Wire a 'check predicate' subcommand into check-command-router.cts and extend
the three generic workflow gate-dispatch sites (execute:wave:post, execute:post,
plan:post) to route check.predicate gates to the new evaluator. The two-step
gate contract (command-failure => onError; block => halt) is unchanged.

- src/gate-predicate-evaluator.cts: pure leaf, KIND_TABLE extensible
- src/check-command-router.cts: cmdCheckPredicate + buildPredicateDeps + parsePredicateFlags
- docs/adr/2008-*, docs/reference/gate-predicates.md, docs/how-to/command-exit-zero-gate.md
- tests: 38 unit + integration tests (exit mapping, timeout, interpolation,
  property-based bijection, malformed-predicate fail-closed, real subprocess e2e)

Closes #2008

* docs(#2008): backfill changeset pr number 2011
2026-07-05 14:04:29 -04:00
Tom Boucher
97730e59a1 fix(#1164): wire external-job config keys + document wave:post choice (#2006)
Refinements A-D to the external-job capability (PR #1998 follow-up):

A. Document why the contribution registers at execute:wave:post: #1164 asks
   for wave:pre, but execute-phase.md only dispatches wave:post today (wave:pre
   is declared in the loop host contract but not rendered). Wiring wave:pre is
   a core-loop change #1164 puts out of scope; the executor honors the
   runtime_budget classification guidance before running any tagged task.

B. external_job.artifact_dir is now consumed (was declared but unused): the
   adapter resolves it via the canonical capability-config seam and surfaces
   the resolved root in submit output.

C. external_job.submit_timeout_ms / poll_timeout_ms are now read from config
   (were shadowed by env-only reads). Precedence: env > config > registry
   default; non-numeric config values fall back (no guessing, no NaN).

D. CLI surface gains unit coverage: parseFlags, findPlanningDir,
   resolveExternalJobSettings, formatShowReport.

Regenerates capability-registry.cjs from the updated capability.json.
2026-07-05 14:04:00 -04:00
Tom Boucher
e2bfe4dc67 fix(#1993): milestone --ws requirements archive header points at workstream path (#2015)
* fix(#1993): milestone --ws requirements archive header points at workstream path

The requirements archive header hardcoded the root .planning/REQUIREMENTS.md
path, so a workstream (--ws) archive pointed readers at the wrong file even
though #1917 fixed the archive LOCATIONS to land inside the workstream.

Derive the display path from the same workstream-aware reqPath the writer
already uses (path.relative(cwd, reqPath)). Root behavior is byte-identical
('.planning/REQUIREMENTS.md'); the --ws case now correctly reads
'.planning/workstreams/<ws>/REQUIREMENTS.md'.

- src/milestone.cts: reqDisplay interpolation in the archive header.
- tests/milestone.test.cjs: #1993 regression in the #1911 --ws block — header
  references the workstream path, not the root literal.

Closes #1993

* docs(#1993): backfill changeset pr 2015

* fix(#1993): use posix separators in archive header (Windows CI) + CRLF-safe test split

- src/milestone.cts: normalize path.relative output to POSIX separators so
  the workstream archive header renders forward slashes on Windows too
  (path.relative yields backslashes there; the original literal was posix).
- tests/milestone.test.cjs: .split(/\r?\n/) for the CRLF-fragile lint rule.
2026-07-05 14:03:43 -04:00
Tom Boucher
ed31e52b67 fix(#1988): exclude stray non-plan *-SUMMARY.md from phase completion count (#2016)
* fix(#1988): exclude stray non-plan *-SUMMARY.md from phase completion count

Stray remediation/gap-closure summaries (30-FIX-CR02-SUMMARY.md,
30-GAPCLOSURE-SUMMARY.md, …) inflated summary_count, and once
summary_count >= plan_count the phase silently flipped to Complete even
though several plans had no summary. A summary now counts toward completion
only if it pairs with a real plan file.

- core-utils.cts: new countMatchedSummaries(planFiles, summaryFiles) —
  layout-agnostic pairing via the PLAN→SUMMARY marker swap (root/nested/bare)
  plus the <stem>-SUMMARY.md form (bare PLAN.md↔PLAN-SUMMARY.md); the swap is
  applied to the basename only so a 'plans/' dir prefix isn't corrupted.
- plan-scan.cts: scanPhasePlans.summaryCount/.completed use the matched count
  (summaryFiles array still holds every summary on disk for listing/reading).
  Fixes roadmap listing, state sync, verification, workstream inventory.
- roadmap.cts: cmdRoadmapUpdatePlanProgress uses the matched count.
- tests/roadmap.test.cjs: countMatchedSummaries unit tests (root/nested/bare/
  stray) + E2E reproducing the exact #1988 report (4 plans, 1 plan summary,
  3 strays → 1/4 In Progress, NOT Complete).

Closes #1988

* docs(#1988): backfill changeset pr 2016

* test(#1988): strengthen countMatchedSummaries unit tests for mutation coverage

Add direct unit tests for the extended (N-PLAN-MM-slug↔N-MM-SUMMARY), bare
(PLAN↔SUMMARY, PLAN↔PLAN-SUMMARY), legacy (N-PLAN-NN↔N-PLAN-NN-SUMMARY), and
stray-exclusion pairings so every branch of countMatchedSummaries is exercised
(Stryker mutation-score coverage).

* test(#1988): move countMatchedSummaries unit tests into core-utils.test.cjs

The Stryker core-utils shard runs ONLY tests/core-utils.test.cjs (per
scripts/mutation-matrix.cjs), so the unit tests for countMatchedSummaries
must live there to be mutation-covered (previously in roadmap.test.cjs, the
shard never ran them → mutants survived → Stryker gate failed). The E2E
#1988 reproduction stays in roadmap.test.cjs. Added an absolute-path case to
guard the lastIndexOf('/') >= 0 boundary.
2026-07-05 14:03:29 -04:00
Tom Boucher
77aa85007a fix(#1581): config-set no longer silently coerces Infinity / project_code (#2023)
* fix(#1581): config-set no longer silently coerces Infinity/project_code

The value parser used !isNaN(Number(val)), which admits Infinity/-Infinity;
JSON.stringify then renders those as null on disk while the CLI echoed the
non-finite value (output ≠ disk). Leading-zero strings like project_code
'007' were also silently number-coerced to 7.

- config.cts parser: Number.isFinite instead of !isNaN, so Infinity falls
  through to the JSON branch (rejected) and stays a string.
- project_code: always persisted as a string (identifier; leading zeros
  matter), bypassing number coercion.
- context_window: new per-key validator — must be a finite positive integer
  (rejects Infinity/0/negatives/non-integers with a non-zero exit).
- tests/config.test.cjs: #1581 regression (Infinity rejected, 0 rejected,
  200000 accepted finite, project_code '007' string-preserved, granularity
  numeric coercion unchanged).

Closes #1581

* docs(#1581): backfill changeset pr 2023
2026-07-05 14:02:53 -04:00
Tom Boucher
8f2ebbe9bf feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity

Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap).

--gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): backfill changeset PR number (#1996)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): drop Gemini CLI from issue templates (review nit)

Removes the sunset Gemini CLI runtime from the two GitHub issue-template
runtime lists that the removal PR missed, per @davesienkowski's review nit:
- feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise
  request a feature for a runtime GSD no longer supports)
- bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json
  retrieval-help line

Leaves the post-removal templates fully consistent with the Antigravity redirect.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 13:32:51 -04:00
Joe Slitzker
4948ad848e Merge remote-tracking branch 'origin/next' into fix/1477-surface-source-marker
# Conflicts:
#	tests/golden-install-parity.test.cjs
2026-07-04 08:22:49 -05:00
Rezolv
55604e9124 fix(#1906): require node-test clean-fixture causation control (#2001)
* fix(#1906): require node-test clean-fixture causation control

The node-test fail-first proof accepted a deceptive content-independent
negative test — one that reds merely because GSD_PROHIB_SUBJECT is set,
ignoring the subject's content — whenever no cleanFixture was supplied,
because #1346's causation control was opt-in. The proof's observed signal
(RED) thus diverged from its target (RED caused by content) by default.

Make the causation control mandatory for the node-test kind: a descriptor
that omits cleanFixture is un-provable (fail-closed), never accepted under
the weaker violation-only proof. When a clean fixture is present, fail-first
is proven exactly as before (RED on violation AND non-vacuous GREEN on clean).
The lint-rule kind is unchanged (its subject IS the linted file; no
GSD_PROHIB_SUBJECT indirection).

Breaking (Hyrum): a previously-green node-test prohibition with no clean
fixture now hard-gates — blast radius is zero in-tree (no node-test
prohibition ships today; only the lint-rule local/no-source-grep dogfood).

Supersedes ADR-1606 Decision 4 / ADR-550 #1346 addendum's opt-in.

Closes #1906

Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ

* docs(#1906): supersede the #1346 opt-in causation control (mandatory for node-test)

Record the node-test mandatory-causation-control supersede across the
governing surfaces:

- ADR-1606 (the enforcement decision-of-record): addendum + Decision 4
  annotated + the "Mandatory causation control — REJECTED" alternative
  flipped to accepted (premise no longer holds: zero in-tree node-test
  consumers).
- ADR-550: the 2026-06-21 #1346 "Why opt-in, not required" paragraph
  marked SUPERSEDED, pointing at ADR-1606.
- spec-phase.md: check_clean_fixture is now REQUIRED for node-test
  (was "optional").
- CONTEXT.md: PROHIB.enforce.causation predicate updated.

Regenerated the shipped-artifact cascade from the spec-phase.md edit
(+149 B, well under the 40960 cap): 16 golden-install-parity fixtures
and the workflow size baseline.

Refs #1906

Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ
2026-07-03 23:21:55 -04:00
Tom Boucher
5ae4ea4c84 feat(#1105): add external-job capability (SLURM scheduler-adapter producer half) (#1998)
* feat(#1105): add external-job capability (SLURM scheduler-adapter producer half)

The async external-job consumer half (#1165) shipped long ago: the core loop
reads .planning/async-jobs/<job>.json manifests and treats a non-terminal one
as the legal external_job_waiting half-state. The PRODUCER half (#1164) was
the remaining unimplemented piece of #1105.

This adds the producer as a default-off capability:

- capabilities/external-job/ — capability.json (execute:wave:post -> executor,
  plan:post -> planner contributions, external_job.* config keys, default-off)
  + fragments teaching runtime-budget classification and externalization.
- src/external-job.cts -> gsd-core/bin/lib/external-job.cjs — pure producer
  module: SLURM state -> manifest-status map (no guessing), manifest
  build/validate (versioned stability contract), sbatch/squeue/sacct parsers,
  and a fail-closed manifest writer (refuses a second non-terminal job for a
  plan_id already in flight; refuses to clobber a malformed manifest). fs/clock
  seams for deterministic tests.
- scripts/slurm-adapter.cjs — operator CLI (submit/poll/show) wrapping bounded
  sbatch/squeue/sacct subprocesses; surfaces manifest commands for confirmation
  and never auto-runs them (trust boundary).
- tests/external-job.test.cjs — 23 behavioral + fast-check property tests.
- docs/reference/long-running-operations.md + docs/how-to/async-external-jobs.md.
- CONTEXT.md glossary entry for the External-job Capability.
- Regenerated capability-registry.cjs; pruned the now-stale test-file-count
  allowlist entry (external-job is at the 2-file cap).

* chore(#1105): backfill PR number in changeset

* fix(#1105): sync capability artifacts + update registry shape-pin tests

gsd-test caught that adding the external-job capability requires its
dependent artifacts regenerated and its registry-shape drift absorbed:

- sync-manifest-versions: stamp 1.7.0-rc.2 into capability.json (was 1.0.0).
- gen-capability-matrix --write: regenerate docs/reference/capability-matrix.md.
- gen-inventory-manifest --write: regenerate docs/INVENTORY-MANIFEST.json.
- check-gap-analysis-plan-post-e2e: plan:post now has 1 contribution
  (external-job planner fragment) instead of 0.
- execute-wave-post-gate-pipeline-e2e: execute:wave:post now has 2
  contributions (mempalace + external-job) instead of 1.

* fix(#1105): regenerate capability-registry after version stamp

sync-manifest-versions re-stamped external-job/capability.json from
1.0.0 to 1.7.0-rc.2 after the last registry regeneration, leaving the
committed capability-registry.cjs stale (CI gen-capability-registry
--check failed). gsd-test masked this because its setup runs the full
'npm run build' (which regenerates the registry); CI's 'npm test'
pretest only runs build:lib.
2026-07-03 19:37:38 -04:00
Tom Boucher
b0bd2f7a48 chore: move committed-generated-artifact freshness checks to lint:ci (#2000)
gsd-test's build leg runs the full 'npm run build' (which regenerates
capability-registry.cjs, loop-host-contract.cjs, package-identity.cjs, etc.),
so committed-freshness guards that lived in the unit suite were masked there:
gsd-test passed a stale-commit that CI's shard-1/3 test then red-flagged
(caught live on PR #1998). The mandated pre-push gate was green on a commit
CI correctly flagged.

Move the committed-state --check guards into a new 'lint:generated-sync'
script wired into lint:ci (the single orchestrated entry point the lint-tests
CI job already runs on a build:lib-only tree, so the committed artifacts are
checked without regeneration). gsd-test no longer contains these guards, so
it can no longer mask them.

- package.json: add lint:generated-sync (7 generators --check); wire into lint:ci.
- generate-package-identity.cjs: add --check mode (was the only generator
  without it); no-arg behaviour unchanged (still writes, as build expects).
- Remove the committed-freshness guards from the unit suite, keeping all
  behavioral/structural tests:
    - capability-registry.test.cjs: drop the --check describe.
    - loop-host-contract.test.cjs: drop the committed-file staleness test
      (keep the normalizeLineEndings unit test).
    - capability-matrix-sync.test.cjs: drop --check + byte-for-byte (keep the
      architectural content invariants: every cap appears, security ship:pre).
    - issue-844-manifest-version-sync.test.cjs: drop describe D (--check).
    - issue-498-package-identity.test.cjs: drop the drift-check test (keep
      behavioral module-export tests); drop the now-unused render import and
      its allow-test-rule exemption (allowlist ratcheted 175 -> 174).
2026-07-03 19:37:14 -04:00
Jeremy McSpadden
e5ef323b15 feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command

Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.

* feat: add /gsd-start smart-entry command

State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.

- src/smart-entry.cts: deterministic situation classifier (no-project,
  paused, blocked, verify-failed, needs-first-phase, planning, executing,
  verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
  ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire  case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
  dispatcher presenting an AskUserQuestion menu (with --text fallback for
  non-Claude runtimes) and dispatching to existing commands. Falls back
  to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
  situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
  (markdown-layer invariants + every emitted command resolves to a real
  slash command).

Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.

* refactor: rename smart-entry command to /gsd:next

Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.

All affected tests (188) pass; lint:ci clean.

* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)

Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.

- detectSignals now reads total_phases + percent from nested progress{}
  first, then scalar fm, then body; current_phase falls back to the
  body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
  body Phase field) covering verify-pending + executing situations.

Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.

* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)

Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.

Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
   layer is down — read .planning/STATE.md directly with the Read tool
   and synthesize a minimal situation + actions menu so /gsd:next stays
   useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
   /gsd:progress as before.

Matches the direct-read resilience the live agent already did by hand.

* docs: add gsd-next skill surface

* chore: trigger no-mistakes validation

* no-mistakes(review): Fix smart-entry phase ordering

* no-mistakes(review): Fix decimal smart-entry phase ordering

* no-mistakes(test): Fix smart-entry next test contracts

* no-mistakes(document): Docs synced for smart entry

* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: shorten next.md description and update golden install parity fixtures

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: trigger no-mistakes validation

* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files

Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.

* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: regenerate golden install parity fixtures for /gsd:next

Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.

* Fix smart-entry verify-failed phase scoping and empty resolve shim step

Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.

* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: recapture all 16 golden fixtures with updated smart-entry.md hash

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: regenerate fixtures + inventory manifest after rebase onto next

Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next

Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).

Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)

The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs

Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo

Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
  recommended action, 1-4 unique-id /gsd:* actions (previously the
  one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout

Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:

- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
  zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
  kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
  consolidation file doing dozens of real installs — 41s even on a fast Mac
  (much worse on the slow Windows I/O path), plus an install-heavy cluster.

Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.

Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 12:18:25 -04:00
Tom Boucher
241b08fa18 ci(#1975): exclude release-tarball-smoke from scoped lane (fixes Windows chunk timeout)
This consolidation PR's breadth (28 changed test files) exposed a scoped-test-lane
capacity limit: ci-test-scope pulls the 3–6 min release-tarball-smoke.install.test.cjs
(npm pack + npm install -g, 10MB/1499 files) into the targeted+windows lane whenever
install files change AND when it is itself a changed file — bundling it with the other
27 files overran the 600s per-chunk timeout on windows-latest-24 (deterministic).

release-tarball-smoke has its OWN dedicated workflow (.github/workflows/install-smoke.yml,
triggered on the production install paths), so its scoped-lane run is redundant. Add a
SCOPED_LANE_EXCLUDE guard that drops it from both targeted_tests and windows_tests however
it entered (matched rule OR changed-file), and remove it from the install rule's tests list.
Update the ci-test-scope.test.cjs assertion accordingly. No coverage lost.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 10:22:11 -04:00
Tom Boucher
b2ed7940c9 test(#1975): scope GSD_TEST_MODE off in real-install folds; fix ci-test-scope fixture
gsd-test surfaced 21 failures:
- 19: folded real-install suites (bug-1834 .sh hooks, enh-2380 --skills-root, fix-1521
  install stamping, bug-2136 .sh hook version) spawn install.js and assert side effects,
  but their host suites (install-minimal-hooks/install.test/managed-hooks) set
  GSD_TEST_MODE=1 at collection time — the install child inherited it and suppressed
  the writes. Clear GSD_TEST_MODE in each of those blocks (before/after; standalone had
  it unset), so the child performs a real install.
- 2: ci-test-scope A1 used deleted tests/bug-1974-context-exhaustion-record.test.cjs as a
  fixture; scopeFor filters nonexistent paths, so it fell back to ['unit']. Repointed to
  its consolidation destination tests/perf-317-context-monitor-fs.test.cjs (an existing test).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 10:22:11 -04:00
Tom Boucher
6d072435d0 test(#1975): consolidate 51 CLI + scripts-tooling regression tests into module suites
Fold 51 issue-named CLI black-box + scripts-tooling regression files into their
canonical module suites (runtime-launcher-parity, worktree-safety, install-*, managed-hooks,
read-guard, capability-registry, etc.), plus a NEW slash-command-namespace.test.cjs grouping
the 4 slash/colon-namespace-leak invariant suites that had no canonical owner. Verbatim
block-scoped describe wrappers; 427 subtests conserved 1:1.

Host-env pre-check (per B2): no CLI-receiving host sets a redirecting GSD_WORKSTREAM/GSD_PROJECT
value. One folded suite (bug-3668 runtime resolver) creates an extension-less PATH gsd-tools
stub + bash -c; co-locating it with the host's chmodSync tripped local/no-unguarded-nonportable-exec,
so it's now Windows-guarded (skip on win32) matching the host suite's own bash -c guard.

Regenerates regression-name allowlist (222->182), ratchets file-count allowlist (graphify 7->6,
docs entry removed), makes 26 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456;
prunes stale ids). Repoints 13 tests/ references across CONTEXT.md, COMMANDS.md/FEATURES.md
(EN + ja/ko/pt/zh) and ADR-0002. lint:ci green.

Part of epic #1969. Closes #1975.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 10:22:11 -04:00
Tom Boucher
697cbb1f05 test(#1977): consolidate 22 misc + repo-invariant regression tests
Final epic-#1969 batch. Fold 22 issue-named files: the 4 genuine repo-wide invariant
scans (551-eslint-bin-lib-coverage, bug-3054 stale /gsd-next, bug-3810 no-gsd-sdk-runtime-refs,
feat-3593 cli-negative-universal) into a NEW shared repo-invariants.test.cjs; the other 18 as
singletons into their nearest module suite (model-resolver, codex-config, runtime-converters,
security, state-transition, worktree-safety, roadmap-parser, etc.). Verbatim block-scoped
describe wrappers; 334 subtests conserved 1:1.

Host-env pre-check (B2+B6): the 6 CLI folds into GSD_TEST_MODE-setting hosts (model-resolver/
codex-config/runtime-converters) are benign — each origin independently sets GSD_TEST_MODE=1
itself (idempotent), unlike the B6 real-install case.

Regenerates regression-name allowlist (222->213), ratchets file-count allowlist (state 17->16),
makes 7 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes stale ids).
Repoints 2 tests/ refs in docs/TESTING-SUITES.md. lint:ci green.

Part of epic #1969. Closes #1977.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 10:06:11 -04:00
Tom Boucher
799230873f test(#1976): scope GSD_TEST_MODE in folded bug-2990 block (adversarial review)
Codex review: the folded bug-2990 block set process.env.GSD_TEST_MODE='1' at
collection time, which persisted into sibling folded suites in agent-frontmatter.test.cjs
(process-isolated when standalone). Scope it to before/after so it no longer leaks.
Assertions unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 09:51:29 -04:00
Tom Boucher
cb7ca4f3dc test(#1976): consolidate 17 agent regression tests into agent-contract suites
Fold 17 issue-named agent-markdown regression files into their canonical agent
suites: agent-frontmatter (shared agent-body-contract catch-all: read-loop guards,
retired-slash-ref bans, sycophancy hardening, doc-writer/code-fixer/review-fix/mapper
contracts, write-truncation), planner-language-regression (planner directive/grep/phase
contract), executor-mvp-tdd-section, verifier-behavior-unverified, intel, and
research-agent-profiles. Verbatim block-scoped describe wrappers; 98 subtests conserved
1:1. All unit (markdown-content) tests — no CLI/env surface. No new files.

Regenerates regression-name allowlist (222->207), ratchets file-count allowlist (intel
entry removed, phase 4->3), makes 13 relocated allow-test-rule exemptions issue-ref-
compliant (ADR-456; prunes stale ids). No CONTEXT.md/ADR refs named the moved files.
lint:ci green.

Part of epic #1969. Closes #1976.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 09:51:29 -04:00
Tom Boucher
0cc7a1a426 test(#1974): consolidate 27 installer/hooks remainder tests into module suites
Fold 27 issue-named installer/hooks/statusline/migration/reapply regression files
into their canonical module suites (installer-migrations, installer-migration-report,
gsd-statusline, reapply-verify-hunks, install-*, gsd-check-update-worker-platform-gate,
etc.). Verbatim block-scoped describe wrappers; 276 subtests conserved 1:1. No new files.

The one subdir origin (tests/installer-migrations/001-legacy-orphan-files) moved up one
level into installer-migrations.test.cjs; its single ../../ module require corrected to
../ so it resolves from tests/ root (verified). Host-env pre-check: no CLI-receiving host
sets a redirecting GSD_WORKSTREAM/GSD_PROJECT value.

Regenerates regression-name allowlist (222->205), ratchets file-count allowlist (verify
11->8, validate entry removed), makes 16 relocated allow-test-rule exemptions issue-ref-
compliant (ADR-456; prunes stale ids). Repoints 15 tests/ references across state-md.md
(EN + ja/ko/pt/zh). lint:ci green.

Part of epic #1969. Closes #1974.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 09:34:39 -04:00
Tom Boucher
2c32e8d890 test(#1973): consolidate 48 workflow regression tests into workflow-aspect suites
Fold 48 issue-named workflow-markdown regression files into the canonical test
that owns each workflow aspect (execute-phase-*, plan-phase-*, quick-*, discuss-*,
worktree-cleanup, secure-phase, verify, update, settings, etc.), across 32 existing
suites. Verbatim block-scoped describe wrappers; 281 subtests conserved 1:1. No new
test files. Host-env pre-check: only gsd-settings-advanced spawns CLI and it sets no
GSD_WORKSTREAM/GSD_PROJECT value — no leak risk.

Regenerates regression-name allowlist (222->181), ratchets file-count allowlist
(verify 11->10), makes 30 relocated allow-test-rule exemptions issue-ref-compliant
(ADR-456; prunes 30 stale ids). lint:ci green.

Part of epic #1969. Closes #1973.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 09:17:32 -04:00
Tom Boucher
85ed50cc4f test(#1972): consolidate 94 command/module regression tests into subject suites
Fold 94 issue-named command/module regression files into the canonical test file
that owns each subject-under-test, across 52 existing suites (state, config, frontmatter,
roadmap-parser, capability-registry, shell-command-projection-dispatch, plan-phase-drift-guard,
health-validation, runtime-converters, commands, etc.). Verbatim block-scoped describe
wrappers; 881 subtests conserved 1:1. No new test files.

Host-env pre-check (per B2): the only GSD_WORKSTREAM/GSD_PROJECT-touching destinations
(intel, planning-workspace) clear those vars hermetically, so folded CLI tests are safe.

Regenerates regression-name allowlist (222->162), ratchets file-count allowlist across
8 buckets (validate entry removed after dropping <=2), makes 34 relocated allow-test-rule
exemptions issue-ref-compliant (ADR-456; prunes 34 stale ids). Repoints CONTEXT.md +
ADR-0002/443/1235/3524 test-file references. lint:ci green.

Part of epic #1969. Closes #1972.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 08:59:23 -04:00
Tom Boucher
7e2f74e9a9 test(#1971): isolate folded CLI blocks from host ambient GSD_ env
gsd-test surfaced 9 failures: folded fix-1437 (phase.list-plans) and bug-1826
(phases clear) tests spawn gsd-tools via runGsdTools, which copies process.env into
the child. Their host suites (phase-command-router / phases-command-router) force
GSD_WORKSTREAM=test-unit at the suite level, redirecting the child's project lookup
away from each test's temp project → plan/dir counts came back 0. Clear GSD_WORKSTREAM
in each folded block's beforeEach (restoring the standalone condition) and restore in
afterEach. Also scope the folded bug-416 GSD_TEST_MODE set to a before/after hook in
health-validation (Codex review) so it no longer leaks into host child processes.
Assertions unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 01:55:48 -04:00
Tom Boucher
de3ba45d00 test(#1971): consolidate 48 gsd-tools CLI regression tests into subcommand suites
Fold 48 issue-named gsd-tools CLI regression files into the canonical test file
that owns each subcommand subject (state, roadmap, phase, milestone, audit, config,
router/dispatch, stats, verify, health, etc.), preserving every assertion and its
origin issue number as provenance (block-scoped describe wrappers, 299 subtests
conserved 1:1). No monolithic gsd-tools.test.cjs created — routes into 18 existing
per-subject suites.

Removes 48 tests/ files. Regenerates regression-name allowlist (271->231), ratchets
the file-count allowlist across 6 buckets (audit/milestone/phase/roadmap/state/verify),
and makes 10 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes
10 stale ids). Repoints one CONTEXT.md symptom ref and ADR-3524's parity-test ref.
lint:ci green.

Part of epic #1969. Closes #1971.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 01:55:48 -04:00
Tom Boucher
c0a8f383db test(#1970): scope HOME/USERPROFILE in folded bug-130/bug-410 blocks
Adversarial review (Codex) surfaced a latent cross-suite leak: the folded
bug-130 and bug-410 install.test.cjs blocks set process.env.HOME/USERPROFILE to a
temp home at collection time and never restored — harmless when each ran as its own
process, but after consolidation it leaked into sibling folded suites in the same
process. Scope both mutations to before()/after() hooks that restore the prior
values, matching the save/restore pattern used elsewhere in the file. Assertions
unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 01:00:09 -04:00
Tom Boucher
4f779eda43 test(#1970): consolidate 61 install-suite regression tests into function suites
Fold 61 issue-named install/codex/runtime regression files into the canonical
test file that owns each subject-under-test, preserving every assertion and its
origin issue number as provenance (block-scoped describe wrappers, zero assertion
loss — 652 subtests conserved 1:1). Routes:

- codex-config.test.cjs        +19 (codex config/toml/hooks/adapter/skill surface)
- install.test.cjs             +18 (node-runner norm, manifest, arg parse, finishInstall)
- install-runtime-artifacts    +13 (per-runtime conversion + emission)
- install-minimal-hooks        +6  (hook-event dialects + guards)
- path-replacement             +2  (opencode absolute pathPrefix)
- install-write-confinement    +2  (pristine dir writes)
- install-regressions          +1  (user-artifact preservation)

Removes 61 tests/ files → 61 fewer CI processes. Regenerates the regression-name
allowlist (271→222), ratchets the file-count allowlist (config 10→9, install 12→9),
and makes 13 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456;
prunes 13 stale allowlist ids). Repoints ADR-0009's moved-test list. lint:ci green.

Part of epic #1969. Closes #1970.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 00:39:32 -04:00
Behruz Nassre Esfahani
8189d2f098 enhance(#1872): document Claude Code advisor inheritance in model profiles (#1922)
* enhance(#1872): document Claude Code advisor inheritance in model profiles

Add an "Advisor Tool (Claude Code)" section to
gsd-core/references/model-profiles.md: session-level advisor is inherited
by all GSD subagents and composes with the per-agent profile/tier system,
candidate executor/advisor pairings per profile (cost/quality/caching
claims attributed to Anthropic's advisor-tool docs, not asserted as GSD
behavior), when it is worth enabling vs not, and the session-level /
no-per-agent-control constraint linking anthropics/claude-code#73072.

Docs-only. Golden-install-parity fixtures recaptured for the edited
reference file (hash-only, one line per runtime).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1872): add changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1872): use Documentation changeset type for docs-only change

Changeset type was `Changed`, which triggers the docs-required lint
(TRIGGERING_TYPES in scripts/lint-docs-required.cjs). This PR only
touches gsd-core/references/model-profiles.md, so there is no docs/
file to pair with and docs-lint failed. `Documentation` is the correct
type for a docs-only enhancement and is exempt from the trigger.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1872): put docs-exempt marker on its own line, revert to type Changed

The prior fix (type: Documentation) was invalid — parse.cjs ALLOWED_TYPES
is {Added, Changed, Deprecated, Removed, Fixed, Security}, so both
changeset-lint and docs-lint failed with invalid_type.

Real root cause of the original docs-lint failure: DOCS_EXEMPT_RE is
anchored to match the `<!-- docs-exempt: ... -->` marker only on its own
line, but the marker was tacked onto the end of the prose line, so it was
never captured (docsExempt: null) and the triggering `Changed` fragment
had no docs/ pairing -> fail_docs_missing.

Fix: keep the valid `type: Changed` and move the marker to its own line.
Verified locally: changeset-lint -> ok_fragment_present,
docs-lint -> ok (own-line marker parses to ok_fragments_exempt).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-02 23:41:39 -04:00