Commit Graph

98 Commits

Author SHA1 Message Date
Tom Boucher
185abe2d66 feat(codex): advance Codex/OpenAI model defaults to GPT-5.6 (Sol/Terra/Luna)
Update runtimeTierDefaults.codex and providerPresets.openai in
model-catalog.json to the GPT-5.6 family (gpt-5.6-sol/terra/luna),
advancing from the superseded GPT-5.4/5.5 generation.

Model IDs verified against OpenAI developer API docs:
- gpt-5.6-sol: flagship, /, reasoning xhigh
- gpt-5.6-terra: balanced, .50/, reasoning medium
- gpt-5.6-luna: fast/cheap, /, reasoning medium

Tier mapping is 1:1 (Sol↔flagship, Terra↔balanced, Luna↔fast),
so profile semantics are unchanged — only the underlying IDs advance.

Updates: catalog JSON, test assertions (catalog defaults), docs
(CONFIGURATION.md + zh-CN/pt-BR translations, workflow settings),
and changeset.

Closes #2122
2026-07-10 12:11:55 -04:00
Tom Boucher
dbc730d8de fix(#2073): capability-probe external killer (timeout/gtimeout) for macOS
Code review (HIGH): a hardcoded 'timeout 600 agy' fails with rc 127 on stock
macOS (no GNU timeout/gtimeout), silently losing the agy reviewer. Probe for
'timeout'/'gtimeout' via command -v and fall back to agy's native --print-timeout
alone when neither exists (mirrors scripts/base64-scan.sh). External cap (600s)
stays >= --print-timeout (540s) so it only backstops a pre-session stall. Factor
the prompt into _AGY_PROMPT to avoid duplicating the long -p string across both
branches. Update the agy + #687 tests to assert the probe + bound + fallback,
regen the 17 goldens + size baseline, refresh the maintainer-note version stamp
to 1.0.16.
2026-07-08 17:53:28 -04:00
Tom Boucher
7abe6e34ac chore(#2073): regen workflow-size baseline for review.md growth
The agy block grew (~file-reference prompt instruction, external-timeout
rationale, --model wiring, richer Step 3 diagnostic) — all load-bearing
content fixing 3 production failure modes (#2073), not bloat. Growth
justified in the PR.
2026-07-08 16:47:02 -04:00
Tom Boucher
1ee00320c7 Merge branch 'next' into fix/2072-thread-model-into-assumptions-and-review-spawns 2026-07-07 22:03:47 -04:00
Tom Boucher
4483300253 fix(#2072): thread resolved model into routed-agent spawns (assumptions-analyzer, code-reviewer, code-fixer)
model_overrides / models.<phaseType> were silently inert for gsd-assumptions-analyzer,
gsd-code-reviewer, and gsd-code-fixer on Claude Code: resolveModelInternal honors them,
but the workflows spawned these agents with no model= param, so the resolved value
never reached the Agent tool and the agents inherited the session model — no warning.

Fix — thread each agent's resolved model at every spawn site (the established
plan-phase pattern; the architecture-consistent Claude mechanism, since 13 other
agents already thread their model):
- discuss-phase-assumptions.md: `resolve-model gsd-assumptions-analyzer --raw`
  → ANALYZER_MODEL, threaded.
- code-review.md + code-review-fix.md (re-review): `resolve-model gsd-code-reviewer --raw`
  → REVIEWER_MODEL, threaded.
- code-review-fix.md (both fixer spawns): `resolve-model gsd-code-fixer --raw`
  → FIXER_MODEL, threaded (same silently-inert bug, same file — folded in per review).
- quick.md review step: was reusing `{executor_model}` for gsd-code-reviewer (so the
  reviewer's own override was ignored); init.quick now resolves `reviewer_model`
  (gsd-code-reviewer) and the spawn threads it.

resolve-model --raw returns the bare model string (resolve-execution --raw would
return effort — wrong). The resolver maps these agents to phaseType discuss /
verification / execution, so models.<phaseType> apply too.

Scope: the three agents reachable from the two issue-named workflows + quick.md. The
wider systemic class (other agents in UNTOUCHED workflows with the same pattern) stays
documented on the issue for a maintainer-scoped structural decision (thread-at-source
vs embed-at-install like #2256), not widened here.

Docs: the stale "discuss — reserved, no subagent today" model-profile tables now list
gsd-assumptions-analyzer and the verification row includes gsd-code-reviewer, across
the English docs, the shipped gsd-core/references/model-profiles.md reference, and the
ja-JP / zh-CN / ko-KR / pt-BR locale mirrors.

Tests:
- tests/model-resolver.test.cjs: #2072 acceptance — model_overrides and
  models.discuss/verification/execution resolve for all three agents.
- tests/model-routing-spawn-threading.test.cjs: every spawn of the three agents threads
  a resolved model (fails pre-fix); a header-precise parity guard fails the suite if a
  new un-threaded spawn of any of them regresses.
All 16 golden-install-parity fixtures + the workflow size baseline regenerated for the
changed shipped files (4 workflows + the reference doc); bin/lib is excluded from parity.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 21:22:48 -04:00
Tom Boucher
6addeccd19 feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates
an external API/SDK/service can no longer seal without a decided coverage matrix.

- src/api-coverage.cts: deterministic detector (compound verb+noun signal +
  <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix
  parse/validate/render with field-length caps.
- check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as
  a token under .planning/phases/ only (traversal-neutralized); validates
  COVERAGE.md or blocks iff a strong integration signal is detected and no
  matrix exists; fail-closed when phases tree exists but phase unresolvable.
- capabilities/ai-integration: workflow.api_coverage_gate config key (default
  true), plan:pre contribution, blocking verify:pre gate. Data-driven.
- gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch.
- Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e.
  Code+security review findings fixed (stopword FP, scope containment, pipe/cap
  rejection, prompt-injection message hygiene).
- Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs.

Closes #1562
2026-07-07 15:11:12 -04:00
Tom Boucher
603593d41d fix(#1857): test gates normalize to one-shot + bounded timeout (no watch-mode hang)
A GSD verification gate resolves a project's test command and runs it. vitest
defaults to WATCH mode in an interactive TTY — exactly where a user runs
`gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never
exited and the orchestrator waited indefinitely. Recovery needed the user to
manually prompt "something blocking?".

Fix — one shared helper + a bounded, surfacing timeout on the test-command gates:
- New pure module src/normalize-test-command.cts + `gsd-tools query
  normalize-test-command` verb: rewrites a resolved command to a best-effort
  one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`;
  a package-manager `test` script whose package.json runner is watch-vitest →
  `CI=true` prefix; handles `--dir`; already-one-shot commands unchanged —
  never double-flagged). Named `normalize-test-command` (not `test-*`) so the
  file does not match node --test's default `test-*` discovery glob.
- The three gates that HUNG or silently-continued route through that ONE helper
  and bound execution with `timeout $(config-get workflow.test_gate_timeout)`
  (new config key, default 600s): the regression gate (extracted to
  execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen —
  it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the
  audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124.
- verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so
  it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint
  on 124, staying under its frozen 40960-byte tier cap.

Security hardening (review): the normalizer only rewrites a runner named as a
standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never
mangled), is length-capped and uses only linear-time split-based scanning (no
super-linear backtracking on an adversarial `workflow.test_command`), and reads
package.json only when it is a regular file (never blocks on a FIFO via `--dir`).

Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the
schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors
workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores,
inventory manifest/index. All 16 golden-install-parity fixtures + workflow size
baseline regenerated for the changed shipped files; bin/lib is excluded from the
parity manifest.

Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security
hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route
through the shared helper + configured timeout + exit-124 watch-mode hint;
verify-phase asserted as normalize-only/already-bounded).
tests/execute-phase-active-flags.test.cjs repointed at the extracted step;
tests/planner-language-regression.test.cjs allowlist comment updated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 12:14:43 -04:00
Tom Boucher
ea4378063f Merge branch 'next' into feat/1820-specless-predicate-rail 2026-07-07 07:49:07 -04:00
Tom Boucher
d7129222c0 Merge branch 'next' into codex/gsd-onboard 2026-07-06 23:49:49 -04:00
Tom Boucher
080bacdb4b Merge branch 'next' into codex/gsd-onboard 2026-07-06 22:43:32 -04:00
Dave
91998dcbca chore(#1820): regen goldens + workflow size baseline after rebase onto next 2026-07-06 14:41:00 -04:00
Dave
7ef834cabc feat(#1820): spec-optional predicate rail — author probe predicates into must_haves when SPEC omits them 2026-07-06 14:40:38 -04:00
Joe Slitzker
7d4fc3a519 test(#1941): regenerate golden fixtures + size baseline after rebase onto next
The rebase onto origin/next pulled in the runtime-launcher preamble resync
(applied repo-wide on next) alongside this branch's quick.md change; both
together shift every runtime's install hashes and workflow sizes, so the
fixtures from the pre-rebase regen were stale.
2026-07-06 13:01:57 -05:00
Joe Slitzker
3fe9a81428 fix(#1941): degrade /gsd-quick worktree dispatch when fork base is stale
Claude Code's isolation="worktree" forks new worktrees from origin/HEAD, not
the live local HEAD. When prior local commits (e.g. an earlier quick task in
the same session, or this task's own Step 5.6 pre-dispatch plan commit)
advance local HEAD without an intervening push, origin/HEAD stays pinned to a
stale ancestor and the executor's worktree_branch_check guard halts with a
base-mismatch fatal that can be many commits behind, not just one.

Port the worktree.base-check auto-degrade pattern already used by
execute-phase (#683/#1369) into quick.md's single-dispatch path, run
immediately before EXPECTED_BASE is captured in Step 6.
2026-07-06 13:01:47 -05:00
Tom Boucher
9f0d785b61 fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows (#2027)
* fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows

gsd-executor.md referenced sdk/src/query/QUERY-HANDLERS.md and reapply-patches.md
referenced sdk/dist/cli.js — both retired with the SDK (ADR-0174). AI runtimes
that locate doc refs via filesystem search ran find /, which on Git Bash for
Windows traverses the whole drive (14h+, orphaned find.exe, 4M+ handles).

- agents/gsd-executor.md: drop dead QUERY-HANDLERS.md ref.
- workflows/reapply-patches.md: drop dead sdk/dist/cli.js clause.
- tests/no-dead-sdk-refs.test.cjs: regression guard — no sdk/src|dist|handlers
  file refs in agents/workflows/references markdown.

Closes #2020

* docs(#2020): backfill changeset pr 2027
2026-07-05 16:43:26 -04:00
Tom Boucher
1bfadec2d0 fix(#1921): preserve verify-work state across gap-closure + defer follow-ups (#2025)
* fix(#1921): preserve verify-work state across gap-closure + defer follow-ups

Resuming /gsd:verify-work after /gsd:execute-phase --gaps-only lost the
verification state: UAT ## Gaps still read status: failed even after their
fix plans executed, so verify-work re-diagnosed them as fresh blockers,
spawned a new gap plan, and reported only the new plan verified. A new
state contract links each gap to its fix plan so fixed gaps are recognized.

- verify-work.md gap YAML: stable gap_id (G-{phase}-{N}) per gap.
- plan_gap_closure planner prompt: each *-PLAN.md tags the gap_ids it
  addresses in its frontmatter (gap_closure: true, gap_ids: [...]).
- new reconcile_gaps step (run at resume_from_file entry): marks a gap
  status: resolved when its plan has a matching *-SUMMARY.md — fixed gaps
  are not re-diagnosed and do not spawn new gap plans; a re-reported break
  is treated as a fresh regression with a new gap_id.
- deferred-follow-up branch: a future-work idea (signals: 'later',
  'next version', 'out of scope', ...) is captured to UAT ## Deferred
  Follow-Ups instead of becoming a blocking gap/plan.
- workflow-size baseline recaptured (verify-work LARGE tier, 38221/61440).

Closes #1921

* docs(#1921): backfill changeset pr 2025

* fix(#1921): balance <step> tags in verify-work.md via 4-backtick outer fence

The plan_gap_closure step nested a ```yaml example inside a ``` block;
same-length fences made stripFencedCode close the outer block early,
swallowing the step's </step> (14 opens / 13 closes). Widen the outer
fence to 4 backticks so the nested yaml is contained. Regenerate golden
fixtures + size baselines.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 15:42:10 -04:00
Codesmith
a979cfd3f4 chore(#1990): regenerate golden fixtures and size baseline after rebase onto next
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-05 19:18:28 +00:00
Tom Boucher
ef8a3e27d4 fix(#1864): balance <step> tags in settings-advanced.md §8 Model Policy (#2014)
* fix(#1864): balance <step> tags in settings-advanced.md §8 Model Policy

§8 Model Policy ended with </step> but had no matching opening tag (5 opens /
6 closes), leaving it as loose inter-step content. Add the missing
<step name="model_policy"> opener so the section is a proper step.

- gsd-core/workflows/settings-advanced.md: add <step name="model_policy">
- tests/workflow-step-tag-balance.test.cjs: regression guard — every top-level
  workflow must have balanced <step>/</step> (fenced code stripped), plus a
  focused assertion that §8 is wrapped in model_policy.
- goldens + workflow-size baseline recaptured.

Closes #1864

* docs(#1864): backfill changeset pr 2014
2026-07-05 14:38:24 -04:00
Tom Boucher
a62079b2da fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR (#2024)
* fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR

The gsd_run preamble resolved the Claude global install only at
$HOME/.claude/gsd-core/bin/, but the installer honors CLAUDE_CONFIG_DIR —
so a global install redirected via CLAUDE_CONFIG_DIR was invisible to every
gsd_run call (every command failed with 'gsd-tools.cjs not found').

The Claude resolver arm now uses ${CLAUDE_CONFIG_DIR:-$HOME/.claude},
matching the installer + the other runtimes' ${VAR:-default} pattern.
Default $HOME/.claude behavior is unchanged.

- _runtime-launcher.snippet.sh: Claude arm honors CLAUDE_CONFIG_DIR.
- sync-runtime-launcher.cjs re-run: 95 workflows/agents re-synced.
- review.md / discuss-phase.md: trimmed to stay under their byte budgets.
- runtime-launcher-parity.test.cjs: (A) substring updated for the new form
  + explicit #1865 assertion that the snippet honors CLAUDE_CONFIG_DIR.
- goldens + size baselines recaptured.

Closes #1865

* docs(#1865): backfill changeset pr 2024
2026-07-05 14:20:52 -04:00
Behruz Nassre Esfahani
23254ca5a7 fix(#1936): reconstruct OpenCode review from JSON events (#1992)
* fix(#1936): reconstruct OpenCode review from JSON events; diagnosable empty-output stub

On a large review prompt, OpenCode's default `build` agent runs a few read
tool calls then ends its turn with zero output tokens (reason:"stop",
output:0), so `opencode run --format default` emits empty stdout. The reviewer
block redirected stderr to /dev/null and wrote a generic "failed or returned
empty output" stub — so the phase silently lost its second independent reviewer
with no diagnostic and no timeout.

Rewrite the OpenCode reviewer block to invoke `--format json` as the primary
call and reconstruct the review from the assistant `text` parts (jq). Capture
stderr to a `.err` sidecar (mirrors the Codex block). When the agent emits no
text, surface the stop reason, output-token count, and stderr so the failure is
diagnosable. Gate the stub on the extracted CONTENT, not the output file size —
an empty jq extraction still prints a lone newline that a `[ -s file ]` check
would treat as populated. Document the wall-clock timeout as a Bash-tool param
(macOS lacks GNU timeout; opencode has no native timeout flag).

review.md was already at the DEFAULT size-tier ceiling (40956/40960), so the
fix cannot fit without reclassifying it into the LARGE tier (it is a
multi-reviewer orchestration file that outgrew "focused single-purpose"; 43.4 KB
sits well under the LARGE high-water mark). Recapture the 16 golden-install
fixtures — the diff is exactly one review.md hash per runtime. Regression block
folded into review-default-reviewers-workflow.test.cjs (new bug-NNNN test files
are not accepted).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1936): add changeset

* test(#1936): property-test the OpenCode review jq reconstruction

Address the re-review's one actionable finding: the jq JSON-event → text
reconstruction had no fast-check property test.

Add tests/opencode-review-reconstruction.property.test.cjs. It extracts the two
shipped jq programs (OPENCODE_REVIEW, OPENCODE_DIAG) verbatim from
gsd-core/workflows/review.md and runs the real jq — not a reimplementation — so
the shipped logic is what gets tested. Properties: the reconstructed review
equals the newline-join of every assistant text part (order preserved); a stream
with no text part reconstructs to empty (drives the #1936 stub); null/absent text
parts are dropped, never rendered as "null". Plus example-based coverage of the
diagnostic edges the reviewer cited: missing .tokens.output and no step_finish
degrade to "?"; non-JSON stdout makes jq fail rather than masquerade as a review.

Verified the invariant has teeth (a comma-join jq fails the property).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1936): skip jq reconstruction property test when jq is absent

The property test shells out to `jq`, which GitHub's windows-latest runners do
not ship (macOS/Linux runners do). `execFileSync('jq')` therefore ENOENT-failed
the whole file on `test (windows-latest, *)`. Probe `jq --version` at load and
skip the suite when jq is not on PATH — the reconstruction logic is
platform-independent, so the assertions still run in full on every jq-present
runner (mirrors how golden-install-parity skips on win32).

Verified: jq present → 7 pass; jq removed from PATH → 7 skipped, 0 fail.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1936): skip jq reconstruction property test on Windows, not just when jq is absent

The prior guard skipped only when `jq` was absent from PATH — but the
windows-latest runners DO ship jq, so the suite still ran there and failed with
`jq: parse error: Invalid numeric literal` (confirmed from the CI job log). Root
cause is Node's child_process argument quoting mangling the jq program (it embeds
double quotes) on Windows, not the shipped review.md logic — the macOS/Linux legs
pass. Gate the suite on `process.platform === 'win32'` (still also skipping when
jq is absent), mirroring golden-install-parity's win32 skip. Logic is
platform-independent and fully asserted on every macOS/Linux CI leg.

Verified: macOS → 7 pass; simulated win32 → skips.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 14:07:07 -04:00
Tom Boucher
8de2ff9121 feat(#2008): generic command-exit-zero gate-predicate evaluator (#2011)
* feat(#2008): add generic command-exit-zero gate-predicate evaluator

Third-party capability gates declared via check.predicate were rendered for
display but never evaluated (only built-in check.query gates fired; the
security capability's gate worked solely via a hard-coded ship.md branch).

Add a generic, deps-injected gate-predicate evaluator (src/gate-predicate-evaluator.cts)
that dispatches by predicate.kind. Built-in kind: command-exit-zero — runs a
bounded sh -c command at the project root (via shell-command-projection.execTool),
inherits env, exit 0 => pass, non-zero => block, timeout => block, fail-closed.

Wire a 'check predicate' subcommand into check-command-router.cts and extend
the three generic workflow gate-dispatch sites (execute:wave:post, execute:post,
plan:post) to route check.predicate gates to the new evaluator. The two-step
gate contract (command-failure => onError; block => halt) is unchanged.

- src/gate-predicate-evaluator.cts: pure leaf, KIND_TABLE extensible
- src/check-command-router.cts: cmdCheckPredicate + buildPredicateDeps + parsePredicateFlags
- docs/adr/2008-*, docs/reference/gate-predicates.md, docs/how-to/command-exit-zero-gate.md
- tests: 38 unit + integration tests (exit mapping, timeout, interpolation,
  property-based bijection, malformed-predicate fail-closed, real subprocess e2e)

Closes #2008

* docs(#2008): backfill changeset pr number 2011
2026-07-05 14:04:29 -04:00
Tom Boucher
8f2ebbe9bf feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity

Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap).

--gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): backfill changeset PR number (#1996)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): drop Gemini CLI from issue templates (review nit)

Removes the sunset Gemini CLI runtime from the two GitHub issue-template
runtime lists that the removal PR missed, per @davesienkowski's review nit:
- feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise
  request a feature for a runtime GSD no longer supports)
- bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json
  retrieval-help line

Leaves the post-removal templates fully consistent with the Antigravity redirect.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 13:32:51 -04:00
Rezolv
55604e9124 fix(#1906): require node-test clean-fixture causation control (#2001)
* fix(#1906): require node-test clean-fixture causation control

The node-test fail-first proof accepted a deceptive content-independent
negative test — one that reds merely because GSD_PROHIB_SUBJECT is set,
ignoring the subject's content — whenever no cleanFixture was supplied,
because #1346's causation control was opt-in. The proof's observed signal
(RED) thus diverged from its target (RED caused by content) by default.

Make the causation control mandatory for the node-test kind: a descriptor
that omits cleanFixture is un-provable (fail-closed), never accepted under
the weaker violation-only proof. When a clean fixture is present, fail-first
is proven exactly as before (RED on violation AND non-vacuous GREEN on clean).
The lint-rule kind is unchanged (its subject IS the linted file; no
GSD_PROHIB_SUBJECT indirection).

Breaking (Hyrum): a previously-green node-test prohibition with no clean
fixture now hard-gates — blast radius is zero in-tree (no node-test
prohibition ships today; only the lint-rule local/no-source-grep dogfood).

Supersedes ADR-1606 Decision 4 / ADR-550 #1346 addendum's opt-in.

Closes #1906

Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ

* docs(#1906): supersede the #1346 opt-in causation control (mandatory for node-test)

Record the node-test mandatory-causation-control supersede across the
governing surfaces:

- ADR-1606 (the enforcement decision-of-record): addendum + Decision 4
  annotated + the "Mandatory causation control — REJECTED" alternative
  flipped to accepted (premise no longer holds: zero in-tree node-test
  consumers).
- ADR-550: the 2026-06-21 #1346 "Why opt-in, not required" paragraph
  marked SUPERSEDED, pointing at ADR-1606.
- spec-phase.md: check_clean_fixture is now REQUIRED for node-test
  (was "optional").
- CONTEXT.md: PROHIB.enforce.causation predicate updated.

Regenerated the shipped-artifact cascade from the spec-phase.md edit
(+149 B, well under the 40960 cap): 16 golden-install-parity fixtures
and the workflow size baseline.

Refs #1906

Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ
2026-07-03 23:21:55 -04:00
Jeremy McSpadden
e5ef323b15 feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command

Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.

* feat: add /gsd-start smart-entry command

State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.

- src/smart-entry.cts: deterministic situation classifier (no-project,
  paused, blocked, verify-failed, needs-first-phase, planning, executing,
  verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
  ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire  case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
  dispatcher presenting an AskUserQuestion menu (with --text fallback for
  non-Claude runtimes) and dispatching to existing commands. Falls back
  to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
  situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
  (markdown-layer invariants + every emitted command resolves to a real
  slash command).

Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.

* refactor: rename smart-entry command to /gsd:next

Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.

All affected tests (188) pass; lint:ci clean.

* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)

Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.

- detectSignals now reads total_phases + percent from nested progress{}
  first, then scalar fm, then body; current_phase falls back to the
  body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
  body Phase field) covering verify-pending + executing situations.

Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.

* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)

Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.

Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
   layer is down — read .planning/STATE.md directly with the Read tool
   and synthesize a minimal situation + actions menu so /gsd:next stays
   useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
   /gsd:progress as before.

Matches the direct-read resilience the live agent already did by hand.

* docs: add gsd-next skill surface

* chore: trigger no-mistakes validation

* no-mistakes(review): Fix smart-entry phase ordering

* no-mistakes(review): Fix decimal smart-entry phase ordering

* no-mistakes(test): Fix smart-entry next test contracts

* no-mistakes(document): Docs synced for smart entry

* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: shorten next.md description and update golden install parity fixtures

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: trigger no-mistakes validation

* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files

Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.

* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: regenerate golden install parity fixtures for /gsd:next

Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.

* Fix smart-entry verify-failed phase scoping and empty resolve shim step

Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.

* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: recapture all 16 golden fixtures with updated smart-entry.md hash

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: regenerate fixtures + inventory manifest after rebase onto next

Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next

Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).

Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)

The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs

Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo

Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
  recommended action, 1-4 unique-id /gsd:* actions (previously the
  one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout

Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:

- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
  zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
  kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
  consolidation file doing dozens of real installs — 41s even on a fast Mac
  (much worse on the slow Windows I/O path), plus an install-heavy cluster.

Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.

Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 12:18:25 -04:00
Rezolv
9d7c046eae refactor(#1852): lazy-split plan-phase.md into steps/ (#1934)
* refactor(#1852): lazy-split plan-phase.md into steps/

Extract 3 self-contained, rarely-hit sections into gsd-core/workflows/plan-phase/steps/
via lazy 'Read and execute' pointers (mirrors execute-phase/steps/, ADR-1610 progressive
disclosure — no eager @-import): closed-phase-gate (1.5), prd-express-path (3.5),
windows-troubleshooting. Byte-invariant: plan-phase.md 94459 -> 89775 (-4684), each step
< 32 KiB anchor, resolved instructions unchanged.

Scope note: only 3 sections were extractable. plan-phase.md is guarded by a dense net of
content-presence tests (plan-bounce/enh-3209/phase6-planning-capabilities assert specific
sections/flags inline) that block extracting the larger blocks without also refactoring
those tests — durable low-80s headroom is deferred pending maintainer re-scope (issue #1852).

Cascade: size:baseline regen, 16 golden-install-parity fixtures regen, INVENTORY in sync.
Full plan-phase test surface green (4129/4130, 0 fail); lint:ci exit 0.

* docs(changeset): Changed fragment for #1934 (plan-phase lazy-split)

* docs(changeset): mark #1934 fragment docs-exempt (internal workflow refactor)

* fix(#1852): add gsd_run launcher preamble to prd-express-path step

The extracted prd-express-path.md calls gsd_run but the canonical launcher
preamble lived in the parent plan-phase.md — runtime-launcher-parity (#373)
walks workflows/ recursively and requires every .md using gsd_run to carry
exactly one preamble + the $HOME/.claude fallback arm (same as the existing
execute-phase/steps/ files). Injected via scripts/sync-runtime-launcher.cjs;
golden fixtures + size baseline regenerated. plan-phase.md unchanged (89775).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-02 23:30:19 -04:00
Behruz Nassre Esfahani
7bef6a6496 fix(#1863): use named flags for state.* calls in executor + workflows (#1873)
* fix(#1863): use named flags for state.* calls in executor + workflows

The named-only state-command router (parseNamedArgs) silently drops
positional args, so state.cjs threw its required-arg error and
metrics/decisions/blockers/session continuity were never recorded.

Convert record-metric / add-decision / add-blocker / record-session in
agents/gsd-executor.md to the named-flag form (mirroring execute-plan.md),
and fix the two remaining positional record-session calls in
gsd-core/workflows/milestone-summary.md and forensics.md. Recapture the
golden-install-parity fixtures and size baselines for the edited files.

Also fix a pre-existing detached-rebuild handle leak in
tests/graphify-auto-update.slow.test.cjs: three dispatch tests returned
after observing only the synchronous "running" status without awaiting the
detached rebuild's terminal state. That leak was latent until the new
#1863 regression block's added runtime shifted --test-force-exit timing
and surfaced it as a non-zero chunk exit. The three tests now await
terminal status via the file's existing waitForBuildStatus helper.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1863): add changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-02 13:01:10 -04:00
jecanore
5657994702 fix(#1716): route resume_from_file to complete_session when no pending tests remain (#1722)
* fix(#1716): route resume_from_file to complete_session when no pending tests remain

When a UAT session has status:partial with blocked_count>0 and pending_count==0 (all remaining tests are blocked, none are pending), resume_from_file found no [pending] test and terminated silently — never routing to complete_session. This blocked the issues==0 auto-transition path even when there were zero code defects.

Guard clause added immediately after the find-pending step: if no [pending] test is found, route to complete_session. complete_session then correctly sets status:partial (because blocked_count>0) without presenting further tests.

Closes #1716

* chore(#1716): add changeset fragment and regenerate golden-install-parity fixtures

Changeset fragment for PR #1722 (type: Fixed).

Golden-install-parity fixtures regenerated for all 16 runtimes — the workflow fix shifts verify-work.md's byte-stable hash in the golden manifest. Regenerated via UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
2026-07-02 11:58:06 -04:00
Tom Boucher
92091d71f2 fix(#1871): wire phase archival end-to-end (phases archive cmd + default + atomic) (#1924)
Follow-up to #1919 (archive-then-remove core). Closes the remaining #1871
acceptance criteria so phase history is preserved across the full milestone
lifecycle, not just at phases.clear:

- #2 src/milestone.cts + src/phases-command-router.cts: extract shared
  archivePhaseDirectories() helper; add cmdPhasesArchive (the previously
  half-wired phases.archive alias now routes instead of erroring Unknown).
- #4 gsd-tools.cjs + src/milestone.cts: milestone complete archives phase
  dirs by default (--no-archive-phases opts out; --archive-phases is now a
  harmless no-op). complete-milestone.md updated to drop the redundant manual
  Yes/Skip archive prompt.
- #3 gsd-core/workflows/new-milestone.md: §6 stages the archive move + source
  removal (git add .planning/milestones/ .planning/phases/) in the same commit
  as the milestone start, so the archive lands atomically — no orphaned
  uncommitted deletions, no un-archived dirs inherited.
- docs/CLI-TOOLS.md (+ ja/zh/ko/pt) + help/modes/full.md: flag accuracy.
- tests: phases archive command (#2) + milestone complete default archive /
  --no-archive-phases opt-out (#4). Goldens + workflow size baseline refreshed.

Closes #1871
2026-07-02 11:14:01 -04:00
Tom Boucher
3c13903dcd feat(#1866): agent-side self-load of configured agent_skills
Each of the 22 consumer agents now self-loads its configured agent_skills
in its mandatory init step, so .planning/config.json agent_skills.<type>
reaches the agent on every runtime — including Cursor and /gsd-autonomous,
where Skill()-delegated workflow bash init did not reliably execute.

- gsd-core/references/agent-skills-bootstrap.md: shared contract
  (query + Read + dedup guard that skips when <agent_skills> is already
  in the prompt, so Claude's orchestrator-side injection never doubles)
- 22 agents/gsd-*.md: one self-load line naming the agent's own type
- gsd-core/workflows/autonomous.md: note that delegated agents self-load
- tests/agent-skills-bootstrap.test.cjs: regression + parity (CONSUMER_AGENTS
  bijection + fast-check property) — Generative-Fix-Divergence guard
- docs: ADR-1866, CONFIGURATION dual-injection How It Works, INVENTORY
  row, Changed changeset

Closes #1866
2026-07-01 20:09:01 -04:00
Tom Boucher
3cc4d1608c test: regenerate golden fixtures + size baselines for the example/doc edits
execute-phase.md and gsd-ai-researcher.md are installed artifacts; their edits
shift install hashes and file sizes. Diff is scoped to those two files' hashes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 23:04:13 -04:00
Tom Boucher
da37986cd0 fix(#1847): resolve standard tier to claude sonnet 5
Point the sonnet/standard tier at Claude Sonnet 5 (`claude-sonnet-5`,
GA 2026-06-30) across the Anthropic-backed runtimes and provider presets,
replacing the superseded `claude-sonnet-4-6`. Mirrors the change into the
CONFIGURATION.md and settings-advanced.md runtime-defaults tables (the
#3229 catalog↔docs parity gate) plus the pt-BR/zh-CN translations, and
updates the tests that pin the old ID. Regenerates the workflow size
baseline for the (smaller) settings-advanced.md.

Scope is Sonnet only — opus/haiku IDs are untouched. Prepared as a 1.6.1
hotfix off the v1.6.0 tag.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 33260555b4)
2026-06-30 22:21:18 -04:00
Tom Boucher
93e5d2dd84 fix(#1525): skip deferred phases on autonomous reruns (#1846)
* fix(#1525): skip deferred phases on autonomous reruns

* chore(#1525): add changeset fragment

* chore(#1525): fix changeset body

* test(#1525): refresh install parity fixtures

* test(#1525): shrink autonomous workflow

* test(#1525): refresh autonomous baselines

* test(#1525): tolerate Windows temp cleanup flake
2026-06-30 21:29:38 -04:00
Behruz Nassre Esfahani
dae7f81482 fix(#1528): drop next-phase guidance from security-blocked verify-work presentation (#1687)
* fix(#1528): drop next-phase guidance from security-blocked verify-work presentation

When security enforcement blocks phase advancement (no SECURITY.md produced),
the verify-work presentation told the user advancement was blocked but still
offered `/gsd:plan-phase {next}` and `/gsd:execute-phase {next}`, competing
with the current-phase fix. Remove those two next-phase lines so the blocked
state routes only to the current-phase resolution (secure-phase, ui-review).
The post-transition presentation — reached only after the completion contract
passes — still offers next-phase planning, which is the correct place for it.

Regression coverage added to tests/ui-review-next-guidance.test.cjs: the
security-blocked block must not offer next-phase actions, and the
post-completion block must still offer them. Regenerated workflow size baseline.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1528): add changeset for security-blocked next-phase fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1528): recapture golden-install-parity fixtures for verify-work.md change

Rebased onto next; verify-work.md's installed hash changed across all 16
runtime fixtures. Diff confined to the single gsd-core/workflows/verify-work.md
key per runtime. Assert mode 16/16 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-30 10:52:30 -04:00
Tom Boucher
1f649838b8 Merge branch 'next' into fix/1698-codex-output-last-message 2026-06-30 10:04:59 -04:00
Rezolv
18995380ce feat(#1154): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1738)
* feat(verify-phase): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1154)

Carry the edge-probe's existing `backstop` (non-inferable) tier through the
plan-phase projection as a structured flat-scalar marker instead of a prose
parenthetical, and make verify-phase abstain -> human_needed (never silent-pass)
on a backstop truth it cannot confirm with explicit evidence. Truth-axis mirror
of #644's prohibition judgment-tier (ADR-550 D4).

Engine (deterministic, CI-tested per ADR-550 D5 — never the LLM verdict):
- src/probe-core.cts: truthStatement/truthVerification normalizers, projectTruths
  (conservative serializer), dispositionForUnverifiableTruth (backstop+no-evidence
  -> unverified/flagged/insufficient_spec; backstop+evidence -> green; inferable
  -> green, the over-abstention guard).
- src/roadmap.cts: coerceTruthToString now reads `statement` first so an object-form
  backstop truth is surfaced, not dropped (Hyrum backward-compat for truth-readers).

Workflow/agent/docs: plan-phase emits the structured marker (flat scalar, ADR-550
#1278); verify-phase + gsd-verifier add the abstain arm; new references/honest-verifier.md;
FEATURES/COMMANDS document insufficient_spec; ADR-550 amended (truth-axis D4 mirror).

Decisions adopted (trek-e review): insufficient_spec feeds existing human_needed with a
distinguishable reason (no new VERIFIER_STATUS); changeset Changed; round-trip parity
test; abstain-on-unconfirmed-backstop regression test red-first.

Implementation notes (deviations from the issue's proposed file list, verified live):
- frontmatter.cts needs no change — its flat parser already round-trips object-form truths.
- verify.cts needs no change — it grades artifacts/key_links structurally; truths are
  LLM-graded at the workflow layer, so consumption lives there + the deterministic helper.
- No CJS<->SDK hand-sync — the SDK seam was retired (ADR-0174); src/*.cts is sole source.

Regenerated artifacts: golden-install-parity fixtures, INVENTORY-MANIFEST, size baselines.

* chore(#1154): add changeset (Changed) for honest verifier

User-facing changelog fragment for #1738. Typed `Changed` (not `Added`) per
trek-e review condition 3 — the verify behavior shifts for backstop-bearing specs
(a confident silent `passed` becomes `human_needed`), which is user-visible even
though the schema marker is additive.

* docs(#1154): score-formula also excludes abstained insufficient_spec truths (review nit-1)

trek-e review nit: the verify-phase score sentence said PRESENT_BEHAVIOR_UNVERIFIED
truths were "the only ones excluded" from verified_truths. Post-#1154 an abstained
`insufficient_spec` backstop truth is also excluded (it is not ✓ VERIFIED and routes
to human_needed). Behavior was already correct; this tightens the wording.
Regenerated golden-install-parity fixtures + workflow-size baseline for the touched
verify-phase.md. (Nit-2 — a dedicated insufficient_spec_items frontmatter list — is
intentionally not taken: the current design is ADR-550-D4-conformant, the abstain
cause rides as a distinguishable report reason, and adding it would exceed the
approved scope.)

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-29 00:14:32 -04:00
Tom Boucher
ac001be49e fix(#1778): use 1.6 named-flag frontmatter.set form in thread workflow (#1816)
* fix(#1778): use 1.6 named-flag frontmatter.set form in thread workflow

The thread workflow's CLOSE and RESUME branches called frontmatter.set with
the pre-1.6 fully-positional shape (frontmatter.set <file> <field> <value>).
Since 1.6 the dispatcher (gsd-tools.cjs) parses the file positionally and
reads field/value from the named flags --field/--value via parseNamedArgs;
the positional form leaves field/value undefined, cmdFrontmatterSet errors
'file, field, and value required', and the status/updated writes are
silently skipped. Closing a thread never marked it status: resolved and
resuming never marked it status: in_progress.

Switch all four sites (CLOSE status+updated, RESUME status+updated) to the
1.6 hybrid form that verify-work.md already uses:
  frontmatter.set <file> --field <field> --value <value>

Add a regression test with three guards: (1) behavioral — the named-flag
form writes the field while the positional form errors with the documented
message and does not mutate the file; (2) workflow parity — no workflow
under gsd-core/workflows/ emits the positional form, so a future edit that
reintroduces it anywhere fails CI; (3) thread-specific — CLOSE writes
status: resolved and RESUME writes status: in_progress via the named flags.

* docs(#1778): add changeset fragment for thread workflow frontmatter fix

* docs(#1778): fix unclosed inline-code backtick in changeset fragment

* fix(#1778): move regression into owning test + regen baselines

lint-regression-test-names rejects new bug-NNNN-*.test.cjs files; move the
#1778 regression (behavioral named-vs-positional + workflow-parity scan +
thread CLOSE/RESUME assertions) into tests/frontmatter-cli.test.cjs, the
canonical home for frontmatter CLI regressions, and delete the standalone
file. frontmatter-cli.test.cjs already carries the allow-test-rule exemption
for workflow .md content tests.

gsd-core/workflows/thread.md ships to every runtime and is size-tracked, so
recapture the 16 golden-install-parity fixtures (thread.md hash) and the
per-file workflow size baseline (thread.md 12400 -> 12464) via UPDATE_GOLDEN=1
and npm run size:baseline.
2026-06-28 22:42:44 -04:00
Behruz Nassre Esfahani
2bada4de1a fix(#1698): capture codex review via --output-last-message, not stdout
The Codex reviewer in review.md captured the review by redirecting codex
exec's stdout to the review file. On Windows, codex writes process-teardown
output to stdout after the final agent message, so that noise was appended
to a non-empty file and slipped past the `[ ! -s ]` empty-output guard as a
silently polluted review (consumed by severity extraction and the
plan-review-convergence gate).

Capture the final message via codex's own `-o/--output-last-message <FILE>`
and discard stdout. The #1115 contract is preserved (stderr to .err,
capability-gated $CODEX_BYPASS_FLAG, --ephemeral, --skip-git-repo-check) and
the empty-output fallback still fires when codex leaves no/empty output.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 08:15:37 -07:00
Tom Boucher
b0d5ca3379 feat(#1517): support custom reviewer instances for /gsd:review (#1766)
* feat(#1517): support custom reviewer instances for /gsd:review

Add a bounded review.reviewer_instances config surface so one model-capable
adapter (e.g. opencode) can run as several independent reviewer identities in a
single /gsd:review pass. Instances participate only via review.default_reviewers,
expand before built-in slugs, are available iff their cli is detected, and a
non-matching entry is a hard error (typo must be loud). >=2 same-cli instances
emit a shared-adapter caveat in REVIEWS.md. Default path with no instances is
byte-for-byte unchanged.

Single-source instance->cli resolution lives in resolveReviewerSelection /
normalizeReviewerInstances (parity-locked in
tests/review-reviewer-instances.test.cjs). cli validated against
KNOWN_REVIEWER_SLUGS only (never arbitrary shell); model/agent opaque, never
shell-interpolated.

Closes #1517

* chore(#1517): backfill changeset pr:1766

---------

Co-authored-by: review-bot <review-bot@gsd>
2026-06-26 23:35:04 -04:00
Tom Boucher
cf2e66b39e feat(#1708): typed documentation-sourced #853 dispatch-flatten (ADR-1239 Phase B) (#1719)
* feat(#1708): typed documentation-sourced #853 dispatch-flatten

Graduate the #853 orchestrator-backgrounding decision from a scattered RUNTIME==='codex' prose check to a typed, documentation-sourced engine decision. Adds a backgroundDispatch dispatch sub-axis (sourced per host: codex+cursor documented true, 9 documented false, 5 undocumented), shouldFlattenDispatch(dispatch) (inline UNLESS background && backgroundDispatch, fail-closed), and a gsd_run query dispatch-should-flatten the plan/execute workflows call. Cursor is newly background-eligible per its docs (inline->background) — a documentation-justified behavior change. No RUNTIME-name residue for this decision.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1708): backgroundDispatch citations in matrix + CONTEXT note

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1708): address review findings on typed dispatch-flatten

Code/adversarial review: convert the manager.md/autonomous.md Compound Action preamble from hardcoded 'On Codex' to FLATTEN-based branching (the handlers already use the query; the preamble contradicted them and was wrong for cursor); make shouldFlattenDispatch null-safe + type-honest (accepts raw 'undocumented' registry values); make backgroundDispatch a required descriptor field (matching its siblings, all 16 carry it); strengthen the config.runtime behavioral test; update the bug-853 prose-pin test + comment. Security review clean; Codex confirmed no fail-open.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1708): backfill backgroundDispatch in role:runtime test fixtures

Making backgroundDispatch a required descriptor field broke role:runtime fixtures in capability-manifest-version/capability-registry/host-integration-descriptors tests that build a dispatch object without it (caught by full gsd-test, not scoped npm test). Backfill backgroundDispatch:false into the well-formed fixtures; the deliberately-malformed 'required-field' test fixture is left malformed by design.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1708): update fix-1521 dispatch-gating assertion to the FLATTEN gate

fix-1521 pinned the codex-specific run_in_background prose that #1708 graduated to the typed dispatch-should-flatten/FLATTEN gate. Update its assertions to verify FLATTEN=false gating (not a runtime name) + that the old RUNTIME===codex gate is gone. Caught by full gsd-test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1708): add changeset for typed dispatch-flatten

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1708): remove stray temp PR-body file

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1708): add issue ref to bug-853 allow-test-rule annotations

ADR-456 requires every allow-test-rule exemption to carry a see #NNN reference; the source-text-is-the-product annotations added when migrating the prose assertions lacked it (lint-tests CI gate). Add (see #1708).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 15:37:33 -04:00
Behruz Nassre Esfahani
47906b052d fix(#1520): randomize mktemp temp paths on BSD/macOS (XXXXXX must be path-final) (#1550)
* fix(#1520): randomize mktemp temp paths on BSD/macOS (XXXXXX must be path-final)

BSD/macOS mktemp only substitutes the XXXXXX template when it is the final
path component. Templates like `...-XXXXXX.json` / `gsd-pr-body.XXXXXX.md`
return a LITERAL `XXXXXX` path (no randomization) on macOS, so concurrent
workflow runs collide on the same temp manifest/body file — one run can
overwrite or consume another's. Reproduced on macOS: the second call to the
suffixed template fails `mkstemp: File exists`.

Fix: use a suffixless `XXXXXX` template (so it IS the final component), then
rename to add the intended extension — portable across BSD + GNU userlands,
no GNU-only `--suffix` flag. Empty-file-then-write semantics are preserved at
every site.

Affected workflow temp files:
- execute-phase.md: gsd-worktree-wave-*.json (wave worktree manifest)
- quick.md:         gsd-quick-worktree-*.json
- spec-phase.md:    edge-probe-reqs-*.json
- ship.md:          gsd-pr-body-*.md
- profile-user.md:  gsd-profile-answers-*.json, gsd-profile-analysis-*.json

The execute-phase.md edit uses a compact intermediate var + trailing comment
to stay under the ADR-857 phase-6 size ceiling (93166); regenerated the
workflow size baseline accordingly. Validated on macOS: 20 concurrent calls
yield 20 unique randomized paths.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1520): add changeset fragment (Fixed)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1520): add fail-first workflow-prose guard for mktemp XXXXXX suffix

Repo-wide scan of gsd-core/workflows/**/*.md that fails on any mktemp
template whose XXXXXX run is followed by a filename suffix (the BSD/macOS
non-randomizing form). Fails on the six pre-fix instances and passes on
the fix, and locks the copy-paste-prone idiom out of future workflows.
Mirrors the bug-637 hardcoded-$HOME workflow guard.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1520): rename regression test to fix- prefix (regression-test-names lint)

New tests/bug-NNNN-*.test.cjs files are banned by the lint-regression-test-names
ratchet; use the fix- prefix (matches the fix-1445 precedent).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1520): add issue ref to allow-test-rule exemption (ADR-456 lint)

lint-allow-test-rule-refs requires every new `allow-test-rule:` comment to
carry a #NNN reference (don't allowlist). Add (#1520) to the source-text
exemption.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1520): abort touched mktemp chains on failure (|| exit 1)

Per review: the VAR=$(mktemp …) && mv … && VAR=… chains dropped the issue's
suggested failure guard. If mktemp fails, $VAR is empty and the subsequent
mv/write lands on an unintended relative path. Add `|| exit 1` to all six
touched chains so a mktemp failure aborts the snippet. Regenerated the
workflow size baseline for the slightly longer lines.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1520): rebase onto next — regen size baseline + describe rename

Resolve the workflow-size-baseline.json conflict from next advancing by
regenerating from the current workflow sizes. Also rename the test describe
from `bug #1520` to `#1520` (the file uses the fix- prefix) per review nit.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1520): harmonize profile-user mktemp to ${TMPDIR:-/tmp} (review nit)

The two profile-user.md temp sites this PR already rewrites kept a hardcoded
/tmp while the four sibling workflows use ${TMPDIR:-/tmp}. Harmonize for
consistency and macOS-correctness (some sandboxes have no writable /tmp).
Regenerated the workflow size baseline.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1520): regen size baseline after rebase onto next

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 17:27:45 -04:00
Tom Boucher
d101daff30 fix(#1516): expose adaptive model_profile in /gsd-new-project AI Models prompt (#1654)
* fix(#1516): expose adaptive model_profile in /gsd-new-project AI Models prompt

Both onboarding paths (Step 2a auto-mode + Step 5 interactive) enumerated only
4 profiles (Balanced/Quality/Budget/Inherit), omitting 'adaptive' even though the
model catalog (model-catalog.json profiles) and docs/CONFIGURATION.md register 5.
Mirrors the proven /gsd:settings two-question split (#3784): Q1 routes between
Adaptive/Standard-tier/Inherit; Q2 (conditional on Q1=Standard) picks
Quality/Balanced/Budget — keeping every AskUserQuestion within the 4-option cap.
Both config-new-project example payloads now list adaptive. Regression cases
folded into the owning tests/new-project-mvp-prompt.test.cjs (per the
lint-regression-test-names ban on new top-level bug-NNNN files): each AI Models
prompt makes adaptive reachable, all 5 profiles reachable, 4-option cap honored,
both example enums include adaptive, brace balance. Workflow size baseline bumped
(new-project.md 62324 -> 66138 bytes; still well under the XL hard cap).

* chore(#1516): backfill changeset pr ref to 1654
2026-06-24 14:46:44 -04:00
Jeremy McSpadden
77c7b4fc9d fix(#1522): enforce canonical verification before phase transition (#1548)
* fix: require fresh phase verification before transition

* no-mistakes(review): Fix canonical verification closeout gates

* no-mistakes(review): Fix verify-work frontmatter promotion command

* no-mistakes(review): Fix stale verification gates

* no-mistakes(review): Fix canonical verification routing gates

* no-mistakes(review): Fix verification dependency and runtime routing gates

* no-mistakes(review): Block stale verification bypasses

* fix: handle large init manager outputs in verification workflows

* chore: update changeset pr number

* fix(verify-work): use fresh verification.status for stale gate

The stale check after UAT used phase_completion.verification_status from
session-start INIT while human_needed promotion already queried fresh
verification.status. Align the stale gate with the canonical query so
mid-session verification refresh is not ignored.

* fix(init): skip roadmap-checked phases when selecting next_phase

Roadmap-only phases without a disk directory were still promoted to
next_phase when their checkbox was already checked. Exclude
checkboxComplete phases so progress routing does not point at work the
roadmap already marks done.

* fix: gaps_found not overridden by stale, transition uses canonical verification

- verification.cts: check gaps_found before stale so gap-closure routing
  is not masked by a newer summary mtime
- phase.cts: remove redundant findStaleVerificationSummary — readVerificationStatus
  already handles stale detection
- transition.md: replace raw grep on file content with verification.status query
  to avoid false-positive blocks from body text matching

* ci: retrigger tests after rebase

* fix(transition): replace gsd_run advisory check with awk frontmatter extraction

The runtime launcher is not defined until the update_roadmap_and_state step
bash block (~line 165). The early verify_completion block used gsd_run to
query verification.status, which violated the runtime-launcher-parity test:
'preamble appears AFTER the first gsd_run reference'.

Replace the gsd_run call with an awk-based frontmatter extractor that reads
only the status: field between the two --- fences. This avoids both the
preamble-ordering constraint and the original false-positive grep bug where
body text like 'previous_status: gaps_found' would match a full-text regex.

The phase.complete gate at update_roadmap_and_state is the canonical
enforcement point; this early check is advisory only.

Also update workflow-size-baseline.json for the updated transition.md size.

Fixes: runtime-launcher-parity test (B)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: re-check verification under planning lock in phase complete

Move readVerificationStatus into withPlanningLock so stale verification
cannot slip through when a SUMMARY.md is written between the gate and
the roadmap/state mutation. Return the blocked status from the lock
callback and emit the error after release to avoid leaving .lock behind.

* fix(transition): gate on canonical verification.status including stale

Replace awk frontmatter read with verification.status query so transition
blocks when summaries are newer than VERIFICATION.md, matching phase.complete
and other workflows (autonomous, progress, verify-work).

* Fix workflow verification gates for yolo transition and stale routing

Require VERIFY_STATUS passed before yolo/interactive transition advance.
Route stale verification recovery to verify-work, matching canonical projection.

* fix(transition): use verification.status query for stale-aware advisory check

The awk-based check read raw frontmatter status: passed, which misses the
stale case where summaries are newer than the VERIFICATION.md file even
though the frontmatter still says passed. The stale status is computed from
file modification times, not stored in frontmatter.

Move the preamble to the verify_completion bash block (the first block with
a gsd_run call) so gsd_run query verification.status can be used for the
advisory check. This gives the full readVerificationStatus logic including
mtime-based staleness detection, matching the enforcement gate at phase.complete.

Capture full JSON (VERIFY_JSON) so next_action can be included in the
advisory output alongside the status.

Also update workflow-size-baseline.json for the updated transition.md size.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* ci: trigger test matrix for 525b946

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix(transition): restore awk frontmatter extraction for pre-shim verification check

The gsd_run launcher shim is not defined until line ~163 of transition.md,
so the verification debt check at line ~80 cannot use gsd_run. Restore the
awk-based frontmatter extraction that correctly reads status without needing
the runtime, and restore the shim at its proper location before
phase.complete.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(#1522): clarify transition verification gate wording

* fix(#1522): update transition workflow size baseline

* fix(#1522): update workflow-size-baseline after rebase onto next

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix(#1522): guard findStaleVerificationSummary FS calls + thread opts.fs seam (review)

Address review blocker B1 on #1548: findStaleVerificationSummary ran fs.readdirSync
and two fs.statSync calls unguarded between readVerificationStatus's try/catch sections,
so a TOCTOU race (a SUMMARY listed by scanPhasePlans then removed before statSync) or any
FS error threw uncaught into callers NOT under the planning lock (init.manager /
init.progress / uat-predicate). Wrap the body in try/catch degrading to 'not stale', and
thread the injectable opts.fs seam (add statSync to FsLike, pass fsImpl from the caller)
for parity with readVerificationStatus's no-throw contract and testability. Also adds the
Verification Module glossary entry to CONTEXT.md (review B3).

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 13:19:10 -04:00
Tom Boucher
207d8f1697 fix(#1626): make the security gate severity-aware via per-threat severity (#1635)
workflow.security_block_on was documented as the minimum threat severity
that blocks advancement, but threats carried no severity and the auditor's
threats_open count (the SECURITY.md gate field) counted every open threat
regardless of severity — so the threshold had no effect, and the auditor's
block_on vocabulary (open/unregistered/none) did not even match the config
enum (critical/high/medium/low/none).

- planner: add a Severity column to the STRIDE threat register; assign
  severity per threat.
- auditor: read severity; reconcile the <config> block_on domain to the
  severity enum; redefine threats_open as the count of OPEN threats whose
  severity is at or above block_on (none => 0). Below-threshold opens are
  reported as non-blocking and excluded from threats_open.
- SECURITY.md template + planning-config.md reconciled.

No gate-check site changed: threats_open == 0 stays the gate everywhere;
only its computation is now severity-filtered.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 19:04:13 -04:00
Tom Boucher
f9d9dfb4bc fix(#1627): scale security rigor by ASVS level (planner disposition + auditor depth) (#1636)
workflow.security_asvs_level was display-only — the planner hardcoded
'mitigate if ASVS L1 requires it' and the auditor only echoed the level,
so L2/L3 behaved identically to L1.

- New reference gsd-core/references/security-asvs-levels.md defines L1
  (opportunistic), L2 (standard), L3 (comprehensive) for both planner
  threat disposition and auditor verification depth (higher = superset).
- planner: disposition now scales with the configured ASVS level (no
  hardcoded L1) + @-pointer to the reference.
- auditor: verification depth scales with asvs_level (L1 grep-presence,
  L2 boundary/vector check, L3 end-to-end trace + bypass check).
- planning-config.md + INVENTORY updated; planner kept under its 48K cap
  by extracting the goal-backward worked example to planner-guidance.md.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 18:48:58 -04:00
Tom Boucher
94be6d5b60 fix(#1625): resolve security config in secure-phase.md before auditor handoff (#1633) 2026-06-23 17:24:59 -04:00
Tom Boucher
c28cccbf85 Merge pull request #1595 from davesienkowski/feat/1592-plan-drift-precheck
feat(#1592): add plan:pre codebase-drift pre-check before planner runs
2026-06-23 10:56:29 -04:00
Tom Boucher
ba96c70b14 feat(#1602): deterministic coverage-metadata UAT routing for verify-work
Add an optional structured `coverage:` block to SUMMARY.md frontmatter and a
deterministic classifier that `verify-work` consumes to route deliverables to
auto-pass vs human-UAT — replacing the rejected #1598/#1599 post-hoc heuristic.

- New `src/coverage.cts` (→ bin/lib/coverage.cjs) parses the nested coverage
  block (extractFrontmatter can't — its `-` items are scalars-only; this is a
  focused parser, sibling of parseMustHavesBlock), validates each entry, and
  classifies into auto_passed vs present. Frozen MODE/PRESENT_REASON/ERROR_CODE
  typed-IR surface. Exposed via `uat classify-coverage --summary <f>`.
- Auto-pass is the narrow proven case only: strict-boolean human_judgment:false
  AND non-empty all-`pass` verification AND zero validation errors. Everything
  else — judgment, empty/failing verification, malformed entry — routes to the
  human (fail-safe). A malformed block falls back to legacy prose extraction and
  surfaces an error; an absent block is byte-identical to pre-#1602.
- execute-plan create_summary populates the block (fail-safe default
  human_judgment:true); verify-work extract_tests consumes it; create_uat_file
  marks auto-passed entries `source: automated`.
- Templates (summary + 3 variants), CONTEXT.md predicate + glossary, INVENTORY,
  eslint/gitignore registration, and Diataxis docs (COMMANDS reference +
  USER-GUIDE explanation) updated.
- Behavioral tests via the CLI (no source-grep); parser-robustness regressions
  for the null-entry/comment-header/mis-indent cases found in adversarial review.

Closes #1602

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 23:43:22 -04:00
Dave
d1f7ba82f2 feat(#1592): add plan:pre codebase-drift pre-check before planner runs
Add a non-blocking, warn-only codebase-drift gate at plan:pre so a stale
STRUCTURE.md is surfaced before /gsd:plan-phase spawns the planner, instead
of being discovered mid-execution by the existing execute:wave:post gate.
Gated on a dedicated workflow.plan_drift_precheck toggle (default true),
independent of schema_drift_gate. Never blocks planning, never spawns the
mapper agent at plan time.

Review feedback (#1595):
- Use a documented conventional-commit type (feat, not enhance) per
  CONTRIBUTING.md / gsd-validate-commit.sh.
- Normalize the plan_drift_precheck command references to the colon prose
  form (/gsd:plan-phase, /gsd:map-codebase) to match plan-phase.md §5.65;
  registry regenerated from capability.json.
- Make the test temp dirs hermetic: drain mkdtemp dirs in an after() hook
  via the helpers.cleanup() budget (local/no-raw-rmsync-in-tests-compliant).

Closes #1592

Claude-Session: https://claude.ai/code/session_016JBiXEAofvB3prJim29XMS
2026-06-22 20:30:59 -04:00
Tom Boucher
248c056532 chore(#1574): regenerate workflow-size baseline for new-project prose growth
The copilot path correction (.github/copilot-instructions.md) lengthened
the new-project.md instruction-file prose by 16 bytes past the prior
baseline. Regenerated; growth is justified by the more accurate path.
2026-06-22 10:09:43 -04:00
Tom Boucher
bf9bd1f4e0 fix(#1529): emit runtime-native instruction file from new-project 2026-06-22 09:32:29 -04:00