579ad30eaec3fd0e1d83ebdbc19e2ce3444f2aba
123 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9f0d785b61 |
fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows (#2027)
* fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows gsd-executor.md referenced sdk/src/query/QUERY-HANDLERS.md and reapply-patches.md referenced sdk/dist/cli.js — both retired with the SDK (ADR-0174). AI runtimes that locate doc refs via filesystem search ran find /, which on Git Bash for Windows traverses the whole drive (14h+, orphaned find.exe, 4M+ handles). - agents/gsd-executor.md: drop dead QUERY-HANDLERS.md ref. - workflows/reapply-patches.md: drop dead sdk/dist/cli.js clause. - tests/no-dead-sdk-refs.test.cjs: regression guard — no sdk/src|dist|handlers file refs in agents/workflows/references markdown. Closes #2020 * docs(#2020): backfill changeset pr 2027 |
||
|
|
1bfadec2d0 |
fix(#1921): preserve verify-work state across gap-closure + defer follow-ups (#2025)
* fix(#1921): preserve verify-work state across gap-closure + defer follow-ups Resuming /gsd:verify-work after /gsd:execute-phase --gaps-only lost the verification state: UAT ## Gaps still read status: failed even after their fix plans executed, so verify-work re-diagnosed them as fresh blockers, spawned a new gap plan, and reported only the new plan verified. A new state contract links each gap to its fix plan so fixed gaps are recognized. - verify-work.md gap YAML: stable gap_id (G-{phase}-{N}) per gap. - plan_gap_closure planner prompt: each *-PLAN.md tags the gap_ids it addresses in its frontmatter (gap_closure: true, gap_ids: [...]). - new reconcile_gaps step (run at resume_from_file entry): marks a gap status: resolved when its plan has a matching *-SUMMARY.md — fixed gaps are not re-diagnosed and do not spawn new gap plans; a re-reported break is treated as a fresh regression with a new gap_id. - deferred-follow-up branch: a future-work idea (signals: 'later', 'next version', 'out of scope', ...) is captured to UAT ## Deferred Follow-Ups instead of becoming a blocking gap/plan. - workflow-size baseline recaptured (verify-work LARGE tier, 38221/61440). Closes #1921 * docs(#1921): backfill changeset pr 2025 * fix(#1921): balance <step> tags in verify-work.md via 4-backtick outer fence The plan_gap_closure step nested a ```yaml example inside a ``` block; same-length fences made stripFencedCode close the outer block early, swallowing the step's </step> (14 opens / 13 closes). Widen the outer fence to 4 backticks so the nested yaml is contained. Regenerate golden fixtures + size baselines. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ef8a3e27d4 |
fix(#1864): balance <step> tags in settings-advanced.md §8 Model Policy (#2014)
* fix(#1864): balance <step> tags in settings-advanced.md §8 Model Policy §8 Model Policy ended with </step> but had no matching opening tag (5 opens / 6 closes), leaving it as loose inter-step content. Add the missing <step name="model_policy"> opener so the section is a proper step. - gsd-core/workflows/settings-advanced.md: add <step name="model_policy"> - tests/workflow-step-tag-balance.test.cjs: regression guard — every top-level workflow must have balanced <step>/</step> (fenced code stripped), plus a focused assertion that §8 is wrapped in model_policy. - goldens + workflow-size baseline recaptured. Closes #1864 * docs(#1864): backfill changeset pr 2014 |
||
|
|
a62079b2da |
fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR (#2024)
* fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR The gsd_run preamble resolved the Claude global install only at $HOME/.claude/gsd-core/bin/, but the installer honors CLAUDE_CONFIG_DIR — so a global install redirected via CLAUDE_CONFIG_DIR was invisible to every gsd_run call (every command failed with 'gsd-tools.cjs not found'). The Claude resolver arm now uses ${CLAUDE_CONFIG_DIR:-$HOME/.claude}, matching the installer + the other runtimes' ${VAR:-default} pattern. Default $HOME/.claude behavior is unchanged. - _runtime-launcher.snippet.sh: Claude arm honors CLAUDE_CONFIG_DIR. - sync-runtime-launcher.cjs re-run: 95 workflows/agents re-synced. - review.md / discuss-phase.md: trimmed to stay under their byte budgets. - runtime-launcher-parity.test.cjs: (A) substring updated for the new form + explicit #1865 assertion that the snippet honors CLAUDE_CONFIG_DIR. - goldens + size baselines recaptured. Closes #1865 * docs(#1865): backfill changeset pr 2024 |
||
|
|
23254ca5a7 |
fix(#1936): reconstruct OpenCode review from JSON events (#1992)
* fix(#1936): reconstruct OpenCode review from JSON events; diagnosable empty-output stub On a large review prompt, OpenCode's default `build` agent runs a few read tool calls then ends its turn with zero output tokens (reason:"stop", output:0), so `opencode run --format default` emits empty stdout. The reviewer block redirected stderr to /dev/null and wrote a generic "failed or returned empty output" stub — so the phase silently lost its second independent reviewer with no diagnostic and no timeout. Rewrite the OpenCode reviewer block to invoke `--format json` as the primary call and reconstruct the review from the assistant `text` parts (jq). Capture stderr to a `.err` sidecar (mirrors the Codex block). When the agent emits no text, surface the stop reason, output-token count, and stderr so the failure is diagnosable. Gate the stub on the extracted CONTENT, not the output file size — an empty jq extraction still prints a lone newline that a `[ -s file ]` check would treat as populated. Document the wall-clock timeout as a Bash-tool param (macOS lacks GNU timeout; opencode has no native timeout flag). review.md was already at the DEFAULT size-tier ceiling (40956/40960), so the fix cannot fit without reclassifying it into the LARGE tier (it is a multi-reviewer orchestration file that outgrew "focused single-purpose"; 43.4 KB sits well under the LARGE high-water mark). Recapture the 16 golden-install fixtures — the diff is exactly one review.md hash per runtime. Regression block folded into review-default-reviewers-workflow.test.cjs (new bug-NNNN test files are not accepted). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1936): add changeset * test(#1936): property-test the OpenCode review jq reconstruction Address the re-review's one actionable finding: the jq JSON-event → text reconstruction had no fast-check property test. Add tests/opencode-review-reconstruction.property.test.cjs. It extracts the two shipped jq programs (OPENCODE_REVIEW, OPENCODE_DIAG) verbatim from gsd-core/workflows/review.md and runs the real jq — not a reimplementation — so the shipped logic is what gets tested. Properties: the reconstructed review equals the newline-join of every assistant text part (order preserved); a stream with no text part reconstructs to empty (drives the #1936 stub); null/absent text parts are dropped, never rendered as "null". Plus example-based coverage of the diagnostic edges the reviewer cited: missing .tokens.output and no step_finish degrade to "?"; non-JSON stdout makes jq fail rather than masquerade as a review. Verified the invariant has teeth (a comma-join jq fails the property). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1936): skip jq reconstruction property test when jq is absent The property test shells out to `jq`, which GitHub's windows-latest runners do not ship (macOS/Linux runners do). `execFileSync('jq')` therefore ENOENT-failed the whole file on `test (windows-latest, *)`. Probe `jq --version` at load and skip the suite when jq is not on PATH — the reconstruction logic is platform-independent, so the assertions still run in full on every jq-present runner (mirrors how golden-install-parity skips on win32). Verified: jq present → 7 pass; jq removed from PATH → 7 skipped, 0 fail. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1936): skip jq reconstruction property test on Windows, not just when jq is absent The prior guard skipped only when `jq` was absent from PATH — but the windows-latest runners DO ship jq, so the suite still ran there and failed with `jq: parse error: Invalid numeric literal` (confirmed from the CI job log). Root cause is Node's child_process argument quoting mangling the jq program (it embeds double quotes) on Windows, not the shipped review.md logic — the macOS/Linux legs pass. Gate the suite on `process.platform === 'win32'` (still also skipping when jq is absent), mirroring golden-install-parity's win32 skip. Logic is platform-independent and fully asserted on every macOS/Linux CI leg. Verified: macOS → 7 pass; simulated win32 → skips. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8de2ff9121 |
feat(#2008): generic command-exit-zero gate-predicate evaluator (#2011)
* feat(#2008): add generic command-exit-zero gate-predicate evaluator Third-party capability gates declared via check.predicate were rendered for display but never evaluated (only built-in check.query gates fired; the security capability's gate worked solely via a hard-coded ship.md branch). Add a generic, deps-injected gate-predicate evaluator (src/gate-predicate-evaluator.cts) that dispatches by predicate.kind. Built-in kind: command-exit-zero — runs a bounded sh -c command at the project root (via shell-command-projection.execTool), inherits env, exit 0 => pass, non-zero => block, timeout => block, fail-closed. Wire a 'check predicate' subcommand into check-command-router.cts and extend the three generic workflow gate-dispatch sites (execute:wave:post, execute:post, plan:post) to route check.predicate gates to the new evaluator. The two-step gate contract (command-failure => onError; block => halt) is unchanged. - src/gate-predicate-evaluator.cts: pure leaf, KIND_TABLE extensible - src/check-command-router.cts: cmdCheckPredicate + buildPredicateDeps + parsePredicateFlags - docs/adr/2008-*, docs/reference/gate-predicates.md, docs/how-to/command-exit-zero-gate.md - tests: 38 unit + integration tests (exit mapping, timeout, interpolation, property-based bijection, malformed-predicate fail-closed, real subprocess e2e) Closes #2008 * docs(#2008): backfill changeset pr number 2011 |
||
|
|
8f2ebbe9bf |
feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap). --gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): backfill changeset PR number (#1996) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): drop Gemini CLI from issue templates (review nit) Removes the sunset Gemini CLI runtime from the two GitHub issue-template runtime lists that the removal PR missed, per @davesienkowski's review nit: - feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise request a feature for a runtime GSD no longer supports) - bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json retrieval-help line Leaves the post-removal templates fully consistent with the Antigravity redirect. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
55604e9124 |
fix(#1906): require node-test clean-fixture causation control (#2001)
* fix(#1906): require node-test clean-fixture causation control The node-test fail-first proof accepted a deceptive content-independent negative test — one that reds merely because GSD_PROHIB_SUBJECT is set, ignoring the subject's content — whenever no cleanFixture was supplied, because #1346's causation control was opt-in. The proof's observed signal (RED) thus diverged from its target (RED caused by content) by default. Make the causation control mandatory for the node-test kind: a descriptor that omits cleanFixture is un-provable (fail-closed), never accepted under the weaker violation-only proof. When a clean fixture is present, fail-first is proven exactly as before (RED on violation AND non-vacuous GREEN on clean). The lint-rule kind is unchanged (its subject IS the linted file; no GSD_PROHIB_SUBJECT indirection). Breaking (Hyrum): a previously-green node-test prohibition with no clean fixture now hard-gates — blast radius is zero in-tree (no node-test prohibition ships today; only the lint-rule local/no-source-grep dogfood). Supersedes ADR-1606 Decision 4 / ADR-550 #1346 addendum's opt-in. Closes #1906 Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ * docs(#1906): supersede the #1346 opt-in causation control (mandatory for node-test) Record the node-test mandatory-causation-control supersede across the governing surfaces: - ADR-1606 (the enforcement decision-of-record): addendum + Decision 4 annotated + the "Mandatory causation control — REJECTED" alternative flipped to accepted (premise no longer holds: zero in-tree node-test consumers). - ADR-550: the 2026-06-21 #1346 "Why opt-in, not required" paragraph marked SUPERSEDED, pointing at ADR-1606. - spec-phase.md: check_clean_fixture is now REQUIRED for node-test (was "optional"). - CONTEXT.md: PROHIB.enforce.causation predicate updated. Regenerated the shipped-artifact cascade from the spec-phase.md edit (+149 B, well under the 40960 cap): 16 golden-install-parity fixtures and the workflow size baseline. Refs #1906 Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ |
||
|
|
e5ef323b15 |
feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command
Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.
* feat: add /gsd-start smart-entry command
State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.
- src/smart-entry.cts: deterministic situation classifier (no-project,
paused, blocked, verify-failed, needs-first-phase, planning, executing,
verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
dispatcher presenting an AskUserQuestion menu (with --text fallback for
non-Claude runtimes) and dispatching to existing commands. Falls back
to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
(markdown-layer invariants + every emitted command resolves to a real
slash command).
Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.
* refactor: rename smart-entry command to /gsd:next
Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.
All affected tests (188) pass; lint:ci clean.
* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)
Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.
- detectSignals now reads total_phases + percent from nested progress{}
first, then scalar fm, then body; current_phase falls back to the
body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
body Phase field) covering verify-pending + executing situations.
Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.
* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)
Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.
Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
layer is down — read .planning/STATE.md directly with the Read tool
and synthesize a minimal situation + actions menu so /gsd:next stays
useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
/gsd:progress as before.
Matches the direct-read resilience the live agent already did by hand.
* docs: add gsd-next skill surface
* chore: trigger no-mistakes validation
* no-mistakes(review): Fix smart-entry phase ordering
* no-mistakes(review): Fix decimal smart-entry phase ordering
* no-mistakes(test): Fix smart-entry next test contracts
* no-mistakes(document): Docs synced for smart entry
* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: shorten next.md description and update golden install parity fixtures
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: trigger no-mistakes validation
* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files
Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.
* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: regenerate golden install parity fixtures for /gsd:next
Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.
* Fix smart-entry verify-failed phase scoping and empty resolve shim step
Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.
* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: recapture all 16 golden fixtures with updated smart-entry.md hash
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: regenerate fixtures + inventory manifest after rebase onto next
Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next
Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).
Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)
The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs
Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo
Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
recommended action, 1-4 unique-id /gsd:* actions (previously the
one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout
Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:
- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
consolidation file doing dozens of real installs — 41s even on a fast Mac
(much worse on the slow Windows I/O path), plus an install-heavy cluster.
Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.
Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
9d7c046eae |
refactor(#1852): lazy-split plan-phase.md into steps/ (#1934)
* refactor(#1852): lazy-split plan-phase.md into steps/ Extract 3 self-contained, rarely-hit sections into gsd-core/workflows/plan-phase/steps/ via lazy 'Read and execute' pointers (mirrors execute-phase/steps/, ADR-1610 progressive disclosure — no eager @-import): closed-phase-gate (1.5), prd-express-path (3.5), windows-troubleshooting. Byte-invariant: plan-phase.md 94459 -> 89775 (-4684), each step < 32 KiB anchor, resolved instructions unchanged. Scope note: only 3 sections were extractable. plan-phase.md is guarded by a dense net of content-presence tests (plan-bounce/enh-3209/phase6-planning-capabilities assert specific sections/flags inline) that block extracting the larger blocks without also refactoring those tests — durable low-80s headroom is deferred pending maintainer re-scope (issue #1852). Cascade: size:baseline regen, 16 golden-install-parity fixtures regen, INVENTORY in sync. Full plan-phase test surface green (4129/4130, 0 fail); lint:ci exit 0. * docs(changeset): Changed fragment for #1934 (plan-phase lazy-split) * docs(changeset): mark #1934 fragment docs-exempt (internal workflow refactor) * fix(#1852): add gsd_run launcher preamble to prd-express-path step The extracted prd-express-path.md calls gsd_run but the canonical launcher preamble lived in the parent plan-phase.md — runtime-launcher-parity (#373) walks workflows/ recursively and requires every .md using gsd_run to carry exactly one preamble + the $HOME/.claude fallback arm (same as the existing execute-phase/steps/ files). Injected via scripts/sync-runtime-launcher.cjs; golden fixtures + size baseline regenerated. plan-phase.md unchanged (89775). --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
7bef6a6496 |
fix(#1863): use named flags for state.* calls in executor + workflows (#1873)
* fix(#1863): use named flags for state.* calls in executor + workflows The named-only state-command router (parseNamedArgs) silently drops positional args, so state.cjs threw its required-arg error and metrics/decisions/blockers/session continuity were never recorded. Convert record-metric / add-decision / add-blocker / record-session in agents/gsd-executor.md to the named-flag form (mirroring execute-plan.md), and fix the two remaining positional record-session calls in gsd-core/workflows/milestone-summary.md and forensics.md. Recapture the golden-install-parity fixtures and size baselines for the edited files. Also fix a pre-existing detached-rebuild handle leak in tests/graphify-auto-update.slow.test.cjs: three dispatch tests returned after observing only the synchronous "running" status without awaiting the detached rebuild's terminal state. That leak was latent until the new #1863 regression block's added runtime shifted --test-force-exit timing and surfaced it as a non-zero chunk exit. The three tests now await terminal status via the file's existing waitForBuildStatus helper. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1863): add changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
5657994702 |
fix(#1716): route resume_from_file to complete_session when no pending tests remain (#1722)
* fix(#1716): route resume_from_file to complete_session when no pending tests remain When a UAT session has status:partial with blocked_count>0 and pending_count==0 (all remaining tests are blocked, none are pending), resume_from_file found no [pending] test and terminated silently — never routing to complete_session. This blocked the issues==0 auto-transition path even when there were zero code defects. Guard clause added immediately after the find-pending step: if no [pending] test is found, route to complete_session. complete_session then correctly sets status:partial (because blocked_count>0) without presenting further tests. Closes #1716 * chore(#1716): add changeset fragment and regenerate golden-install-parity fixtures Changeset fragment for PR #1722 (type: Fixed). Golden-install-parity fixtures regenerated for all 16 runtimes — the workflow fix shifts verify-work.md's byte-stable hash in the golden manifest. Regenerated via UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs. |
||
|
|
92091d71f2 |
fix(#1871): wire phase archival end-to-end (phases archive cmd + default + atomic) (#1924)
Follow-up to #1919 (archive-then-remove core). Closes the remaining #1871 acceptance criteria so phase history is preserved across the full milestone lifecycle, not just at phases.clear: - #2 src/milestone.cts + src/phases-command-router.cts: extract shared archivePhaseDirectories() helper; add cmdPhasesArchive (the previously half-wired phases.archive alias now routes instead of erroring Unknown). - #4 gsd-tools.cjs + src/milestone.cts: milestone complete archives phase dirs by default (--no-archive-phases opts out; --archive-phases is now a harmless no-op). complete-milestone.md updated to drop the redundant manual Yes/Skip archive prompt. - #3 gsd-core/workflows/new-milestone.md: §6 stages the archive move + source removal (git add .planning/milestones/ .planning/phases/) in the same commit as the milestone start, so the archive lands atomically — no orphaned uncommitted deletions, no un-archived dirs inherited. - docs/CLI-TOOLS.md (+ ja/zh/ko/pt) + help/modes/full.md: flag accuracy. - tests: phases archive command (#2) + milestone complete default archive / --no-archive-phases opt-out (#4). Goldens + workflow size baseline refreshed. Closes #1871 |
||
|
|
3c13903dcd |
feat(#1866): agent-side self-load of configured agent_skills
Each of the 22 consumer agents now self-loads its configured agent_skills in its mandatory init step, so .planning/config.json agent_skills.<type> reaches the agent on every runtime — including Cursor and /gsd-autonomous, where Skill()-delegated workflow bash init did not reliably execute. - gsd-core/references/agent-skills-bootstrap.md: shared contract (query + Read + dedup guard that skips when <agent_skills> is already in the prompt, so Claude's orchestrator-side injection never doubles) - 22 agents/gsd-*.md: one self-load line naming the agent's own type - gsd-core/workflows/autonomous.md: note that delegated agents self-load - tests/agent-skills-bootstrap.test.cjs: regression + parity (CONSUMER_AGENTS bijection + fast-check property) — Generative-Fix-Divergence guard - docs: ADR-1866, CONFIGURATION dual-injection How It Works, INVENTORY row, Changed changeset Closes #1866 |
||
|
|
1bd04e1565 |
docs(#1847): add Claude Sonnet 5 changeset for next-line changelog + refresh stale model examples
The 1.6.1 forward-port (#1851) used the no-changelog opt-out instead of carrying a changeset, so Sonnet 5 — unlike every other 1.6.1 fix (#1580/#1591/#1693 whose fragments live on next) — had no fragment and would be MISSING from the 1.7.0 changelog. Add the fragment so the release render reflects current shipping code. Also note the bold-checklist form in the #1591 fragment, and refresh two stale claude-sonnet-4-6 illustrative examples to current IDs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
da37986cd0 |
fix(#1847): resolve standard tier to claude sonnet 5
Point the sonnet/standard tier at Claude Sonnet 5 (`claude-sonnet-5`,
GA 2026-06-30) across the Anthropic-backed runtimes and provider presets,
replacing the superseded `claude-sonnet-4-6`. Mirrors the change into the
CONFIGURATION.md and settings-advanced.md runtime-defaults tables (the
#3229 catalog↔docs parity gate) plus the pt-BR/zh-CN translations, and
updates the tests that pin the old ID. Regenerates the workflow size
baseline for the (smaller) settings-advanced.md.
Scope is Sonnet only — opus/haiku IDs are untouched. Prepared as a 1.6.1
hotfix off the v1.6.0 tag.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit
|
||
|
|
93e5d2dd84 |
fix(#1525): skip deferred phases on autonomous reruns (#1846)
* fix(#1525): skip deferred phases on autonomous reruns * chore(#1525): add changeset fragment * chore(#1525): fix changeset body * test(#1525): refresh install parity fixtures * test(#1525): shrink autonomous workflow * test(#1525): refresh autonomous baselines * test(#1525): tolerate Windows temp cleanup flake |
||
|
|
dae7f81482 |
fix(#1528): drop next-phase guidance from security-blocked verify-work presentation (#1687)
* fix(#1528): drop next-phase guidance from security-blocked verify-work presentation When security enforcement blocks phase advancement (no SECURITY.md produced), the verify-work presentation told the user advancement was blocked but still offered `/gsd:plan-phase {next}` and `/gsd:execute-phase {next}`, competing with the current-phase fix. Remove those two next-phase lines so the blocked state routes only to the current-phase resolution (secure-phase, ui-review). The post-transition presentation — reached only after the completion contract passes — still offers next-phase planning, which is the correct place for it. Regression coverage added to tests/ui-review-next-guidance.test.cjs: the security-blocked block must not offer next-phase actions, and the post-completion block must still offer them. Regenerated workflow size baseline. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1528): add changeset for security-blocked next-phase fix Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1528): recapture golden-install-parity fixtures for verify-work.md change Rebased onto next; verify-work.md's installed hash changed across all 16 runtime fixtures. Diff confined to the single gsd-core/workflows/verify-work.md key per runtime. Assert mode 16/16 green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
1f649838b8 | Merge branch 'next' into fix/1698-codex-output-last-message | ||
|
|
18995380ce |
feat(#1154): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1738)
* feat(verify-phase): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1154) Carry the edge-probe's existing `backstop` (non-inferable) tier through the plan-phase projection as a structured flat-scalar marker instead of a prose parenthetical, and make verify-phase abstain -> human_needed (never silent-pass) on a backstop truth it cannot confirm with explicit evidence. Truth-axis mirror of #644's prohibition judgment-tier (ADR-550 D4). Engine (deterministic, CI-tested per ADR-550 D5 — never the LLM verdict): - src/probe-core.cts: truthStatement/truthVerification normalizers, projectTruths (conservative serializer), dispositionForUnverifiableTruth (backstop+no-evidence -> unverified/flagged/insufficient_spec; backstop+evidence -> green; inferable -> green, the over-abstention guard). - src/roadmap.cts: coerceTruthToString now reads `statement` first so an object-form backstop truth is surfaced, not dropped (Hyrum backward-compat for truth-readers). Workflow/agent/docs: plan-phase emits the structured marker (flat scalar, ADR-550 #1278); verify-phase + gsd-verifier add the abstain arm; new references/honest-verifier.md; FEATURES/COMMANDS document insufficient_spec; ADR-550 amended (truth-axis D4 mirror). Decisions adopted (trek-e review): insufficient_spec feeds existing human_needed with a distinguishable reason (no new VERIFIER_STATUS); changeset Changed; round-trip parity test; abstain-on-unconfirmed-backstop regression test red-first. Implementation notes (deviations from the issue's proposed file list, verified live): - frontmatter.cts needs no change — its flat parser already round-trips object-form truths. - verify.cts needs no change — it grades artifacts/key_links structurally; truths are LLM-graded at the workflow layer, so consumption lives there + the deterministic helper. - No CJS<->SDK hand-sync — the SDK seam was retired (ADR-0174); src/*.cts is sole source. Regenerated artifacts: golden-install-parity fixtures, INVENTORY-MANIFEST, size baselines. * chore(#1154): add changeset (Changed) for honest verifier User-facing changelog fragment for #1738. Typed `Changed` (not `Added`) per trek-e review condition 3 — the verify behavior shifts for backstop-bearing specs (a confident silent `passed` becomes `human_needed`), which is user-visible even though the schema marker is additive. * docs(#1154): score-formula also excludes abstained insufficient_spec truths (review nit-1) trek-e review nit: the verify-phase score sentence said PRESENT_BEHAVIOR_UNVERIFIED truths were "the only ones excluded" from verified_truths. Post-#1154 an abstained `insufficient_spec` backstop truth is also excluded (it is not ✓ VERIFIED and routes to human_needed). Behavior was already correct; this tightens the wording. Regenerated golden-install-parity fixtures + workflow-size baseline for the touched verify-phase.md. (Nit-2 — a dedicated insufficient_spec_items frontmatter list — is intentionally not taken: the current design is ADR-550-D4-conformant, the abstain cause rides as a distinguishable report reason, and adding it would exceed the approved scope.) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
ac001be49e |
fix(#1778): use 1.6 named-flag frontmatter.set form in thread workflow (#1816)
* fix(#1778): use 1.6 named-flag frontmatter.set form in thread workflow The thread workflow's CLOSE and RESUME branches called frontmatter.set with the pre-1.6 fully-positional shape (frontmatter.set <file> <field> <value>). Since 1.6 the dispatcher (gsd-tools.cjs) parses the file positionally and reads field/value from the named flags --field/--value via parseNamedArgs; the positional form leaves field/value undefined, cmdFrontmatterSet errors 'file, field, and value required', and the status/updated writes are silently skipped. Closing a thread never marked it status: resolved and resuming never marked it status: in_progress. Switch all four sites (CLOSE status+updated, RESUME status+updated) to the 1.6 hybrid form that verify-work.md already uses: frontmatter.set <file> --field <field> --value <value> Add a regression test with three guards: (1) behavioral — the named-flag form writes the field while the positional form errors with the documented message and does not mutate the file; (2) workflow parity — no workflow under gsd-core/workflows/ emits the positional form, so a future edit that reintroduces it anywhere fails CI; (3) thread-specific — CLOSE writes status: resolved and RESUME writes status: in_progress via the named flags. * docs(#1778): add changeset fragment for thread workflow frontmatter fix * docs(#1778): fix unclosed inline-code backtick in changeset fragment * fix(#1778): move regression into owning test + regen baselines lint-regression-test-names rejects new bug-NNNN-*.test.cjs files; move the #1778 regression (behavioral named-vs-positional + workflow-parity scan + thread CLOSE/RESUME assertions) into tests/frontmatter-cli.test.cjs, the canonical home for frontmatter CLI regressions, and delete the standalone file. frontmatter-cli.test.cjs already carries the allow-test-rule exemption for workflow .md content tests. gsd-core/workflows/thread.md ships to every runtime and is size-tracked, so recapture the 16 golden-install-parity fixtures (thread.md hash) and the per-file workflow size baseline (thread.md 12400 -> 12464) via UPDATE_GOLDEN=1 and npm run size:baseline. |
||
|
|
2bada4de1a |
fix(#1698): capture codex review via --output-last-message, not stdout
The Codex reviewer in review.md captured the review by redirecting codex exec's stdout to the review file. On Windows, codex writes process-teardown output to stdout after the final agent message, so that noise was appended to a non-empty file and slipped past the `[ ! -s ]` empty-output guard as a silently polluted review (consumed by severity extraction and the plan-review-convergence gate). Capture the final message via codex's own `-o/--output-last-message <FILE>` and discard stdout. The #1115 contract is preserved (stderr to .err, capability-gated $CODEX_BYPASS_FLAG, --ephemeral, --skip-git-repo-check) and the empty-output fallback still fires when codex leaves no/empty output. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b0d5ca3379 |
feat(#1517): support custom reviewer instances for /gsd:review (#1766)
* feat(#1517): support custom reviewer instances for /gsd:review Add a bounded review.reviewer_instances config surface so one model-capable adapter (e.g. opencode) can run as several independent reviewer identities in a single /gsd:review pass. Instances participate only via review.default_reviewers, expand before built-in slugs, are available iff their cli is detected, and a non-matching entry is a hard error (typo must be loud). >=2 same-cli instances emit a shared-adapter caveat in REVIEWS.md. Default path with no instances is byte-for-byte unchanged. Single-source instance->cli resolution lives in resolveReviewerSelection / normalizeReviewerInstances (parity-locked in tests/review-reviewer-instances.test.cjs). cli validated against KNOWN_REVIEWER_SLUGS only (never arbitrary shell); model/agent opaque, never shell-interpolated. Closes #1517 * chore(#1517): backfill changeset pr:1766 --------- Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
cf2e66b39e |
feat(#1708): typed documentation-sourced #853 dispatch-flatten (ADR-1239 Phase B) (#1719)
* feat(#1708): typed documentation-sourced #853 dispatch-flatten Graduate the #853 orchestrator-backgrounding decision from a scattered RUNTIME==='codex' prose check to a typed, documentation-sourced engine decision. Adds a backgroundDispatch dispatch sub-axis (sourced per host: codex+cursor documented true, 9 documented false, 5 undocumented), shouldFlattenDispatch(dispatch) (inline UNLESS background && backgroundDispatch, fail-closed), and a gsd_run query dispatch-should-flatten the plan/execute workflows call. Cursor is newly background-eligible per its docs (inline->background) — a documentation-justified behavior change. No RUNTIME-name residue for this decision. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1708): backgroundDispatch citations in matrix + CONTEXT note Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1708): address review findings on typed dispatch-flatten Code/adversarial review: convert the manager.md/autonomous.md Compound Action preamble from hardcoded 'On Codex' to FLATTEN-based branching (the handlers already use the query; the preamble contradicted them and was wrong for cursor); make shouldFlattenDispatch null-safe + type-honest (accepts raw 'undocumented' registry values); make backgroundDispatch a required descriptor field (matching its siblings, all 16 carry it); strengthen the config.runtime behavioral test; update the bug-853 prose-pin test + comment. Security review clean; Codex confirmed no fail-open. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1708): backfill backgroundDispatch in role:runtime test fixtures Making backgroundDispatch a required descriptor field broke role:runtime fixtures in capability-manifest-version/capability-registry/host-integration-descriptors tests that build a dispatch object without it (caught by full gsd-test, not scoped npm test). Backfill backgroundDispatch:false into the well-formed fixtures; the deliberately-malformed 'required-field' test fixture is left malformed by design. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1708): update fix-1521 dispatch-gating assertion to the FLATTEN gate fix-1521 pinned the codex-specific run_in_background prose that #1708 graduated to the typed dispatch-should-flatten/FLATTEN gate. Update its assertions to verify FLATTEN=false gating (not a runtime name) + that the old RUNTIME===codex gate is gone. Caught by full gsd-test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1708): add changeset for typed dispatch-flatten Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1708): remove stray temp PR-body file Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1708): add issue ref to bug-853 allow-test-rule annotations ADR-456 requires every allow-test-rule exemption to carry a see #NNN reference; the source-text-is-the-product annotations added when migrating the prose assertions lacked it (lint-tests CI gate). Add (see #1708). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
47906b052d |
fix(#1520): randomize mktemp temp paths on BSD/macOS (XXXXXX must be path-final) (#1550)
* fix(#1520): randomize mktemp temp paths on BSD/macOS (XXXXXX must be path-final) BSD/macOS mktemp only substitutes the XXXXXX template when it is the final path component. Templates like `...-XXXXXX.json` / `gsd-pr-body.XXXXXX.md` return a LITERAL `XXXXXX` path (no randomization) on macOS, so concurrent workflow runs collide on the same temp manifest/body file — one run can overwrite or consume another's. Reproduced on macOS: the second call to the suffixed template fails `mkstemp: File exists`. Fix: use a suffixless `XXXXXX` template (so it IS the final component), then rename to add the intended extension — portable across BSD + GNU userlands, no GNU-only `--suffix` flag. Empty-file-then-write semantics are preserved at every site. Affected workflow temp files: - execute-phase.md: gsd-worktree-wave-*.json (wave worktree manifest) - quick.md: gsd-quick-worktree-*.json - spec-phase.md: edge-probe-reqs-*.json - ship.md: gsd-pr-body-*.md - profile-user.md: gsd-profile-answers-*.json, gsd-profile-analysis-*.json The execute-phase.md edit uses a compact intermediate var + trailing comment to stay under the ADR-857 phase-6 size ceiling (93166); regenerated the workflow size baseline accordingly. Validated on macOS: 20 concurrent calls yield 20 unique randomized paths. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1520): add changeset fragment (Fixed) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1520): add fail-first workflow-prose guard for mktemp XXXXXX suffix Repo-wide scan of gsd-core/workflows/**/*.md that fails on any mktemp template whose XXXXXX run is followed by a filename suffix (the BSD/macOS non-randomizing form). Fails on the six pre-fix instances and passes on the fix, and locks the copy-paste-prone idiom out of future workflows. Mirrors the bug-637 hardcoded-$HOME workflow guard. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1520): rename regression test to fix- prefix (regression-test-names lint) New tests/bug-NNNN-*.test.cjs files are banned by the lint-regression-test-names ratchet; use the fix- prefix (matches the fix-1445 precedent). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1520): add issue ref to allow-test-rule exemption (ADR-456 lint) lint-allow-test-rule-refs requires every new `allow-test-rule:` comment to carry a #NNN reference (don't allowlist). Add (#1520) to the source-text exemption. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#1520): abort touched mktemp chains on failure (|| exit 1) Per review: the VAR=$(mktemp …) && mv … && VAR=… chains dropped the issue's suggested failure guard. If mktemp fails, $VAR is empty and the subsequent mv/write lands on an unintended relative path. Add `|| exit 1` to all six touched chains so a mktemp failure aborts the snippet. Regenerated the workflow size baseline for the slightly longer lines. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1520): rebase onto next — regen size baseline + describe rename Resolve the workflow-size-baseline.json conflict from next advancing by regenerating from the current workflow sizes. Also rename the test describe from `bug #1520` to `#1520` (the file uses the fix- prefix) per review nit. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#1520): harmonize profile-user mktemp to ${TMPDIR:-/tmp} (review nit) The two profile-user.md temp sites this PR already rewrites kept a hardcoded /tmp while the four sibling workflows use ${TMPDIR:-/tmp}. Harmonize for consistency and macOS-correctness (some sandboxes have no writable /tmp). Regenerated the workflow size baseline. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1520): regen size baseline after rebase onto next --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
d101daff30 |
fix(#1516): expose adaptive model_profile in /gsd-new-project AI Models prompt (#1654)
* fix(#1516): expose adaptive model_profile in /gsd-new-project AI Models prompt Both onboarding paths (Step 2a auto-mode + Step 5 interactive) enumerated only 4 profiles (Balanced/Quality/Budget/Inherit), omitting 'adaptive' even though the model catalog (model-catalog.json profiles) and docs/CONFIGURATION.md register 5. Mirrors the proven /gsd:settings two-question split (#3784): Q1 routes between Adaptive/Standard-tier/Inherit; Q2 (conditional on Q1=Standard) picks Quality/Balanced/Budget — keeping every AskUserQuestion within the 4-option cap. Both config-new-project example payloads now list adaptive. Regression cases folded into the owning tests/new-project-mvp-prompt.test.cjs (per the lint-regression-test-names ban on new top-level bug-NNNN files): each AI Models prompt makes adaptive reachable, all 5 profiles reachable, 4-option cap honored, both example enums include adaptive, brace balance. Workflow size baseline bumped (new-project.md 62324 -> 66138 bytes; still well under the XL hard cap). * chore(#1516): backfill changeset pr ref to 1654 |
||
|
|
77c7b4fc9d |
fix(#1522): enforce canonical verification before phase transition (#1548)
* fix: require fresh phase verification before transition * no-mistakes(review): Fix canonical verification closeout gates * no-mistakes(review): Fix verify-work frontmatter promotion command * no-mistakes(review): Fix stale verification gates * no-mistakes(review): Fix canonical verification routing gates * no-mistakes(review): Fix verification dependency and runtime routing gates * no-mistakes(review): Block stale verification bypasses * fix: handle large init manager outputs in verification workflows * chore: update changeset pr number * fix(verify-work): use fresh verification.status for stale gate The stale check after UAT used phase_completion.verification_status from session-start INIT while human_needed promotion already queried fresh verification.status. Align the stale gate with the canonical query so mid-session verification refresh is not ignored. * fix(init): skip roadmap-checked phases when selecting next_phase Roadmap-only phases without a disk directory were still promoted to next_phase when their checkbox was already checked. Exclude checkboxComplete phases so progress routing does not point at work the roadmap already marks done. * fix: gaps_found not overridden by stale, transition uses canonical verification - verification.cts: check gaps_found before stale so gap-closure routing is not masked by a newer summary mtime - phase.cts: remove redundant findStaleVerificationSummary — readVerificationStatus already handles stale detection - transition.md: replace raw grep on file content with verification.status query to avoid false-positive blocks from body text matching * ci: retrigger tests after rebase * fix(transition): replace gsd_run advisory check with awk frontmatter extraction The runtime launcher is not defined until the update_roadmap_and_state step bash block (~line 165). The early verify_completion block used gsd_run to query verification.status, which violated the runtime-launcher-parity test: 'preamble appears AFTER the first gsd_run reference'. Replace the gsd_run call with an awk-based frontmatter extractor that reads only the status: field between the two --- fences. This avoids both the preamble-ordering constraint and the original false-positive grep bug where body text like 'previous_status: gaps_found' would match a full-text regex. The phase.complete gate at update_roadmap_and_state is the canonical enforcement point; this early check is advisory only. Also update workflow-size-baseline.json for the updated transition.md size. Fixes: runtime-launcher-parity test (B) Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> * fix: re-check verification under planning lock in phase complete Move readVerificationStatus into withPlanningLock so stale verification cannot slip through when a SUMMARY.md is written between the gate and the roadmap/state mutation. Return the blocked status from the lock callback and emit the error after release to avoid leaving .lock behind. * fix(transition): gate on canonical verification.status including stale Replace awk frontmatter read with verification.status query so transition blocks when summaries are newer than VERIFICATION.md, matching phase.complete and other workflows (autonomous, progress, verify-work). * Fix workflow verification gates for yolo transition and stale routing Require VERIFY_STATUS passed before yolo/interactive transition advance. Route stale verification recovery to verify-work, matching canonical projection. * fix(transition): use verification.status query for stale-aware advisory check The awk-based check read raw frontmatter status: passed, which misses the stale case where summaries are newer than the VERIFICATION.md file even though the frontmatter still says passed. The stale status is computed from file modification times, not stored in frontmatter. Move the preamble to the verify_completion bash block (the first block with a gsd_run call) so gsd_run query verification.status can be used for the advisory check. This gives the full readVerificationStatus logic including mtime-based staleness detection, matching the enforcement gate at phase.complete. Capture full JSON (VERIFY_JSON) so next_action can be included in the advisory output alongside the status. Also update workflow-size-baseline.json for the updated transition.md size. Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> * ci: trigger test matrix for 525b946 Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> * fix(transition): restore awk frontmatter extraction for pre-shim verification check The gsd_run launcher shim is not defined until line ~163 of transition.md, so the verification debt check at line ~80 cannot use gsd_run. Restore the awk-based frontmatter extraction that correctly reads status without needing the runtime, and restore the shim at its proper location before phase.complete. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(#1522): clarify transition verification gate wording * fix(#1522): update transition workflow size baseline * fix(#1522): update workflow-size-baseline after rebase onto next Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> * fix(#1522): guard findStaleVerificationSummary FS calls + thread opts.fs seam (review) Address review blocker B1 on #1548: findStaleVerificationSummary ran fs.readdirSync and two fs.statSync calls unguarded between readVerificationStatus's try/catch sections, so a TOCTOU race (a SUMMARY listed by scanPhasePlans then removed before statSync) or any FS error threw uncaught into callers NOT under the planning lock (init.manager / init.progress / uat-predicate). Wrap the body in try/catch degrading to 'not stale', and thread the injectable opts.fs seam (add statSync to FsLike, pass fsImpl from the caller) for parity with readVerificationStatus's no-throw contract and testability. Also adds the Verification Module glossary entry to CONTEXT.md (review B3). --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
207d8f1697 |
fix(#1626): make the security gate severity-aware via per-threat severity (#1635)
workflow.security_block_on was documented as the minimum threat severity that blocks advancement, but threats carried no severity and the auditor's threats_open count (the SECURITY.md gate field) counted every open threat regardless of severity — so the threshold had no effect, and the auditor's block_on vocabulary (open/unregistered/none) did not even match the config enum (critical/high/medium/low/none). - planner: add a Severity column to the STRIDE threat register; assign severity per threat. - auditor: read severity; reconcile the <config> block_on domain to the severity enum; redefine threats_open as the count of OPEN threats whose severity is at or above block_on (none => 0). Below-threshold opens are reported as non-blocking and excluded from threats_open. - SECURITY.md template + planning-config.md reconciled. No gate-check site changed: threats_open == 0 stays the gate everywhere; only its computation is now severity-filtered. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f9d9dfb4bc |
fix(#1627): scale security rigor by ASVS level (planner disposition + auditor depth) (#1636)
workflow.security_asvs_level was display-only — the planner hardcoded 'mitigate if ASVS L1 requires it' and the auditor only echoed the level, so L2/L3 behaved identically to L1. - New reference gsd-core/references/security-asvs-levels.md defines L1 (opportunistic), L2 (standard), L3 (comprehensive) for both planner threat disposition and auditor verification depth (higher = superset). - planner: disposition now scales with the configured ASVS level (no hardcoded L1) + @-pointer to the reference. - auditor: verification depth scales with asvs_level (L1 grep-presence, L2 boundary/vector check, L3 end-to-end trace + bypass check). - planning-config.md + INVENTORY updated; planner kept under its 48K cap by extracting the goal-backward worked example to planner-guidance.md. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
94be6d5b60 | fix(#1625): resolve security config in secure-phase.md before auditor handoff (#1633) | ||
|
|
c28cccbf85 |
Merge pull request #1595 from davesienkowski/feat/1592-plan-drift-precheck
feat(#1592): add plan:pre codebase-drift pre-check before planner runs |
||
|
|
ba96c70b14 |
feat(#1602): deterministic coverage-metadata UAT routing for verify-work
Add an optional structured `coverage:` block to SUMMARY.md frontmatter and a deterministic classifier that `verify-work` consumes to route deliverables to auto-pass vs human-UAT — replacing the rejected #1598/#1599 post-hoc heuristic. - New `src/coverage.cts` (→ bin/lib/coverage.cjs) parses the nested coverage block (extractFrontmatter can't — its `-` items are scalars-only; this is a focused parser, sibling of parseMustHavesBlock), validates each entry, and classifies into auto_passed vs present. Frozen MODE/PRESENT_REASON/ERROR_CODE typed-IR surface. Exposed via `uat classify-coverage --summary <f>`. - Auto-pass is the narrow proven case only: strict-boolean human_judgment:false AND non-empty all-`pass` verification AND zero validation errors. Everything else — judgment, empty/failing verification, malformed entry — routes to the human (fail-safe). A malformed block falls back to legacy prose extraction and surfaces an error; an absent block is byte-identical to pre-#1602. - execute-plan create_summary populates the block (fail-safe default human_judgment:true); verify-work extract_tests consumes it; create_uat_file marks auto-passed entries `source: automated`. - Templates (summary + 3 variants), CONTEXT.md predicate + glossary, INVENTORY, eslint/gitignore registration, and Diataxis docs (COMMANDS reference + USER-GUIDE explanation) updated. - Behavioral tests via the CLI (no source-grep); parser-robustness regressions for the null-entry/comment-header/mis-indent cases found in adversarial review. Closes #1602 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d1f7ba82f2 |
feat(#1592): add plan:pre codebase-drift pre-check before planner runs
Add a non-blocking, warn-only codebase-drift gate at plan:pre so a stale STRUCTURE.md is surfaced before /gsd:plan-phase spawns the planner, instead of being discovered mid-execution by the existing execute:wave:post gate. Gated on a dedicated workflow.plan_drift_precheck toggle (default true), independent of schema_drift_gate. Never blocks planning, never spawns the mapper agent at plan time. Review feedback (#1595): - Use a documented conventional-commit type (feat, not enhance) per CONTRIBUTING.md / gsd-validate-commit.sh. - Normalize the plan_drift_precheck command references to the colon prose form (/gsd:plan-phase, /gsd:map-codebase) to match plan-phase.md §5.65; registry regenerated from capability.json. - Make the test temp dirs hermetic: drain mkdtemp dirs in an after() hook via the helpers.cleanup() budget (local/no-raw-rmsync-in-tests-compliant). Closes #1592 Claude-Session: https://claude.ai/code/session_016JBiXEAofvB3prJim29XMS |
||
|
|
b2c0086c1b |
fix(#1574): resolve review — copilot instruction file is .github/copilot-instructions.md
GitHub Copilot reads repository-wide instructions only from .github/copilot-instructions.md (confirmed via GitHub Docs), not a root copilot-instructions.md. Aligns getProjectInstructionFile with the installer (runtime-config-adapter-registry installSurface 'copilot-instructions') and cites the docs source in the doc-comment. |
||
|
|
bf9bd1f4e0 | fix(#1529): emit runtime-native instruction file from new-project | ||
|
|
e12a2abfd8 |
feat(#441): add /gsd-capture --list-seeds for seed listing and audit (#722)
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed), enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way to browse or audit parked seeds on demand. This adds a read-only listing, following the established --list → workflow pattern (per the approved scope on - gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the seeds dir, returns { count, seeds[], summary } JSON with each seed's id, slug, status, scope, trigger_when, planted, title. Optional case-insensitive status filter. User-controlled content is sanitized (sanitizeForDisplay) and every path validated (requireSafePath); read-only. Independent of audit.scanSeeds, which only returns unimplemented seeds for the milestone surface. - /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that renders the seed table. Closes #441 * chore(#441): point changeset fragment at PR #722 * test(#441): allowlist list-seeds test in prompt-injection scan The test asserts that list-seeds neutralizes injection payloads (<system>, [INST]) embedded in seed content, so the fixtures legitimately contain those patterns — same as the sibling security tests already on the allowlist. * fix(#441): use canonical /gsd:capture colon form in list-seeds workflow Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new list-seeds workflow used the hyphen form. * docs(#441): sync help full.md + INVENTORY for --list-seeds Adds the --list-seeds entry to the help reference (help/modes/full.md, per bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json. * docs(#441): add --list-seeds how-to + drop phantom statuses Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers): - USER-GUIDE.md Seeds section (how-to): extend the task to cover auditing parked seeds on demand via --list-seeds, including the status filter — kept task-oriented per Diataxis how-to mode. - CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected from the list-seeds filter vocabulary; the system only produces dormant|active|triggered (src/audit.cts scanSeeds). Reference must be factually accurate and complete. * fix(#441): guard non-scalar status frontmatter in cmdListSeeds A seed with a bare `status:` line (extractFrontmatter yields {}) or a `status: [a, b]` value (yields an array) crashed the whole audit list: `(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string. Coerce every frontmatter read through a `fmStr` helper (mirrors the existing `typeof fm.id === 'string'` guard), so a non-scalar status falls back to dormant and non-scalar scope/trigger_when/title can no longer leak a raw array/object into the JSON contract. Title is now capped symmetrically. Adds regression coverage for empty and array `status:` and non-scalar fields. Refs #441 * docs(#441): align list-seeds workflow status vocabulary The load_seeds step listed `implemented` as an example status filter, but the real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds); `implemented` has no producer. Matches the earlier CLI-TOOLS.md correction. Refs #441 * refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested in-process (review minor #1). No behavior change. Filter comparison now matches the raw lowercased status (both sides already normalized) instead of sanitizeForDisplay(status); sanitization is for output, not matching (review nit #3). * test(#441): add fast-check property coverage and count=1 boundary for list-seeds Adds tests/list-seeds.property.test.cjs with four fast-check properties over deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing (review minor #1). Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2). * chore(#441): sync runtime launcher snippet into list-seeds workflow Propagate the current _runtime-launcher.snippet.sh (with non-Claude runtime home probes) into the new list-seeds.md workflow via scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation. * test(#441): record list-seeds.md in workflow size baseline (#1074) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
ac40f070ef |
feat(#1318): require external reviewers to verify plan claims against source (#1421)
* feat(#1318): require external reviewers to verify plan claims against source /gsd-review built its external-reviewer prompt from plan text only and never asked reviewers to open the repo and verify claims, so a grounded HIGH could be outvoted by ungrounded LOWs. Add a concise, generic source-grounding block to build_prompt's Review Instructions: treat yourself as running in the working tree, open referenced files, cite path:line + mechanism, trace asserted mechanisms, downgrade to an open question if you have no file access, and know that grounded findings are weighted more heavily. Also clarify that CodeRabbit (a diff-only reviewer that never receives the prompt) must not be weighted as a grounded plan-level verdict in consensus synthesis. Workflow stays under its size cap (baseline bumped deliberately). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1318): add changeset for reviewer source-grounding Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1318): mark changeset docs-exempt (internal reviewer-prompt wording) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#1318): document reviewer source-grounding in COMMANDS.md; drop docs-exempt Review: a user-visible behavioral Changed warrants a docs touch, not a docs-exempt. Add a sentence to the /gsd-review entry in docs/COMMANDS.md (reviewers verify against source, cite file:line, grounded findings weighted higher) and remove the changeset docs-exempt marker so lint:docs passes via docs-updated. Also note the literal build_prompt test anchor. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1318): harden build_prompt fence extraction to be fence-run-aware Addresses maintainer review on PR #1421 (required-before-merge). The buildPromptReviewInstructions() test helper located the closing fence with `src.indexOf('\n```')`, which terminates at the FIRST triple-backtick line — so a build_prompt ```markdown block whose body embeds a fenced code example would truncate mid-content (dropping the `## Review Instructions` section) and give a spurious failure or false pass. Since this feature feeds source/plan content (which routinely contains code fences) to reviewers, that is a live fragility. Rewrite the extraction to be fence-run-aware, mirroring the CommonMark close rule in src/markdown-sectionizer.cts stripFencedCode: parse the opener's backtick run length, then close on the first line with >= that many backticks and only trailing whitespace — so a shorter nested fence is treated as content. Add a fail-first regression test (a 4-backtick outer fence wrapping a nested ```bash block) asserting the trailing `## Review Instructions` still extracts. Test-only change; no production .cts touched. Verified: test file 7/7, empirical fail-first proof the old indexOf logic truncated, full suite 4236/4236, eslint clean. Codex review: approve. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
8748e95ed1 |
fix(#666): pr-branch silently ignored planning.sub_repos (#667)
* fix(pr-branch): handle sub_repos from config with git -C (#666) Adds a `handle_sub_repos` step between `detect_state` and `analyze_commits`. When `planning.sub_repos` is set in config, the workflow now: - Reads sub-repo paths via `gsd_run query config-get sub_repos` - Skips the step entirely when the list is empty/null/[] - Scans each repo with `git -C "$REPO" status --porcelain` - Offers the user all/select/skip choices - For selected repos: creates a PR branch, commits all staged/unstaged changes, pushes, and opens a companion PR via `gh pr create` All git commands use `git -C "$REPO"` — never `cd "$REPO"` — because shell state does not persist between agent-executed commands. Closes #666 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore: update changeset pr number to 667 * fix(pr-branch): address maintainer review — correct seam, behavioral tests, robustness Resolves all three blockers and seven robustness issues raised in PR #667 review: Blockers: - Use `planning.sub_repos` (not top-level `sub_repos`) so config-get actually resolves - Replace prose grep test with behavioral fixture tests using runGsdTools + local bare repo - Extract sub-repo git work into new `cmdPrSubrepo` seam in src/commands.cts; never uses git add -A — stages explicit files only (universal-anti-patterns.md:44) Robustness: - Dirty-repo list persisted via mktemp/cat, not bash arrays (cross-block safe) - Branch name embeds repo slug (${CURRENT_BRANCH}-${REPO_SAFE}-pr) to avoid collision - push --set-upstream so gh pr create finds the branch - Sub-repo base branch resolved via ls-remote with fallback to repo's default branch - Remote slug parsed with /github\.com[:/]/ (handles SSH + HTTPS + .git-less URLs) - rollback() cleans up branch on any mid-sequence failure - node -e replaces jq (always available, no undeclared hard dep) Refs: #666 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(pr-branch): security guard, push timeout, rollback fix, porcelain fix Security (Blocker 1): - Use security.cjs validatePath() in cmdPrSubrepo for symlink-safe workspace containment check — rejects ../escape, absolute paths, and symlink traversal - Add negative regression test: '../escape' repo path must be rejected Robustness: - Push uses timeout: 60_000 ms (network op needs more than the 10 s default) - Capture prevBranchName before checkout -b so rollback uses explicit name instead of git checkout - (fails on fresh single-branch repos) - Porcelain path parse: line.trimStart().slice(2).trim() handles all XY combinations and the execGit global-trim edge case uniformly Tests: 17/17 pass, lint: 0 errors Refs: #666 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(pr-branch): move regression tests to commands.test.cjs, add core.quotePath=false - Move cmdPrSubrepo behavioral + workflow source-invariant tests from standalone bug-666-*.test.cjs into tests/commands.test.cjs under describe('pr-subrepo') per TESTING-SUITES.md policy (no new bug-* files). Adds allow-test-rule: source-text-is-the-product see #666 for the workflow-source-invariant suite. - Add -c core.quotePath=false to git status --porcelain call so non-ASCII filenames (e.g. café) are not C-escaped, keeping slice(2) parse correct. * fix(pr-branch): remove obsolete regression tests for sub-repos handling * fix(pr-branch): update workflow-size-baseline, add dirty-scan timeout - Regenerate tests/workflow-size-baseline.json for pr-branch.md growth (+handle_sub_repos step, +timeout addition). - Add { timeout: 10_000 } to the execFileSync git status --porcelain call in the handle_sub_repos dirty-scan (repo convention: every git subprocess is bounded, never hangs). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore: regenerate INVENTORY-MANIFEST after rebase onto next Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#666): handle rename staging and split changedFiles from filesToStage For git mv renames, the old path no longer exists in the worktree after the move — staging it with git add fails. Split parsing into changedFiles (both paths, for result.files) and filesToStage (new path only for renames; old is already staged by git mv). Also adds porcelain tests for staged renames, non-ASCII filenames, and a fast-check property test. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#666): rollback on push failure in cmdPrSubrepo If push fails the branch only exists locally; rollback cleans it up so the sub-repo is not left in a half-committed state. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#666): do not rollback after commit on push failure; add push-fail regression test Post-commit push failures are network/auth/policy issues — the user's work is already committed on the local branch. Calling rollback() at that point force-deletes the only ref holding the commit (data loss). Leave the branch in place and emit a retry instruction instead. Adds a regression test (pre-receive hook that rejects all pushes) asserting the branch and commit survive a push rejection so the failure path stays covered going forward. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore: regenerate INVENTORY-MANIFEST after rebase onto next Rebased onto current next (#1267 retired core.cjs). Stale tsbuildinfo and a leftover bin/lib/core.cjs build artifact were masking the drift — wiped both, rebuilt clean, and regenerated the manifest. gen-inventory-manifest --check now exits 0. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#666): validate sub-repo paths before git invocation in pr-branch.md The handle_sub_repos workflow ran git -C on raw planning.sub_repos config values at two points before the pr-subrepo seam's validatePath guard ever ran: the dirty-scan detection (git status) and the base-branch resolution (git ls-remote / remote show). A traversal entry could point git outside the workspace; an embedded newline could inject a spurious record into the newline-joined dirty-file output and into the shell-interpolated commit message. Adds a containment check + character allowlist to the dirty-scan node script (reject before any execFileSync), and a defense-in-depth shell case guard on the same value before the second, independent git -C invocation in the base-branch resolution block. Adds a behavioral test that extracts and executes the actual shipped node script from pr-branch.md (not a mirror) against a real traversal target and an embedded-newline entry, asserting neither reaches git or the dirty-file output. Also updates the stale cmdPrSubrepo doc comment: push failures no longer delete the branch (see prior commit), only stage/commit failures do. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#666): make sub-repo traversal scan test genuinely fail-first The outside repo's only change was an untracked file, which the ?? filter excludes — so the repo looked clean even with the guard removed, making the traversal assertion vacuous (it passed against a neutered guard). Commit the file first, then modify it, so the outside repo has a tracked dirty change: without the path guard it WOULD be reported dirty, so the test now fails-first. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#666): symlink-safe (realpath) sub-repo containment in pr-branch.md Finding A from re-review: the workflow guard used path.resolve, which only normalizes '..' textually and does not follow symlinks — so an in-tree symlink whose name has no '..' or '/' (e.g. "evil" -> /outside) passed both the charset filter and the resolve+startsWith check, letting git status / ls-remote / remote show run against a directory outside the workspace. The pr-subrepo seam already used fs.realpathSync (validatePath); this brings the workflow layer to parity. - dirty-scan: realpathSync the root once, and realpathSync each candidate before the containment check; skip on throw. - base-branch resolution: replace the weak `case *..*|/*` guard with a realpath containment check that yields a validated absolute SUB_REPO_DIR, and run git -C against that instead of re-concatenating $ROOT/$REPO_REL. - security test: add a symlink-escape entry and a positive control (legit in-root backend must still be reported). Confirmed fails-first — regressing the scan to path.resolve makes the symlink case leak. Also fixes a misleading-fallback minor: the workflow now checks the seam's exit status and skips the companion-PR step on failure, instead of printing "branch pushed, open PR manually" after a real stage/commit/push failure. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#666): harden pr-branch sub-repo flow against round-12 edge cases Pre-emptive hardening of the workflow changes from the symlink fix: - continue-outside-loop: the "skip companion PR on seam failure" block used a bash `continue`, but the per-sub-repo iteration is prose-driven (the agent loops, not a literal `for`), so `continue` would warn and no-op. Reframed as prose-gated control flow keyed on $SUBREPO_EXIT — no bash loop assumption. - Windows portability: the new symlink security case now degrades gracefully (try/catch around fs.symlinkSync; skip just the symlink assertion when symlink creation lacks privileges) so it doesn't hard-fail on Windows CI. Verified: seam exits 1 on error / 0 on success (error() → process.exit(1), propagated through the shim), so the $SUBREPO_EXIT check is meaningful; bash -n clean on the touched blocks; commands 156/156; lint:ci green; manifest in sync. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
195d356d7c | Merge branch 'next' into feat/1346-enhance-verify-phase-project-a-check-vio | ||
|
|
faac9331f2 |
feat(#1298): add validated worktree record-agent writer verb for wave manifests (#1448)
Closes #1298 |
||
|
|
2c718bf972 |
fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs (#1537)
* fix(#1521): resolve own runtime + worktrees-off for all non-Claude installs Generalizes the Codex-only #1515/#1519 fix to every non-Claude runtime, and wires it into the real install path (where it was previously dead-on-arrival). Root causes: 1. The runtime-default stamping lived only in `_applyRuntimeRewrites`, but the installer emits `gsd-core/workflows/*.md` via `copyWithPathReplacement`, which never calls it — so a real `--codex`/`--cursor`/etc. install emitted `--default claude` and worktrees-on. RUNTIME mis-resolved to claude and the workflow ran executors unisolated against the main checkout. (#1515/#1519 were also dead-on-arrival in real installs; this repairs them.) 2. Only `case 'codex'` was stamped; every other non-Claude runtime kept the Claude default. Fix: - New `_stampNonClaudeRuntimeDefaults(content, runtime)` (single shared helper) stamps `--default <runtime>` + `use_worktrees=false` for every `runtime != claude`; called from both `_applyRuntimeRewrites` and, crucially, `copyWithPathReplacement` in bin/install.js (the real workflow emit path). - Generalize the fail-closed worktree guard `= codex` -> `!= claude` in execute-phase/quick/diagnose-issues (worktree isolation is Claude-Code-only). - Flip manager/autonomous inline-vs-background gating to `codex -> background, everything-else -> inline` (research: only Codex can background-nest the pipeline's subagents; all others run inline, which they support). Worktree-capability determination is research-backed (official docs for all 14 non-Claude runtimes: none honor GSD's isolation="worktree" mechanism, only Codex background-nests). New end-to-end real-install test asserts the EMITTED workflow is stamped — the regression guard that would have caught the dead-on-arrival bug. Closes #1521 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r * chore(#1521): backfill changeset PR number (#1537) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
2436b76980 |
fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees (#1519)
* fix(#1515): make Codex installs resolve their own runtime and fail closed on worktrees A Codex install with a runtime-neutral .planning/config.json resolved RUNTIME=claude and enabled git worktree isolation, which Codex's spawn_agent cannot honor. Two root causes: 1. Workflows read `config-get runtime` / `config-get workflow.use_worktrees` without `--raw`, so config-get's JSON-quoted output ("codex") was captured verbatim into the bash var and broke every `[ "$RUNTIME" = ... ]` check — the Codex fail-closed guard was dead even when runtime:codex was explicit, and Claude's own worktree degrade-check was dead too. Add `--raw` to those reads across execute-phase, autonomous, manager, diagnose-issues, quick. 2. The conversion engine emitted `--default claude` for every runtime. Stamp the codex-emitted workflows to `--default codex` (runtime) and `--default false` (use_worktrees) in _applyRuntimeRewrites case 'codex', so a neutral config on a Codex install resolves runtime=codex / worktrees off. Also extend the Codex fail-closed worktree guard to quick.md and diagnose-issues.md (they spawned isolation="worktree" with no runtime guard). Regression test asserts source<->engine parity across all five workflows (DEFECT.GENERATIVE-FIX) plus fast-check property coverage of the stamping. Closes #1515 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r * chore(#1515): backfill changeset PR number (#1519) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013vX5eUtWa2wsZEyeMf5i3r --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c09c13f295 |
enhance(verify-phase): node-test causation control — prove the RED is content-caused (#1346)
The #1279 node-test machine-proof confirmed a known-bad subject drives the negative test RED, but could not distinguish a genuine content-violation from a deceptive test that reds merely because GSD_PROHIB_SUBJECT is set. Add an optional fifth flat scalar `check_clean_fixture` (-> CheckDescriptor.cleanFixture) threading a KNOWN-CLEAN control subject through projectProhibitions + descriptorFromProjection. When present, the prover also runs the check against the clean subject and requires GREEN, so fail-first is proven only when the check is RED on the violation AND GREEN on the clean subject (content-dependent). Opt-in and additive: absent a clean fixture the prover behaves exactly as post-#1314 (no control, documented residual), preserving the zero-authoring compose path; the lint-rule kind needs no analog (its subject IS the linted file, no env indirection). Coverage: RED-first deceptive case, positive, missing-clean fail-closed, round-trip read-back/emit, fast-check property extended to the 5th scalar, and an end-to-end COMPOSE capstone (honest vs deceptive). Docs: ADR-550 dated addendum, prohibition-probe reference, spec-phase + verify-phase workflows. Closes #1346 Claude-Session: https://claude.ai/code/session_01GsPRb8zvpcT7Eat6vZw8PX |
||
|
|
b63a500fad |
fix(#1505): extract context_guard step to reference file; fix allow-test-rule see ref
- Extract execute-phase.md context_guard step prose to gsd-core/references/execute-phase-context-guard.md (@-ref lazy load), bringing execute-phase.md back under the ADR-857 phase-6 size ceiling (92914 < 93166 bytes) - Fix allow-test-rule comment: add `see #1452` per ADR-456 lint rule - Update feat-1452 tests to check the reference file for extracted content - Register execute-phase-context-guard.md in INVENTORY-MANIFEST.json and INVENTORY.md Workflow References section - Regenerate workflow-size-baseline.json after file shrinkage Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
4eca5ac96c |
feat(#1452): add workflow.context_guard_mode to guard execute-phase against context exhaustion
Proactive checkpoint guard fires at each wave boundary before spawning agents. Self-assesses context pressure against context-budget.md degradation tiers and warns (warn, default) or auto-invokes /gsd:pause-work (auto) when POOR tier (70%+) is detected. Config key validated; defaults to \"warn\". Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
f5276b36b3 |
fix(#1369): refresh wave manifest and re-check base before each wave in execute-phase (#1492)
* fix(#1369): refresh wave manifest and re-check base before each wave in execute-phase Two compounding issues caused wave N+1 worktrees to fork from the stale pre-wave-N commit, immediately tripping the worktree_branch_check FATAL guard in every executor: 1. worktree.base-check auto-degrade only ran once at initialize time. After wave N merges advanced orchestrator HEAD past origin/HEAD, new worktrees were still forked from origin/HEAD (Claude Code "fresh" base). 2. WAVE_WORKTREE_MANIFEST was never unset between waves, so wave N+1 reused the consumed wave-N manifest file, which would have blocked the step 5.5 manifest guard (#3384) on subsequent waves. Fix: add two safeguards in execute-phase.md — - Step 0.5 (start of each wave): re-runs worktree.base-check; auto-degrades USE_WORKTREES=false for that wave when HEAD has diverged from origin/HEAD. - Step 7c (end of each wave): unsets WAVE_WORKTREE_MANIFEST so wave N+1 creates a fresh per-wave manifest; re-asserts worktree.set-baseref (idempotent) and re-evaluates base degradation after wave merges land. 17 regression tests added in tests/bug-1369-wave-stale-base.test.cjs. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1369): rename test to fix-NNN convention; update workflow size baseline Rename tests/bug-1369-wave-stale-base.test.cjs → tests/fix-1369-wave-stale-base.test.cjs to satisfy the lint-regression-test-names gate (new files cannot use bug-NNN prefix). Update tests/workflow-size-baseline.json for execute-phase.md: 93157 → 97393 (LF-normalized byte count after adding step 0.5 inter-wave base re-check and step 7c between-wave manifest reset). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1369): add issue reference to allow-test-rule comment Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1369): extract new execute-phase steps to references; satisfy ADR-857 cap Step 0.5 (inter-wave worktree base re-check) and steps 7b–7c (pre-wave dependency check + between-wave manifest reset/base refresh) added by this PR grew execute-phase.md to 97393 bytes, violating the ADR-857 phase-6 architectural mandate that host-loop bodies remain strictly below the pre-phase-6 baseline of 93166 bytes. Extract both new blocks into dedicated reference files: - gsd-core/references/execute-phase-wave-guard.md (step 0.5) - gsd-core/references/execute-phase-between-wave-reset.md (steps 7b + 7c) Replace inline prose with @-reference pointers. File now measures 92851 bytes (LF-normalized), satisfying the ADR-857 capstone conformance gate. Also update tests/workflow-size-baseline.json to 92851 and add both new reference files to docs/INVENTORY-MANIFEST.json. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1369): update regression tests to read from extracted reference files Steps 0.5 and 7b+7c were moved to reference files to satisfy the ADR-857 size cap on execute-phase.md. Tests now check @-reference pointers in the workflow for ordering and read content assertions from the reference files. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
3a3b2135c2 |
chore(#1073): purge phantom pre-migration issue refs from source, tests, docs (#1471)
#2551/#3182/#2361 are pre-migration get-shit-done-redux issue numbers with no equivalent in open-gsd/gsd-core; they mislead triage and manufacture phantom blockers. Repoint to real successors (#717 byte-budget rework, #720) or rewrite as prose referencing the discuss-phase/modes progressive-disclosure split. Correct co-located 'line budget'/'<500 lines' framing to the byte-based reality (#717). Add a CI guard (tests/no-phantom-issue-refs.test.cjs) that fails if a phantom ref is reintroduced. SSH-key patterns (id_ed25519) left untouched. No user-facing runtime behavior change. Closes #1073 |
||
|
|
120f85164b |
feat(#1355): detect-and-warn guard for claude-code agent-teams (#1371)
* feat(#1355): detect-and-warn guard for claude-code agent-teams GSD's multi-agent orchestration can stall under claude-code's experimental agent-teams (a subagent's completion fails to route to the orchestrator). Per the maintainer decision, the accepted scope is a read-only detector + one non-fatal warning — NOT the declined run_in_background/TaskOutput conversion. - New Teams Status Module (src/teams-status.cts → gsd-core/bin/lib/teams-status.cjs): pure resolveTeamsStatus({runtime, env}) + thin CLI cmdTeamsStatus reusing resolveRuntime. active = strictly-truthy env flag AND runtime === 'claude'. - Wire `gsd-tools query teams-status [--active]` (read-only; no capability registration needed — conformance gates govern features, not query commands). - One non-fatal warning in plan-phase.md before the first Agent spawn, gated on `query teams-status --active`; zero behavior change on non-claude/teams-off. - Hermeticity: clear CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS in run-tests.cjs + SESSION_ENV_KEYS. Docs reference + CONTEXT.md glossary. Built lib gitignored. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1355): add changeset for teams-detect guard Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1355): bump plan-phase.md workflow size baseline (+407B for teams warning) The non-fatal agent-teams warning block added to plan-phase.md grew it 92759 → 93166 bytes, past its committed per-file baseline ratchet. The growth is small, deliberate, and still well under the workflow tier hard cap. Regenerate the baseline via `npm run size:baseline` (only plan-phase.md changed). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1355): register teams-status.cjs in the inventory manifest The new teams-status CLI module is a tracked surface; regenerate docs/INVENTORY-MANIFEST.json (cli_modules family) via gen-inventory-manifest.cjs --write so the inventory-manifest-sync gate passes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1a678eb0e9 |
fix(#1359): migrate workflows off deprecated TaskOutput to Read(outputFile) (#1362)
* fix(#1359): migrate workflows off deprecated TaskOutput to Read(outputFile) The map-codebase and docs-update workflows collected background sub-agent results with the deprecated Claude Code `TaskOutput` tool using `block: true`, which has a confirmed main-session hang after the agent completes (anthropics/claude-code#20236). Migrate the collection steps to the upstream-recommended pattern: keep `run_in_background=true` on the Agent spawn, then `Read` each agent's `outputFile` (from the `async_launched` result) once it reports completion. Completion-marker contracts and on-disk verification are unchanged, and the non-Claude runtime fallbacks (sequential_mapping / sequential_generation) are preserved byte-for-byte. docs-update's timeout note no longer references the unwired `workflow.subagent_timeout` key (it kept a literal before). Regression coverage folded into tests/subagent-timeout.test.cjs (the owning module for background-subagent collection). Workflow size baseline regenerated for the justified prose growth. Refs #1355 (same latent hang surface). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1359): add changeset for TaskOutput migration Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
00c05eb717 | Merge branch 'next' into feat/1279-fail-first-prover |