Commit Graph

51 Commits

Author SHA1 Message Date
Tom Boucher
ef8a3e27d4 fix(#1864): balance <step> tags in settings-advanced.md §8 Model Policy (#2014)
* fix(#1864): balance <step> tags in settings-advanced.md §8 Model Policy

§8 Model Policy ended with </step> but had no matching opening tag (5 opens /
6 closes), leaving it as loose inter-step content. Add the missing
<step name="model_policy"> opener so the section is a proper step.

- gsd-core/workflows/settings-advanced.md: add <step name="model_policy">
- tests/workflow-step-tag-balance.test.cjs: regression guard — every top-level
  workflow must have balanced <step>/</step> (fenced code stripped), plus a
  focused assertion that §8 is wrapped in model_policy.
- goldens + workflow-size baseline recaptured.

Closes #1864

* docs(#1864): backfill changeset pr 2014
2026-07-05 14:38:24 -04:00
Tom Boucher
a62079b2da fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR (#2024)
* fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR

The gsd_run preamble resolved the Claude global install only at
$HOME/.claude/gsd-core/bin/, but the installer honors CLAUDE_CONFIG_DIR —
so a global install redirected via CLAUDE_CONFIG_DIR was invisible to every
gsd_run call (every command failed with 'gsd-tools.cjs not found').

The Claude resolver arm now uses ${CLAUDE_CONFIG_DIR:-$HOME/.claude},
matching the installer + the other runtimes' ${VAR:-default} pattern.
Default $HOME/.claude behavior is unchanged.

- _runtime-launcher.snippet.sh: Claude arm honors CLAUDE_CONFIG_DIR.
- sync-runtime-launcher.cjs re-run: 95 workflows/agents re-synced.
- review.md / discuss-phase.md: trimmed to stay under their byte budgets.
- runtime-launcher-parity.test.cjs: (A) substring updated for the new form
  + explicit #1865 assertion that the snippet honors CLAUDE_CONFIG_DIR.
- goldens + size baselines recaptured.

Closes #1865

* docs(#1865): backfill changeset pr 2024
2026-07-05 14:20:52 -04:00
Behruz Nassre Esfahani
23254ca5a7 fix(#1936): reconstruct OpenCode review from JSON events (#1992)
* fix(#1936): reconstruct OpenCode review from JSON events; diagnosable empty-output stub

On a large review prompt, OpenCode's default `build` agent runs a few read
tool calls then ends its turn with zero output tokens (reason:"stop",
output:0), so `opencode run --format default` emits empty stdout. The reviewer
block redirected stderr to /dev/null and wrote a generic "failed or returned
empty output" stub — so the phase silently lost its second independent reviewer
with no diagnostic and no timeout.

Rewrite the OpenCode reviewer block to invoke `--format json` as the primary
call and reconstruct the review from the assistant `text` parts (jq). Capture
stderr to a `.err` sidecar (mirrors the Codex block). When the agent emits no
text, surface the stop reason, output-token count, and stderr so the failure is
diagnosable. Gate the stub on the extracted CONTENT, not the output file size —
an empty jq extraction still prints a lone newline that a `[ -s file ]` check
would treat as populated. Document the wall-clock timeout as a Bash-tool param
(macOS lacks GNU timeout; opencode has no native timeout flag).

review.md was already at the DEFAULT size-tier ceiling (40956/40960), so the
fix cannot fit without reclassifying it into the LARGE tier (it is a
multi-reviewer orchestration file that outgrew "focused single-purpose"; 43.4 KB
sits well under the LARGE high-water mark). Recapture the 16 golden-install
fixtures — the diff is exactly one review.md hash per runtime. Regression block
folded into review-default-reviewers-workflow.test.cjs (new bug-NNNN test files
are not accepted).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1936): add changeset

* test(#1936): property-test the OpenCode review jq reconstruction

Address the re-review's one actionable finding: the jq JSON-event → text
reconstruction had no fast-check property test.

Add tests/opencode-review-reconstruction.property.test.cjs. It extracts the two
shipped jq programs (OPENCODE_REVIEW, OPENCODE_DIAG) verbatim from
gsd-core/workflows/review.md and runs the real jq — not a reimplementation — so
the shipped logic is what gets tested. Properties: the reconstructed review
equals the newline-join of every assistant text part (order preserved); a stream
with no text part reconstructs to empty (drives the #1936 stub); null/absent text
parts are dropped, never rendered as "null". Plus example-based coverage of the
diagnostic edges the reviewer cited: missing .tokens.output and no step_finish
degrade to "?"; non-JSON stdout makes jq fail rather than masquerade as a review.

Verified the invariant has teeth (a comma-join jq fails the property).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1936): skip jq reconstruction property test when jq is absent

The property test shells out to `jq`, which GitHub's windows-latest runners do
not ship (macOS/Linux runners do). `execFileSync('jq')` therefore ENOENT-failed
the whole file on `test (windows-latest, *)`. Probe `jq --version` at load and
skip the suite when jq is not on PATH — the reconstruction logic is
platform-independent, so the assertions still run in full on every jq-present
runner (mirrors how golden-install-parity skips on win32).

Verified: jq present → 7 pass; jq removed from PATH → 7 skipped, 0 fail.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1936): skip jq reconstruction property test on Windows, not just when jq is absent

The prior guard skipped only when `jq` was absent from PATH — but the
windows-latest runners DO ship jq, so the suite still ran there and failed with
`jq: parse error: Invalid numeric literal` (confirmed from the CI job log). Root
cause is Node's child_process argument quoting mangling the jq program (it embeds
double quotes) on Windows, not the shipped review.md logic — the macOS/Linux legs
pass. Gate the suite on `process.platform === 'win32'` (still also skipping when
jq is absent), mirroring golden-install-parity's win32 skip. Logic is
platform-independent and fully asserted on every macOS/Linux CI leg.

Verified: macOS → 7 pass; simulated win32 → skips.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 14:07:07 -04:00
Tom Boucher
8de2ff9121 feat(#2008): generic command-exit-zero gate-predicate evaluator (#2011)
* feat(#2008): add generic command-exit-zero gate-predicate evaluator

Third-party capability gates declared via check.predicate were rendered for
display but never evaluated (only built-in check.query gates fired; the
security capability's gate worked solely via a hard-coded ship.md branch).

Add a generic, deps-injected gate-predicate evaluator (src/gate-predicate-evaluator.cts)
that dispatches by predicate.kind. Built-in kind: command-exit-zero — runs a
bounded sh -c command at the project root (via shell-command-projection.execTool),
inherits env, exit 0 => pass, non-zero => block, timeout => block, fail-closed.

Wire a 'check predicate' subcommand into check-command-router.cts and extend
the three generic workflow gate-dispatch sites (execute:wave:post, execute:post,
plan:post) to route check.predicate gates to the new evaluator. The two-step
gate contract (command-failure => onError; block => halt) is unchanged.

- src/gate-predicate-evaluator.cts: pure leaf, KIND_TABLE extensible
- src/check-command-router.cts: cmdCheckPredicate + buildPredicateDeps + parsePredicateFlags
- docs/adr/2008-*, docs/reference/gate-predicates.md, docs/how-to/command-exit-zero-gate.md
- tests: 38 unit + integration tests (exit mapping, timeout, interpolation,
  property-based bijection, malformed-predicate fail-closed, real subprocess e2e)

Closes #2008

* docs(#2008): backfill changeset pr number 2011
2026-07-05 14:04:29 -04:00
Tom Boucher
8f2ebbe9bf feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity

Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap).

--gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): backfill changeset PR number (#1996)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): drop Gemini CLI from issue templates (review nit)

Removes the sunset Gemini CLI runtime from the two GitHub issue-template
runtime lists that the removal PR missed, per @davesienkowski's review nit:
- feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise
  request a feature for a runtime GSD no longer supports)
- bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json
  retrieval-help line

Leaves the post-removal templates fully consistent with the Antigravity redirect.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 13:32:51 -04:00
Rezolv
55604e9124 fix(#1906): require node-test clean-fixture causation control (#2001)
* fix(#1906): require node-test clean-fixture causation control

The node-test fail-first proof accepted a deceptive content-independent
negative test — one that reds merely because GSD_PROHIB_SUBJECT is set,
ignoring the subject's content — whenever no cleanFixture was supplied,
because #1346's causation control was opt-in. The proof's observed signal
(RED) thus diverged from its target (RED caused by content) by default.

Make the causation control mandatory for the node-test kind: a descriptor
that omits cleanFixture is un-provable (fail-closed), never accepted under
the weaker violation-only proof. When a clean fixture is present, fail-first
is proven exactly as before (RED on violation AND non-vacuous GREEN on clean).
The lint-rule kind is unchanged (its subject IS the linted file; no
GSD_PROHIB_SUBJECT indirection).

Breaking (Hyrum): a previously-green node-test prohibition with no clean
fixture now hard-gates — blast radius is zero in-tree (no node-test
prohibition ships today; only the lint-rule local/no-source-grep dogfood).

Supersedes ADR-1606 Decision 4 / ADR-550 #1346 addendum's opt-in.

Closes #1906

Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ

* docs(#1906): supersede the #1346 opt-in causation control (mandatory for node-test)

Record the node-test mandatory-causation-control supersede across the
governing surfaces:

- ADR-1606 (the enforcement decision-of-record): addendum + Decision 4
  annotated + the "Mandatory causation control — REJECTED" alternative
  flipped to accepted (premise no longer holds: zero in-tree node-test
  consumers).
- ADR-550: the 2026-06-21 #1346 "Why opt-in, not required" paragraph
  marked SUPERSEDED, pointing at ADR-1606.
- spec-phase.md: check_clean_fixture is now REQUIRED for node-test
  (was "optional").
- CONTEXT.md: PROHIB.enforce.causation predicate updated.

Regenerated the shipped-artifact cascade from the spec-phase.md edit
(+149 B, well under the 40960 cap): 16 golden-install-parity fixtures
and the workflow size baseline.

Refs #1906

Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ
2026-07-03 23:21:55 -04:00
Jeremy McSpadden
e5ef323b15 feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command

Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.

* feat: add /gsd-start smart-entry command

State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.

- src/smart-entry.cts: deterministic situation classifier (no-project,
  paused, blocked, verify-failed, needs-first-phase, planning, executing,
  verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
  ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire  case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
  dispatcher presenting an AskUserQuestion menu (with --text fallback for
  non-Claude runtimes) and dispatching to existing commands. Falls back
  to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
  situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
  (markdown-layer invariants + every emitted command resolves to a real
  slash command).

Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.

* refactor: rename smart-entry command to /gsd:next

Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.

All affected tests (188) pass; lint:ci clean.

* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)

Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.

- detectSignals now reads total_phases + percent from nested progress{}
  first, then scalar fm, then body; current_phase falls back to the
  body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
  body Phase field) covering verify-pending + executing situations.

Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.

* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)

Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.

Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
   layer is down — read .planning/STATE.md directly with the Read tool
   and synthesize a minimal situation + actions menu so /gsd:next stays
   useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
   /gsd:progress as before.

Matches the direct-read resilience the live agent already did by hand.

* docs: add gsd-next skill surface

* chore: trigger no-mistakes validation

* no-mistakes(review): Fix smart-entry phase ordering

* no-mistakes(review): Fix decimal smart-entry phase ordering

* no-mistakes(test): Fix smart-entry next test contracts

* no-mistakes(document): Docs synced for smart entry

* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: shorten next.md description and update golden install parity fixtures

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: trigger no-mistakes validation

* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files

Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.

* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: regenerate golden install parity fixtures for /gsd:next

Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.

* Fix smart-entry verify-failed phase scoping and empty resolve shim step

Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.

* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: recapture all 16 golden fixtures with updated smart-entry.md hash

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: regenerate fixtures + inventory manifest after rebase onto next

Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next

Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).

Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)

The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs

Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo

Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
  recommended action, 1-4 unique-id /gsd:* actions (previously the
  one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout

Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:

- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
  zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
  kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
  consolidation file doing dozens of real installs — 41s even on a fast Mac
  (much worse on the slow Windows I/O path), plus an install-heavy cluster.

Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.

Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 12:18:25 -04:00
Behruz Nassre Esfahani
8189d2f098 enhance(#1872): document Claude Code advisor inheritance in model profiles (#1922)
* enhance(#1872): document Claude Code advisor inheritance in model profiles

Add an "Advisor Tool (Claude Code)" section to
gsd-core/references/model-profiles.md: session-level advisor is inherited
by all GSD subagents and composes with the per-agent profile/tier system,
candidate executor/advisor pairings per profile (cost/quality/caching
claims attributed to Anthropic's advisor-tool docs, not asserted as GSD
behavior), when it is worth enabling vs not, and the session-level /
no-per-agent-control constraint linking anthropics/claude-code#73072.

Docs-only. Golden-install-parity fixtures recaptured for the edited
reference file (hash-only, one line per runtime).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1872): add changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1872): use Documentation changeset type for docs-only change

Changeset type was `Changed`, which triggers the docs-required lint
(TRIGGERING_TYPES in scripts/lint-docs-required.cjs). This PR only
touches gsd-core/references/model-profiles.md, so there is no docs/
file to pair with and docs-lint failed. `Documentation` is the correct
type for a docs-only enhancement and is exempt from the trigger.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1872): put docs-exempt marker on its own line, revert to type Changed

The prior fix (type: Documentation) was invalid — parse.cjs ALLOWED_TYPES
is {Added, Changed, Deprecated, Removed, Fixed, Security}, so both
changeset-lint and docs-lint failed with invalid_type.

Real root cause of the original docs-lint failure: DOCS_EXEMPT_RE is
anchored to match the `<!-- docs-exempt: ... -->` marker only on its own
line, but the marker was tacked onto the end of the prose line, so it was
never captured (docsExempt: null) and the triggering `Changed` fragment
had no docs/ pairing -> fail_docs_missing.

Fix: keep the valid `type: Changed` and move the marker to its own line.
Verified locally: changeset-lint -> ok_fragment_present,
docs-lint -> ok (own-line marker parses to ok_fragments_exempt).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-02 23:41:39 -04:00
Rezolv
9d7c046eae refactor(#1852): lazy-split plan-phase.md into steps/ (#1934)
* refactor(#1852): lazy-split plan-phase.md into steps/

Extract 3 self-contained, rarely-hit sections into gsd-core/workflows/plan-phase/steps/
via lazy 'Read and execute' pointers (mirrors execute-phase/steps/, ADR-1610 progressive
disclosure — no eager @-import): closed-phase-gate (1.5), prd-express-path (3.5),
windows-troubleshooting. Byte-invariant: plan-phase.md 94459 -> 89775 (-4684), each step
< 32 KiB anchor, resolved instructions unchanged.

Scope note: only 3 sections were extractable. plan-phase.md is guarded by a dense net of
content-presence tests (plan-bounce/enh-3209/phase6-planning-capabilities assert specific
sections/flags inline) that block extracting the larger blocks without also refactoring
those tests — durable low-80s headroom is deferred pending maintainer re-scope (issue #1852).

Cascade: size:baseline regen, 16 golden-install-parity fixtures regen, INVENTORY in sync.
Full plan-phase test surface green (4129/4130, 0 fail); lint:ci exit 0.

* docs(changeset): Changed fragment for #1934 (plan-phase lazy-split)

* docs(changeset): mark #1934 fragment docs-exempt (internal workflow refactor)

* fix(#1852): add gsd_run launcher preamble to prd-express-path step

The extracted prd-express-path.md calls gsd_run but the canonical launcher
preamble lived in the parent plan-phase.md — runtime-launcher-parity (#373)
walks workflows/ recursively and requires every .md using gsd_run to carry
exactly one preamble + the $HOME/.claude fallback arm (same as the existing
execute-phase/steps/ files). Injected via scripts/sync-runtime-launcher.cjs;
golden fixtures + size baseline regenerated. plan-phase.md unchanged (89775).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-02 23:30:19 -04:00
Tom Boucher
d3d689a5e9 fix(#1920): resolve host version from gsd-core/VERSION + ship capability generators (#1938)
The flattened install layout broke the third-party capability ecosystem in two
ways. Both are fixed at the host-version / installer boundary.

Gap 1 — host version read as 0.0.0. The running GSD version was resolved via
require('../../../package.json') (loader/source) and require('../../package.json')
(the gsd-tools CLI), which in the installed layout is the versionless CommonJS
marker → the fail-closed fallback reported 0.0.0, so `capability install` rejected
any manifest with a real engines.gsd range as "incompatible with GSD 0.0.0". For
runtimes that get no marker, and for local installs, that walked-up package.json
could even be the USER's own project, reporting a wrong version. Fix:
readHostVersion() (capability-loader.cts, capability-source.cts) and capHostVersion()
(gsd-tools.cjs) now prefer the authoritative gsd-core/VERSION the installer already
writes for EVERY runtime (mirrors resolveVersionFrom(), #1383), falling back to the
runtime-root package.json for the dev/source tree, then fail-closing. This fixes
every runtime and the actual `capability install` CLI path without touching the
marker package.json, so uninstall is unchanged (no data-loss surface).

Gap 2 — scripts/gen-capability-registry.cjs (+ its sibling
gen-loop-host-contract.cjs) were never copied by the installer, so the loader's
never-crash invariant discarded every overlay and fell back to the frozen
first-party registry (installed capabilities silently inert). Now copied,
uninstalled, and manifest-tracked exactly like fix-slash-commands.cjs (#1223).

Regenerates the 16 golden-install-parity fixtures to capture the two newly shipped
generator scripts and the gsd-tools.cjs change.

Regression tests (RED→GREEN): readHostVersion VERSION-first / fallback / fail-closed
resolution; an end-to-end `capability install` against a REAL installed layout
proving the engines gate sees the real host version, not 0.0.0; and a real-install
check that both generators are shipped and manifest-tracked.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-02 20:34:30 -04:00
Tom Boucher
f13db4cbf5 feat(#1933): VS Code IDE reference host binding — completes Phase 5 IDE profile (#1935)
* feat(#1933): VS Code IDE reference host binding — completes Phase 5 IDE profile

The reference VS Code host binding composes the Phase-3 engine seams for the
ide profile (host-integration.cts PROFILE_BASELINES): active model via vscode.lm
(createModelAdapter active + sendRequest), engine-owned hook bus
(createHookBus engine — VS Code has no host bus), sandboxed-storage stateIO
(createStateIO sandboxed-storage + host backend — no fs / no child_process),
and the imperative adapter (engine-as-library). Command surface: palette/chat.

VS Code is extension-distributed (Marketplace), not file-projected onto a config
dir, so it intentionally has NO runtime descriptor / --vscode installer entry —
the extension IS the host (architecturally N/A for the descriptor/installer
model, not a deferral). Mock-friendly binding (vscode.lm + hostStorage injected)
so it is behaviorally testable without a live VS Code host.

Tests: IDE profileOf classification; full seam composition (active model routes
to vscode.lm; engine bus pub/sub; sandboxed state routes to backend; imperative
adapter; command surface); fail-closed construction.

* fix(#1933): correct vscode fail-closed test case matrix
2026-07-02 17:25:05 -04:00
Tom Boucher
21a9af4048 feat(#1682): pi ExtensionAPI imperative reference host-plugin — Slice 3 (#1932)
Proves the Programmatic-CLI reference binding for pi via the ExtensionAPI
imperative adapter (#1682 AC): createImperativeAdapter({runtime:'pi'}) classifies
as imperative + composes the registry; pi axes (imperative + bun) classify as
'programmatic-cli'; the reference pi host-plugin binds GSD via ExtensionAPI
(registerCommand /gsd + registerTool gsd_invoke + on tool_call).

Reference plugin (tests/fixtures/pi-host-plugin.cjs) is CJS + mock-friendly so
it is behaviorally testable without a live pi runtime. Full --pi installable-
runtime integration (descriptor + installer + golden parity 16→17) is a larger
follow-up — intentionally NOT added to the runtime registry here.
2026-07-02 16:49:50 -04:00
Tom Boucher
325e9fad4b feat(#1682): OpenCode session.idle + opencode-subset dialect + Claude parity — Slice 1b/c (#1930)
* feat(#1682): OpenCode session.idle + opencode-subset dialect + Claude parity — Slice 1b/c

- Plugin (.opencode/plugins/gsd-core.js): handle session.idle (↔ Claude Stop
  lifecycle point; no-op sentinel — state already persisted to .planning/).
  Completes the compaction/idle pair (#1914 shipped compaction).
- Declare hookEvents: 'opencode-subset' in the OpenCode descriptor — the
  reserved dialect now has a real consumer (no longer zero-consumer).
- host-integration.cts: add HOOK_EVENT_SURFACES + hookEventSurfaceFor() — the
  pure consumer that resolves a dialect to its host-fireable event surface.
  opencode-subset = session/tool/file subset with NO workflow-phase events
  (engine owns phase sequencing; ADR-1239 §OpenCode binding).
- Tests: hookEventSurfaceFor unit tests (claude/gemini/opencode-subset/null);
  plugin session.idle no-throw; compaction breadcrumb; opencode-subset surface
  parity vs the plugin's handlers (Claude parity).

* fix(#1682): don't declare hookEvents on opencode (hooksSurface:none invariant)

The repo invariant couples runtime.hookEvents to the managed settings.json hook
surface: hooksSurface:'none' runtimes (opencode — plugin owns hooks) must NOT
declare hookEvents. Declaring 'opencode-subset' there violated 3 capability-
registry invariants + 2 install-plan golden masters + opencode golden parity.

The opencode-subset dialect is still IMPLEMENTED — just not via the legacy
descriptor field: hookEventSurfaceFor() (host-integration.cts) is its consumer,
and the OpenCode plugin consumes the subset events at runtime (session.idle
added here; compaction shipped in #1914). Declaring it on the descriptor would
require weakening the hooksSurface:none ⇔ no-hookEvents invariant (flagged for
decision).

* fix(#1682): refresh opencode golden parity for plugins/gsd-core.js (session.idle)

* docs(changeset): OpenCode session.idle + opencode-subset dialect (#1682)

* docs(changeset): backfill PR #1930
2026-07-02 16:16:00 -04:00
Tom Boucher
541de6894f feat(#1682): OpenCode companion-MCP binding (mcp.gsd) — Phase 5 Slice 1a (#1929)
* docs: align PR-FLOW push gate from gsd-test-summary to gsd-test

gsd-test is the application (open-gsd/gsd-test-runner); gsd-test-summary is
the legacy local wrapper. RULESET.PR-FLOW.docker-before-push now names gsd-test
as the pre-push gate (exit 0 / verdict outcome 'passed').

* docs: drop WORKTREE.SEAM local node --test rule; align PROC dispatch to gsd-test

- Remove WORKTREE.SEAM.execution-rule (prefer local node --test) — contradicts
  CLAUDE.md 'NEVER run node --test locally'; CLAUDE.md wins.
- PROC.PARALLEL-FIX-DISPATCH: gsd-test-summary --both -> gsd-test (app; --both
  was a legacy wrapper flag).

* feat(#1682): OpenCode companion-MCP binding (mcp.gsd) — Phase 5 Slice 1a

configureOpencodePermissions registers the Phase-4 companion MCP server
(gsd-mcp-server) as opencode mcp.gsd, so OpenCode connects to GSD's command
(point 1) + state-IO (point 5) with no bespoke plugin (ADR-1239 Phase D).
Idempotent + non-clobbering (add-if-absent; respects a user-defined mcp.gsd).
Local-stdio schema per OpenCode config (packages/core/src/config/mcp.ts);
`-p @opengsd/gsd-core` resolves the bin (name != package) under npx.

Tests: registers-on-object-config; does-not-clobber-user-entry.

* docs(changeset): OpenCode companion-MCP binding (#1682)

* fix(#1682): use PACKAGE_NAME single-source (#516) + refresh opencode golden parity

- mcp.gsd command: replace hardcoded '@opengsd/gsd-core' literal with
  PACKAGE_NAME from gsd-core/bin/lib/package-identity.cjs (#516 single-source).
- opencode golden-install-parity fixture: refresh opencode.json hash for the
  added mcp.gsd block (configureOpencodePermissions output change).

* docs(changeset): add docs-exempt marker (Phase 5 slice)

* docs(changeset): backfill PR #1929
2026-07-02 14:57:44 -04:00
Tom Boucher
ff6b5cb024 feat(#1914): OpenCode native plugin integration (Option 1 file-copy) (#1923)
* feat(#1914): OpenCode native plugin integration (Option 1 file-copy)

Ship a native OpenCode plugin (.opencode/plugins/gsd-core.js) plus the
installer step that delivers it, so GSD's lifecycle hooks run on OpenCode.
OpenCode declares hooksSurface:'none', so GSD's hook scripts already ship to
<configDir>/hooks/ but nothing invokes them; the plugin bridges OpenCode's
event bus onto those scripts as subprocesses (prompt/read/worktree/workflow
guards, injection scanner, context monitor).

Distribution is Option 1 (file copy) per the #1914 triage decision: no
scripts.build rename, no prepare/prepack removal. package.json gains
main + .opencode in files[] for discovery.

Corrected against OpenCode's docs + loader source (not the reference branch):
- Auto-discovery globs {plugin,plugins}/*.{ts,js} — .cjs is never matched, so
  the installed adapter must be .js (config dir carries {"type":"commonjs"}).
- No opencode.json plugin-array patch — that array is npm-only; local files
  are auto-discovered.
- REPO_ROOT is resolved by walking up to the dir holding hooks/ + gsd-core/,
  correct for package tree, global install, and local install.
- Config-hook registration is gated (IS_PACKAGE_TREE) so it never
  double-registers commands/agents/skills already delivered by native copy.

Also fixes an incidental .gitignore drift: 8 ADR-1239 .cts-generated .cjs
artifacts were untracked-and-not-ignored (leak risk) — now ignored.

Tests: tests/opencode-plugin-adapter.test.cjs (14, pure helpers + real
subprocess bridge against stub hooks); golden-install-parity regenerated.
Green on Mac + Linux (gsd-test): 0 failures.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1914): harden OpenCode plugin export shape + advisory accumulation (adversarial findings)

Address Codex adversarial-review findings:
- HIGH: export `{ id, server }` could trip OpenCode's loader
  (`for (entry of Object.values(mod)) getServerPlugin(entry)` throws on a
  non-extractable value). Make `id` NON-ENUMERABLE and assign module.exports
  from a variable (not a literal) so no string `id` is ever iterated — verified
  loader-safe under real import(pathToFileURL) (default + module.exports alias,
  both objects with .server; no bare id string).
- MEDIUM: sequential advisory hooks clobbered output.metadata._gsdAdvisory;
  now accumulate into an array.
- LOW: resolveRepoRoot fallback returned ".." while the comment said "../.." —
  aligned to "../.." (package-tree depth).

Tests: added a faithful loader-loop emulation (raw CJS + ESM namespace views),
advisory-accumulation, and a real bin/install.js copy→manifest→uninstall
integration test. golden regenerated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1914): add changeset fragment for OpenCode plugin integration (PR #1923)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1914): Windows — assert rewritten path with string include, not path-regex

The Read content-rewrite test built a RegExp from `path.join(root,'gsd-core')`.
On Windows the backslashes in the path are interpreted as regex escapes, so the
assertion never matched and `test (windows-latest, 24)` failed — even though the
adapter rewrote the path correctly. Replace the RegExp with a separator-agnostic
`String.includes` check (the repo's no-path-literal-in-assert concern).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-02 13:01:51 -04:00
Behruz Nassre Esfahani
7bef6a6496 fix(#1863): use named flags for state.* calls in executor + workflows (#1873)
* fix(#1863): use named flags for state.* calls in executor + workflows

The named-only state-command router (parseNamedArgs) silently drops
positional args, so state.cjs threw its required-arg error and
metrics/decisions/blockers/session continuity were never recorded.

Convert record-metric / add-decision / add-blocker / record-session in
agents/gsd-executor.md to the named-flag form (mirroring execute-plan.md),
and fix the two remaining positional record-session calls in
gsd-core/workflows/milestone-summary.md and forensics.md. Recapture the
golden-install-parity fixtures and size baselines for the edited files.

Also fix a pre-existing detached-rebuild handle leak in
tests/graphify-auto-update.slow.test.cjs: three dispatch tests returned
after observing only the synchronous "running" status without awaiting the
detached rebuild's terminal state. That leak was latent until the new
#1863 regression block's added runtime shifted --test-force-exit timing
and surfaced it as a non-zero chunk exit. The three tests now await
terminal status via the file's existing waitForBuildStatus helper.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1863): add changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-02 13:01:10 -04:00
jecanore
5657994702 fix(#1716): route resume_from_file to complete_session when no pending tests remain (#1722)
* fix(#1716): route resume_from_file to complete_session when no pending tests remain

When a UAT session has status:partial with blocked_count>0 and pending_count==0 (all remaining tests are blocked, none are pending), resume_from_file found no [pending] test and terminated silently — never routing to complete_session. This blocked the issues==0 auto-transition path even when there were zero code defects.

Guard clause added immediately after the find-pending step: if no [pending] test is found, route to complete_session. complete_session then correctly sets status:partial (because blocked_count>0) without presenting further tests.

Closes #1716

* chore(#1716): add changeset fragment and regenerate golden-install-parity fixtures

Changeset fragment for PR #1722 (type: Fixed).

Golden-install-parity fixtures regenerated for all 16 runtimes — the workflow fix shifts verify-work.md's byte-stable hash in the golden manifest. Regenerated via UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
2026-07-02 11:58:06 -04:00
Tom Boucher
92091d71f2 fix(#1871): wire phase archival end-to-end (phases archive cmd + default + atomic) (#1924)
Follow-up to #1919 (archive-then-remove core). Closes the remaining #1871
acceptance criteria so phase history is preserved across the full milestone
lifecycle, not just at phases.clear:

- #2 src/milestone.cts + src/phases-command-router.cts: extract shared
  archivePhaseDirectories() helper; add cmdPhasesArchive (the previously
  half-wired phases.archive alias now routes instead of erroring Unknown).
- #4 gsd-tools.cjs + src/milestone.cts: milestone complete archives phase
  dirs by default (--no-archive-phases opts out; --archive-phases is now a
  harmless no-op). complete-milestone.md updated to drop the redundant manual
  Yes/Skip archive prompt.
- #3 gsd-core/workflows/new-milestone.md: §6 stages the archive move + source
  removal (git add .planning/milestones/ .planning/phases/) in the same commit
  as the milestone start, so the archive lands atomically — no orphaned
  uncommitted deletions, no un-archived dirs inherited.
- docs/CLI-TOOLS.md (+ ja/zh/ko/pt) + help/modes/full.md: flag accuracy.
- tests: phases archive command (#2) + milestone complete default archive /
  --no-archive-phases opt-out (#4). Goldens + workflow size baseline refreshed.

Closes #1871
2026-07-02 11:14:01 -04:00
Tom Boucher
3c13903dcd feat(#1866): agent-side self-load of configured agent_skills
Each of the 22 consumer agents now self-loads its configured agent_skills
in its mandatory init step, so .planning/config.json agent_skills.<type>
reaches the agent on every runtime — including Cursor and /gsd-autonomous,
where Skill()-delegated workflow bash init did not reliably execute.

- gsd-core/references/agent-skills-bootstrap.md: shared contract
  (query + Read + dedup guard that skips when <agent_skills> is already
  in the prompt, so Claude's orchestrator-side injection never doubles)
- 22 agents/gsd-*.md: one self-load line naming the agent's own type
- gsd-core/workflows/autonomous.md: note that delegated agents self-load
- tests/agent-skills-bootstrap.test.cjs: regression + parity (CONSUMER_AGENTS
  bijection + fast-check property) — Generative-Fix-Divergence guard
- docs: ADR-1866, CONFIGURATION dual-injection How It Works, INVENTORY
  row, Changed changeset

Closes #1866
2026-07-01 20:09:01 -04:00
Tom Boucher
3cc4d1608c test: regenerate golden fixtures + size baselines for the example/doc edits
execute-phase.md and gsd-ai-researcher.md are installed artifacts; their edits
shift install hashes and file sizes. Diff is scoped to those two files' hashes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 23:04:13 -04:00
Tom Boucher
a2a9d38884 test(#1847): regenerate golden-install-parity fixtures for claude-sonnet-5
The sonnet-5 catalog change alters model-catalog.json and settings-advanced.md,
so the per-runtime install-hash fixtures are recaptured (UPDATE_GOLDEN=1). Only
those two file hashes change per runtime.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 22:21:18 -04:00
Tom Boucher
93e5d2dd84 fix(#1525): skip deferred phases on autonomous reruns (#1846)
* fix(#1525): skip deferred phases on autonomous reruns

* chore(#1525): add changeset fragment

* chore(#1525): fix changeset body

* test(#1525): refresh install parity fixtures

* test(#1525): shrink autonomous workflow

* test(#1525): refresh autonomous baselines

* test(#1525): tolerate Windows temp cleanup flake
2026-06-30 21:29:38 -04:00
Behruz Nassre Esfahani
dae7f81482 fix(#1528): drop next-phase guidance from security-blocked verify-work presentation (#1687)
* fix(#1528): drop next-phase guidance from security-blocked verify-work presentation

When security enforcement blocks phase advancement (no SECURITY.md produced),
the verify-work presentation told the user advancement was blocked but still
offered `/gsd:plan-phase {next}` and `/gsd:execute-phase {next}`, competing
with the current-phase fix. Remove those two next-phase lines so the blocked
state routes only to the current-phase resolution (secure-phase, ui-review).
The post-transition presentation — reached only after the completion contract
passes — still offers next-phase planning, which is the correct place for it.

Regression coverage added to tests/ui-review-next-guidance.test.cjs: the
security-blocked block must not offer next-phase actions, and the
post-completion block must still offer them. Regenerated workflow size baseline.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1528): add changeset for security-blocked next-phase fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1528): recapture golden-install-parity fixtures for verify-work.md change

Rebased onto next; verify-work.md's installed hash changed across all 16
runtime fixtures. Diff confined to the single gsd-core/workflows/verify-work.md
key per runtime. Assert mode 16/16 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-30 10:52:30 -04:00
Tom Boucher
1f649838b8 Merge branch 'next' into fix/1698-codex-output-last-message 2026-06-30 10:04:59 -04:00
Tom Boucher
b31b562dd2 fix: exclude CHANGELOG.md from golden-install-parity hash manifest (#1840)
CHANGELOG.md contains historical version strings from prior releases.
The PKG_VERSION normalization applied to all files only replaces the
*current* package version, so locally (PKG_VERSION=1.6.0) the normalization
mutates CHANGELOG.md content (1.6.0 appears in old entries), producing a
different hash than in CI (PKG_VERSION=1.7.0-rc.1, which doesn't appear
in CHANGELOG.md). The hash can never match across build contexts.

Add gsd-core/CHANGELOG.md to VOLATILE_FILES so it is excluded from the
parity manifest. It's release documentation — not a functional install
artifact — and changes with every release anyway.

Regenerate all 16 golden fixtures to remove the stale CHANGELOG.md entry
and establish the new baseline.

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-30 00:55:09 -04:00
Tom Boucher
4bdf0fd1cd fix: normalize package version in golden-install-parity hashes (#1837)
The rc release step runs `npm version X.Y.Z-rc.N` before tests, which
rebakes the current version string into hook files and `gsd-core/VERSION`.
Without normalization, every golden parity hash differed post-bump and the
entire test suite failed with 16 golden failures — even though no files
actually changed in a semantically meaningful way.

Add `PKG_VERSION` normalization (`.split(PKG_VERSION).join('<VERSION>')`)
alongside the existing `<HOME>` root normalization so the goldens are
stable across version bumps. Regenerate all 16 fixtures with the new
normalization to establish the new baseline.

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-30 00:11:34 -04:00
Rezolv
18995380ce feat(#1154): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1738)
* feat(verify-phase): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1154)

Carry the edge-probe's existing `backstop` (non-inferable) tier through the
plan-phase projection as a structured flat-scalar marker instead of a prose
parenthetical, and make verify-phase abstain -> human_needed (never silent-pass)
on a backstop truth it cannot confirm with explicit evidence. Truth-axis mirror
of #644's prohibition judgment-tier (ADR-550 D4).

Engine (deterministic, CI-tested per ADR-550 D5 — never the LLM verdict):
- src/probe-core.cts: truthStatement/truthVerification normalizers, projectTruths
  (conservative serializer), dispositionForUnverifiableTruth (backstop+no-evidence
  -> unverified/flagged/insufficient_spec; backstop+evidence -> green; inferable
  -> green, the over-abstention guard).
- src/roadmap.cts: coerceTruthToString now reads `statement` first so an object-form
  backstop truth is surfaced, not dropped (Hyrum backward-compat for truth-readers).

Workflow/agent/docs: plan-phase emits the structured marker (flat scalar, ADR-550
#1278); verify-phase + gsd-verifier add the abstain arm; new references/honest-verifier.md;
FEATURES/COMMANDS document insufficient_spec; ADR-550 amended (truth-axis D4 mirror).

Decisions adopted (trek-e review): insufficient_spec feeds existing human_needed with a
distinguishable reason (no new VERIFIER_STATUS); changeset Changed; round-trip parity
test; abstain-on-unconfirmed-backstop regression test red-first.

Implementation notes (deviations from the issue's proposed file list, verified live):
- frontmatter.cts needs no change — its flat parser already round-trips object-form truths.
- verify.cts needs no change — it grades artifacts/key_links structurally; truths are
  LLM-graded at the workflow layer, so consumption lives there + the deterministic helper.
- No CJS<->SDK hand-sync — the SDK seam was retired (ADR-0174); src/*.cts is sole source.

Regenerated artifacts: golden-install-parity fixtures, INVENTORY-MANIFEST, size baselines.

* chore(#1154): add changeset (Changed) for honest verifier

User-facing changelog fragment for #1738. Typed `Changed` (not `Added`) per
trek-e review condition 3 — the verify behavior shifts for backstop-bearing specs
(a confident silent `passed` becomes `human_needed`), which is user-visible even
though the schema marker is additive.

* docs(#1154): score-formula also excludes abstained insufficient_spec truths (review nit-1)

trek-e review nit: the verify-phase score sentence said PRESENT_BEHAVIOR_UNVERIFIED
truths were "the only ones excluded" from verified_truths. Post-#1154 an abstained
`insufficient_spec` backstop truth is also excluded (it is not ✓ VERIFIED and routes
to human_needed). Behavior was already correct; this tightens the wording.
Regenerated golden-install-parity fixtures + workflow-size baseline for the touched
verify-phase.md. (Nit-2 — a dedicated insufficient_spec_items frontmatter list — is
intentionally not taken: the current design is ADR-550-D4-conformant, the abstain
cause rides as a distinguishable report reason, and adding it would exceed the
approved scope.)

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-29 00:14:32 -04:00
Tom Boucher
ac001be49e fix(#1778): use 1.6 named-flag frontmatter.set form in thread workflow (#1816)
* fix(#1778): use 1.6 named-flag frontmatter.set form in thread workflow

The thread workflow's CLOSE and RESUME branches called frontmatter.set with
the pre-1.6 fully-positional shape (frontmatter.set <file> <field> <value>).
Since 1.6 the dispatcher (gsd-tools.cjs) parses the file positionally and
reads field/value from the named flags --field/--value via parseNamedArgs;
the positional form leaves field/value undefined, cmdFrontmatterSet errors
'file, field, and value required', and the status/updated writes are
silently skipped. Closing a thread never marked it status: resolved and
resuming never marked it status: in_progress.

Switch all four sites (CLOSE status+updated, RESUME status+updated) to the
1.6 hybrid form that verify-work.md already uses:
  frontmatter.set <file> --field <field> --value <value>

Add a regression test with three guards: (1) behavioral — the named-flag
form writes the field while the positional form errors with the documented
message and does not mutate the file; (2) workflow parity — no workflow
under gsd-core/workflows/ emits the positional form, so a future edit that
reintroduces it anywhere fails CI; (3) thread-specific — CLOSE writes
status: resolved and RESUME writes status: in_progress via the named flags.

* docs(#1778): add changeset fragment for thread workflow frontmatter fix

* docs(#1778): fix unclosed inline-code backtick in changeset fragment

* fix(#1778): move regression into owning test + regen baselines

lint-regression-test-names rejects new bug-NNNN-*.test.cjs files; move the
#1778 regression (behavioral named-vs-positional + workflow-parity scan +
thread CLOSE/RESUME assertions) into tests/frontmatter-cli.test.cjs, the
canonical home for frontmatter CLI regressions, and delete the standalone
file. frontmatter-cli.test.cjs already carries the allow-test-rule exemption
for workflow .md content tests.

gsd-core/workflows/thread.md ships to every runtime and is size-tracked, so
recapture the 16 golden-install-parity fixtures (thread.md hash) and the
per-file workflow size baseline (thread.md 12400 -> 12464) via UPDATE_GOLDEN=1
and npm run size:baseline.
2026-06-28 22:42:44 -04:00
Tom Boucher
fd576528a7 fix(#1747): register four search-provider keys in the config schema (#1814)
* fix(#1747): register four search-provider keys in the config schema

buildNewProjectConfig emits seven search-provider availability flags and
research-provider.cts providerAvailability() consumes all seven, but only
three were registered in VALID_CONFIG_KEYS (config-schema.manifest.json).
config-loader.cts then printed an 'unknown config key(s)' warning for the
four unregistered keys (tavily_search, ref_search, perplexity, jina) on
every freshly generated .planning/config.json.

Register the four missing keys in the schema manifest and document them
alongside brave/exa/firecrawl in CONFIGURATION.md. Add a regression test
plus a structural drift guard that requires every config-driven
research-provider flag to be in VALID_CONFIG_KEYS, so a future provider
addition cannot silently reintroduce the drift.

* fix(#1747): move regression into owning test file + add changeset

lint-regression-test-names rejects new bug-NNNN-*.test.cjs files; move the
#1747 regression (four provider keys in VALID_CONFIG_KEYS + provider-flag
drift guard) into tests/bug-2530-valid-config-keys.test.cjs, the canonical
home for VALID_CONFIG_KEYS regressions, and delete the standalone file.

Add the missing .changeset fragment — config-schema.manifest.json lives
under gsd-core/ (user-facing), so changeset-lint requires a fragment.

* test(#1747): regenerate golden-install-parity fixtures for schema change

Adding four provider keys to config-schema.manifest.json shifts its shipped
content hash (65dea848 -> 7d398e94); recapture all 16 runtime fixtures via
UPDATE_GOLDEN=1. Each fixture changes exactly one line — the manifest hash.
2026-06-28 22:42:36 -04:00
Tom Boucher
2d314c3a28 fix(#1772): read full multi-line command in graphify-update hook Gate 2 (#1815)
* fix(#1772): read full multi-line command in graphify-update hook Gate 2

The PostToolUse hook joined tool_name + newline + tool_input.command and
extracted the command with sed -n '2p' — line 2 only. Agent runtimes
(Claude Code's Bash tool among them) routinely emit HEAD-advancing commits
as multi-line scripts ('cd /path', then 'git add', then 'git commit …'), so
line 2 is the 'cd', Gate 2's *"git commit"* match failed, and the rebuild
silently no-op'd on real commits despite graphify.auto_update: true.

Capture line 2 through EOF (sed -n '2,$p') so the case glob sees the full
multi-line command string. Single-line behavior is unchanged (the match
only widens); non-HEAD-advancing multi-line commands still no-op cleanly.

Regression tests cover multi-line commit/merge/pull dispatch plus a
multi-line no-op no-regression guard.

* docs(#1772): add changeset fragment for graphify-update multi-line fix

* test(#1772): regenerate golden-install-parity fixtures for hook change

gsd-graphify-update.sh ships to 9 graphify-aware runtimes; widening the
sed range (2p -> 2,$p) shifts its shipped hash. Recapture the 9 affected
fixtures via UPDATE_GOLDEN=1 — each changes exactly one line (the hook hash).
2026-06-28 22:42:23 -04:00
Behruz Nassre Esfahani
32b8c836c3 test(#1698): recapture golden-install-parity fixtures for review.md change
Rebased onto next; review.md's installed hash changed across all 16
runtime fixtures. Diff confined to the single gsd-core/workflows/review.md
key per runtime. Assert mode 16/16 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 08:17:47 -07:00
Tom Boucher
4ced0a64cc feat(#1561): assumption-delta advisory checkpoint (#1767)
* feat(#1561): assumption-delta advisory checkpoint

* chore(#1561): backfill changeset PR number (#1767)

---------

Co-authored-by: review-bot <review-bot@gsd>
2026-06-27 08:43:47 -04:00
Tom Boucher
b0d5ca3379 feat(#1517): support custom reviewer instances for /gsd:review (#1766)
* feat(#1517): support custom reviewer instances for /gsd:review

Add a bounded review.reviewer_instances config surface so one model-capable
adapter (e.g. opencode) can run as several independent reviewer identities in a
single /gsd:review pass. Instances participate only via review.default_reviewers,
expand before built-in slugs, are available iff their cli is detected, and a
non-matching entry is a hard error (typo must be loud). >=2 same-cli instances
emit a shared-adapter caveat in REVIEWS.md. Default path with no instances is
byte-for-byte unchanged.

Single-source instance->cli resolution lives in resolveReviewerSelection /
normalizeReviewerInstances (parity-locked in
tests/review-reviewer-instances.test.cjs). cli validated against
KNOWN_REVIEWER_SLUGS only (never arbitrary shell); model/agent opaque, never
shell-interpolated.

Closes #1517

* chore(#1517): backfill changeset pr:1766

---------

Co-authored-by: review-bot <review-bot@gsd>
2026-06-26 23:35:04 -04:00
Tom Boucher
e075a41c86 feat(#1754): CLI version-skew detection — warn when a global install shadows project-local GSD (#1755)
* feat(#1754): CLI version-skew detection — warn when a global install shadows project-local GSD

Addresses #1754 (approved-enhancement). Detects when the running gsd-tools.cjs
is outside the project root while a project-local install exists — the shadowing
scenario from #1748 where a stale global canary CLI (retired @gsd-build/sdk)
silently overrides project-local GSD.

Implementation (Node CLI entry-point, not shell snippet — avoids bloating 93
workflow files past their size caps):

- src/cli-skew-check.cts: pure function checkCliSkew({resolvedPath, projectRoot,
  projectLocalExists}) → string|null. Compares paths via path.relative; returns
  a warning when the resolved CLI is outside the project root AND a project-local
  install exists. Includes @gsd-build/sdk removal hint when the path matches.
  No I/O (pure), no gsd-sdk literal (avoids bug-2801 lint).
- gsd-core/bin/gsd-tools.cjs: wired at startup via the existing findProjectRoot
  resolver. Non-blocking (try/catch; advisory stderr warning, never gates).
- eslint.config.mjs: registers the new ADR-457 generated artifact in the ignores.
- tests: 6-case suite (skew/no-skew/legacy/normalization); all green.
- Golden fixtures regenerated (UPDATE_GOLDEN=1) for the new compiled artifact.
- docs/how-to/update-gsd.md: Diátaxis reference note for the skew warning.

Full suite: 3354 pass, 0 regressions (1 pre-existing local AGENTS.md failure).
lint:ci green.

Closes #1754

* chore(#1754): backfill changeset pr placeholder

* chore(#1754): regenerate INVENTORY-MANIFEST for the new cli-skew-check source module

---------

Co-authored-by: review-bot <review-bot@gsd>
2026-06-26 12:19:39 -04:00
Tom Boucher
b307c4cfde refactor(#1734): extract install engine from bin/install.js (ADR-1239 Phase B deep move) (#1735)
* refactor(#1734): extract install engine from bin/install.js (ADR-1239 Phase B deep move)

Relocate the runtime-artifact install cluster out of the 12,490-line
bin/install.js into a dedicated src/install-engine.cts -> install-engine.cjs:
installRuntimeArtifacts, uninstallRuntimeArtifacts, installOpencodeFamilySkills,
and their cluster helpers (_copyStaged, snapshot/restore, legacy migration,
GSD-entry pruning, preserve/restoreUserArtifacts, OpenCode-family converters,
USER_OWNED_ARTIFACTS).

- bin/install.js imports the engine and re-exports the moved symbols for
  back-compat; getCommitAttribution STAYS in install.js (impure config I/O +
  argv explicitConfigDir global) and is injected via a resolveAttribution param.
- 17 test files migrated to import the moved symbols from the engine.
- Bookkeeping: eslint built-artifact ignore, .gitignore, INVENTORY manifest+row,
  CONTEXT.md Install Engine Module glossary seam.

Behaviour-preserving: install output is byte-identical for all 16 runtimes
(golden-parity harness #1730) — the only delta is the new install-engine.cjs
file shipping in the installed gsd-core/bin/lib/ tree.

Closes #1734

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1734): backfill changeset PR number (#1735)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: review-bot <review-bot@gsd>
2026-06-25 21:16:47 -04:00
Tom Boucher
07ec1b91d5 test(#1730): golden-parity install harness for all 16 runtimes (#1732)
ADR-1239 Phase B (parent #1679). Safety net for the upcoming engine deep-move:
captures the COMPLETE emitted install output of all 16 runtimes as a golden
baseline so the move PR can prove byte-identical parity.

tests/golden-install-parity.test.cjs spawns the real installer per runtime into
a temp HOME, normalizes the temp path to <HOME>, excludes the two volatile
metadata files (gsd-file-manifest.json, gsd-install-state.json — the only
run-to-run variance after normalization, empirically), SHA-256s every remaining
file, and asserts the manifest matches tests/fixtures/golden-install-parity/<rt>.json
(8,957 hashes total). UPDATE_GOLDEN=1 regenerates; mismatches list
added/removed/changed paths. Non-vacuous (corrupting a hash fails the run).

Closes #1730

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 19:00:32 -04:00
Tom Boucher
5fa4dcd78c fix: recover silently-excluded test dirs + test-architecture audit hardening (#1195)
* fix: recurse test discovery so subdir test suites actually run

scripts/run-tests.cjs discovered tests with a flat readdirSync(testDir),
silently excluding tests/observability/ (4 files), tests/dispatch/ (1) and
tests/installer-migrations/ (1) — 94 passing tests — from `npm test` and all
CI lanes. Walk the tree recursively (relative subpaths preserved), classify
suites by basename, and add a fail-on-zero-executed guard for suite/default
runs (escape hatch GSD_ALLOW_EMPTY_SUITE=1) while preserving the empty
--files/--files-from path the CI inert lane relies on.

Unit suite 735 -> 741 files; surfaces ADR-227's observability/dispatch seam.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: retire 5 verified-worthless tests

Adversarial verification confirmed these 5 prove nothing — their coverage is
provided more strictly elsewhere:
- enh-2790 'has a name: field' spot-checks (command-contract enforces /^gsd[:-]/)
- command-routing-hub duplicate construct + duplicate ERROR_KINDS assertions
- no-cjs-sdk-handsync-tooling (guarded files that never existed on main; bug-190
  covers the real retired SDK artifacts)
- runtime-artifact-layout cline edge case (subsumed by the explicit-global test
  and bug-782-cline-skills-emission)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: add ADR-218 release version-validation coverage

ADR-218 (reject leading-zero versions like 1.01.0; npm duplicate pre-check) had
zero tests — the logic lived only in release.yml bash. Add a test that extracts
the actual rejection regexes from the workflow and exercises them against a
boundary table (leading-zero/malformed rejected, valid accepted) plus structural
wiring assertions. Goes red if the regex is reverted to [0-9]+.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: redesign weak tests into behavioral, deterministic assertions

Per the ADR test audit, rewrite 27 weak test files (test-only, no source
changes) so each can go red for the defect it guards:
- kill pass-always assert.ok(true) placeholders (research-cli, worktree-baseref,
  bug-260 security guard, eslint-rules x24, clusters '|| true')
- replace source-text grep with behavioral calls (install Kilo, sh-hook-paths,
  plan-review-convergence) and add a repo-layout governance test
- de-flake real-clock/Math.random coupling (phase last_updated, bug-3707 mtime,
  context-utilization property, feat-3594)
- fix independence/shared-state violations (bug-492 singleton, issue-844 tmpRoot,
  core reapStaleTempFiles, active-workstream TTY, feat-488 GSD_HOME)
- strengthen property/shape-only tests (research-provider/store classification +
  collision) and unconditional plugin.json schema validation (issue-766)

Verified: all 28 files run together 1220 pass / 0 fail / 1 skip.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: add no-tautological-assert lint rule, error in test suite

New custom ESLint rule (eslint-rules/no-tautological-assert.cjs) bans asserts
that can never fail: assert(true)/assert.ok(<always-truthy literal>),
'cond || true' inside an assert, and equality asserts comparing two identical
literals. Wired as error on tests/**; full sweep confirmed zero existing
violations so the suite stays green. Prevents the placeholder-assert regressions
the audit redesigns just removed. RuleTester coverage added (6 valid, 8 invalid).

Note: no-only-tests was already enforced via eslint-plugin-no-only-tests, so no
duplicate rule was added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: gate new allow-test-rule exemptions to require an issue ref

ADR-456 requires any allow-test-rule exemption added after the ADR to carry a
tracking issue number, but nothing enforced it. New ratchet gate
(scripts/lint-allow-test-rule-refs.cjs, wired into lint:ci) fails when a NEW
allow-test-rule comment lacks a #NNN/URL reference; the 323 existing untracked
exemptions are grandfathered in an allowlist that ratchets down as they gain
refs. Red-green verified (novel untracked offender fails; compliant passes).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: add ADR test-audit evidence report (#1192)

Full risk-first qa-test-architect audit of the ADR portfolio (37 ADRs + 4
platform lenses, adversarial verification of retire verdicts) that drove the
P0 discovery fix, ADR-218 coverage, 5 retires, 27 redesigns, and the two new
lint gates. Filed as point-in-time evidence under docs/issueevidence/, named
for tracking issue #1192.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: replace pre-existing raw NUL byte with escape in feat-3594 fixture

feat-3594's null-byte parser fixture contained a literal NUL byte (pre-existing
on next at b10e5681 — confirmed: base blob has 1 NUL, this fix has 0), which
made git treat the file as binary and would break grep/editors. Switch to the
\x00 escape; the runtime string value (a real NUL in the parser input) is
unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: address adversarial-review findings

Codex adversarial pass over the branch:
- capability-registry drift test no longer mutates the committed generated
  capability-registry.cjs in place (concurrency hazard) — uses in-memory
  checkPipeline comparison instead.
- allow-test-rule ratchet now detects exemptions in ALL comment forms (block
  /* */ too, matching no-source-grep) so a block comment can't bypass it;
  one newly-surfaced pre-existing offender grandfathered (323->324).
- install.test Kilo case asserts on what install(false,'kilo') actually writes
  rather than manually calling configureKiloPermissions (masked the call site).
- issue-766 drops the undeclared transitive ajv dep for explicit structural
  assertions from the schema fixture.
- adr-218 test notes the hotfix leading-zero gap is tracked in #1186.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: address code-review findings (subdir discovery, rule + test gaps)

xhigh code review surfaced 15 confirmed issues, all fixed:
- run-tests.cjs --files now resolves subdir tests by bare basename + handles
  Windows backslash paths (ambiguous basenames error clearly).
- affected-tests-lib.cjs listTestFiles made recursive — the targeted CI lane was
  silently dropping changed subdir tests (same false-green class the audit fixed).
- no-tautological-assert now catches 'true || cond' and empty []/{}  equality.
- verify-test-quality: restore provenance-classification coverage, tighten the
  writeFile circular-detection check, guard the module-level file read.
- sh-hook-paths: cover the global-install .sh delegation branch (#2045 guard).
- active-workstream null-guard runs deterministically (no longer skipped on TTY).
- adr-218 structural guards tightened (major/minor leading-zero; needs: membership).
- repo-layout AGENTS.md guard no longer false-alarms on equivalent refactors.
- cross-ai ordering guard fails red when the step is missing.
- issue-766 parses required fields from the schema fixture (auto-enforced).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: stub USERPROFILE alongside HOME in feat-488 (Windows parity)

The feat-488 redesign stubbed process.env.HOME but not USERPROFILE; os.homedir()
resolves from USERPROFILE on Windows, so the home stub was not hermetic there —
caught by windows-test-parity-guard (stubsHomeNoUserProfile). Save/set/restore
USERPROFILE symmetrically with HOME (delete-if-originally-undefined).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: reconcile allow-test-rule allowlist after rebase onto next

Rebasing onto current next pulled in merged PR #1170, which added
inventory-headings-countfree.test.cjs (a baseline allow-test-rule exemption) and
deleted inventory-counts.test.cjs. Grandfather the former and prune the latter so
the ratchet matches the merged tree. No new debt from this PR.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 23:35:08 -04:00
Tom Boucher
e4dfa6b9ea fix(#1012): invoke fallow with its real CLI and wire the report normalizer (#1044)
* fix(#1012): invoke fallow with its real CLI and wire the report normalizer

The /gsd-code-review structural pre-pass invoked fallow with flags no published
fallow version accepts (--json, --profile, --stdin-files), so it failed on every
run and degraded silently per REQ-FALLOW-02 — the feature never delivered on any
fallow version. Three compounding defects:

1. Invalid flags. Real fallow audit uses --format json (not --json), -q/--quiet,
   --changed-since/--base for changed-files scoping (no file-list input), and
   --max-crap for thresholds. There is no --profile or --stdin-files.
2. Exit-code handling. fallow audit exits 1 when it FINDS issues (verdict=fail),
   0 when clean. The pre-pass treated any non-zero exit as a crash and discarded
   the output — i.e. it threw away exactly the findings it exists to surface.
   Success is now decided by whether a valid fallow JSON report was produced,
   not by the exit code.
3. Schema mismatch. normalizeFallowReport parsed a fictional top-level schema
   (unusedExports/duplicates/circularDependencies) fallow never shipped, and was
   dead code (the workflow embedded raw JSON; its tests asserted the fictional
   schema, one even calling a non-existent runFallowAudit and passing vacuously).

Fixes: align the invocation to fallow's documented agent-facing pattern; map the
profile preset (minimal/standard/strict) to --max-crap (50/30/15); scope phase
runs via --changed-since with a repo-scope fallback; rewrite the normalizer to
fallow's real schema (dead_code.unused_exports/unused_files/circular_dependencies
+ duplication.clone_groups) and wire it into the workflow so the reviewer
receives normalized findings; replace the fictional-schema fixtures and tests
with real-schema ones and delete the vacuous runFallowAudit test.

Closes #1012

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1012): backfill changeset PR number to 1044

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:36:50 -04:00
Colin
cd5db1f8db test(suites): seed security/slow/integration suites via measured retags
Renames (git mv) with all references updated (ci-test-scope RULES,
windows-parity allowlist, test-file-count allowlist, docs in 6 locales):

- 5 scanner tests -> *.security.test.cjs — the 'Run security tests' CI step
  ran zero files since the suite taxonomy landed; it is now honest.
- graphify-auto-update -> *.slow.test.cjs (36s, slowest file in the suite;
  e2e gsd-tools spawns) — runs on full-matrix lanes and push to next.
- installer-migration-install-integration -> *.integration.test.cjs
  (13s; an integration test by its own name).

Coverage gate measured after retags: 88.55% lines (gate 70%).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:50:41 -04:00
Tom Boucher
cdd78bd2aa fix(#670): self-healing recovery for installer-migration checksum drift (#675)
Editing the body of an already-released installer migration drifts its computed
checksum (it hashes plan.toString()). The integrity guard then hard-aborted
every prior install on upgrade with "applied migration checksum changed" — a
100% reproducible blocker (v1.3.0, all platforms).

Already-applied migrations are filtered out of `pending` and never re-run, so
a drifted checksum is functionally inert. ADR-0008 anticipates checksum-mismatch
state as something the install-state layer must handle gracefully (plan -> apply
-> recover/report), not abort on.

This supersedes the published-checksum allowlist merged in #674 (per-release
maintenance debt — every historical checksum hand-pinned, still throws for any
unregistered value) with a general, self-healing recovery:

- Replace the throwing guard with non-fatal `collectAppliedChecksumDrift`,
  surfaced on `plan.checksumDrift`.
- Reconcile drifted stored checksums durably on the next state write
  (`reconcileDriftedChecksums`), idempotently (no perpetual writes).
- Relocate the "shipped migration bodies are immutable" rule to a CI baseline
  test that locks every shipped migration's checksum and fails on body drift —
  where #615 should have been caught, instead of blocking users.

Removes #674's legacyChecksums field, per-migration checksum pins,
published-checksums.json fixture, and compat test.

Fixes #670

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 13:44:05 -04:00
Colin
b8e15c7a98 fix: accept published installer migration checksums 2026-06-04 12:30:13 -04:00
Tom Boucher
463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00
Tom Boucher
06cf7826e3 test: deepen fixture module v2 for reliability (#329)
* test: deepen fixture seam for init/state/workstream suites

* test: deepen fixture module v2 with declarative builders
2026-05-26 11:23:42 -04:00
Tom Boucher
33ffc647e2 feat(117): reproducible npm environment bootstrap + check-env validator (#136)
* test(117): add failing tests for env validator (check-env.sh)

RED phase. Six tests for scripts/check-env.sh — none pass because the
script does not exist yet. Fixtures:
  good/             — engines.node >=22, .nvmrc 26, synced lockfile
  bad-node-version/ — engines.node <14.0.0 (current Node v26 fails)
  missing-lockfile/ — no package-lock.json
  bad-nvmrc/        — .nvmrc says 22, current Node is v26

Tests cover:
  1. Happy path exits 0
  2. engines.node constraint failure exits 1
  3. Missing lockfile exits 1
  4. .nvmrc major mismatch exits 1
  5. --json flag emits {pass: boolean, checks: array}
  6. Integration smoke: exits 0 on live worktree root

Sources:
  npm engines: https://docs.npmjs.com/cli/v10/configuring-npm/package-json#engines
  npm ci docs: https://docs.npmjs.com/cli/v10/commands/npm-ci

Closes #117

* feat(117): add scripts/check-env.sh with Node/npm/lockfile/version-manager checks

GREEN phase. Implements the five-check environment validator:

  1. Node version vs engines.node (semver constraint — >=, >, <=, <, =)
  2. npm version vs engines.npm (skipped if field absent)
  3. package-lock.json presence
  4. Lockfile sync via `npm ci --dry-run` (exits non-zero when drift detected)
  5. Version-manager pin (.nvmrc / .node-version / .tool-versions) vs active Node major

Exit codes: 0 = all green; 1 = at least one failure; 2 = tool error.
Flags: --json (structured report), --help.

All 7 tests pass. shellcheck clean. bash -n syntax check clean.

Sources:
  npm engines:         https://docs.npmjs.com/cli/v10/configuring-npm/package-json#engines
  Reproducible builds: https://reproducible-builds.org/docs/source-tree/
  npm ci docs:         https://docs.npmjs.com/cli/v10/commands/npm-ci

Closes #117

* chore(117): pin Node engines + .nvmrc; add check:env npm script

- Add engines.npm: ">=10.0.0" (npm 10 ships with Node 22, the CI floor).
  Source: https://docs.npmjs.com/cli/v10/configuring-npm/package-json#engines
- Add .nvmrc pinning Node 22 (lowest supported version per CI matrix in
  .github/workflows/test.yml; node-version: [22, 24]).
- Add "check:env": "./scripts/check-env.sh" script to package.json.

No generator is involved (not a .generated. file). The test update in this
commit adjusts the integration smoke: it now asserts on --json structured
output rather than raw text, and accepts exit 0 or 1 (version-manager pin
mismatch is expected when developer runs Node 26 against a .nvmrc of 22).

Closes #117

* ci(117): wire environment check into test workflow

Add "Environment check" step to .github/workflows/test.yml in the `test`
job. Positioned AFTER actions/setup-node and BEFORE npm ci so that env
mismatches (wrong Node version, missing npm version, absent lockfile) are
caught before the install step obscures the root cause.

Runs `npm run check:env` (./scripts/check-env.sh) on every matrix lane
(ubuntu, macos, windows) × (Node 22, 24).

Source: https://docs.npmjs.com/cli/v10/configuring-npm/package-json#engines

Closes #117

* docs(117): publish docs/contributing/bootstrap.md + link from CONTRIBUTING.md

Adds docs/contributing/bootstrap.md with:
  1. Prerequisites (nvm, fnm, asdf, mise; gh CLI)
  2. One-time setup (clone, nvm use, check:env, npm ci)
  3. Daily commands table
  4. Validation guide (check table, exit codes, --json usage)
  5. Troubleshooting (node-version, npm-version, lockfile-present,
     lockfile-sync, version-manager-pin, missing modules, locale errors)
  6. Alternative: Docker via gsd-test-runner
     (https://github.com/open-gsd/gsd-test-runner)

Adds "Bootstrap your environment" section to CONTRIBUTING.md pointing to
the new doc. No content duplication — CONTRIBUTING.md links only.

Adds .changeset/117-npm-bootstrap.md (type: Added) for changelog.

Sources:
  npm engines:         https://docs.npmjs.com/cli/v10/configuring-npm/package-json#engines
  Reproducible builds: https://reproducible-builds.org/docs/source-tree/
  npm ci docs:         https://docs.npmjs.com/cli/v10/commands/npm-ci
  gsd-test-runner:     https://github.com/open-gsd/gsd-test-runner

Closes #117

* fix(#117): make check:env script run on Windows runners

Invoke check-env.sh via `bash` instead of a bare POSIX path so
Windows CI runners (which have Git Bash on PATH) execute the script
without requiring a POSIX shell shebang dispatcher.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(117): make check-env.sh fixture .nvmrc adapt to active Node major (cross-platform fix)

Before() hook writes good/.nvmrc = activeNodeMajor and bad-nvmrc/.nvmrc = activeNodeMajor+99
at test-run time. Hardcoded .nvmrc=26 failed on every CI matrix row except Node 26.
After() restores originals so the checked-in files stay stable.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(117): exempt gsd-test-runner URL path from slash-command registry check

docs/contributing/bootstrap.md links to https://github.com/open-gsd/gsd-test-runner.
The parity-test regex captures /gsd-test-runner from the URL path component and
flags it as an unregistered slash command. Add 'test-runner' to INTERNAL_COMPONENT_SLUGS
(mirrors the existing 'build' entry for GitHub org URLs) with an explanatory comment.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(117): fix Node-24 and Windows-22 CI failures in check-env

Two root causes:

1. version-manager-pin on Node 24 (mac/ubuntu/win):
   The project root .nvmrc pins major 22 for local dev. When the CI
   matrix runs Node 24, check-env.sh fails the version-manager-pin
   check and exits 1, blocking the entire test job before any test
   runs. Fix: skip the pin check when CI=true (GitHub Actions always
   sets this). The pin is a local dev guard, not a gate for multi-
   version matrix CI.

2. engines.node appears missing on Windows-22 (pkg_field backslash):
   pkg_field() embedded PACKAGE_JSON directly into a node -e string
   literal using require(). On Windows, the path uses backslashes
   (D:\a\...) which are silently interpreted as JS escape sequences
   inside the string, causing require() to fail silently (2>/dev/null
   || true). engines.node returns empty, triggering a spurious FAIL.
   Fix: switch to fs.readFileSync + JSON.parse and normalise
   backslashes to forward-slashes before embedding in the JS literal.

Also pass { CI: '' } from the bad-nvmrc unit test so the
version-manager-pin fixture test still exercises the mismatch path
even when running inside CI runners.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(117): use relative ./package.json path in pkg_field to fix Windows CI

On Windows, Git Bash exposes \$PWD as a POSIX path (/d/a/…) which
node.exe cannot resolve via fs.readFileSync. The previous fix embedded
the absolute PACKAGE_JSON path in the node -e string after converting
backslashes to forward-slashes, but the POSIX form produced by Git Bash
(/d/a/…) has no backslashes — so the conversion was a no-op and node
received an unresolvable path. The silent catch(e) { process.exit(0) }
swallowed the ENOENT, returning empty string for every engines.* field.

Fix: use './package.json' (relative to CWD). pkg_field() is always
called before any cd in the script so CWD === PROJECT_ROOT at call time.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 18:38:49 -04:00
Tom Boucher
89886d90b4 feat(114): npm dependency integrity gate (npm ls invalid/extraneous) (#135)
* test(114): add failing regression tests for npm dependency integrity gate

Adds tests/npm-integrity-gate.test.cjs and four fixture directories under
tests/fixtures/npm-integrity/ covering:
  - clean: matching lockfile and node_modules (expects exit 0)
  - drift: declared vs installed version mismatch (expects exit 1)
    Reproduces the ws 8.20.1 declared / 8.20.0 installed incident shape
    using stable-dep@8.20.1 (package.json) vs stable-dep@8.20.0 (node_modules).
  - extraneous: unlisted package in node_modules (exits 1; exits 0 with --ignore-extraneous)
  - missing: declared package absent from node_modules (exits 1 regardless of flags)

Each test spawns scripts/check-npm-integrity.sh as a subprocess and asserts
on exit code first, then stderr content. Tests are RED at this commit because
the script does not yet exist.

Sources:
  npm ls docs: https://docs.npmjs.com/cli/v10/commands/npm-ls
  NIST SSDF PW.4.1: https://csrc.nist.gov/publications/detail/sp/800-218/final

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(114): add check-npm-integrity.sh + workspace-aware drift detection

Adds scripts/check-npm-integrity.sh, a Bash script that:
  1. Runs `npm ls --all --json` at the invocation directory
  2. Parses JSON output for invalid, missing, and extraneous package flags
  3. Exits 1 on any finding; emits a structured report to stderr listing offenders
     with both declared and installed versions for invalid packages
  4. Exits 2 on tool error (npm/node not found, JSON parse failure)
  5. Accepts --ignore-extraneous to suppress extraneous-only failures
  6. Documents behaviour in --help output including remediation path

Workspace behaviour: the root package.json in this repo has no "workspaces"
field. npm ls runs at the invocation root and covers that tree only. The sdk/
sub-package is a separate, non-workspace package and is out of scope for a
single invocation. If workspaces are added in future, npm ls will traverse
them automatically (npm >=7).

The drift scenario (ws 8.20.1 declared vs 8.20.0 installed) is reproduced by
using an exact version pin in package.json combined with a mismatched
node_modules/package.json -- npm ls marks this as "invalid" and exits 1.

npm exits 0 for extraneous packages even though they appear in the JSON
"problems" array; this script detects them via JSON parsing regardless of
the npm exit code.

Sources:
  npm ls docs: https://docs.npmjs.com/cli/v10/commands/npm-ls
  NIST SSDF PW.4.1: https://csrc.nist.gov/publications/detail/sp/800-218/final
  OpenSSF Scorecard Pinned-Dependencies:
    https://github.com/ossf/scorecard/blob/main/docs/checks.md#pinned-dependencies

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* ci(114): wire dependency integrity gate into CI/release/security workflows

Adds a "Dependency integrity gate" step invoking
scripts/check-npm-integrity.sh to three workflows, always after `npm ci`
and before any test or build step:

  .github/workflows/test.yml
    - matrix job: after "Install dependencies" / before "Build SDK dist"
    - coverage job: after "Install dependencies" / before "Build SDK dist"

  .github/workflows/release.yml
    - rc job "Install and test": after npm ci, before npm run test:coverage
    - finalize job "Install and test": after npm ci, before npm run test:coverage

  .github/workflows/security-scan.yml
    - Added setup-node + npm ci + gate before existing source-scan steps
    - Bumped timeout-minutes from 5 to 10 to accommodate the install step

Also adds "check:integrity": "./scripts/check-npm-integrity.sh" to root
package.json scripts for local contributor invocation.

No new workflow files created. All edits extend existing workflows.

Sources:
  npm ls docs: https://docs.npmjs.com/cli/v10/commands/npm-ls
  NIST SSDF PW.4.1: https://csrc.nist.gov/publications/detail/sp/800-218/final

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(114): document dependency integrity gate in audit runbook

Appends a "Dependency Integrity Verification" section to SECURITY.md
(no docs/runbooks/ directory exists in this repo). Covers:
  - The three detection classes: invalid, missing, extraneous
  - Local invocation: ./scripts/check-npm-integrity.sh + npm run check:integrity
  - Remediation: rm -rf node_modules && npm ci
  - Bypass policy: no flag; commit-message documentation required if skipped
  - Scope: root package only (sdk/ is a non-workspace package, out of scope)
  - CI coverage listing

Sources cited:
  NIST SSDF PW.4.1: https://csrc.nist.gov/publications/detail/sp/800-218/final
  OpenSSF Scorecard Pinned-Dependencies:
    https://github.com/ossf/scorecard/blob/main/docs/checks.md#pinned-dependencies

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#114): npm integrity gate satisfies its own clean/drift/extraneous fixtures

Replace npm-ls-based analysis with pure package-lock.json parsing so the
script runs correctly in CI and test environments where node_modules is not
installed. Key changes:

- Rewrite check-npm-integrity.sh parser to read package-lock.json directly
  instead of spawning `npm ls --all --json`, which required node_modules on
  disk and incorrectly flagged clean/drift fixtures as MISSING.
- Implement a self-contained semver satisfies() covering exact, caret, tilde,
  comparison-operator, and compound ranges — no external semver package needed.
- Update extraneous fixture package-lock.json to include ghost-pkg with
  "extraneous: true" so the lockfile-based detector can identify it.
- Update missing fixture package-lock.json to omit the node_modules/absent-dep
  entry, making the absent-dep MISSING condition derivable from lockfile alone.

All 13 tests (clean ×2, drift ×3, extraneous ×3, missing ×3, help ×2) pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#114): treat transitive deps as non-extraneous in integrity gate

The extraneous check was comparing all lockfile packages against root
package.json declarations only. This caused every transitive dependency
(e.g. hono, ajv, @anthropic-ai/claude-agent-sdk-darwin-arm64) to be
flagged as EXTRANEOUS, producing false-positive failures in CI.

Only packages that npm itself marks with "extraneous: true" in the
lockfile represent genuinely unwanted packages. Transitive dependencies
installed by parent packages are valid and should be skipped.

All 13 existing tests continue to pass; the extraneous fixture still
works because it uses "extraneous: true" explicitly (npm's own marker).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* ci: retrigger checks after transient git-auth runner failure

The original run for this PR had a single CI job fail with:
"fatal: could not read Username for 'https://github.com': terminal prompts disabled"
That is a hosted-runner infrastructure flake — no code defect. The run
cannot be retried via gh CLI (too old). This empty commit kicks a fresh
full CI cycle.

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 15:50:17 -04:00
Tom Boucher
e32a53b974 feat(113): detect javascript:/data:/userinfo/token-in-query in markdown links (#133)
* test(113): add per-rule failing tests + hostile fixture for markdown link payloads

RED phase for issue #113 — scanForInjection() currently returns { clean: true }
for markdown links containing javascript:, data:text/html, userinfo credentials,
and token-in-query payloads.

Changes:
- tests/fixtures/adversarial/security/context-malicious-markdown-link.md:
  Extended to contain one hostile example per rule class (MD-LINK-JS-SCHEME,
  MD-LINK-DATA-SCHEME, MD-LINK-USERINFO, MD-LINK-TOKEN-IN-QUERY) plus benign
  negative controls (data:image/png, mailto:, https://github.com, port-only URL).
- tests/security-prompt-injection.test.cjs:
  - Flipped PINNED "malicious-markdown-link fixture is NOT flagged" assertion
    to "malicious-markdown-link fixture is flagged by scanner" (forward-looking).
  - Added 4×positive + 4×negative per-rule unit tests asserting structuredFindings
    with ruleId, file, line, match fields.
  - Added parity guard: every MARKDOWN_LINK_PATTERNS source string from
    security.cjs must appear in gsd-read-injection-scanner.js hook source.

D3 false-positive grep: 0 legitimate matches — no allowlist entries needed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(113): detect javascript:/data:/userinfo/token-in-query in markdown links (security.cjs + hook)

GREEN phase for issue #113.

Rule details (all with primary source citations):

  MD-LINK-JS-SCHEME
    Flags ](javascript:...) regardless of case.
    Source: OWASP XSS Prevention Cheat Sheet
    https://cheatsheetseries.owasp.org/cheatsheets/Cross_Site_Scripting_Prevention_Cheat_Sheet.html

  MD-LINK-DATA-SCHEME
    Flags data: URIs NOT in the explicit safe-list.
    Safe-list: image/(png|jpeg|gif|webp|bmp|ico|avif|heic) and font/(woff2?|otf|ttf).
    data:image/svg+xml is intentionally BLOCKED — SVG can host <script>.
    Source: OWASP File Upload Cheat Sheet — SVG Files
    https://cheatsheetseries.owasp.org/cheatsheets/File_Upload_Cheat_Sheet.html#svg-files

  MD-LINK-USERINFO
    Flags https?://user:pass@host in markdown link targets.
    Does NOT fire on: mailto:user@host (no :// before user) or https://host:443/path (port, not userinfo).
    Source: RFC 3986 §3.2.1 (userinfo syntax)
    https://www.rfc-editor.org/rfc/rfc3986#section-3.2.1
    RFC 9110 §4.2.4 (HTTP deprecates userinfo)
    https://www.rfc-editor.org/rfc/rfc9110#section-4.2.4

  MD-LINK-TOKEN-IN-QUERY
    Flags key NAMES: token, access_token, id_token, refresh_token, api_key, apikey,
    secret, password, client_secret, code — regardless of value.
    Source: RFC 9700 OAuth 2.0 Security BCP §4.3.1
    https://www.rfc-editor.org/rfc/rfc9700#section-4.3.1
    D3 false-positive grep: 0 legitimate matches in codebase — no allowlist needed.

Architecture:
- scripts/security.cjs: canonical MARKDOWN_LINK_PATTERNS export, scanForInjection()
  extended with structuredFindings (ruleId, file, line, match) via opts.file.
- hooks/gsd-read-injection-scanner.js: patterns inlined for hook independence
  (same pattern sources, verified by parity test).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(113): flip PINNED malicious-markdown-link assertion and add parity guard

REFACTOR phase — tightening test rigor after test-rigor skill review:

1. Fixture assertion now enumerates all 4 expected ruleIds explicitly:
   [MD-LINK-JS-SCHEME, MD-LINK-DATA-SCHEME, MD-LINK-USERINFO, MD-LINK-TOKEN-IN-QUERY].
   Previously findings.length > 0 would pass even if 3 of 4 rules were broken.

2. line field assertions tightened: `f.line >= 1` (meaningful lower bound for
   1-based line numbers) instead of `typeof f.line === 'number'` (vacuous).

3. match field assertions tightened to check the hostile content is present:
   - MD-LINK-JS-SCHEME: /javascript:/i in match
   - MD-LINK-DATA-SCHEME: /data:/i in match
   - MD-LINK-USERINFO: /@/ in match (the @ character is the definitive userinfo marker)
   - MD-LINK-TOKEN-IN-QUERY: /token=/i in match

4. Parity test checks actual RegExp .source strings (not just lengths), verifying
   the hook contains the exact canonical pattern sources character-for-character.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#113): add changeset fragment + Windows/Node 24 state.test compatibility

1. .changeset/113-malicious-markdown-links.md — required Security fragment
   for the user-facing markdown-link scanner changes in this PR (changeset-lint
   was failing with FAIL_MISSING_FRAGMENT).

2. get-shit-done/bin/lib/state-command-router.cjs — add OUTPUT_ON_SDK_ERROR
   set for mutation state subcommands whose CJS contract is always exit-0.
   On Windows/Node 24 the SDK bridge returns result.ok===false for validation
   failures (e.g. state record-metric --phase 1 with no --plan/--duration),
   causing dispatchViaSdk() to call error() (exit 1) instead of output({error})
   (exit 0). The fix maps SDK non-ok results to JSON output for the affected
   mutation commands (record-metric, advance-plan, record-session, add-decision,
   add-blocker, resolve-blocker, update-progress), restoring the exit-0 CJS
   contract on all platforms.

tests/state.test.cjs:1161 "returns error when required fields missing" passes
locally (104/104 pass).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 10:15:03 -04:00
Tom Boucher
7bd77d6268 fix(116): locale-safe base64-scan with portable timeout and partial-scan signaling (#132)
* test(116): reproduce base64-scan illegal byte sequence on non-UTF8 fixtures

Adds regression fixtures and failing tests for #116. Empirically verified
on macOS 26.5 (BSD tr) that `tr -cd '[:print:]'` under LC_CTYPE=en_US.UTF-8
exits non-zero with "Illegal byte sequence" when its input contains bytes
that are not valid UTF-8 start sequences (e.g. lone continuation bytes 0x80–0x9F).

The base64-scan.sh root cause is a known bash pitfall: the assignment
  `local printable_count=$(... | tr -cd '[:print:]' | ...)`
uses `local` on the same line, which always returns exit 0, masking the tr
failure. Result: tr errors surface only on stderr; the scan exits 0 with
incomplete coverage (false-clean signal).

Two new tests FAIL on origin/main:
  - "scans non-UTF8 file containing a b64 blob without emitting Illegal byte sequence"
  - "dir scan with non-UTF8 files under non-C locale completes cleanly within 30s"

Fixtures in tests/fixtures/base64-locale/:
  utf8-with-injection.md       — UTF-8 + base64-encoded injection (positive control)
  non-utf8-with-b64blob.bin    — raw 0x80-0x9F bytes + b64 blob that decodes to
                                 binary (this is the reproducer that triggers tr error)
  mixed-encoding.txt           — valid UTF-8 + lone continuation bytes
  clean-text.md                — negative control (must not be flagged)

Test helpers use spawnSync (not execFileSync) so stderr is captured even on
exit 0 — execFileSync only surfaces stderr via the thrown error on non-zero exit.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(116): locale-safe base64-scan with portable timeout and partial-scan signaling

Root cause: BSD tr(1) on macOS rejects input bytes that are not valid UTF-8
start sequences with "Illegal byte sequence" when LC_CTYPE is set to any
UTF-8 locale (e.g. en_US.UTF-8). Empirically verified on macOS 26.5 using
`man tr` (ENVIRONMENT section) and direct testing:
  printf '\x80\x81hello' | LC_ALL=en_US.UTF-8 tr -cd '[:print:]'
  → tr: Illegal byte sequence (exit 1)

The error is silently masked because base64-scan.sh uses `local` on the same
line as the tr assignment. Bash's `local` built-in always returns 0 regardless
of the subshell's exit code — so the tr failure never propagates under
`set -euo pipefail`. Result: the script exits 0 with truncated printable_count,
causing binary-decoded blobs to be skipped (false-clean, security gap).

Fix: `export LC_ALL=C` at script level (line 33).
  - LC_ALL=C forces the POSIX C locale throughout: tr treats every byte 0x00–0xFF
    as a valid character, never rejects high bytes.
  - Safe for all script operations: all injection patterns are ASCII, grep POSIX
    classes ([:space:], [:print:]) behave correctly in C locale, base64 -d is
    locale-independent.
  - Script-level export is appropriate because all operations in this script are
    byte-level; no multi-byte character handling is needed.

Additional hardening:
  - MAX_LINE_BYTES=1048576 guard in extract_and_check_blobs: lines longer than
    1 MiB are skipped with an explicit "partial scan" warning to stderr. This
    bounds grep -oE cost on pathological inputs (minified JS, single-line binary
    blobs) and satisfies the "partial-scan failure signaling" requirement.
  - Portable run_with_timeout + is_timeout_exit: probes for GNU timeout,
    gtimeout (homebrew), and falls back to perl alarm(N)+exec. Defined for
    future use guarding external sub-commands. Verified: no timeout binary on
    this macOS host, perl alarm fallback works correctly (exit 142 on SIGALRM).

Test-rigor fixes applied per test-rigor skill review:
  - Fixture validity check: assert `isInvalidUtf8` (round-trip length difference)
    rather than checking for a specific byte range — the property that matters is
    "file is not valid UTF-8", not "file has bytes in 0x80–0x9F".
  - FAIL assertions: assert `FAIL: ${INJECTION_FIXTURE}` (specific filepath) not
    `result.stdout.includes('FAIL')` — rules out false-positives on other fixtures.
  - Test name: renamed "mixed-encoding file does not cause scan to abort or hang"
    to accurately describe what is tested (no extractable blobs → exits clean).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(116): fix shellcheck warnings in base64-scan.sh

Address SC2034 (unused variables) and SC2329 (functions never invoked)
warnings flagged by shellcheck after the locale-hardening changes.

SC2034 fixes (pre-existing):
  - Remove unused SCRIPT_DIR variable (set but never referenced)
  - Remove unused printable_ratio local (declared but no assignment or use)

SC2329 fixes (new functions from this PR):
  - Add shellcheck disable=SC2329 annotations on run_with_timeout,
    _init_timeout_cmd, and is_timeout_exit — these are intentionally
    defined as infrastructure helpers, not called from the main loop.
    The line-length guard (MAX_LINE_BYTES) is the primary runtime
    protection; the timeout helpers are available for future use.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#116): exclude scanner fixtures from base64-scan diff mode

Add tests/fixtures/* to should_skip_file() so deliberate prompt-injection
samples in scanner fixture directories are never flagged in CI diff-mode.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 10:14:48 -04:00
Tom Boucher
835dd6ab44 test(3596): adversarial security/prompt-injection abuse suite (#3654)
* test(3596): adversarial security/prompt-injection abuse suite

Adds `tests/security-prompt-injection.test.cjs` and a fixtures
directory at `tests/fixtures/adversarial/security/` covering the
attack classes enumerated in #3596:

  - Command substitution / backticks / heredoc payloads in workstream
    names — sentinel-file probes prove no shell is spawned, slugifier
    neutralises the input.
  - Path traversal through `--ws` and slash-bearing workstream names —
    rejected with structured `--json-errors` payload, no stack trace,
    no filesystem mutation outside the project root.
  - Fake `<system>` / `[SYSTEM]` / `<<SYS>>` / `[INST]` boundary tags —
    sanitizeForPrompt neutralises every form; structural negative
    property locked across all six styles in one place.
  - Zero-width / bidi-override codepoints — stripped per the documented
    codepoint set; asserted via codePoint inspection, not regex
    literals.
  - Hostile read of CONTEXT.md / PLAN.md / ROADMAP.md fixtures —
    `gsd-read-injection-scanner.js` surfaces the advisory; excluded
    paths and non-Read tools stay silent; malformed JSON does not
    crash the hook.
  - Hostile write of `.planning/` files — `gsd-prompt-guard.js` emits
    a `PreToolUse` advisory; non-Write/Edit tools stay silent.
  - Fake `ghp_*` / `sk-*` env tokens — never echoed in CLI stdout or
    stderr under hostile inputs; covered under
    `// allow-test-rule: structural-regression-guard` because the only
    way to assert byte-level absence is `.includes(token)` against the
    captured streams.
  - `validatePath`, `validateShellArg`, `validatePhaseNumber`,
    `validateFieldName` — focused negative-input contract pins.

Pinned behavior gaps (documented, NOT fixed in this PR):

  - `<instructions>` is intentionally whitelisted by both the scanner
    and the sanitiser (GSD's own prompt scaffolding). Two REGRESSION
    GUARD tests lock that contract.
  - The current `scanForInjection` does NOT flag malicious markdown
    links (javascript:/data:/embedded-credentials URLs). PINNED with
    negative-proof so any future scope extension fails the assertion
    and forces a deliberate update to the acceptance map.
  - `prompt-builder.ts` does not yet wrap plan/context markdown in an
    "untrusted data" envelope. That seam lives on the TS side and is
    covered by `sdk/src/prompt-builder.test.ts`; out of scope for a
    CJS test file. Mentioned in the file header.

Verification:

  - `node --test tests/security-prompt-injection.test.cjs` → 73 tests
    pass.
  - `node scripts/lint-no-source-grep.cjs` → 0 violations across
    546 test files (one `allow-test-rule: structural-regression-guard`
    annotation on this file for the token-absence assertions).
  - `node scripts/run-tests.cjs` → 9730 tests pass, 0 fail.

Refs #3596

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(3596): allow adversarial fixtures in scan + harden graphify status parse

* fix(3596): skip adversarial security fixtures in secret scan

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-17 00:34:53 -04:00
Tom Boucher
b1317633db test(3594): adversarial parser fixtures + frontmatter/roadmap matrix + property-style suite (#3633)
* test(3594): adversarial parser fixtures + frontmatter/roadmap matrix + property-style suite

Lands the adversarial parser-input corpus that CONTRIBUTING.md
§"QA Matrix Requirements / Parser and project-file inputs" and
TEST-EXAMPLES.md §"Parser Adversarial Fixtures" describe.

New tests/fixtures/adversarial/ layout:

  frontmatter/
    duplicate-keys.md            — same key twice (collapses last-wins)
    crlf-mixed.md                — CRLF endings throughout
    unclosed-block.md            — `---` open with no close
    unicode-keys-and-values.md   — non-ASCII + emoji + Greek
    null-byte-value.md           — U+0000 in a value
    huge-bounded.md              — 2000-item array, ~30KB

  roadmap/
    phase-heading-inside-fenced-code.md   — #2787 fence shadowing
    nested-fenced-code.md                 — outer + inner ``` blocks
    unicode-phase-titles.md               — JP / Greek / emoji titles
    repeated-phase-ids.md                 — phase 1 declared twice
    decimal-phase-mixed.md                — 2 vs 2.1 vs 2.10 vs 21
    markdown-headings-inside-html-comment.md — comment shadowing

Test files (all node:test, no try/finally in test bodies, no source-grep,
no raw-text matching on stdout/file content):

  tests/feat-3594-parser-adversarial-frontmatter.test.cjs (12 tests)
    Loads each fixture, pins parser invariants on extractFrontmatter()
    return shape. Cross-corpus "does not throw on any fixture" sweep.

  tests/feat-3594-parser-adversarial-roadmap.test.cjs (18 tests)
    Loads each fixture into a temp project's .planning/ROADMAP.md and
    drives `gsd-tools roadmap get-phase <N>` via the runCli harness
    introduced by #3593. Asserts on the typed JSON payload.

  tests/feat-3594-parser-property-style.test.cjs (2 tests)
    Deterministic mulberry32 PRNG generates 500 malformed-ish
    frontmatter inputs per test. Pins (a) extractFrontmatter is total
    over the corpus (no null-deref TypeError, always returns a plain
    object on success), (b) the suite completes well under 2 seconds
    (quadratic-regression guard).

Known-open bugs surfaced and pinned (intentionally NOT fixed in this
PR — separate issues warranted):

  - CJS roadmap parser matches `## Phase N:` headings inside fenced
    code blocks (the SDK parser tracks fences per the #2787 comment in
    sdk/src/query/roadmap.ts but the CJS path has not caught up).
  - CJS roadmap parser matches `## Phase N:` headings inside HTML
    comments.

Both are documented in-test with the "currently STILL matches it
(open: needs <fix>)" naming pattern so the day the production fix
lands, flipping the assertion from `found: true` to `found: false` is
the regression guard.

Test totals:
  - 32 new feat-3594-* tests (12 frontmatter + 18 roadmap + 2 property)
  - 108/108 pass when running together with the pre-existing
    frontmatter.test.cjs + roadmap.test.cjs suites (76 of theirs).

Closes #3594

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(3594): use Fisher-Yates shuffle for deterministic seeded inputs

Replaces `arr.sort(() => rng() - 0.5)` with a Fisher-Yates shuffle
driven by the supplied PRNG. The sort-based shuffle is non-transitive:
V8's TimSort behavior on non-transitive comparators is engine-defined,
so the same seed produced different orderings across Node versions —
undermining the test's stated reproducibility guarantee.

Fisher-Yates is O(n), transitive (no comparator at all), and consumes
exactly n-1 RNG values in a fixed order. The mulberry32 seed now
determines the input sequence end-to-end.

Codex review on PR #3633.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-16 00:23:27 -04:00
Tom Boucher
75d5ca5875 feat(code-review): integrate fallow structural pre-pass for /gsd-code-review (#3424)
* feat(code-review): add optional fallow structural pre-pass

* fix(ci): sync lockfile for fallow optional binaries

* fix(test): make fallow integration tests cross-platform

* fix(review): require executable fallow binary paths

* docs(review): clarify structural findings usage and size guard

* fix(fallow): preserve line:0, prefer node_modules/.bin, sync SDK twin (H1, M2, N1 from #3424 review)

* fix(workflow): harden fallow pre-pass — exit check, timeout, atomic write, size-guard order (B1, H2-H4, M1, M3 from #3424 review)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(deps): pin fallow floor to ^2.70.0 matching lockfile (H7 from #3424 review)

* fix(config): enum-validate fallow.scope/profile + group code_quality.* contiguously (H5, N3 from #3424 review)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(fallow): label mcp gate reserved, version-pin install, expand context schema (B3, H8, M4, M8, L1 from #3424 review)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(fallow): replace source-grep with behavioral tests, expand fixtures, fail-loud tmpdir (B4, H6, L2, L3, M5, M6, N2 from #3424 review)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(workflow): escape closing structural_findings tag in JSON payload (CR #3424 inline finding)

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-13 21:19:47 -04:00