Commit Graph

4646 Commits

Author SHA1 Message Date
Tom Boucher
b2f4aa9435 docs(#2420): clean stale get-shit-done/ path refs in translated docs (#2421)
After the package/repo rename in #604, the English docs were updated to
use gsd-core/... paths, but the four translated doc trees (ja-JP, zh-CN,
ko-KR, pt-BR) and .changeset/README.md were never updated and still
referenced the pre-rename get-shit-done/ runtime directory, which no
longer exists.

This commit brings the translations in line with the English docs:

  - docs/{ja-JP,zh-CN,ko-KR,pt-BR}/**/*.md (57 files):
      get-shit-done/ -> gsd-core/  (path references)
      #references-get-shit-donereferencesmd -> #references-gsd-corereferencesmd
                                            (anchor in INVENTORY -> ARCHITECTURE links)
  - .changeset/README.md:9 issue URL:
      open-gsd/get-shit-done-redux -> open-gsd/gsd-core

Legacy references intentionally preserved (historical record):
  - CHANGELOG.md, .changeset/archived/*, docs/RELEASE-NOTES-LEGACY.md
  - docs/cleanup-get-shit-done-cc.md, docs/adr/*, docs/research/*
  - docs/{ja-JP,ko-KR}/superpowers/plans/2026-03-18-* (developer's local paths)
  - docs/{INVENTORY,README,FEATURES,installer-migrations}.md (rename-history
    descriptions, some tagged <!-- gsd-allow-legacy-name -->)
  - Code/tests implementing or testing legacy-cleanup logic
    (bin/install.js, gsd-core/bin/lib/legacy-cleanup.cjs,
    scripts/lint-legacy-dir-name.cjs, migration sources/tests)

No source code changes — documentation only.

Fixes #2420
2026-07-18 23:11:15 -04:00
Tom Boucher
873bdf51e5 fix(#2352): expand tilde paths in review scope before the deleted-file filter (#2419)
* fix(#2352): tilde-expand SUMMARY.md key-files paths before deleted-file filter

compute_file_scope's "Filter deleted files" step tested the literal `~/...`
value from SUMMARY.md key-files entries with `[ -f "$file" ]`, which bash
never tilde-expands (only a literal `~` in source text expands, not one
arriving as an already-expanded variable value). Real files recorded with a
`~/...` path were silently misclassified as deleted and dropped from
REVIEW_FILES, and a phase whose every recorded file used a tilde path hit the
empty-scope skip as a false negative.

Adds a tilde-normalization loop as step 1 of post-processing (all tiers),
before the deleted-file filter, rewriting a leading `~/` to `${HOME}/...` so
downstream existence checks, the empty-scope short-circuit, and the
FILES_TO_READ/CONFIG_FILES construction all see a real, openable path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2352): regenerate fixtures + lint gate-prep

* chore(#2352): add Fixed changeset fragment (pr 2419)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 21:29:58 -04:00
Tom Boucher
dd5a2211c9 enhance(#1964): semantic knowledge-base recall via MemPalace (keyword fallback) (#2416)
* test(#1964): add failing-first semantic-recall contract tests

Epic #1957 Phase 3C (final). Source-text-is-the-product contract tests:
semantic recall via MemPalace (top-k meaning-similar prior resolutions, catches
same-root-cause/different-wording cases), indexing resolved sessions at archive,
graceful degradation to keyword matching when MemPalace is absent,
knowledge-base.md stays the durable plain-text source of truth, agent Phase 0 /
Matching Logic is semantic-first (the stale 'keyword overlap, not semantic
similarity' claim must go), and no new embedding/vector infra (reuse MemPalace).

Failing-first: reference, the Matching Logic reframe, the Phase 0 consolidation,
and the archive indexing step do not yet exist.

* feat(#1964): semantic knowledge-base recall via MemPalace (keyword fallback)

Epic #1957 Phase 3C (FINAL). Replaces keyword-overlap matching with semantic
recall: at Phase 0 the debugger queries MemPalace with the current symptoms
and surfaces the top-k meaning-similar prior resolutions, catching the
same-root-cause/different-wording cases keyword overlap missed (the self-noted
'keyword overlap, not semantic similarity' limitation). Resolved sessions are
indexed into MemPalace at archive (symptoms + root_cause(s) + fix + recurrence
guard). knowledge-base.md remains the durable plain-text source of truth; when
MemPalace is absent the debugger falls back to keyword-overlap matching
(logged, never a silent skip). No new embedding/vector infrastructure —
MemPalace is reused.

Size-neutral agent edits: the Matching Logic section reframed (keyword-only ->
semantic-first + keyword-fallback + @-include); Phase 0's three keyword bullets
consolidated into one semantic-first bullet; one MemPalace-indexing step added
at archive. Agent at 57222 B (122 B headroom — final phase). Full rules in
gsd-core/references/debugger-semantic-recall.md. INVENTORY + manifest +
agent-size baseline + install-parity goldens + AGENTS.md updated.

* fix(#1964): address orthogonal review (invocation mechanism, index Resolution-not-symptoms + redaction, fallback detail)

- HIGH: the 'query MemPalace' instruction was WHAT-level only; the agent has
  no MCP tools. Added an Invocation section naming the Bash CLI
  (mempalace search --wing <wing>) + MCP-when-registered + wing resolution
  (config.mempalace.wing -> project_code -> project dir), matching every other
  MemPalace integration. Without this the feature silently degraded to keyword
  matching even when MemPalace was present.
- MEDIUM (security x2): index the agent-authored Resolution summary
  (root_cause + fix + recurrence_guard), NOT raw user-supplied Symptoms —
  excludes attacker-controlled prose from the cross-session index AND reduces
  secret/PII leakage. Redact secret-shaped values before indexing. Stated the
  write order (KB append + commit MUST succeed before indexing).
- LOW: restored 'identifiers' + 'case-insensitive' to the keyword fallback;
  added a test asserting the fallback mechanics survived the Phase 0
  consolidation (Error patterns field, 2+ token overlap, identifiers,
  case-insensitive).

* chore(#1964): ratchet agent-size baseline downward (leaner archive bullet shrank gsd-debugger.md 57222->57197)

* chore(#1964): backfill changeset pr number (PR #2416)
2026-07-18 19:01:04 -04:00
Tom Boucher
c67f301867 feat(#1963): emit blameless-postmortem Prevention block at resolution (#2410)
* test(#1963): add failing-first prevention/postmortem contract tests

Epic #1957 Phase 3B. Source-text-is-the-product contract tests: blameless
5-Whys that BRANCHES per Phase 2A RCA (not a single-cause chain; treats agent
error as 'why was that possible?'), the 'why wasn't this caught?' question,
the recurrence-guard taxonomy (regression test / assertion / lint rule / KB
pattern), the KB-entry why_not_caught + recurrence_guard fields with backward
compat, the session-manager prevention summary line, and the Zawinski
scope-boundary (a block, not a subsystem).

Failing-first: reference, archive_session edit, KB schema extension, and
session-manager summary do not yet exist.

* feat(#1963): emit blameless-postmortem Prevention block at resolution

Epic #1957 Phase 3B. At archive_session the debugger now produces a
Prevention block with three blame-free components: a branching 5-Whys causal
chain (branches per Phase 2A RCA, not a single chain; 'agent error' prompts
'why was that possible?', never blame), a 'why wasn't this caught?' answer
naming the missed gate (test/typecheck/lint/review/verify), and a concrete
recurrence guard (regression test / assertion / lint rule / KB pattern).

The knowledge-base entry gains two structured fields (why_not_caught +
recurrence_guard) so future Phase-0 recall surfaces the prior prevention, not
just the prior fix. Additive: old entries without the fields still load. The
session-manager compact summary surfaces a one-line prevention summary.

Full rules extracted to gsd-core/references/debugger-prevention.md (slim
archive_session step + 2 KB fields kept in the agent). INVENTORY + manifest +
agent-size baseline + install-parity goldens + AGENTS.md updated.

* fix(#1963): address orthogonal review (CRITICAL append-template drift + Phase-0 consumption + parity test)

- CRITICAL: the archive_session KB append template omitted Why not caught +
  Recurrence guard (only the Entry Format had them) — the feature's core
  deliverable silently did not happen. Added both fields to the append template
  the agent actually follows (nearest-instruction wins).
- HIGH: Phase 0 (KB read) only surfaced root_cause + fix; the new fields were
  dead data. Extended the Phase 0 Evidence line to consume why_not_caught +
  recurrence_guard when present (absent on old entries — backward compat holds).
- MEDIUM: added a cross-section parity test (every Entry-Format field must also
  appear in the append template — the guard that would have caught the
  Critical) + a Phase-0-consumption assertion.
- MEDIUM: the 'branches per Phase 2A' claim is now wired — reuses
  reasoning_checkpoint.candidate_causes across the four categories.
- MEDIUM: recurrence-guard taxonomy gains type refinement + config-default
  change; LOW: added 'build' gate to both surfaces for parity.
- NIT: compact-summary fallback shape ('no gate existed'); verify the guard
  artifact exists before recording it.

* test(#1963): anchor Phase-0 consumption test on the specific heading

The regex /Phase 0[\s\S]{0,1200}/ matched the first 'Phase 0' in the file
(in knowledge_base_protocol prose), not the Phase 0 block in investigation_loop.
Anchor on '**Phase 0: Check knowledge base**' and widen to 1500 chars.

* chore(#1963): backfill changeset pr number (PR #2410)
2026-07-18 17:38:54 -04:00
Tom Boucher
36a311c5bb enhance(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries) (#2409)
* test(#1962): add failing-first repro-hardening contract tests

Epic #1957 Phase 3A. Source-text-is-the-product contract tests: PBT shrinking
(fast-check/Hypothesis, minimized seed, manual-minimization degradation), the
four oracle types (specified/derived/metamorphic/implicit with implicit flagged
weakest), boundary neighbors (off-by-one/min-max/empty-singleton tied to the
equivalence class), oracle_type in DEBUG Resolution, and the Phase 1A tie-in
(minimized seed + real oracle => the mutation guardrail bites).

Failing-first: reference, agent cross-refs, and template field do not yet exist.

* feat(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries)

Epic #1957 Phase 3A. Extends Minimal Reproduction (shrinking) and Test-First
Debugging (oracle classification + boundary neighbors):
- Shrinking: wrap an input-space failing input in a property (fast-check JS/TS,
  Hypothesis Python) and store the MINIMIZED counterexample as the regression
  seed; degrade to manual minimization when no PBT framework is present.
- Oracle classification: state specified / derived (contract/model) /
  metamorphic / implicit (crash, weakest) before writing the assertion; record
  under Resolution.oracle_type; never default to implicit silently.
- Boundary neighbors: off-by-one, min/max, empty/singleton around the fixed
  defect's equivalence class.

Together they turn the regression test into a root-cause check — what the Phase
1A mutation guardrail needs to bite. Full rules extracted to gsd-core/references/
debugger-repro-hardening.md. INVENTORY + manifest + agent-size baseline +
install-parity goldens + AGENTS.md + DEBUG template updated.

* fix(#1962): address orthogonal review (bounding, provenance, oracle scope, sufficient-triple)

- HIGH: added a 'Bound the property/shrink run' section (60s timeout, degrade-
  to-manual on timeout, do-not-raise-default-run-limits, argv-not-shell) —
  the gauntlet violation the sibling references already honored.
- Medium: test-provenance caveat (the failing input often comes from the bug
  report — author the generator from a sanitized description, cross-ref
  debugger-fix-acceptance.md).
- Medium: oracle scope note — the 4 types cover deterministic bugs; non-
  deterministic failures re-route to stability-stress per bug-taxonomy.
- Medium: Phase 1A tie-in corrected — seed+oracle is necessary not sufficient;
  boundary neighbors close the adjacent-input escape; the sufficient triple is
  seed+oracle+neighbors.
- Low: preserve the original noisy repro as a secondary reference; operationalize
  'equivalence class' (the predicate the fix draws). Nit: degradation reworded.

* chore(#1962): backfill changeset pr number (PR #2409)

---------

Co-authored-by: sim <sim@local>
2026-07-18 15:46:42 -04:00
Tom Boucher
6baa2a8182 feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger (#2407)
* test(#1961): add failing-first bug-taxonomy routing contract tests

Epic #1957 Phase 2B. Source-text-is-the-product contract tests (3 taxonomy
classes, explicit class->technique routing table, Bohrbug->repro+SBFL+bisect,
Heisenbug->record-replay/stability+SKIP-SBFL, Concurrency->atomicity/order/
deadlock checklist, bug_class in DEBUG Current Focus, supersede-not-append)
plus a routing-table specification object pinning the documented decisions
(SBFL forbidden on Heisenbug is the load-bearing 1B/2B seam).

Failing-first: reference, Phase 1.75, and routing-table reframe do not yet exist.

* feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger

Epic #1957 Phase 2B (reliability-critical). Adds Phase 1.75: classify the
failure as Bohrbug / Heisenbug-Mandelbug / Concurrency, then route the
investigation technique via an explicit class->technique table (Kernighan: no
opaque heuristic). Bohrbug -> reproduction + SBFL (Phase 1.25) + git bisect;
Heisenbug/Mandelbug -> record-replay (rr) + stability-stress + statistical
sampling, with SBFL explicitly SKIPPED (a flaky spectrum poisons the Ochiai
ranking — the load-bearing 1B/2B seam); Concurrency -> the
atomicity/order/deadlock checklist first.

Reframes (supersedes, not appends — Zawinski) the flat 'Technique Selection by
situation' table into a class-routed table; the 11 techniques remain as routed
targets. bug_class recorded in Current Focus (DEBUG template); common-bug-
patterns catalog cross-referenced to the taxonomy.

Full rules extracted to gsd-core/references/debugger-bug-taxonomy.md. INVENTORY
+ manifest + agent-size baseline + install-parity goldens + AGENTS.md updated.

* fix(#1961): address orthogonal review (phase-name drift, General lane, revoke framing, row-scoped tests, bounding)

- HIGH: reference said 'Phase 1B' (epic shorthand); corrected to the deployed
  'Phase 1.25' (matches the agent + SBFL reference).
- HIGH: 6 of 11 techniques (Rubber duck, Delta, Working backwards,
  Differential, Comment-out, Follow-the-indirection) were orphaned by the
  situation-table reframe. Added a 'General (any class, situation-cued)'
  lane to BOTH the reference routing table and the agent's Technique
  Selection table that re-homes them — supersede-not-append now holds.
- MEDIUM: the SBFL-skip is structurally retroactive (Phase 1.25 runs before
  Phase 1.75 classification), so reframed the table column from 'Do NOT use'
  to 'Revoke if already run' + an explicit 'retroactive revocation, not
  proactive skip' note stating the ordering honestly.
- MEDIUM: contract tests are now row-scoped (parse the table by class, assert
  per-row) instead of presence-only; added a guard that the previously-
  orphaned techniques now have a General-lane route.
- LOW: pinned the canonical bug_class value form (lowercase-kebab:
  bohrbug|heisenbug-mandelbug|concurrency; prose may use title-case).
- NIT: added a 'Bound the Heisenbug-chase runs' note (rr/stability/sampling
  timeouts) per the unbounded-subprocess gauntlet.

* chore(#1961): backfill changeset pr number (PR #2407)
2026-07-18 14:42:29 -04:00
Tom Boucher
f8b16d1874 enhance(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger (#2405)
* test(#1960): add failing-first RCA-branching contract + schema-invariant tests

Epic #1957 Phase 2A. Source-text-is-the-product contract tests (fishbone
>=2 categories, AND-gate, multi-cause root_cause, backward compat, reasoning
checkpoint candidate_causes+and_gate fields, debugger-philosophy single-cause
note, DEBUG template) plus behavioral schema-invariant checks on two fixtures:
two contributing causes (AND-gate yes) -> both recorded; single-cause
(AND-gate no) -> one root_cause, identical to today.

Failing-first: reference, agent edits, and template note do not yet exist.

* feat(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger

Epic #1957 Phase 2A. Guards against 5-Whys single-cause bias: before committing
root_cause, the debugger enumerates candidate causes across >=2 Ishikawa
categories (code/config/environment/data) and explicitly answers an AND-gate
question. When the AND-gate fires, every contributing cause is recorded, so a
multi-cause fix no longer recurs via the unaddressed second cause.
Resolution.root_cause may hold one OR a small set (additive; single-cause
sessions are byte-identical to today). The Structured Reasoning Checkpoint gains
candidate_causes + and_gate fields; debugger-philosophy.md adds the
single-cause-bias trap.

Full rules extracted to gsd-core/references/debugger-rca-branching.md (slim
Phase 2 routing + 2 checkpoint fields kept in the agent). INVENTORY + manifest +
agent-size baseline + install-parity goldens + AGENTS.md + DEBUG template updated.

* fix(#1960): address orthogonal review (AND-gate self-consistency, parity guard, narrowed claim, ripples)

- Reference: the collapse rule now enforces AND-gate self-consistency —
  and_gate=yes with a single confirmed cause is flagged as incomplete
  (return to Phase 3); a race/timing note clarifies such bugs bridge
  categories; the 'byte-identical' backward-compat claim narrowed to
  'root_cause shape unchanged; reasoning_checkpoint gains 2 fields in every
  session'.
- DEBUG.md: stale 'five-field' mirror prose -> seven-field (parallel-surface
  drift the reviewer flagged); new debug-session-management parity test pins
  the field-count claim to the gsd-debugger.md YAML keys (CRLF-safe).
- Scalar-assuming consumers of set-valued root_cause updated: session-manager
  compact summaries (319/332), diagnose-only return (1062), archive entry
  (1216), ROOT CAUSE FOUND return (1322).
- Test: added the AND-gate-yes/single-cause invariant + fixture; rephrased the
  fixture describe block honestly as a schema-invariant specification.
- Phase 2 bullet phrasing clarified ('at hypothesis formation, before the
  Phase 4 commit').

* test(#1960): parity regex accepts word-form count ('seven-field' or '7-field')

* test(#1960): parity regex counts array-valued YAML keys (no inline value)

* chore(#1960): backfill changeset pr number (PR #2405)
2026-07-18 13:42:58 -04:00
0xdhx
50efae13ce fix(#2305): stage the shared guard hooks Kilo's native plugin spawns (#2327)
* fix(#2305): stage shared guard hooks for Kilo — drop skipSharedHooksInstall

Kilo's capability descriptor declared BOTH hostBehaviors.nativePlugin (a
plugin that spawns the shared PreToolUse guard scripts as subprocesses)
AND hostBehaviors.skipSharedHooksInstall:true, which suppresses staging
of hooks/*.js into the Kilo config dir. The plugin's runHook treats an
absent hook script as a silent allow, so every guard it spawned
(gsd-prompt-guard, gsd-read-guard, gsd-worktree-path-guard) no-opped on
every Kilo install. OpenCode uses the byte-identical plugin with hook
staging on and is unaffected — it is the reference shape.

The skip flag predates Kilo's plugin surface: it dates to #1821 (hooks
were dead weight for a runtime with no hook consumer), and #2093 added
the hooks-dependent nativePlugin without revisiting it.

- capabilities/kilo/capability.json: remove skipSharedHooksInstall
  (regenerated gsd-core/bin/lib/capability-registry.cjs accordingly)
- bin/install.js: correct the stale #1821 comments claiming Kilo has no
  plugin surface
- tests/kilo-upgrades.test.cjs: install-fixture tests (global + local)
  asserting the guard scripts land where the plugin's walk-up resolves
  them; an end-to-end test driving a disallowed out-of-worktree write
  through the REAL installed Kilo tree and asserting the guard rejects
  it; a cross-runtime descriptor invariant (nativePlugin and
  skipSharedHooksInstall:true must never coexist)
- tests/kilo-imperative-reference.test.cjs: flip the pinned assertion
- golden fixtures regenerated (kilo now stages the 24 hook files, same
  set as OpenCode)

Fixes #2305

* fix(#2305): warn loudly when a guard hook script is missing (runHook)

runHook's absent-file branch returned a silent exit-0 allow — the
mechanism that let #2305 ship undetected: with the hooks bundle never
staged on Kilo, every PreToolUse guard the plugin spawned resolved to
"file not found → allow" with zero signal anywhere.

Keep the adapter's design contract (a missing hook must never break the
tool call — pinned by the existing adapter test) but make the absence
loud: console.error once per hook file, naming the unresolved path and
the remediation. Applied identically to .kilo/ and .opencode/ plugin
copies (byte-parity guard). Golden parity fixtures regenerated (the
installed plugin file's hash changed).

Fixes #2305

* chore(#2305): add changeset fragment

* test(#2305): include gsd-workflow-guard.js in the staged-guards regression list

The native plugin spawns four guards on write-like tool calls — the
regression test's PLUGIN_GUARD_HOOKS list covered three. Staging itself
was already asserted via the golden fixtures (the full bundle), but the
named per-guard assertion should cover every guard the plugin actually
dispatches. Surfaced by cross-AI review of PR #2327.

* test(#2305): update the #1821 tests that encoded Kilo's false no-plugin premise

The #1821 hook-copy test asserted Kilo must receive no staged hooks — the
exact behavior this PR reverses (and the cause of all 8 CI failures). Kilo
moves from the ZCode "no dead hooks" loop to the OpenCode group, with
positive assertions on the new contract: the three guard hooks the plugin
spawns, hooks/lib/git-cmd.js, and plugins/gsd-core.js all staged. The
integration runtime contract flips kilo packageJson to true (the CommonJS
marker ships with the bundle), and the pi contract comment no longer cites
Kilo as a no-plugin runtime.

* chore(#2305): scope the queued #1821 changeset fragment to ZCode only

The fragment still claimed the installer skips hooks for Kilo — rendering
both it and this PR's fragment into the same release would ship two
contradictory statements about Kilo's install behavior. It now claims
ZCode only and notes that #2327 reverses the Kilo half.

* chore(#2305): rename changeset fragment to the generator naming convention

2305-kilo-stage-guard-hooks.md -> loud-guard-hooks.md, matching the
<adjective>-<noun>-<noun> shape npm run changeset generates (review nit).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-18 12:20:32 -04:00
Tom Boucher
13d181aedf fix(#2349): exclude status: superseded plans from phase completion counts (#2404)
Adds a status: superseded plan-frontmatter marker that scanPhasePlans excludes from both plan and summary counts, so a phase with a deliberately-unexecuted plan no longer reads incomplete forever (the plan-level analogue of #1514). Includes all-superseded completion handling and a bounded, symlink-safe frontmatter read. Fixes #2349.
2026-07-18 08:08:03 -04:00
Tom Boucher
56a5c6404c feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger (#2403)
* test(#1959): add failing-first SBFL contract + Ochiai correctness tests

Epic #1957 Phase 1B. Source-text-is-the-product contract tests (Ochiai
formula documented, Tarantula fallback, top-N seeding, no-coverage skip
logged, ranking->Evidence, Bohrbug gating) plus a behavioral Ochiai
formula-correctness section: bound [0,1], max-score invariant, a known-fault
fixture proving the fault ranks #1 (criterion 2), clean degradation on
zero failing tests, and two fast-check properties.

Failing-first: reference file and agent routing do not yet exist.

* feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger

Epic #1957 Phase 1B. When a runnable test suite with per-test coverage exists
(>=1 failing AND >=1 passing test), the debugger computes an Ochiai
suspiciousness ranking over the coverage spectrum and seeds the top-N
suspicious locations into Evidence as first-class hypothesis candidates,
narrowing the search space deterministically before LLM reasoning. Tarantula
documented as fallback. Degrades cleanly (logged, never silent) when there is
no test suite, no failing tests, or no per-test coverage, and is explicitly
not trusted on flaky/Heisenbug spectra (pairs with Phase 2B bug-taxonomy).

Full rules extracted to gsd-core/references/debugger-sbfl.md (slim Phase 1.25
routing kept in the agent to respect the size cap). No new coverage framework
— reuses the project's existing test/coverage runner. INVENTORY + manifest +
agent-size baseline + install-parity goldens + AGENTS.md updated.

* test(#1959): bound property generators to valid coverage counts

The [0,1] property generated failedExec independently of totalFailed, but
Ochiai's score is only bounded by 1 under the coverage invariant
failedExec <= totalFailed (a failing test that executed s is one of the
totalFailed failing tests). Out-of-domain inputs (failedExec=100, totalFailed=5)
make the formula correctly return >1. Bound failedExec by totalFailed via
fc.chain so the property tests the real domain. Also cleaned up the ranking
property (removed dead code).

* fix(#1959): address orthogonal review (monotonicity property, degradation row, coverage bounding)

- Replace vacuous ranking property (true-by-sort-construction) with a
  non-trivial monotonicity property: holding totalFailed + passedExec fixed,
  ochiai is non-decreasing in failedExec. An inverted formula would fail it.
- Add the missing 'no passing tests' degradation row (preconditions require
  >=1 passing test; Tarantula would divide by totalPassed=0).
- Bound the coverage subprocess (CLAUDE.md gauntlet): cap the coverage run,
  degrade-to-skip on timeout, never hang the debug session.
- Reword 'discard the ranking' -> 'mark the Evidence entry as revoked (do not
  delete)' per Kernighan auditability.

* test(#1959): bound monotonicity-property generator to valid coverage (failedExecA <= totalFailed)

* chore(#1959): backfill changeset pr number (PR #2403)
2026-07-18 07:55:28 -04:00
Tom Boucher
863a54ec82 fix(#2350): pass --raw to config-get in every build/test gate (#2399)
Adds --raw to config-get workflow.build_command|test_command reads in the post-merge, regression, verify-phase, and audit-fix gates so an unset key is a genuinely empty string, not the literal "" — restoring the auto-detect cascade and graceful skip instead of a false exit-127 failure. Regression guard sweeps all four gate files. Fixes #2350.
2026-07-18 01:14:45 -04:00
Tom Boucher
5e52350736 feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger (#2396)
* test(#1958): add failing-first guardrail contract tests

Epic #1957 Phase 1A. Adds source-text-is-the-product tests asserting the
5-signal fix-acceptance guardrail contract (target test, mutation check,
no-op/deletion detector, adjacent tests, revert-and-reconfirm), graceful
degradation, FIX REJECTED BY GUARDRAIL return path, per-signal debug-file
recording, and subprocess bounding.

Failing-first: reference file and agent sections do not yet exist.

* feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger

Epic #1957 Phase 1A. Prevents accepting a fix that merely greens the test
(Goodhart defense / APR overfitting). Adds a 5-signal gate run before fix
acceptance: target test, mutation check (Stryker), no-op/behavior-deleting
detector, adjacent/held-out tests, revert-and-reconfirm. Degrades gracefully
when Stryker or a test suite is absent (each skip logged, never a silent pass),
records per-signal results under Resolution.verification, and returns a
FIX REJECTED BY GUARDRAIL outcome the session-manager surfaces for
revise / accept-as-debt / abandon.

Full rules extracted to gsd-core/references/debugger-fix-acceptance.md (slim
routing kept in the agent to respect the agent-size cap). Debug template +
INVENTORY + manifest + agent-size baseline + AGENTS.md updated.

* test(#1958): correct newline-tolerant assertion + regen install-parity goldens

The revert-and-reconfirm assertion collapsed whitespace before matching so
markdown line-wrapping does not break it. Regenerated the golden-install-parity
and install-tree fixtures (npm run gen:golden) to absorb the intentional
gsd-debugger.md / gsd-debug-session-manager.md / DEBUG.md / new reference-file
changes to the installed artifact tree.

* fix(#1958): tighten guardrail per orthogonal review

Addresses the isolated reviewer's findings:
- signal 5 now states its recorded-repro dependency and routes the no-repro
  case to the degradation row; revert mechanism specified (git stash / git
  revert -n); minimality flag tied to diff structure, not revert-ability.
- bounded-subprocesses section now bounds the git subprocess (5-30s) too,
  requires argv-array argument passing, and scopes Stryker to the driving
  regression test (a mutant killed only by a non-driving test is a finding).
- new test-provenance (security) clause: the driving test must be
  agent-authored; bug-report repro scripts are DATA, never executed verbatim.
- tightened 3 contract assertions to bind to specific clauses
  (guardrail_verdict field, deletion-reject-unless-RCA, 60s+git bounding).
- Goodhart framing softened to 'partially-independent'; DEBUG.md template
  verification field notes the nested map shape.

* chore(#1958): backfill changeset pr number (PR #2396)

* fix(#1958): add issue ref to allow-test-rule annotation (ADR-456)

CI lint-allow-test-rule-refs requires every allow-test-rule exemption to
carry a 'see #NNN' issue ref per ADR-456. The new test file's annotation
lacked it; this adds (see #1958).
2026-07-18 00:39:28 -04:00
Tom Boucher
2c54f219c9 fix(#2348): derive verification staleness from git commit time, not mtime (#2394)
readVerificationStatus() decided a phase's verification was `stale` (a
*-SUMMARY.md newer than the *-VERIFICATION.md) by comparing filesystem
mtimes. mtimes are assigned at checkout time and are not preserved by
`git clone` / `cp -R`, and any unrelated `touch` / reformat / editor-save
re-stales a valid report — so a committed phase declaring `status: passed`
could silently read `stale` on a fresh clone purely from checkout order,
falsely rewriting a ROADMAP row and blocking milestone close (#2022 gate).

Each file's effective "last changed" time is now its git commit time when
the file is committed AND clean, and its mtime otherwise (uncommitted or
working-tree-dirty). Both are real wall-clock change times, so a summary
committed after — or edited after — the verification reads stale, while a
clean fresh clone stays passed. Git commit time is content-tied and clone-
stable; mtime is retained only where it is the true last-changed signal.

Implementation:
- Two bounded git calls per phase (never one-per-file): `git log
  --first-parent --format=%ct --name-only` for commit times, and `git diff
  --name-only HEAD` to drop dirty files. readVerificationStatus runs
  per-phase in the init/roadmap listing loops, so per-file spawning would
  fan out to P×(S+1) git processes ("Unbounded Subprocesses").
- `--first-parent` so merge commits report their file lists (plain
  `--name-only` omits merge diffs and would under-date merge-landed content).
- The dirty-check fails SAFE: if `git diff` is inconclusive (errors / exits
  non-zero) the commit times are discarded so every file falls back to mtime,
  never trusting a possibly-stale commit time (no false "not stale").
- Paths matched back by `/`-bounded suffix (root vs nested `plans/` can't
  collide) and passed after `--` (dash-named files can't be read as flags).
- A phase with no summaries skips git entirely; the scan short-circuits on
  the first stale summary.

A `phaseCleanCommitTimesMs` seam keeps the unit tests hermetic (no git
spawn); the resolver's two-call error handling is unit-tested via an
injected execGit; two real-git integration tests lock the end-to-end path,
the committed-then-edited (dirty) regression, and the `--` argv guard.
2026-07-17 21:58:02 -04:00
Tom Boucher
f2c077df38 chore(#2387): refactor CONTEXT.md legacy content + add glossary drift gate (#2391)
* chore(#2387): refactor CONTEXT.md legacy content + add glossary drift gate

Apply the audit-and-enforce concept from the ADR index (#2356) to CONTEXT.md:
correct stale facts, and add a CI gate so the machine-verifiable claims can't
silently re-rot.

CONTEXT.md was entirely hand-maintained with nothing checking its claims against
the shipped tree, so it had rotted. An audit against live code (Memtrace +
filesystem + gh), each finding adversarially re-verified, drove 38 factual
corrections + 1 surfaced by the new gate:

- Dead references: Package Identity named @opengsd/get-shit-done-redux (package
  is @opengsd/gsd-core); Shell Command Projection named run-git/run-npm/run-tool
  (real exports execGit/execNpm/execTool); a partial docs/adr/1606 ref; retired
  sdk/ framing.
- Superseded facts: allRuntimes 15 -> 17 (pi #2102, zcode); "seven nested-loader
  runtimes" -> five (claude reverted flat #924, antigravity flat); stacked-PR
  examples rebasing onto main -> next; QUOTA_SENTINELS precedence corrected to
  match src/agent-command-router.cts.
- Drifted CONTRIBUTING.md line citations refreshed.

Per CONTRIBUTING.md:179, only stale FACTS were corrected -- no maintainer intent,
lesson, or opinion was rewritten, and the append-only session log is untouched
except one dated in-place superseding note. The three tests that assert on
CONTEXT.md content (phase6-capstone-conformance, tracer-bullet,
external-job-waiting) keep all their anchors.

New scripts/check-glossary-refs.cjs (--check, wired into lint:generated-sync):
- Check A: every backticked file reference under a TRACKED_PREFIXES allowlist
  resolves on disk. Generated gsd-core/bin/lib/*.cjs (77 refs, gitignored),
  ~/-paths, .planning/, and bare filenames are deliberately skipped so a clean
  CI checkout never false-fails.
- Check B: the allRuntimes count + member set in the glossary prose match
  bin/install.js's allRuntimes literal (drifts on every runtime addition).
tests/check-glossary-refs.test.cjs covers both, including the false-positive
guard that a missing bin/lib/*.cjs ref does NOT trip the gate.

Closes #2387

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2387): confine glossary-gate file refs to ROOT (no `..` traversal)

Pre-PR security review finding (low): extractTrackedRefs fed tokens straight to
fs.existsSync(path.join(ROOT, token)), and PATH_TOKEN_RE admits `.` in a segment,
so a CONTEXT.md token like `src/../../../etc/passwd` passed the `src/` prefix
check and normalized to an out-of-tree absolute path — turning the doc lint into
a filesystem-existence oracle on the CI host (existsSync only; CONTEXT.md is a
trusted committed file, hence low severity, but a defense-in-depth gap).

Add isWithinRoot() confinement in extractTrackedRefs: a token is dropped unless
path.resolve(ROOT, token) stays within ROOT. A CONTEXT.md reference is always a
plain in-repo path, so a `..` escape is never legitimate. Regression test asserts
a `..`-bearing token is skipped and never named in output.

Refs #2387

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2387): drop legacy `get-shit-done` name from a CONTEXT.md defect entry

CI lint-legacy-dir-name failed: the line-928 upstream-issue re-point I applied
wrote the historical provenance as "gsd-build/get-shit-done#3545", and
scripts/lint-legacy-dir-name.cjs forbids the legacy `get-shit-done` name. Reword
to "moved from #3545 in the predecessor repo" — same provenance, no legacy name.

Caught by `npm run lint:ci` (the CI lint chain), which I had not run locally —
lint:generated-sync + eslint do not include lint-legacy-dir-name.

Refs #2387

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 19:24:28 -04:00
Tom Boucher
81f7ab4df1 refactor(#2384): leaf dispatch table + runCommand collapse (ADR-2346 P4) (#2392)
* refactor(#2384): leaf dispatch table + runCommand collapse (ADR-2346 P4)

Cutover all 55 remaining case arms from runCommand's switch to
HOST_COMMAND_ROUTERS. runCommand now contains only its default case
(~40 lines): the three-layer dispatch (capability → overlay → host table)
plus the unknown-command diagnostic. The 73-case switch is dissolved.

Each case body was relocated verbatim to a module-scope route*Command
function via a brace-matching extractor; inner break; statements (from
_dispatchNonFamily early-exit patterns) were converted to return;
(5 arms affected); loop break; statements preserved.

Closes #2384

* chore: retrigger CI
2026-07-17 18:29:28 -04:00
Tom Boucher
d91e32b3ce fix(#2383): untrack node_modules — accidentally committed as a hardcoded absolute-path symlink (#2385)
* fix(#2383): untrack node_modules — accidentally committed as a hardcoded absolute-path symlink

cf004df67 (#2360/#2364) swept node_modules into git as a tracked
120000 (symlink) blob pointing at /Users/trekkie/projects/gsd-core/node_modules
— a path specific to one contributor's machine. .gitignore already
lists node_modules/, so this was almost certainly a broad `git add`
run while node_modules happened to be a symlink at that path, not
intentional (git add on an explicitly-added path isn't blocked by
.gitignore).

Two concrete problems this caused: (1) anyone else cloning the repo,
or any CI runner, checks out a symlink pointing at a path that does
not exist on their machine; (2) it silently self-heals for most
people (npm ci detects the checked-out symlink is "not a directory"
and replaces it), but anyone who runs a tool directly against
node_modules/.bin/* before ever running npm ci hits ENOENT/ELOOP
failures that read as environment corruption and are expensive to
diagnose — exactly what happened while preparing PR #2380 before this
tracked entry was found to be the actual root cause.

git rm --cached only, no working-tree content touched. .gitignore
already covers node_modules/ going forward; confirmed via
`git show cf004df67 --stat` that no other file was swept into that
same commit by the same mistake.

Closes #2383

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2383): gitignore node_modules regardless of file type

node_modules/ (trailing slash) only matches directories, so it never
suppressed the worktree-sharing symlink some worktrees use to point
node_modules back at the main checkout — every such worktree showed a
perpetual, un-ignorable "?? node_modules" in git status, exactly the
noise that trains people to stop reading git status output. Dropping
the trailing slash matches node_modules regardless of whether it's a
real directory, a file, or a symlink, which is what every other repo's
node_modules ignore rule actually needs to do. Found while directly
verifying #2383's untrack fix was complete, not assumed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: regenerate cursor golden-install-parity fixture after rebase

next advanced again during rebase — #2386 (fix #2341, "de-dup Cursor
menu by marking skills user-invocable:false") landed and legitimately
changed every cursor SKILL.md's content. Confirmed via git log that
this is the explanation before committing: all 71 changed hash entries
are isolated to cursor.json, matching a runtime-specific skill-output
change, not noise.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: regenerate cursor golden fixture from a clean clone (worktree was stale)

Local worktree regeneration didn't match CI's clean-room result despite
multiple attempts; a fresh clone + npm ci + regenerate in isolation
produced a different, correct result. Using that.

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 17:56:28 -04:00
Tom Boucher
f15eb5f5c9 fix(#2347): make the decision-shape evidence test format-agnostic (#2389)
#1365's fail-loud guard reused the parser's own D- grammar as its evidence test, so a populated <decisions> block using any other ID prefix (e.g. D5-01) was invisible to both parser and guard, collapsing could-not-parse into a clean none-present pass. Add an ID-shaped bold-lead-in probe as format-agnostic evidence on both parse paths; empty/prose scaffolds stay none-present. Graduates the #2371 d5-prefix representative fixture to its expected* assertion.

Closes #2347. Admin-merged (self-review bypass) with full green CI.
2026-07-17 16:25:13 -04:00
Tom Boucher
062f3fda90 chore(#2371): representative gate-fixture corpus + document-shaped property test (#2380)
* chore(#2371): representative gate-fixture corpus + document-shaped property test

Adds tests/fixtures/representative/ — a permanent corpus of verbatim,
incident-sourced fixtures (never author-invented) from #2286, #2347,
#2365, #2366, each labeled with its expected gate verdict in a
MANIFEST.json and driven through the real CLI gate entrypoint via
tests/representative-corpus.test.cjs.

Adds a document-shaped fast-check property test alongside the existing
writer-seeded bijection test in tests/api-coverage.test.cjs: the existing
generator produces rows and renders them through the writer, so the
document shape is a constant and it cannot fail against a decoy table;
the new one generates the document space instead.

Two gates (#2365, #2347) are still open, so their corpus/property
assertions are marked with node:test's official `todo` option — the test
executes and reports its failure without affecting the process exit code
(https://nodejs.org/api/test.html#test-options). The audit-uat corpus
(#2286, fixed by #2317) is a normal passing assertion, proving the
methodology works end to end and not just cataloguing gaps.

Records the fixture-provenance rule in CONTRIBUTING.md: a gate's fixtures
may not be derived from the gate's own writer, grammar, or docstring
examples; a negative fixture must come from a source that doesn't know
the gate exists.

No production src/*.cts changes — validation only, per #2371's scope.

Closes #2371

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2371): address orthogonal-review findings — dedup generators, fix field naming, wire dead fields

Standards-axis review findings, all fixed:

- Deduplicated the row-shape generators (capabilityGen/rowGen/validRowGen)
  that were copy-pasted between the parse/render bijection test and the
  new document-shaped property test in tests/api-coverage.test.cjs — a
  future edit to one could have silently desynced the two properties.
  Hoisted to a single module-scope declaration both tests reference.

- Renamed decision-coverage-guard/MANIFEST.json's expectedOutcome ->
  expectedReason. It asserted against the gate's `reason` field, but this
  codebase already has a real, different `outcome` field at parser
  altitude (extractDecisions' DecisionOutcome) — naming the manifest
  field after the wrong altitude's term was exactly the ambiguity the
  "Fixture provenance" rule this PR adds exists to eliminate.

- Removed the unused `role` field from three MANIFEST.json files (never
  read by any test) and wired the previously-dead per-fixture
  `expectedMinItems` in audit-uat/MANIFEST.json into a real per-file
  assertion in tests/representative-corpus.test.cjs, using cmdAuditUat's
  `results` array — catches a regression that moves items between the
  two fixture files while preserving the aggregate total, which the
  existing total_items check alone would miss.

No changes to test intent or coverage — same assertions, correctly named
and fully wired.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: remove dead capabilityState/capabilityWriter requires from gsd-tools.cjs

Surfaced by the mandatory pre-PR lint gate (no-unused-vars) while
preparing this PR — unrelated to #2371's own changes, but a defect
found while working is fixed in place rather than deferred.

Leftover from #2368/#2370 (merged just before this branch rebased onto
it): the case 'capability' arm that needed these two requires was
relocated to bin/lib/capability-command-router.cjs, which already
requires both modules directly (lines 24-25) and is their only real
consumer (cmdCapabilityState, resolveCapabilityRuntimeState,
cmdCapabilitySet). The two requires left behind in gsd-tools.cjs had
zero other references in the file and were never re-exported —
confirmed via grep across the file and its module.exports.

Behavior-preserving: Node's require cache means the underlying modules
still load exactly once via capability-command-router.cjs's own
requires; gsd-tools.cjs never used its now-removed local bindings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2371): replace todo-marked assertions with characterization tests

gsd-test's own JSONL result parser (gsd-test-runner's
internal/pipeline/parse.go, verified directly against that repo's
source) has no concept of node:test's `todo` option — it only
recognizes kind:"pass"|"fail" and hard-errors on anything else. A
{ todo: true } test whose body throws is counted as a real failure in
gsd-test's own verdict, exactly as if it weren't marked todo — proven
by an actual gsd-test run against this branch, which reported
outcome:"failed" with all six todo-marked assertions (the property
test plus five representative-corpus fixtures) in the failure list,
each carrying the correct raw node:test `todo` field the tool's parser
simply doesn't read.

Replaces todo with characterization: MANIFEST.json now carries both
the correct target verdict (expected*) and the exact current observed
verdict (currentBuggyOutput, directly verified against live CLI
output for all five fixtures). Tests assert currentBuggyOutput — an
honest, non-vacuous pin of today's known-broken reality that passes
today and will fail loudly the moment the referenced fix changes the
observed output, at which point the assertion should be flipped to
expected* and currentBuggyOutput deleted.

The document-shaped property test switches from throwing fc.assert to
non-throwing fc.check (returns RunDetails per fast-check's own docs)
and asserts report.failed === true directly, for the same reason.

Updates all prose (CONTRIBUTING.md, the fixture READMEs) that
previously claimed todo would be respected — that claim was
factually wrong for this repo's actual tooling and must not ship.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: regenerate golden-install-parity fixtures after rebase onto next

Rebasing onto the current next (which now includes #2381's
todo-severity changes to gsd-core/bin/gsd-tools.cjs) produced real
conflicts in all 18 golden-install-parity fixtures — expected, since
both branches changed the same gsd-tools.cjs hash entry. Resolved by
taking one side to unblock the rebase, then regenerating fresh from
source via npm run gen:golden and verifying the result; every file's
diff is exactly the one hash line for gsd-core/bin/gsd-tools.cjs,
correcting a stale intermediate hash from the arbitrary conflict pick.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 15:23:09 -04:00
Tom Boucher
58028eaf56 fix(#2341): de-dup Cursor / menu by marking skills user-invocable:false (#2386)
Cursor installs both a skills and a commands surface and shows both in '/', duplicating every /gsd-*. Extend the #789 CodeBuddy de-dup to Cursor: convertClaudeCommandToCursorSkill (in both src and the live bin/install.js) now emits user-invocable:false, so the skill stays model-invocable while the commands surface is the single '/' entry point.

Closes #2341. Admin-merged (self-review bypass) with full green CI.
2026-07-17 15:20:19 -04:00
Tom Boucher
e6c16efa6d refactor(#2373): cutover resolve/git/config/research host routers (ADR-2346 P3) (#2382)
* refactor(#2373): cutover resolve/git/config/research host routers (ADR-2346 P3)

Relocate 13 case arms from runCommand's switch to HOST_COMMAND_ROUTERS:
- resolve: resolve-model, resolve-granularity, resolve-execution (3)
- git: git (1)
- config: config-ensure-section, config-set, config-set-model-profile,
  config-get, config-new-project, config-path, migrate-config (7)
- research: research-store, research-plan (2)

Each case body moved verbatim to a module-scope route*Command function
(closures over config/commands/output/_dispatchNonFamily preserved).
dispatchHostCommand extended to pass defaultValue + workstreamContext
(needed by config-get and config-path). No new files, no logic change.

Closes #2373

* chore: retrigger CI
2026-07-17 14:26:35 -04:00
Tom Boucher
b0f672f88c fix(#2337): capture and surface todo severity (#2381)
add-todo.md gains a confirm-based infer_severity step (infer from the blocker/major/minor/cosmetic taxonomy, confirm via AskUserQuestion with TEXT_MODE fallback, before writing) and a severity frontmatter field. cmdListTodos and cmdInitTodos now surface severity, backward-compatible (key omitted when absent), in parity.

Closes #2337. Admin-merged (self-review bypass) with full green CI.
2026-07-17 14:01:38 -04:00
Tom Boucher
dcb4954131 fix(#2335): normalize volta node image paths to the stable shim (#2375)
Adds a volta branch to normalizeNodePath() that rewrites the version-pinned node image path to volta's stable shim, so managed hooks survive a volta node prune (the fnm/Homebrew/mise class, now covered for volta). Also unifies the release-smoke install timeout into one shared 600s constant across before() and runSmoke() so slow benches no longer spuriously time out.

Closes #2335. Admin-merged (self-review bypass) with full green CI.
2026-07-17 12:38:37 -04:00
Tom Boucher
b302f53ee6 refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2) (#2370)
* refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2)

Behavior-preserving relocation of the 706-line case 'capability': arm from
gsd-tools.cjs into a new hand-authored bin/lib/capability-command-router.cjs
(sibling of ensure-runtime-build.cjs). The 15 bin/-relative require paths are
rewritten to sibling-relative (correct for bin/lib/). dispatchHostCommand is
now async (capability's install/upgrade/consent ops await the lifecycle); sync
routers (state/phase/…) pass through await unchanged. case 'capability':
removed; capability dispatches via HOST_COMMAND_ROUTERS.

Validated by the existing capability-lifecycle / -consent / -trust / -loader
test suites (no logic changed). Golden install-parity fixtures regenerated.

Closes #2368 (Slice 1 — relocation). Probe consolidation (capHostVersion→
readHostVersion, capReadStrict dedup) deferred to a follow-up slice.

* fix(#2368): add capabilityState/capabilityWriter requires + INVENTORY row

The relocated capability arm references capabilityState (cmdCapabilityState,
resolveCapabilityRuntimeState) and capabilityWriter (cmdCapabilitySet) — both
module-scope requires in gsd-tools.cjs (L288/289) that the initial closure-dep
scan missed. Added as sibling requires to capability-command-router.cjs. Also
adds the new cli module to docs/INVENTORY.md + regenerates the manifest.

* fix(#2368): correct capHostVersion __dirname depth for bin/lib/ relocation

capHostVersion's VERSION/package.json paths were bin/-relative ('..' and
'..','..'); on relocation to bin/lib/ they resolved one level too deep,
so capHostVersion returned 0.0.0 and capability install failed the
engines.gsd gate (#1920). Added one more '..' to each (now resolves
gsd-core/VERSION and repo-root package.json correctly).

* test(#2368): drop capability from the invocation loop (async/FS vs /fake/cwd)

capability is async and does FS/config reads, so invoking it against the
unit test's /fake/cwd is fragile. The 6 sync Tier-1 routers stay in the
invocation loop; capability is covered by the non-invoking registry-
ownership assertion + the dedicated capability-* test suites.

* chore: retrigger CI (no-changelog label now present)
2026-07-17 11:04:47 -04:00
Tom Boucher
67a9243cf1 chore(#2356): make the ADR index a generated artifact and enforce ADR lifecycle invariants (#2367)
* chore: rebuild ADR index as a generated artifact and enforce lifecycle invariants

The ADR index in docs/adr/README.md was hand-maintained with nothing checking
it, and had drifted to 40 of 65 ADRs. The absent rows included the entire
capability family (857/894/959/1016/1143/1213/1244) and ADR-1239 (EoS) itself,
so the decisions a reader most needed were the ones they could not find.

Make the index a derived artifact, matching the repo's existing generated-file
idiom (lint:generated-sync), and enforce the corpus' lifecycle invariants:

- scripts/gen-adr-index.cjs generates the index between markers and validates
  the status vocabulary (Accepted/Proposed/Superseded/Legacy/Retired),
  successor links, id/filename agreement, and supersession symmetry.
- Wire --check into lint:generated-sync so drift fails CI.

Correct the lifecycle metadata the gate surfaced, without flipping any status:

- ADR-1239 (EoS) declared it subsumed ADR-1016/58/3660/894; none recorded it.
  Add reciprocal "Subsumed by" pointers + dated amendments. Subsumption keeps
  the target Accepted -- these are live adapters, not dead decisions.
- ADR-857/894 carry dated status caveats: they read Proposed while the
  capability system shipped and epic #857 is closed. Ratification is a
  maintainer act and is deliberately left open.
- Link ADR-0005/0007/0012/3524 -> ADR-0174 and ADR-0010 -> ADR-0009; record
  the reciprocal Supersedes on ADR-0009.
- ADR-218 declared itself "ADR-0175" -- an unfinished rename.
- The 0011 PRD moves from the non-canonical "Draft" to "Legacy".

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: capture stderr via spawnSync; record ADR-0010 draft supersession

Two fixes surfaced by the first gsd-test run and by regenerating the index:

- tests/adr-index-gate.test.cjs used execFileSync, which only surfaces stderr
  through the thrown error on non-zero exit. The `--write` path exits 0 while
  reporting outstanding violations on stderr, so the helper always saw ''.
  spawnSync captures both streams on both outcomes.
- The hand-maintained index recorded 0010-skill-surface-budget-module.md as
  "earlier draft superseded by ADR-0011" while the file itself still said
  Proposed. Deriving the index from the files would have dropped that
  assertion and resurrected a superseded draft as a live decision, so it is
  recorded at its source, with the reciprocal Supersedes on ADR-0011.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: drop the dead sdk/ model-catalog candidate retired by ADR-0174

src/model-catalog.cts resolved model-catalog.json through three candidates, the
second being sdk/shared/model-catalog.json three levels up. That was the legacy
source-repo fallback kept by the #3288 fix ("check the co-located path FIRST,
before the legacy source-repo path").

ADR-0174 then retired the @opengsd/gsd-sdk package boundary and deleted the sdk/
tree (11918dcc3), so the candidate can no longer resolve in any layout: a source
repo has no sdk/, and an install layout points it at ~/.claude/sdk/shared/, which
the installer never writes -- the original #3288 bug. It was dead weight implying
a package boundary this repo no longer has.

No test depends on it: the #3288 regression tests in tests/install.test.cjs write
their own synthetic old-path fixture and assert it throws.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: ratify nine shipped ADRs; record why ten others stay Proposed

The corpus carried 19 Proposed ADRs, most describing architecture that had
already shipped. A Proposed label on live architecture tells contributors and
agents the decision is an unbuilt idea -- the capability system and EoS were
both being misread that way.

Audited all 19 against the shipped tree and GitHub. Each candidate flip then had
to survive two independent reviewers instructed to refute it.

Ratified Proposed -> Accepted, each with a dated Ratification section carrying
the verified evidence (file:line, symbols, tests, issue state):

  857  capability system      894  declaration format   1244 capability ecosystem
  1577 injection boundary     1610 size-budget ratchet  1990 existing-code onboarding
  15   cross-AI convergence   22   plan-drift guard     0011 default reviewers

Held ten, each now carrying a "Why this is still Proposed" section naming the
blocker and its unblock condition, so the audit is not repeated:

  2264 its own headline acceptance criterion is unmet in the tree
  230  live branch protection contradicts the decided spec (1 approval, not 2)
  660  the namesake release/<version> re-cut is manual, not automated
  959  issue #2346 is approved and plans its graduation as its own ADR
  1213 the shipped writer's return shape differs from the decided interface
  443  the orchestrator override path has no live caller
  1143 / 1606 each states its own bar for acceptance; neither is met
  612 / 1671 legitimately open

Shipped code proved necessary but not sufficient: eight ADRs had every named
module, symbol, and test present with their epics closed, and still failed the
bar. That lesson is written into README.md's ratification procedure.

Also corrected ADR-857's "Supersedes (generalizes)" to "Subsumes": taken
literally it would have marked two live seams dead -- ADR-0011 (surface.cts:348)
and ADR-58 (runtime-artifact-install-plan.cts:82). Both keep Accepted status and
gain Subsumed-by pointers.

Index: Active 39->48, Proposed 19->10, Superseded/Legacy 7. 65 total.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: harden gen-adr-index against hostile titles and non-ADR filenames (#2356)

Three findings from the pre-PR orthogonal security review, all confirmed:

- An ADR title containing the literal ADR-INDEX:END marker was emitted verbatim
  into its table cell, relocating the splice boundary so the NEXT --write
  spliced against the wrong marker and truncated README.md. Titles now render
  through cellText(), which escapes pipes and angle brackets -- making an HTML
  comment (and any other HTML) unformable from ADR-authored text.
- A docs/adr/*.md without a numeric prefix crashed on match(...)[1] of null.
  Such a file is also invisible to the index -- the very failure this gate
  exists to prevent -- so it is now reported as a naming-convention violation
  naming the file and the fix.
- Tests leaked their mkdtemp dirs. They now use helpers.createTempDir/cleanup
  via t.after(); helpers.cleanup carries the Windows-EBUSY retry budget that a
  raw fs.rmSync lacks (caught by local/no-raw-rmsync-in-tests).

Adds five regression tests: marker hijack, HTML injection, pipe cell-break,
non-conforming filename, and splice stability across repeated writes.

Refs #2356

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: close two gate false-passes; read ## Supersedes sections (#2356)

Second round of confirmed findings from the pre-PR orthogonal code review. Both
false-passes matter more than a false-fail: a gate that silently misses a
violation is worse than no gate, because it is trusted.

- A relation field mixing a link with a bare id silently dropped the bare claim:
  the check tested `rel.links.length` (does this field have ANY link?) instead
  of whether THAT id was linked. `Supersedes: [ADR-0001](...), ADR-0011` passed
  clean -- accepting exactly the ambiguous bare reference the rule forbids. Now
  each bare id is checked against the ids actually linked in the same field, so
  a repeat in trailing prose stays quiet while an unlinked claim is flagged.
- The ratification guard (`statusToken !== 'Accepted'`) skipped BOTH relation
  directions, which killed the IN check entirely: `supersedes.in` is only ever
  populated on an ADR whose status IS `Superseded`, so a dangling `Superseded by
  X` where X never claims it always passed. The guard now applies to OUT only --
  a prospective claim must not obligate its target, but an ADR's statement about
  ITSELF is always owed a reciprocal.
- Fixing that surfaced a parser gap: ADR-0174 declares its supersessions in a
  `## Supersedes` table SECTION, not a header field, and headerBlock() stops at
  the first `##`. The repo's best-documented supersession was invisible. Section
  form is now parsed for both relations.
- Replaced a vacuous test: the em-dash negation case passed whether or not
  NEGATED_RELATION_RE matched (a mutation to /$^/ survived). It now carries a
  link that would create a failing asymmetric relation if negation did not fire.

Also removes docs/adr/9401-test-target.md -- a synthetic fixture a reviewer
created in the worktree while reproducing a finding, swept in by `git add -A`.

Adds regression tests for each: mixed link+bare, linked-and-repeated-in-prose,
dangling superseded-by from a non-Accepted ADR, and the ADR-0174 section shape.

Refs #2356

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: escape backslashes before pipes in the ADR index cell renderer (#2356)

CodeQL js/incomplete-sanitization (high) on scripts/gen-adr-index.cjs: cellText()
escaped `|` -> `\|` without first escaping the backslash. Markdown's escape
character is the backslash, so the input `\|` became `\\|`, which renders as a
literal backslash followed by an UNESCAPED pipe -- re-opening the cell break the
pipe escape exists to prevent. Order is load-bearing: escape the escape
character first, then everything that emits one.

Same class as the index-marker hijack fixed earlier: ADR-authored text breaking
out of the cell it is rendered into.

Adds a regression test asserting a `\|`-bearing title leaves exactly the row's
own 5 unescaped delimiters and cannot forge a Status cell. Uses split(/\r?\n/)
per local/no-crlf-fragile-split -- a literal "\n" split is CRLF-fragile on the
Windows CI leg.

Refs #2356

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 10:51:58 -04:00
Tom Boucher
cf004df678 refactor(#2360): host dispatch table + state cutover pilot (ADR-2346 P1) (#2364)
* refactor(#2360): host dispatch table + state cutover pilot (ADR-2346 P1)

Pilot cutover for ADR-2346 Phase 1 (epic #2345). Introduces the Layer-2 host
dispatch table — dispatchHostCommand + HOST_COMMAND_ROUTERS, consulted in
runCommand's default case after capability/overlay dispatch, before the
unknown-command error. Migrates 'state' as the pilot: removes the hardcoded
case 'state': arm; state now dispatches default -> dispatchHostCommand ->
routeStateCommand, byte-identical to the old path (proven by the new
state-command-cutover equivalence test, 5-category template).

Host commands are NOT capabilities (core, non-toggleable, no tier/activationKey)
— the capability registry stays reserved for toggleable feature bundles per
ADR-959. This is the host-vs-capability distinction the merged ADR-2346 lacked;
the ADR is corrected here alongside the code that realizes it.

- gsd-core/bin/gsd-tools.cjs: HOST_COMMAND_ROUTERS + dispatchHostCommand
  (prototype-pollution-safe); wired into default case; case 'state': removed;
  dispatchHostCommand + HOST_COMMAND_ROUTERS exported for tests.
- tests/state-command-cutover.test.cjs: UNIT/DISPATCH/BEHAVIOR/REGISTRY
  equivalence (recording-mock + runGsdTools end-to-end + pollution guard).
- docs/adr/2346-*.md: refine Decision 1/2 to the host-table vs capability-
  registry model (correction that did not land in the merged #2355).

Behavior-preserving. Subsequent P1b/c PRs migrate phase/init/roadmap/validate/
verify using this proven template.

Closes #2360.

* test(#2360): regenerate golden fixtures + allowlist for state cutover

Bookkeeping for the gsd-tools.cjs change: npm run gen:golden regenerates the
install-parity fixtures (gsd-tools.cjs content hash changed), and the new
tests/state-command-cutover.test.cjs is added to the lint-test-file-count
allowlist under the 'state' prefix.

* refactor(#2360): migrate remaining Tier-1 routers (phase/init/roadmap/validate/verify)

Completes P1: all 6 Tier-1 host routers now dispatch via HOST_COMMAND_ROUTERS
(state landed in the pilot commit). init preserves its #1688 warnIfStaleBake
pre-hook; validate binds the output emitter. Cutover test extended to assert
all 6 are consumed + owned. Golden install-parity fixtures regenerated.
2026-07-17 09:19:28 -04:00
Tom Boucher
f1a91072b6 docs(#2357): fix Registry Discussions category name and state its format (#2361)
The submission process told an admin to create a `Registry` Discussions
category and never stated its format. Both were wrong in a way a correct
reading of the docs could not catch.

Name: the category is `EoS Registry`, and it carries threads for both
registries — `discussion` is a required field on Capability entries as well as
EoS entries. A reader following the old text would create a second, duplicate
category.

Format: `discussion` being required means the thread must exist before the
entry's PR, opened by the entry's author — an outside contributor holding
neither `maintain` nor `admin`. GitHub's Announcement format restricts starting
discussions to those two levels, so an admin could pick it, block every external
submission at the first step, and see nothing wrong. Records open-ended as the
required format, and why Announcement and Question/Answer do not work.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 08:20:02 -04:00
Tom Boucher
ed06b6a4b9 fix(#2329): write opencode slash commands to commands/ (plural), migrate legacy command/ (#2354)
* test(#2329): fail-first tests for opencode commands/ (plural) command dir

Red phase, empirically probed: global/local install lands in command/ (singular)
with 71 gsd-*.md files and no commands/; the manifest records 71 keys under
command/ and zero under commands/; all four declaring sites report 'command'.
Migration coverage is black-box (two sequential install runs against one
configDir) so it holds regardless of how the fix implements cleanup.

The Kilo guard passes today by design — a forward-looking no-collateral check.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2329): write opencode commands to commands/ (plural), migrate legacy command/

OpenCode discovers slash commands from commands/ (plural); the installer wrote
them to command/ (singular), so none of the ~71 /gsd-* commands appeared in the
TUI. Five sites declared the directory and all had to agree:

- capabilities/opencode/capability.json: both artifactLayout destSubpath entries
  (global + local) and hostBehaviors.flatCommandDir
- bin/install.js: the manifest prefix was a SEPARATE hardcoded 'command/' literal,
  so the manifest would have diverged from the descriptor even after a rename. It
  now derives from _hostBehaviors(runtime).flatCommandDir.
- src/install-engine.cts installOpencodeFamilyArtifacts: the actual write target,
  which bypasses resolveRuntimeArtifactLayout via combinedFamilyInstall. This was
  a fifth site the issue did not list — without it the descriptor change alone
  would not have moved a single file.

Migration: an upgrade over a pre-fix install removes only manifest-proven
GSD-managed files from the legacy command/ dir and rmdirs it once empty.
Unmanifested user files are preserved, never deleted.

Kilo shares the opencode family install path and is explicitly unaffected —
pinned by a no-collateral test.

Note on the tests: the migration cases originally built their legacy fixture by
running the installer and relying on it to produce command/ — i.e. they depended
on the bug to set up the fixture, and became unsatisfiable the moment it was
fixed (block 1 requires command/ to be absent after a fresh install). They now
fabricate the legacy layout explicitly, including rewriting the manifest keys to
the command/ prefix — which is load-bearing, since the migration only removes
manifest-proven files and an unrewritten fixture would silently no-op and pass
even against a broken migration.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* test(#2329): regenerate opencode install golden after rebase onto next

The golden conflicted on rebase because #2322 also regenerated it. Resolved by
regenerating from the merged source rather than hand-merging a generated file;
the only delta is the 71 command/gsd-*.md -> commands/gsd-*.md key renames.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* test(#2329): update stale tests that pinned opencode's singular command/ dir

Seven tests encoded the old contract (opencode: command/gsd-help.md exists, the
descriptor's flatCommandDir, the install-integration contract, and the
resolveRuntimeArtifactLayout golden). They passed in the red phase precisely
because they pinned the buggy singular dir; the fix intentionally changes that
contract, so these are stale-test corrections, not regressions.

Kilo shares the opencode family install path and is deliberately NOT changing —
it stays on command/ (singular). The shared opencode/kilo test is now split via
an explicit per-runtime dir map so the two cannot be conflated, and Kilo's own
layout test is untouched. tests/opencode-command-dir-plural.test.cjs
independently pins Kilo unchanged end-to-end.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2329): changeset for opencode commands/ dir fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2329): backfill PR number 2354 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2329): correct the changeset — do not assert opencode ignores command/

The changeset repeated the issue's stated mechanism ("OpenCode discovers them
from commands/ ... a clean install produced no usable commands in the TUI at
all"). OpenCode's source contradicts that: packages/core/src/v1/config/command.ts
globs {command,commands}/**/*.md, so BOTH names resolve, and its own skill doc
still calls .opencode/command/ typical. Shipping that claim as a release note
would document a mechanism that does not exist.

The change is still right, for the stronger reason: OpenCode's config docs list
plural as the convention and singular as backwards compatibility, so GSD was
shipping on the alias the vendor may withdraw. Reworded to describe it as the
alignment it is, decided on OpenCode's source and docs rather than on bug reports
in either repo.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2329): baseline opencode's commands/ surface — closes a data-loss path this PR opened

Not a bookkeeping gap. Moving opencode's command dir to commands/ moved the
install destination to a surface the first-time baseline scan does not cover:
000-first-time-baseline's RUNTIME_SURFACES.opencode lists ['gsd-core','command',
'skills','agents'] — no 'commands'.

installOpencodeFamilyCommands unconditionally unlinks every gsd-*.md under its
destination before writing the fresh set (install-engine.cts:870-873), with zero
manifest or migration involvement. The only thing that protects a pre-existing
file is assertInstallerMigrationsUnblocked, which runs before materialization and
halts when the baseline scan flags an unknown file at a KNOWN surface.

Probed: a pre-existing commands/gsd-plan.md is silently destroyed (install exits
0). The identical file under the legacy, already-baselined command/ surface
correctly halts the install with "installer migration blocked pending user
choice". So this PR would have traded a protected surface for an unprotected one.

Fixed with a NEW fix-forward migration rather than editing 000, per
docs/installer-migrations.md:131-134 — an applied migration never re-runs, so
editing 000 would only protect fresh installs and leave every existing machine
exposed. A new id runs for both populations and drifts no shipped checksum;
adding its entry to EXPECTED_CHECKSUMS is the case that test explicitly sanctions.
All five pre-existing shipped checksums verified byte-identical.

Kilo is excluded by the migration's runtimes filter and keeps command/.

This was previously deferred as a PR-body note claiming "low impact — nothing
else acts on baseline-scan misses". That claim was never probed and was wrong.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2329): drop the parenthetical product description from the changeset

The product-name purity guard (#1777) rejects "Kilo (which still uses
command/)" — fragment prose renders verbatim into CHANGELOG.md, so a product
name must not carry a parenthetical. Reworded to a plain sentence; the meaning
is unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 08:14:59 -04:00
Tom Boucher
15b3cc8690 docs(#2346): Command Dispatch Completion ADR + graduate ADR-959 to Accepted (#2355)
Records the decision (ADR-2346) to dissolve runCommand's 73-case switch into a
two-layer dispatch (registry families + leaf-verb table filling the prepared
_dispatchNonFamily seam), collapsing it to ~15 lines. Covers the four decisions
ADR-959 leaves open: full dissolution, family/leaf classification rule, shared
parseFamilyArgs, and the capability-arm extraction shape. Phased under epic
#2345 (P1-P4). Behavior-preserving; each cutover proven by the
audit-command-cutover equivalence template.

- docs/adr/2346-command-dispatch-completion.md (new)
- docs/adr/959-*.md: Status Proposed -> Accepted + amendment section
- docs/adr/README.md: index rows for 959 + 2346
- docs/ARCHITECTURE.md: forward-reference note under Command Routing Hub
- CONTEXT.md: seed glossary entry

Closes #2346 (docs-only; no production code).
2026-07-17 07:19:17 -04:00
Tom Boucher
23a65c4a3d fix(#2322): materialize installed third-party capability skills (#2340)
* test(#2322): fail-first tests for third-party capability skill materialization

Red phase: tests (1) and (6) fail — resolveSurface reports the third-party stem
surfaced (#2045) but no SKILL.md is ever written to disk. The other four are
controls that must keep holding: first-party-wins collision, profile-tier filter,
nested-router layout unperturbed, and absent/malformed capability must not throw.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2322): materialize installed third-party capability skills

A capability could report installed:true, surfaced:true, active:true and still
never exist as an invocable command. #2045 fixed the registry layer —
resolveSurface unions registry.capabilityClusters into the resolved skill set —
but the materialization layer never got the matching fix.
stageSkillsForRuntimeAsSkills only ever read gsd-core's own bundled
commands/gsd/*.md and silently skipped any stem it couldn't find there, so a
third-party skill living at <GSD_HOME>/.gsd/capabilities/<id>/skills/<stem>/
was never copied. Registry said surfaced; disk had nothing.

Installed capability skills are now staged alongside the first-party ones, copied
verbatim (they are authored complete for their target runtime and need no
converter). First-party stems always win a collision, the profile filter still
applies, and an absent or malformed capability degrades rather than throwing.

Security: capability.json's skills[] entries are validated only as non-empty
non-reserved strings (capability-validator.cjs:503-514) — no path shape is
enforced upstream — so stems are sanitized (rejecting separators, '..', absolute
paths, NUL) with an independent isPathConfined check on both the read and write
paths. A '../../evil' stem writes nothing outside the capability's own dir.

Also fixes a defect this surfaced in pruneSkillDirs: a materialized capability
skill dir has no first-party manifest entry, so every apply logged
"preserving (user-owned or unknown)" for a live GSD-managed dir. The retained
check now precedes the manifest gate; no deletion outcome changes, and genuinely
unknown gsd-* dirs still warn and are preserved.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2322): address security review — bind skills to declaring capability, fix full profile

An independent security review BLOCKED the first pass. Both blockers were mine.

BLOCKER 1 (security): readInstalledCapabilitySkill scanned every capability dir
and returned the first sorted match, never checking that a capability DECLARES
the stem — ownership was inferred from attacker-controlled filesystem layout.
Since install copies the whole bundle and the validator only checks DECLARED
entries, a capability declaring `skills: []` could ship an undeclared
skills/deploy/SKILL.md and win the `deploy` stem on sort order, supplying the
agent-invocable instructions the user believed came from the registered
capability. Stems are now bound to their owning capId via
registry.capabilityClusters, and only that capability's dir is read.

BLOCKER 2: the fill-in pass was gated `skills !== '*'` on the premise that
applySurface materializes `full` into a concrete Set. True for applySurface —
false for the installer, which is the default path: resolveProfile returns the
'*' sentinel and bin/install.js passes it straight to staging. So #2322 survived
on the default `full` profile, i.e. the fix didn't fix the reported bug. The
registry is now plumbed to staging, and '*' stages all capability-cluster stems.
Wiring this surfaced a second gap: the ADR-1239 imperative adapter (the primary
install path) never threaded its registry either, which would have silently
defeated the fix on the real default install.

HIGH: staged capability skills were never prunable — pruneSkillDirs gates on the
first-party manifest, so uninstalling a capability left its instructions live in
the agent's context forever. Staged skills now carry a marker making them
GSD-owned and prunable; genuinely unknown gsd-* dirs still warn and are preserved.

MEDIUM: the "staged verbatim" claim was false — applySurface rewrites bodies over
the whole stage dir. The tests asserted byte-equality and passed only because
their fixtures contained no rewrite triggers. Claim dropped; tests now assert the
rewrite against triggering content.

LOW: isPathConfined is lexical, not realpath (symlink-defeatable, currently
unreachable because install rejects symlinks) — comment corrected. The validator
does not enforce non-empty, so isSafeCapabilitySkillStem is the sole defense, not
a second layer — comment corrected and it now has traversal/NUL/absolute/empty
test coverage (previously mutating it to `return true` left every test green).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* test(#2322): pin that the imperative adapter forwards a capability registry

The delegation-args test deep-equalled the exact argv to
installRuntimeArtifacts, so threading the composed capability registry through
the ADR-1239 imperative adapter (required for #2322 — without it the default
`full` install path never materializes third-party capability skills) failed it.

The contract legitimately gained a parameter, so this is a stale-test
correction, not a regression. Rather than deep-equalling the whole composed
registry (brittle — it embeds the full agent/profile map), the test pins the
leading args exactly and asserts only that a registry-shaped value is forwarded.
That still fails if the adapter stops threading it, which is the regression the
test exists to catch.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2322): backfill PR number 2340 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 06:50:06 -04:00
Tom Boucher
19c1f54a2a fix(#2316): stop phase complete silently dropping ghost requirement IDs (#2339)
* test(#2316): fail-first tests for ghost REQ-IDs, v-heading over-match, all-orphan gap check

Red phase: 4 of 10 fail against current source (#2316-1 ghost-ID warning,
-3 requirements_updated honesty, -4a v1-heading suppression, -6b all-orphan gap
rows). The other 6 are controls/boundaries that must keep passing — including the
#1159 deferred-heading guard and the literal "TBD" placeholder boundary.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2316): stop phase complete silently dropping ghost requirement IDs

phase complete parses a phase's `**Requirements**:` line from ROADMAP.md and
reconciles it into REQUIREMENTS.md. When a cited ID was registered nowhere, every
branch degraded to a no-op and the report was indistinguishable from a run that
applied every update: `requirements_updated: true, warnings: [], has_warnings:
false`, file byte-for-byte unchanged.

Four defects on that path, all long-standing (traced to 2bc295b32d, v1.2.0):

1. The only cross-check compared REQUIREMENTS.md's own body against its own
   Traceability table, so an ID cited by ROADMAP but defined in neither was
   invisible to it. `citedReqIds` is now hoisted out of the `if (reqMatch)` block
   and cross-checked; ghost IDs raise a warning through the existing
   warnings/has_warnings surface that execute-phase.md already prints.
2. `if (reqUpdate.ok)` had no `else`, so a Traceability-row write matching
   nothing was discarded silently. Misses are now recorded.
3. `requirements_updated` was set unconditionally inside `if (existsSync)`,
   reporting "the file was in the transaction" rather than "a write landed". It
   now reflects whether the content actually changed.
4. DEFERRED_HEADING_RE's bare `v\d+` alternative treated an active `## v1
   Requirements` heading as deferred, zeroing the body scan for that section.
   Dropped; `deferred|backlog|future` still match, so #1159's suite still holds.

Also fixes the same class in gap-checker: the `items.length === 0` early return
fired before ghost rows were folded in, so an all-unregistered phase reported
LESS than a partially-unregistered one. It is now gated on ghostReqIds too.

Follows the #2140 precedent in milestone.cts, which hardened this exact class for
`requirements mark-complete`; adapted to this function's single `warnings` array
rather than adding fields execute-phase.md never reads.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2316): address review — milestone-aware deferred headings, probe-based ghost check

Independent review found the first pass green but wrong — the tests were built to
miss the shape that breaks.

BLOCKER: dropping the `v\d+` alternative from DEFERRED_HEADING_RE regressed #1159
against GSD's OWN shipped template. gsd-core/templates/requirements.md:35 ships
`## v2 Requirements` with the deferred-ness in the BODY PROSE ("Deferred to future
release"), not the heading — so `v\d+` was the only alternative matching it.
Deleting it made every scaffolded project emit a false traceability warning on
every phase complete, forever. The real problem is that `v\d+` matched BOTH the
active v1 and the deferred v2: deleting it swings from suppress-everything to
suppress-nothing. The heading check is now milestone-aware — a `## v<N>` heading
is deferred only when <N> is not the current milestone's major, resolved through
the existing state.cjs seam. Unresolvable milestone fails safe to the old
always-deferred behavior, since suppressing a warning beats spamming every project.

HIGH: the ghost check re-derived membership from two lossy index sets that
disagreed with the case-insensitive write paths, so it warned "not registered
anywhere" about an ID whose checkbox it had just ticked, and about a
case-mismatched ID whose write landed. It now probes the same two write surfaces
the writes use, mirroring milestone.cts's #2140 precedent (probe the write, don't
re-derive it) — which is what the original brief asked for and the first pass
didn't do.

HIGH: the cited-ID tokenizer split the rest of the line on whitespace, so every
trailing word became a "cited REQ-ID". The shipped roadmap.md template puts an
inline HTML comment on that exact line, so a fully correct run told users to
register `<!--` and `-->` as requirements. Cited IDs are now filtered to the
REQ-ID shape, matching what bodyReqIds/tableReqIds already require.

HIGH: gap-checker had the identical early-return defect 34 lines above the one
fixed — the could-not-parse branch gated ghost rows behind the same unguarded
items.length. One malformed <decisions> line made all ghost rows vanish.

Tests: the #2316-5 guard only covered `## Deferred`/`## Backlog`/`## Future` —
headings that trivially still match — omitting `## v2 Requirements`, the only
shape the change altered. Added the shipped-template fixture plus guards for each
finding above.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2316): backfill PR number 2339 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 06:49:58 -04:00
Tom Boucher
ada79bee97 fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md (#2338)
* fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md

Step 4 rewrote the `## Current Milestone` heading in the shared root PROJECT.md
unconditionally. references/workstream-flag.md marks PROJECT.md `# Shared`, and
per-workstream milestone state already lives in the workstream's own STATE.md /
ROADMAP.md / REQUIREMENTS.md. With parallel milestones — the sanctioned design —
whichever workstream ran new-milestone last silently won the shared heading.
Step 4 is now skipped when a workstream is active; step 6 no longer stages
PROJECT.md in that mode (cmdCommit returns nothing_to_commit rather than failing
when a staged path is unchanged).

Also fixes a second defect found while diagnosing this, same root cause (the
workflow was workstream-unaware): step 1 parsed only --reset-phase-numbers and
the milestone name, so GSD_WS was never set — yet ${GSD_WS} was interpolated at
the routing lines. It always expanded to empty, so `/gsd:new-milestone --ws x`
suggested `/gsd:discuss-phase [N]` with the workstream scope silently dropped,
violating the routing-propagation contract. Step 1 now parses --ws using the
established idiom from verify-work.md.

Guard is keyed on GSD_WS, not $GSD_WORKSTREAM: the runtime launcher does not
export the latter and it is only priority 2 of 5 in resolution, so it would miss
the --ws flag case that is the actual repro.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* test(#2308): regenerate install goldens for the new-milestone workflow change

gsd-core/workflows/ ships as an installed artifact, so new-milestone.md's content
hash is pinned in all 18 runtime golden fixtures. Only that hash changed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2308): address review — inert step-6 guard, dropped Evolution repair, tautological tests

Independent review found the first pass was partly cosmetic:

1. The step-6 `if [ -n "$GSD_WS" ]` branch was INERT. GSD_WS is assigned in
   step 1's shell and each step's bash block runs in its own shell — this file
   already proves it, since step 5 round-trips OUTGOING_MILESTONE through a file
   for exactly that reason (#2288). The guard read an unset variable, always took
   the flat branch, and staged PROJECT.md anyway. Rather than re-deriving GSD_WS
   in step 6, the branch is removed entirely: step 4 Part A's guard is what
   protects the shared heading, so post-guard the only content PROJECT.md can
   carry is Part B's idempotent Evolution backfill — which must be staged, not
   stranded. A regression test now asserts no cross-step GSD_WS branch returns.

2. Skipping ALL of step 4 also dropped the `## Evolution` structural repair — a
   shared, idempotent backfill that is not workstream state. A pre-Evolution
   project running only `--ws` would never get the section that transition and
   complete-milestone expect. Step 4 is now split: Part A (milestone-state write)
   is workstream-guarded; Part B (Evolution) always runs.

3. The tests were tautological prose-pinning — including one asserting a comment
   mentions "#2308". The step-6 test asserted the guard's TEXT was present, so it
   passed on the inert guard it existed to catch. Replaced with executable tests
   that extract the step-1 and step-6 fences and run them under bash with stubbed
   gsd_run, asserting real parse and --files behavior.

4. --ws is now stripped from the milestone name (step 1 previously left
   "--ws search" in the remaining text), and documented in argument-hint,
   help/modes/full.md, and docs/COMMANDS.md.

5. Changeset no longer overstates: --ws reaches the prose guard and routing hints
   only, not the SDK calls (state.milestone-switch/phases.clear/init.new-milestone
   still take no ${GSD_WS} — out of scope here).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* chore(#2308): regenerate SKILL.md, goldens, and size baseline for the argument-hint change

skills/gsd-new-milestone/SKILL.md is generated from commands/gsd/new-milestone.md,
so documenting --ws in the argument-hint made it stale (caught by lint:ci's
gen-plugin-skills --check). Regenerated it plus the install goldens and workflow
size baseline, since commands/, skills/, and gsd-core/workflows/ all ship.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2308): backfill PR number 2338 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 06:48:13 -04:00
Tom Boucher
b2961c3f69 fix(#2070): accept adaptive model_profile in validate health; warn on invalid models tiers (W022) (#2336)
* test(#2070): fail-first tests for adaptive model_profile and models tier validation

Encodes the three acceptance criteria from #2070 plus the boundary cases the
resolver silently ignores today (non-string values, empty string, mistyped
phase-type key), and pins VALID_TIERS to a catalog-derived set.

Red phase: these fail against current src/ by design.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2070): accept adaptive model_profile in validate health; warn on invalid models tiers (W022)

W004 sourced its profile list from a hand-maintained literal that predated the
adaptive profile, so `"model_profile": "adaptive"` was false-flagged. It now
reads VALID_PROFILES, which model-catalog.cts derives from model-catalog.json.

models.<phase_type> was validated nowhere: the resolver's tier gate silently
drops unknown values, so a typo like `"planning": "opuss"` was an undiagnosable
no-op. A new W022 flags unknown phase-type keys and invalid tier values
(including non-string values, which the same gate also drops).

VALID_TIERS moves from a function-local literal in model-resolver.cts to a
catalog-derived export, so health and the resolver cannot disagree by
construction rather than by parity test. Object.values(adaptiveTierMap) is
['opus','sonnet','haiku'] plus 'inherit' — identical to the previous literal,
so resolution behavior is unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2070): changeset for validate health adaptive profile + W022

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2070): close review findings — malformed models, tier-list duplication, changeset gate

Review of the initial fix surfaced three real defects, folded in per the
no-defer rule:

1. verify.cts: the W022 guard skipped a top-level `models` that is present but
   not a plain object (`[]`, `"opus"`, `5`, `true`). The resolver ignores those
   identically, so they were the same undiagnosable no-op #2070 targets — just
   one level up. They now warn; absent/null/{} stay silent.

2. config-loader.cts: RUNTIME_OVERRIDE_TIERS was a second hardcoded copy of the
   tier vocabulary this change had just de-hardcoded elsewhere. It now derives
   from the catalog via ADAPTIVE_TIER_VALUES (no 'inherit' — runtime overrides
   resolve to a concrete tier). Byte-equivalent to the old literal.

3. scripts/changeset/lint.cjs: USER_FACING_PREFIXES omitted `src/`. Post-ADR-457
   the product source is src/*.cts compiled to a gitignored gsd-core/bin/lib,
   so the `gsd-core/` prefix is dead coverage for library code and a src/-only
   PR could merge with no release note — including this one. Adding `src/`
   closes the gate; tests/ stays non-user-facing.

Also corrects a false docstring in the VALID_TIERS test: value-equality cannot
detect a re-hardcoded literal, so the test no longer claims it does.

Two review findings were rejected with evidence rather than actioned:
- W021 double-allocation is governed by ADR-612 ("W021 renumber -> void ...
  kept, message-disambiguated"), not a defect.
- Global-defaults validation would be a false-positive generator: config-loader
  reads ~/.gsd/defaults.json only on the "no .planning/" branch, and health
  early-returns E001 without .planning/, so those values provably never affect
  resolution in any context health can run.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* test(#2070): regenerate install goldens for the changeset-lint change

scripts/ ships as an installed artifact, so scripts/changeset/lint.cjs's content
hash is pinned in all 18 runtime golden fixtures. Adding 'src/' to
USER_FACING_PREFIXES changed that hash and tripped every golden parity check.
Regenerated via `npm run gen:golden`; the only delta is the lint.cjs hash.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2070): backfill PR number 2336 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 06:48:04 -04:00
Tom Boucher
a30fb75b51 fix(#2068): dynamic routing escalates the model per --attempt (#2334)
* fix(#2068): dynamic routing escalates the model per --attempt, not just effort

cmdResolveExecution resolved the model via resolveModelInternal (which ignores
dynamic_routing), so retries escalated effort but the model stayed pinned to the
default tier. Resolve the model via resolveModelForTier when --attempt is given,
gated identically to the effort resolution so model and effort stay symmetric —
an omitted --attempt keeps the classic profile model (unchanged for everyone,
incl. dynamic_routing users who don't pass --attempt). Escalation is capped at
max_escalations. Falls back to resolveModelInternal when dynamic_routing is off.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#2068): backfill PR number 2334 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 15:25:19 -04:00
Tom Boucher
1794acb255 chore(#2331): trigger PR-policy workflows on pull_request_target so fork PRs get the verdict (#2333)
* chore(#2331): trigger PR-policy workflows on pull_request_target

Three PR-policy workflows (pr-title-validator, pr-target-validator,
require-issue-link) triggered on plain `pull_request`, so a fork PR's
GITHUB_TOKEN was downgraded to read-only regardless of the declared
`permissions:`. Each one comments on the PR and THEN emits its verdict, so the
createComment 403 killed the github-script step before core.setFailed ran: the
contributor saw an API stack trace instead of the instructions the comment
exists to deliver. Confirmed on PR #2084 (job 86573823878), whose title has
been non-compliant since 2026-07-08 while the explanatory comment 403'd on
every run.

Switches all three to pull_request_target (base-repo context, write-capable
token), matching the three siblings that already do this correctly
(pr-template-format, close-draft-prs, auto-close-unsolicited-prs). Safe: the
only checkouts are BASE-branch with persist-credentials: false, and every
PR-controlled input is read as data — no head code executes. Also wraps each
comment in try/catch so a comment failure can never again suppress the verdict.

Extends tests/workflow-maintainer-skip.test.cjs with the trigger lock already
applied to close-draft-prs.yml (:32-42) for this same defect class.

Closes #2331

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2331): strip backticks before echoing untrusted text into bot comments

Found by the orthogonal security review of this change.

pr-title-validator and pr-target-validator echo attacker-controlled text (the PR
title; the fork's branch name) into an inline-code span in a comment posted by
github-actions[bot]. A single backtick closes the span early and the remainder
renders as live Markdown — GFM autolinks a bare URL — so a fork author could
make our own bot post an arbitrary clickable link into a PR thread, borrowing
the bot's credibility for phishing.

This interpolation is unchanged from next, but it was NOT previously reachable
from forks: the createComment call 403'd and the comment was never posted. The
trigger switch in the parent commit is what makes it reachable by untrusted
authors for the first time, using the write token it grants — so it is in scope
here and fixed here rather than deferred.

A PR title has no charset restriction, so that vector is fully exploitable. The
branch-name vector is weaker (check-ref-format forbids space, ':', '[' and '*',
so no bare URL, link or emphasis is expressible) but is the same class and is
stripped identically rather than left to the charset to police. Stripping the
backtick is complete: it is the only character that can break out of an
inline-code span. Only the rendered body needs this — core.warning/setFailed go
to the job log, where @actions/core already escapes workflow commands.

Also strengthens the try/catch test to assert core.setFailed sits AFTER the
catch block rather than merely existing, so moving the verdict inside the try
(the exact inversion #2331 fixes) fails the test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2331): assert verdict ordering on code, not on comment prose

The first cut of the verdict-ordering guard failed against correct code. It used
indexOf('core.setFailed') on raw source, and these workflows name core.setFailed
in their own comments while explaining the bug — at lines 29/118/147, 16 and 6,
all BEFORE the catch block. So the assertion compared a comment to the call and
reported the inversion it was written to catch. gsd-test caught it: 4 unique
failures across linux-node22/24.

The code was right; the test was measuring the wrong text. Fixes:

- readWorkflowCode() strips whole-line YAML/JS comments so positional
  assertions see only executable text.
- The ordering check is extracted to verdictSurvivesCommentFailure() and
  exercised against BOTH a good and an inverted sample, so the guard is proven
  non-vacuous rather than merely passing.
- A test pins the trap itself: raw source really does mention core.setFailed
  before the catch, while the stripped view puts the real call after it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2331): drop the unnecessary permission widening; the trigger was the whole bug

Both orthogonal review passes flagged the permissions block, from opposite
directions — one said issues:write was dead surface on require-issue-link, the
other said it was the load-bearing scope the two validators lacked. Neither is
right, and the repo's own history settles it:

- pr-title-validator declares pull-requests:write ONLY, and its sticky comment
  has posted 26 times.
- pr-target-validator declares pull-requests:write ONLY — posted 8 times.
- require-issue-link declares issues:write ONLY — posted on same-repo PRs
  #106, #164, #232, #259.

So GitHub accepts EITHER scope for issues.createComment when the target is a
PR, and all three files already declared a sufficient one. The 403 was purely
the fork token downgrade. My added scopes fixed nothing and widened privilege
on precisely the workflows now running as pull_request_target — the context
where surplus scope matters most. Reverted: permissions are byte-identical to
next, and the diff is now trigger + try/catch + sanitizer only.

The permission test previously used an (issues|pull-requests) alternation, so
it passed on the pre-fix tree and would not have caught removal of the scope
that matters. It now asserts each file's SPECIFIC scope and, more usefully,
asserts the absence of the other — locking the least-privilege property against
a future 'add it to be safe' regression. It is a forward lock, not a #2331
fails-first test; the trigger assertion is the fails-first one.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 13:55:15 -04:00
Tom Boucher
9ad2bab4be fix(#2297): scope resolve_model_ids:"omit" to the resolving runtime (#2332)
* fix(#2297): scope resolve_model_ids:"omit" to the resolving runtime

The installer writes resolve_model_ids:"omit" for non-alias runtimes into the
machine-wide ~/.gsd/defaults.json (#1156); any runtime read it back, so install
order silently flipped Claude's adaptive tier aliases (executor->sonnet,
planner->opus) to '' in no-project sessions.

Resolution is now scoped to the runtime actually resolving, identified by a new
per-install <install>/gsd-core/.gsd-runtime marker (installer writes it beside
VERSION). The "omit" branch returns '' only when the PROJECT explicitly set omit
(honored for all runtimes, #2517 finding #4) OR the active runtime lacks native
aliases. Claude ignores a global-defaults-only omit and keeps its aliases; the
active runtime is canonicalized (GSD_RUNTIME -> config.runtime -> marker ->
claude) so alias/case spellings can't defeat the check; explicit project
omit is workstream/project-scope aware; explicit true still materializes IDs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#2297): backfill PR number 2332 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 13:33:36 -04:00
Tom Boucher
1bb724048a fix(#2293): recognize --agy/--antigravity in plan-review-convergence whitelist (#2325)
* fix(#2293): recognize --agy/--antigravity in plan-review-convergence whitelist

The convergence reviewer-flag whitelist predated the 1.7.0 Antigravity CLI
adapter and silently dropped --agy/--antigravity, so convergence fell back to
--codex only and the working adapter was unreachable (worse after Gemini CLI's
upstream shutdown). Add both flags to the workflow grep whitelist, the command
argument-hint + flag docs, and the regenerated SKILL.md; they pass through to
/gsd-review unchanged. --gemini behavior is untouched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#2293): backfill PR number 2325 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 01:55:45 -04:00
Tom Boucher
5ea3401d4f fix(#2289): context-monitor emits injection envelope only for supported events (#2324)
* fix(#2289): context-monitor emits injection envelope only for supported events

gsd-context-monitor is wired to Codex Stop/SubagentStart/SubagentStop/PreCompact
(#772), but it emitted a hookSpecificOutput.additionalContext envelope for every
event. Codex's Stop schema rejects that shape ("hook returned invalid stop hook
JSON output") exactly when context is low.

Use a positive allowlist: emit only for context-injection events (PostToolUse,
AfterTool, and the pre-existing Gemini missing-name fallback); exit 0 silently
for Stop and every other event. Debounce and critical-session bookkeeping still
run on silenced events. Behavioral regression tests drive the real hook.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#2289): backfill PR number 2324 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#2285): de-flake per-plan use_worktree property test on duplicate briefs

The fast-check property in claude-orchestration.test.cjs located each plan's
agent() call via indexOf on the brief, but only plan IDs were unique — briefs
could collide (fast-check shrinks toward short strings). On a colliding seed the
lookup found the first duplicate's line and misattributed its isolation, failing
"use_worktree:true must carry isolation" intermittently (surfaced on macOS CI).
emitWorkflowScript is correct for duplicate briefs (verified); suffix the unique
id onto each agent() label so the test probe is unambiguous.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 00:56:07 -04:00
Tom Boucher
52fab7d9d7 fix(#2288): archive phase history under the outgoing milestone version (#2323)
* fix(#2288): archive phase history under the outgoing milestone version

phases.clear derived its archive directory from a live getMilestoneInfo()
read, but new-milestone.md switches the milestone BEFORE phases.clear runs,
so phase history was filed under the NEW milestone's <version>-phases/ dir.

Add a --archive-version override (threaded from new-milestone.md, captured
before the switch) with precedence override -> live read -> dated label.
Harden the version label against path traversal on both phases.clear and the
sibling milestone-complete sink (the label is a moved directory name), and
persist the outgoing version via a file + quoted shell expansion so untrusted
STATE.md content is never re-parsed by the shell.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#2288): backfill PR number 2323 into changesets

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 22:51:19 -04:00
Tom Boucher
b041f101fb fix(#2287): surface unresolved deferred-items.md entries in progress + audit-uat (#2318)
The executor SCOPE BOUNDARY convention (agents/gsd-executor.md) logs
out-of-scope discoveries to a phase directory's deferred-items.md, but no
reader ever consumed it — forensic_audit, cmdAuditUat, and capture --list
all skipped it — so deferred items were permanently invisible.

cmdAuditUat (src/uat.cts) now scans each phase dir's deferred-items.md via
a new parseDeferredItems (reusing the collectSection/splitGapsEntries/
extractGapEntryFields seams) and surfaces entries whose status != resolved
(fail-safe: a missing/garbled status surfaces rather than hides, matching
the false-negative-averse posture of #2286). forensic_audit
(gsd-core/workflows/progress.md) gains Check 7 that globs
.planning/phases/*/deferred-items.md and reports unresolved entries.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 21:40:59 -04:00
Tom Boucher
a22333034b fix(#2310): guard Codex agent model_overrides so Anthropic aliases never leak into .toml (#2312)
* fix(#2310): guard Codex agent model_overrides so Anthropic aliases never leak into .toml

generateCodexAgentToml embedded a per-agent `model_overrides` value verbatim as the
Codex `.toml` `model`, leaking GSD/Claude tier aliases (opus/sonnet/haiku/fable) and
`claude-*` ids. Codex/ChatGPT rejects those (400 "The 'sonnet' model is not supported
when using Codex with a ChatGPT account"), and since spawn_agent has no inline model
param, the model is baked into the .toml at install time — so the orchestrator could
not recover and fell back to the non-equivalent generic-agent workaround.

Translate a GSD tier alias through the Codex tier map (sonnet -> gpt-5.6-terra); drop
with a deduped warning any Anthropic-flavored value with no Codex mapping (fable) or a
`claude-*` id, so emission falls through to the runtime-aware resolver or Codex's
default. A final safety gate blocks an Anthropic-flavored model from the runtime-
resolver path too (runtime/target mismatch). Mirrors the Claude-side override guard
(#2041). Real Codex/OpenAI model ids in model_overrides still pass through verbatim
(#2256 preserved); runtime:"codex" tier resolution unchanged (#2517).

Adds regression + fast-check property tests in tests/codex-config.test.cjs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2310): backfill changeset PR number to #2312

* fix(#2310): Codex passive-model posture — omit Anthropic-flavored model (all namespacings)

Adopt ADR-1239's passive/session-only posture for Codex model handling: a Codex
agent .toml `model` is embedded ONLY for an explicit real-Codex model_overrides
pin; any Anthropic-flavored value is omitted so the agent inherits the always-
available session model (never a 400).

- model_overrides tier alias (opus/sonnet/haiku/fable) or a Claude model id →
  omit (was: translate to gpt-*); an explicit real-Codex model id → embed
  verbatim (#2256 preserved).
- Detect ALL Anthropic namespacings, not just `claude-*`: single-source the
  canonical CLAUDE_AGENT_ALIASES from model-resolver.cts and treat any id whose
  value contains "claude" (case-insensitive) as Anthropic-flavored — catching
  `anthropic/claude-*` and `us.anthropic.claude-*` (the forms the catalog assigns
  to opencode/hermes/kilo), which reach a Codex .toml via the runtime-resolver
  path on a mixed-runtime + Codex install.
- The final safety gate applies to the runtime-resolver path too.

The full passive posture (removing #2517's runtime-resolver per-tier embedding +
a correctness health-check + a Codex TOML sync path) is tracked as the ADR-2310
epic #2313.

Regression + fast-check property tests in tests/codex-config.test.cjs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 21:39:38 -04:00
Tom Boucher
636316f720 fix(#2286): audit-uat surfaces Gaps section + frontmatter/heading verification items (#2317)
parseUatItems only scanned '### N.' expected/result blocks and
parseVerificationItems only recognized table/bullet/numbered shapes, so
audit-uat returned a false-clean total_items:0 when a file recorded open
findings in a '## Gaps' section, declared items in a frontmatter
human_verification: array, or used the '### N. <label>'+bold-paragraph
verification shape.

parseUatItems now also scans '## Gaps' (via collectSection + iterateBullets)
and surfaces any entry whose status != resolved. parseVerificationItems
now treats the frontmatter human_verification: array (via extractFrontmatter)
as the primary source when present, and adds a tokenizeHeadings fallback
for the '### N.'+bold-paragraph shape, preserving the existing
table/bullet/numbered recognition (no double-count, no regression).

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 20:42:03 -04:00
Tom Boucher
ff9cb6069f fix(#2285): wire claude-orchestration Workflow backend into execute-phase (#2314)
The claude-orchestration capability (#1143) shipped registered 'active'
but fully inert: detectWorkflowBackend/emitWorkflowScript had no caller
outside their own CLI router, and execute-phase.md declared an
execute:wave:pre hook point that the workflow body never rendered — so
claude_orchestration.enabled:true had zero effect on real runs.

Approach B (maintainer-chosen):
- execute-phase.md now renders the execute:wave:pre hook
  (gsd_run loop render-hooks execute:wave:pre) at a new step 2.75,
  immediately before each wave's Agent() dispatch — fixing the latent
  dead-hook gap for any pre-wave capability.
- Move the claude-orchestration contribution execute:wave:post ->
  execute:wave:pre (a pre-wave backend selector belongs before dispatch,
  not after); rename fragments/execute-wave-post.md -> execute-wave-pre.md
  with prose instructing the orchestrator to call resolve-wave-dispatch
  before step 3. Unrelated wave:post contributions (ui.safety-gate, drift,
  external-job, mempalace) untouched.
- New .cts seam resolveWaveDispatch(input) composes detectWorkflowBackend
  + emitWorkflowScript into one {backend:'inline'|'workflow', ...} result;
  exposed as gsd-tools claude-orchestration resolve-wave-dispatch. This is
  a real non-CLI-router, non-test caller of both functions.

Fail-closed: any gate miss (disabled, non-Claude runtime, Workflow tool
absent, SDK below floor, execution_backend:inline, malformed input) or an
emit failure resolves to inline with a byte-identical result shape — no
regression to the default-off execute-phase path.

Regression tests (tests/fix-2285-*) cover happy-path activation + SDK-floor
BVA, the fail-closed gate-miss table with detectWorkflowBackend parity, a
fast-check composition property, capability.json contribution assertions,
and a source-contract guard that execute:wave:pre is now actually rendered.
Dependent registry-shape assertions updated in-scope.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 19:18:28 -04:00
Tom Boucher
75bedf16fd fix(#2284): project Hermes named dispatch onto delegate_task; protect comparison tables (#2309)
Hermes installs brand-swapped "Claude Code" -> "Hermes Agent" in shipped
workflows/*.md but never projected the Agent(...) dispatch calls onto
Hermes's delegate_task contract, so installed workflows kept literal
Agent(...) syntax and falsely asserted "The Agent tool IS available"
(Hermes exposes delegate_task, not Agent).

Dispatch projection: a generic named-dispatch engine
(projectNamedDispatchToStructuralDelegate) wired into the per-runtime
RUNTIME_CONTENT_DISPATCH.hermes.md converter, branching entirely on the
documentation-sourced hostIntegration.dispatch facts read via
_hostIntegrationDispatch (capability.json unchanged): namedDispatch:false
-> resolve the gsd-* role and embed a load-its-prompt instruction in the
payload; background:true -> map onto delegate_task background; read-only /
maxDepth:1 -> no nested delegation to leaf roles; per-call model dropped.
Span detection uses literal Agent( scanning + local balanced paren/quote
matching (immune to upstream document quote imbalance) and handles all
three corpus call forms (multi-line, object-literal, single-line compact).
An independent, mask-free post-projection guard fails the install loud on
any residual Agent(/subagent_type/leaked model. Fail-closed: install
throws if a literal gsd-* role reference cannot be resolved. commands ->
skill path untouched. No literal Agent( survives in installed Hermes
workflows.

Folded in (maintainer-directed) a pre-existing cross-cutting branding
defect: the "Claude Code" -> brand swap corrupted <runtime_compatibility>
comparison tables (where "Claude Code" is a compared-runtime label) for
every branding runtime. New shared applyClaudeCodeBrandSwap helper
protects <runtime_compatibility> regions via split-and-rejoin (no sentinel
token) while still rebranding genuine self-references; adopted by all six
branding .md converters. Also a surgical prose-consistency fix so
plan-review-convergence.md's dispatch-adjacent terminology is coherent
post-projection (no broad bare-word rename).

Golden install-parity regenerated for the six branding runtimes
(dispatch/branding scope only); other runtimes unchanged.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 17:10:06 -04:00
Cody Anderson
612fcb00f7 fix(#2232): cap phase-token continuation segments at exactly 2 digits (all sites) (#2254)
* fix(#2232): cap phase-token continuation segments at exactly 2 digits (all sites)

A phase whose slug's first word is a ≥2-digit number (dir
14-2026-photos-performance, roadmap phase "2026 Photos & Performance" →
slug 2026-photos-…) had its phase token over-collected as "14-2026"
instead of "14", so every phase-locating verb (init.plan-phase,
init.execute-phase, phase-plan-index, state.planned-phase,
roadmap.annotate-dependencies) resolved phase_dir=null / plan_count=0
while the directory existed. This is the residual case #2043 explicitly
scoped out: its ≥2-digit continuation gate (\d{2,}) distinguishes
single-digit slug words but not multi-digit ones (years, counts).

The structural distinguisher: getPhaseDirFromPhaseId writes sub-phase and
plan continuation segments zero-padded to EXACTLY 2 digits, so a genuine
continuation's digit run is exactly 2 — \d{2}(?!\d). The (?!\d) guard
caps the run without anchoring what follows, so each call site keeps its
own trailing grammar (letter suffixes, dotted sub-phases, boundaries).

Shared-source, not hand-synced: the grammar lives once in phase-id.cts as
PHASE_CONTINUATION_SEGMENT_SOURCE / isPhaseContinuationSegment (the #2121
single-owner seam), consumed by all five #2043 sites:
- phase-id.cts extractPhaseToken (the reported repro)
- validate.cts PHASE_TOKEN_FROM_DIR_RE + canonicalPlanStem
- roadmap-parser.cts isDirInMilestone numericRe (hyphenated mode)
- core-utils.cts + phase.cts extractCanonicalPlanId (paired plan
  component only — the LEADING phase component keeps unbounded \d{2,};
  phase numbers ≥100 are legitimate)

Digit-width policy, resolved per triage and locked by boundary tests at
1/2/3/4-digit continuation widths across all sites: sub-phase/plan
numbers ≥100 are out of the dir-token grammar. validate.cts
phaseDirNameRe's leading \d{2,} is intentionally untouched — it encodes
the write-side padding of the leading dir number, not the continuation
heuristic, and has no year collision.

Fixes #2232

Claude-Session: https://claude.ai/code/session_017KaYUJnfzV3JVVuQnhkcjg

* chore(#2232): add changeset for PR #2254

Claude-Session: https://claude.ai/code/session_017KaYUJnfzV3JVVuQnhkcjg

* test(#2232): parity gate + fast-check properties for the continuation cap

Addresses trek-e's review on PR #2254 (M1, M2, B1). Test-only — the fix
itself was verified as a true root-cause fix, so no source changes.

M1 — drift/parity enforcement for the new shared constant.
scripts/lint-phase-id-drift.cjs guards PHASE_NUMBER_TOKEN_SOURCE only; its
TOKEN_DRIFT_RE cannot match a bare \d{2,} re-derivation, so a future edit
reintroducing a raw digit-cap at a consuming site would pass lint + CI
silently. Extending the lint was rejected: \d{2,} legitimately appears at
the intentionally-unbounded LEADING-token sites (validate phaseDirNameRe,
core-utils/phase tokenRe), so a textual guard would need sanctions on
correct code and would flag by spelling rather than by behaviour.

Instead, per the repo's *-parity.test.cjs precedent, added
tests/phase-continuation-parity.test.cjs: a shared digit-width corpus
(1/2/3/4/5) asserting every consuming surface's notion of "is this segment
absorbed" equals isPhaseContinuationSegment(). Covers all five #2043 sites:
extractPhaseToken, PHASE_TOKEN_FROM_DIR_RE, canonicalPlanStem,
extractCanonicalPlanId (paired component), and roadmap isDirInMilestone
(hyphenated mode, on a real ROADMAP fixture). The corpus states the policy
independently of the regex, so it fails on divergence rather than mirroring
whatever the code does.

Failing-first verified: reverting PHASE_TOKEN_FROM_DIR_RE to \d{2,} fails 3
parity tests; reverting the owner constant itself fails 11 across parity +
properties + examples.

M2 — fast-check properties for the changed parser (4 added to
phase-id.test.cjs, following its existing inline fc precedent):
- biconditional: a segment is absorbed IFF its digit run is exactly 2
- the owner agrees with observable extraction for every digit run
- metamorphic: a write-side getPhaseDirFromPhaseId dir round-trips to its
  own normalizePhaseName id — ties the cap to the zero-padding convention
  it mirrors, so a change to the write-side width fails loudly
- metamorphic: the round-trip holds when the phase name leads with a year
  (the #2232 bug itself, generatively)
Digit runs are generated as digit strings (not String(int)) so leading-zero
forms like "02" — the whole point of the rule — are actually exercised.

B1 — GitGuardian red. The session-trailer hypothesis is disproven: the same
Claude-Session trailer rides 3 commits now merged to next via #2173, whose
GitGuardian check PASSED. GitGuardian's own comment names
tests/phase-id.test.cjs:260 — the synthetic dir literal 'M1-14-2026-photos'
tripping the generic high-entropy detector. Composed it from parts; the
assertion is unchanged, only the source spelling.

Refs #2232

Claude-Session: https://claude.ai/code/session_019SkiJk38YWAbmxHrGxEmuU

* test(#2232): name the parity gate after the invariant, not the phase module

CI caught two failures from the new parity test, both one root cause:
lint-test-file-count caps each production module at 2 test files (primary +
one integration, per the #3740 consolidation). The file was named
phase-continuation-parity.test.cjs, and the linter clusters a test to a
production module by name prefix — "phase-*" bound it to src/phase.cts,
whose cluster (phase.test.cjs + phase-dependency-levels.test.cjs) was
already at the cap, making 3. That tripped the lint-tests job AND the
ubuntu-24 unit lane, where tests/lint-test-file-count.test.cjs is a
meta-test asserting the linter exits 0 against the real repo.

Renamed to continuation-grammar-parity.test.cjs, matching the convention
the repo's other cross-cutting parity gates already follow: they are named
after the INVARIANT, not a module — capability-precedence-parity,
agent-classification-parity, and runtime-launcher-parity all have no
corresponding src/*.cts, so they cluster to nothing. The gate tests a
grammar shared ACROSS phase-id/validate/core-utils/roadmap-parser rather
than the phase module specifically, so the invariant-name is also the
semantically correct home. Not allowlisted: a novel offender belongs under
the cap, not ratcheted into the exemption list.

Content unchanged — same 12 assertions across the same 5 surfaces.

Refs #2232

Claude-Session: https://claude.ai/code/session_019SkiJk38YWAbmxHrGxEmuU

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-15 15:33:58 -04:00
Tom Boucher
4a9833d3e3 fix(#2278): use Edit() not Write() for Claude allow-permissions + migrate legacy (#2302)
GSD_CLAUDE_ALLOW_PERMISSIONS pre-populated Claude Code settings.json
with Write(.planning/*) and Write(STATE.md). Claude Code has no
standalone Write permission gate — file-editing tools are gated
collectively via Edit(pattern) — so those rules never matched, fresh
installs still hit first-run approval prompts for .planning/* and
STATE.md, and Claude Code emitted a session-start warning about the
unmatched rules.

Swap the two entries to Edit(.planning/*) / Edit(STATE.md). Add a
GSD_CLAUDE_LEGACY_ALLOW_PERMISSIONS list of the retired Write(...) forms,
consulted by mergeClaudePermissions (actively remove stale entries when
adding current ones, idempotent, user entries preserved) and by the
uninstall cleanup filter (still removes the legacy form). Sample
settings.json in docs/USER-GUIDE.md corrected to match.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 13:46:26 -04:00
Tom Boucher
f74442310d fix(#2257): auto-resume debug on non-terminal session-manager return (#2300)
The /gsd-debug orchestrator handled the gsd-debug-session-manager return
with only two literal-string checks (DEBUG SESSION COMPLETE, ABANDONED)
and no else branch, so a usable-but-non-terminal progress summary (the
manager's own turn/context budget exhausted mid-loop, with a valid
on-disk checkpoint) fell through to the user as if the debug were
complete. Same gap at the continue subcommand.

Callee side (agents/gsd-debug-session-manager.md): add an explicit
non-terminal CONTINUE_REQUIRED return marker, distinct from the two
terminal shapes and from a genuine user-input checkpoint.

Orchestrator (gsd-core/workflows/debug.md Sections 4 and 1c): classify
returns exhaustively — recognized terminal markers behave as before,
anything else is non-terminal and auto-resumes by re-spawning the
session manager from the same slug/checkpoint. Anti-loop guard: after
two consecutive no-progress resumes (unchanged next_action/updated),
emit a blocker report instead of looping.

Regression test (source-text contract guard, fix-2196 idiom) asserts
both sections' non-terminal/auto-resume branch, the CONTINUE_REQUIRED
marker, and the anti-loop bound.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 12:55:38 -04:00
Tom Boucher
9c65a2ea02 fix(#2256): resolve capability-registry configSchema defaults in config-get (#2299)
cmdConfigGet resolved absent keys through only the 4-key SCHEMA_DEFAULTS
map, so the ~42 registry-declared configSchema defaults (including the
workflow.security_enforcement security gate, default true) returned
'Key not found' (rc=1) — diverging from the runtime's own
resolveConfigKey Level-4 resolver and letting '... || echo false'
guards silently read the gate as disabled.

Add a resolveSchemaDefault helper that layers SCHEMA_DEFAULTS over the
already-imported getCapabilityConfigSchema(cwd) accessor, wired into all
three absent-key branches. --default flag precedence, the legacy 4 keys,
and 'Key not found' for genuinely unknown keys are preserved.

Two pre-existing, security-relevant defects in the same surface, found
while writing the regression tests, are fixed inline (no-defer policy):
- The --default fallback path never masked secret-named keys, printing
  e.g. 'config-get brave_search --default <secret>' in plaintext. All
  six default-emission sites now route through emitResolvedDefault,
  which applies the same isSecretKey/maskSecret masking the found-key
  path uses.
- Dotted-key traversal used raw bracket access, so 'config-get __proto__'
  / 'constructor' walked the JS prototype chain and returned internals
  at rc=0 instead of erroring. Each segment is now own-property-gated.

Regression tests folded into tests/config-get-default.test.cjs cover
registry defaults (boolean/enum/number, read live from the registry),
the no-file/mid-traversal/final-undefined branches, --default and legacy
precedence, prototype-pollution keys, secret masking, and the
Key-not-found vs No-config-file negative cases.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 11:47:15 -04:00
Tom Boucher
315d94f6d4 feat(#1945): tracer-first planning default + executor feedback gate (#2294)
* feat(#1945): tracer-first planning default + executor feedback gate

Make "thin end-to-end slice first, verify, then expand" the default planning + execution discipline instead of the opt-in --mvp mode.

- gsd-planner: first-class `type="tracer"` task; every plan LEADS with one production-quality end-to-end tracer slice by default; --no-tracer restores horizontal layers; --mvp/--tdd compose on top.
- gsd-executor + execute-plan: post-tracer feedback gate — autonomous runs halt-on-fail before expansion, interactive runs emit checkpoint:human-verify after the tracer.
- --no-tracer flag wired through plan-phase workflow/command/help/skill.
- CONTEXT.md glossary defines tracer bullet vs prototype; docs + references reconciled.
- tests/tracer-bullet.test.cjs: prose-contract + behavioral (verify plan-structure accepts tracer) coverage.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1945): backfill changeset PR number to 2294

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 09:41:36 -04:00
Tom Boucher
a68f1be10e ci(#2280): fix release-pipeline workflow defects (finalize timeout + auto-backmerge build:lib) (#2281)
Closes #2280

- release.yml: finalize timeout 10 -> 30 (match rc)
- auto-backmerge.yml: npm ci + build:lib before version-sync so the version hook can require the gitignored capability-ledger.cjs
2026-07-14 22:49:50 -04:00
Tom Boucher
c4237df8e6 docs(#2276): 1.7.0 release documentation — what's-new, EoS explanation, feature index (#2282)
Add a curated 1.7.0 release-highlights page (docs/whats-new-1.7.0.md) and a
conceptual Embeddable Orchestration System (EoS) explanation
(docs/explanation/embeddable-orchestration-system.md), extend docs/FEATURES.md
with a v1.7.0 feature section, and wire both new docs into the docs index
(docs/README.md) and the root README.

Covers the release's marquee changes: the ADR-1239 Host-Integration Interface /
EoS (Embeddable Orchestration System) runtime expansion, the Capability + EoS
discoverability registries, the gsd-mcp-server companion, model-catalog advances
(GPT-5.6, (1M) badge), statusline enhancements, the compact GSD-state format,
plus a themed summary of the 100 fixes and 4 security hardenings.

Also corrects a stale CONTEXT.md glossary entry: the Capability Registry Overlay
now documents the #2009 fail-open behavior for a load-failed gate-declaring
capability (previously described as fail-closed).

American house style; no parity-gated reference docs hand-edited.

Refs #2276, #1678

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 22:49:44 -04:00