Commit Graph

429 Commits

Author SHA1 Message Date
Tom Boucher
f80505aca3 fix(#2090): restore src/state.cts to origin/next (unrelated contamination) 2026-07-09 19:41:31 -04:00
Tom Boucher
fdfff96d52 fix(#2090): restore unrelated files accidentally deleted/modified by subagent
The implementation subagent cross-contaminated the branch with changes from
PR #2121 (phase-identifier parsing consolidation):
- Deleted docs/adr/2121-phase-identifier-parsing-consolidation.md (restored)
- Deleted src/phase-id.cts (restored)
- Deleted tests/phase-id.test.cjs (restored)
- Modified src/roadmap-parser.cts (restored to origin/next)
- Modified src/state.cts (restored to origin/next)

None of these are related to the Cline EoS migration.
2026-07-09 19:40:58 -04:00
Tom Boucher
9e7f80ea69 fix(#2090): resolve lint — drop unnecessary type assertion + unused vars 2026-07-09 19:38:27 -04:00
Tom Boucher
760eb71b6b feat(#2090): migrate cline onto imperative adapter + beforeTool/createAgentModel upgrades
Fold all hardcoded runtime === 'cline' / isCline branches in bin/install.js
into descriptor-driven runtime.hostBehaviors lookups (reapplyCommand,
frontmatterDialect, skipSharedHooksInstall, localTargetIsProjectRoot,
clineRulesSurface, localCommandsViaRules). Add cline-sdk-binding adapter
(ADR-1239 Phase D): UPGRADE 1 re-implements the .clinerules/hooks/PreToolUse
guard as a real AgentPlugin.hooks.beforeTool handler (fail-open, same
semantics); UPGRADE 2 wires DefaultGateway.createAgentModel params from
model_overrides/model_profile_overrides resolution (modelMode: active).
Install output is byte-identical (golden parity asserted for cline +
claude/cursor/codex/opencode).
2026-07-09 19:38:27 -04:00
Tom Boucher
bd5fcfceeb fix(#2125): route complete-phase resolver through canonical parser (review)
Orthogonal review surfaced that resolvePhaseIdForCompletePhase (state.cts) and
cmdStateCompletePhase's idempotency check still used an unanchored
/(\d+[A-Z]?(?:\.\d+)*)/i — even more permissive than the parseProsePhaseField
regex this phase fixes. Reachable corruption: after `milestone complete v0.5`,
`state complete-phase` (no --phase) mined "0.5" from the body line
"Phase: Milestone v0.5 complete" and rewrote STATE.md as "Phase 0.5 complete".

Both sites now delegate to phase-id.cts:parsePhaseFromProse (the same anchored
parser), so a milestone-closure line yields no token and the existing
"unable to resolve" guard fires instead of corrupting. Canonical tokens
(3, 03, 3A, 3.3, "3 of 5", "1 — Setup") are preserved unchanged.

Regression (tests/state.test.cjs, complete-phase suite): `state complete-phase`
on a "Milestone v0.5 complete" STATE.md now rejects and does not mine "0.5".
Demonstrated fail-first.

Refs #2125, #2121

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-09 18:11:55 -04:00
Tom Boucher
e7ff7be1a4 fix(#2125): migrate state.cts prose parsing to canonical parser (drives #2111)
Phase 2 of epic #2121. state.cts:parseProsePhaseField now delegates to the
anchored phase-id.cts:parsePhaseFromProse (built in Phase 1), removing this
module's independent prose phase-id regex.

Drives #2111: `milestone complete vX.Y` no longer corrupts current_phase. The
body line "Phase: Milestone v0.5 complete" previously had "5" mined from it by
the unanchored regex; the anchored parser returns { phase: null }, so
syncStateFrontmatter's #905 guard preserves the real current_phase. This also
fixes the broader family the review surfaced — every milestone completion
(e.g. v1.0 -> "0") was silently corrupting current_phase, not just .5-versions.

Regression (tests/milestone.test.cjs, in the milestone-complete suite, #2111):
`milestone complete v0.5` on a project with current_phase: "19" now preserves
"19". Demonstrated fail-first end-to-end: reverting the migration reproduces
current_phase = "5".

Closes #2125
Refs #2121

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-09 17:57:47 -04:00
Tom Boucher
80923244b1 fix(#2124): harden parsePhaseFromProse per orthogonal review (ReDoS + coercion)
Orthogonal security review of the Phase 1 surface found two issues; both fixed
and regression-tested:

- MEDIUM ReDoS: the name-extraction regexes /\(([^)]+)\)/ and
  /—\s*([^(\n]+?).../ backtrack O(n^2) on a crafted STATE.md field value with a
  long unterminated "(" / "—" run (reviewer measured ~38s at 320k chars).
  Length-bound both quantifiers to {1,200} -> linear (320k now ~100ms). A real
  phase name is far shorter than the cap.
- LOW: parsePhaseFromProse threw on non-string truthy input, unlike its three
  sibling #2121 functions. Coerce via String(value) up front.

The identical ReDoS regexes are copied verbatim from the pre-existing
state.cts:parseProsePhaseField; per the no-defer rule that surfaced defect is
fixed inline there too (Phase 2 / #2125 later supersedes that function by
delegating to the bounded phase-id.cts parser).

Adds a behavioral bound-guard regression test (a >200-char parenthetical is not
extracted) and a non-string-coercion test.

Refs #2124, #2121

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-09 16:33:39 -04:00
Tom Boucher
bf25e0b435 chore(#2124): build canonical phase-id.cts surface (Phase 1 of #2121)
Add the ADR-2121-locked canonical functions to src/phase-id.cts. No consumer
behavior changes — Phases 2-4 migrate the divergent call sites against them.

- parsePhaseFromProse: anchored prose parser. A phase is returned only when the
  STATE.md "Phase:" field VALUE begins with a phase token, so
  "Milestone v0.5 complete" yields { phase: null } instead of "5" (the #2111
  root cause: the old unanchored \b(\d+..)\b mined the minor-version digit).
  Name extraction (parenthetical / em-dash tail, minus status words) unchanged.
- stripConfiguredProjectCodePrefix / isForeignPrefixedPhaseQuery: config-aware
  prefix policy. A foreign prefix (MEM-01 when the configured code is LKML) is
  preserved rather than collapsed to a bare numeric phase — the #2104 fix's
  canonical home (consumed later, outside this epic's critical path).
- roadmapPhaseLookupSources: moved from roadmap-parser.cts so phase-id.cts is
  the single owner of the exact -> numeric -> prefix-tolerant ordering.
  roadmap-parser.cts now imports it (behavior-identical); its two now-unused
  imports (phaseMarkdownRegexSourceExact, OPTIONAL_PROJECT_CODE_PREFIX_SOURCE)
  are dropped.

Tests: subject-named suites in tests/phase-id.test.cjs covering the ADR
boundary set (v0.5, v1.0, MEM-01, AB-29, bare 29, zero-padded 029) plus two
fast-check properties: the #2111 "Milestone vX.Y complete never yields a phase"
invariant and a parse/normalize property.

Extend-never-mutate: the 12 pre-existing phase-id.cts exports are unchanged.

Closes #2124
Refs #2121

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-09 16:13:06 -04:00
Tom Boucher
45f3a2a5f9 fix(#2089): wire adapter into install path + address all review findings
MEDIUM fixes (code review):
- Wire resolveManagedHookEvents + resolveHookScripts + buildHookBusEntries
  from imperative-hook-bus.cts into writeCursorHooksJson — the install path
  is now truly descriptor-driven (reads hostBehaviors.managedHookEvents),
  not a hardcoded constant that happens to match the descriptor. bin/install.js
  passes the descriptor list via opts.managedHookEvents.
- buildHookBusEntries is now consumed (was dead code); entry-building is no
  longer duplicated inline.
- Remove try/finally from cursor-hook-bus-upgrade.test.cjs test bodies
  (violated CONTRIBUTING.md L342; redundant with t.after cleanup).

LOW fixes:
- Remove dead require('fs')/require('path') from gsd-cursor-pre-tool.js
- Fix resolveManagedHookEvents docstring (all-invalid fallback behavior)
- Add src/runtime-hooks-surface.cts to the AC2 source-guard file list

Security review: no CRITICAL/HIGH/MEDIUM findings (3 LOW are pre-existing
#777 baseline patterns, not regressions).
2026-07-09 13:17:04 -04:00
Tom Boucher
24896ddac7 fix(#2089): register 4 new cursor hook scripts in build + managed-hooks whitelists 2026-07-09 00:58:09 -04:00
Tom Boucher
b0d985ccb3 feat(#2089): migrate cursor onto imperative adapter + hook-bus/dispatch upgrades 2026-07-09 00:21:39 -04:00
Tom Boucher
6e773d97df feat(#2088): migrate Codex onto the Embeddable Orchestration System (ADR-1239)
Drive Codex install/uninstall through the descriptor-driven Host-Integration
Interface (declarative embedding adapter → engine surface dispatch) and fold
every positive `runtime === 'codex'` / `isCodex` projection into descriptor-driven
`runtime.hostBehaviors`. Install/uninstall output stays byte-parity-gated
(tests/fixtures/golden-install-parity/codex.json); no other runtime changes.

Three Context7-verified upgrades, each with a test on the user-reachable surface:
- Skill root → canonical $HOME/.agents/skills via a skills-kind `home` override,
  with pre-move migration cleanup (stale ~/.codex/skills/gsd-* removed on install
  and uninstall; user content preserved). Fixes getGlobalSkillsBase, writeManifest,
  and the skill-manifest inventory to honor the override so --skills-root /
  sync-skills / the manifest report the real location.
- Six new hooks.json lifecycle events (PreToolUse, PermissionRequest, PreCompact,
  PostCompact, SubagentStop, UserPromptSubmit) shared by install + uninstall;
  extendedHookEvents reconciled [] -> the schema-valid wired subset.
- Explicit `[agents] max_depth = 1` in the managed config.toml block, pinning the
  negotiated dispatch.maxDepth:1 axis. validateCodexConfigSchema now permits a
  known-scalar-only bare `[agents]` AgentsToml table (still rejects [[agents]] and
  unknown-key break-forms, #2760); mergeCodexConfig preserves the user's own
  AgentsToml scalars (max_threads etc.) instead of dropping them.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 21:42:11 -04:00
Tom Boucher
396f44bd0b feat(architecture): [EoS/opencode] Migrate OpenCode onto the Embeddable Orchestration System (ADR-1239, #2087)
Route OpenCode (and its Kilo sibling) through the public Host-Integration Interface and
land two Context7-verified capability upgrades. Byte-identical install output for all 16
runtimes (golden parity asserted).

Through the interface (AC2):
- OpenCode/Kilo's bespoke commands+skills+plugin install (the inline
  `else if (isOpencode || isKilo)` block) moves into the engine
  (installOpencodeFamilyCommands/Artifacts in src/install-engine.cts), dispatched by
  installRuntimeArtifacts when the descriptor declares hostBehaviors.combinedFamilyInstall.
  opencode/kilo now flow CLI -> _runtimeAdapter -> installRuntimeArtifacts like the skills
  runtimes. _isSkillsRuntime no longer excludes them; the bespoke block + dead
  copyFlattenedCommands are removed.
- Every hardcoded `runtime === 'opencode'`/`isOpencode` branch is folded into
  descriptor-driven runtime.hostBehaviors. ZERO `runtime === 'opencode'`/`'kilo'`
  string-equality remain in bin/install.js / install-engine.cts / runtime-artifact-conversion.cts.

Upgrades (AC4):
- Background dispatch: OpenCode shipped experimental background subagents in v1.15 and
  made them default-on in v1.17 -> dispatch.background/backgroundDispatch flip to true;
  shouldFlattenDispatch(opencode) now returns false (behavioral change; type: Changed).
- Expanded event surface: the OpenCode plugin subscribes permission.asked/replied +
  session.error.

Tests: opencode-imperative-reference (adapter/profile, shouldFlattenDispatch pin,
fail-closed negotiate, hostBehaviors, AC2 source-guard) + extended plugin surface test.
Docs: capability matrix v1.15/v1.17 citations. Changeset (Changed). gitignore .memdb//.memtrace/.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 17:17:05 -04:00
Tom Boucher
8c18d032a5 fix: restrict phase-complete checkbox regex gap (#2067)
The checkbox regex in cmdPhaseComplete used a greedy .* between ] and
'Phase N', so completing an already-checked phase (idempotent re-run)
matched a LATER phase whose description merely mentioned the target phase
number — checking the wrong phase's box. Restrict the gap to whitespace /
optional markdown bold emphasis, mirroring the tight pattern already used
by phase-insert.
2026-07-08 10:52:55 -04:00
Tom Boucher
8ffb1261d1 Merge branch 'next' into fix/2028-phase-complete-milestone-end-and-workstream-guard 2026-07-07 23:26:39 -04:00
Tom Boucher
5aa7771dd3 Merge branch 'next' into fix/2009-load-failed-capability-injects-blocking- 2026-07-07 23:12:09 -04:00
Tom Boucher
3f82929cc2 Merge pull request #2074 from open-gsd/fix/2072-thread-model-into-assumptions-and-review-spawns
fix(#2072): thread resolved model into routed-agent spawns (assumptions-analyzer, code-reviewer, code-fixer)
2026-07-07 23:11:53 -04:00
Tom Boucher
809f07b2f6 Merge branch 'next' into fix/2028-phase-complete-milestone-end-and-workstream-guard 2026-07-07 23:05:24 -04:00
Tom Boucher
015c3a7fda fix(#2071): extract install-time effort resolvers so effort sync stops requiring the un-shipped bin/install.js
`gsd-tools effort sync` crashed in every installed runtime (e.g. ~/.claude/gsd-core/)
with `Cannot find module '../../../bin/install.js'`: cmdEffortSync (src/commands.cts)
required the package-root bin/install.js for its install-time effort resolvers, but the
installer only copies the gsd-core/ subtree into a runtime home — bin/install.js is never
present there. So `effort` config changes silently never reached installed agents without
a full reinstall (exactly the gap #488 was meant to close). 4th instance of the recurring
"runtime code under gsd-core/ requires a file outside the shipped subtree via ../../../"
anti-pattern (#1223/#1920/#1383 were the prior three, all already mitigated).

Fix (ADR-457 direction — extract, single source): move readGsdEffectiveEffortConfig +
resolveInstallTimeEffort (with their _getGsdEffortCatalog + _readGsdConfigFile helpers)
out of the hand-authored bin/install.js into a new src/install-effort-resolver.cts that
compiles into the shipped gsd-core/bin/lib/install-effort-resolver.cjs. commands.cts now
requires it as a sibling (`./install-effort-resolver.cjs`) — always present in the
installed tree — instead of `../../../bin/install.js`. bin/install.js imports the same four
symbols back from the new module (it still calls them + re-exports them), so there is one
source of truth and no duplication/drift. The lazy manifest read is repointed from the
package-root layout (`.., gsd-core, bin, shared`) to the bin/lib layout (`.., shared`).

Scope note: this is one of four instances of the anti-pattern; the other three are already
shipped/guarded. A build-time guard rejecting new cross-boundary requires whose target isn't
in the installer copy manifest (to prevent instance #5) is recommended on the issue but kept
out of this fix.

Tests: tests/effort-sync-installed-runtime.test.cjs does a real minimal install into a temp
home (the golden-parity helper) and runs the issue's exact repro
(`gsd-tools effort sync --config-dir <temp>`), asserting no MODULE_NOT_FOUND for
bin/install.js. Fail-first verified: against pristine next the same test throws
`Cannot find module '../../../bin/install.js'` at cmdEffortSync; post-fix it syncs cleanly.

New module registered in .gitignore (ADR-457), eslint ignores, docs/INVENTORY.md +
INVENTORY-MANIFEST.json. bin/install.js is not shipped and the new module is under bin/lib
(excluded from golden parity), so no golden fixtures change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 22:21:45 -04:00
Tom Boucher
1ee00320c7 Merge branch 'next' into fix/2072-thread-model-into-assumptions-and-review-spawns 2026-07-07 22:03:47 -04:00
Tom Boucher
fad205e961 Merge branch 'next' into fix/2028-phase-complete-milestone-end-and-workstream-guard 2026-07-07 21:38:12 -04:00
Tom Boucher
4483300253 fix(#2072): thread resolved model into routed-agent spawns (assumptions-analyzer, code-reviewer, code-fixer)
model_overrides / models.<phaseType> were silently inert for gsd-assumptions-analyzer,
gsd-code-reviewer, and gsd-code-fixer on Claude Code: resolveModelInternal honors them,
but the workflows spawned these agents with no model= param, so the resolved value
never reached the Agent tool and the agents inherited the session model — no warning.

Fix — thread each agent's resolved model at every spawn site (the established
plan-phase pattern; the architecture-consistent Claude mechanism, since 13 other
agents already thread their model):
- discuss-phase-assumptions.md: `resolve-model gsd-assumptions-analyzer --raw`
  → ANALYZER_MODEL, threaded.
- code-review.md + code-review-fix.md (re-review): `resolve-model gsd-code-reviewer --raw`
  → REVIEWER_MODEL, threaded.
- code-review-fix.md (both fixer spawns): `resolve-model gsd-code-fixer --raw`
  → FIXER_MODEL, threaded (same silently-inert bug, same file — folded in per review).
- quick.md review step: was reusing `{executor_model}` for gsd-code-reviewer (so the
  reviewer's own override was ignored); init.quick now resolves `reviewer_model`
  (gsd-code-reviewer) and the spawn threads it.

resolve-model --raw returns the bare model string (resolve-execution --raw would
return effort — wrong). The resolver maps these agents to phaseType discuss /
verification / execution, so models.<phaseType> apply too.

Scope: the three agents reachable from the two issue-named workflows + quick.md. The
wider systemic class (other agents in UNTOUCHED workflows with the same pattern) stays
documented on the issue for a maintainer-scoped structural decision (thread-at-source
vs embed-at-install like #2256), not widened here.

Docs: the stale "discuss — reserved, no subagent today" model-profile tables now list
gsd-assumptions-analyzer and the verification row includes gsd-code-reviewer, across
the English docs, the shipped gsd-core/references/model-profiles.md reference, and the
ja-JP / zh-CN / ko-KR / pt-BR locale mirrors.

Tests:
- tests/model-resolver.test.cjs: #2072 acceptance — model_overrides and
  models.discuss/verification/execution resolve for all three agents.
- tests/model-routing-spawn-threading.test.cjs: every spawn of the three agents threads
  a resolved model (fails pre-fix); a header-precise parity guard fails the suite if a
  new un-threaded spawn of any of them regresses.
All 16 golden-install-parity fixtures + the workflow size baseline regenerated for the
changed shipped files (4 workflows + the reference doc); bin/lib is excluded from parity.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 21:22:48 -04:00
Tom Boucher
1092a71554 Merge branch 'next' into fix/2028-phase-complete-milestone-end-and-workstream-guard 2026-07-07 19:42:01 -04:00
Tom Boucher
46f8d3814c fix(#2009): load-failed capability gates fail open with a loud warning
Previously a capability that failed to LOAD (e.g. incompatible engines.gsd) but
declared a gate-kind loop hook caused the loop resolver to inject a BLOCKING
synthetic gate (blocking:true, onError:halt) at every declared point, halting
every ship:pre / verify:post project-wide over an unrelated load error, with no
remediation surfaced.

Per maintainer decision (#2009) it now fails OPEN: no gate is injected (the loop
proceeds; --active-cap correctly reports the failed cap inactive) and a loud
warning is emitted — to stderr (the channel host workflows/agents actually see)
and in the envelope 'warnings' array — naming the load reason and the exact
'gsd capability remove <id>' remediation. The loader still records blockedGates;
only the consequence changes from block to warn.

Security (review): capId and reason originate from a third-party manifest /
directory name. capId is validated against the canonical kebab-case id shape
before it is placed in the runnable remediation command (withheld otherwise);
reason is stripped of control chars and backticks. This closes an argument/
prompt-injection vector in the surfaced message.

Docs updated to the fail-open-warning posture (ARCHITECTURE, INVENTORY,
CONFIGURATION, README, capability-overlay-model). Also removes a dead 'before'
import surfaced by lint in the issue-2045 test.
2026-07-07 19:38:01 -04:00
Tom Boucher
6addeccd19 feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates
an external API/SDK/service can no longer seal without a decided coverage matrix.

- src/api-coverage.cts: deterministic detector (compound verb+noun signal +
  <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix
  parse/validate/render with field-length caps.
- check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as
  a token under .planning/phases/ only (traversal-neutralized); validates
  COVERAGE.md or blocks iff a strong integration signal is detected and no
  matrix exists; fail-closed when phases tree exists but phase unresolvable.
- capabilities/ai-integration: workflow.api_coverage_gate config key (default
  true), plan:pre contribution, blocking verify:pre gate. Data-driven.
- gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch.
- Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e.
  Code+security review findings fixed (stopword FP, scope containment, pipe/cap
  rejection, prompt-injection message hygiene).
- Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs.

Closes #1562
2026-07-07 15:11:12 -04:00
Tom Boucher
51dfa683d4 fix(#2028): phase.complete milestone-end out-of-order + workstream root-fallback guard
Two code-confirmed defects in `gsd-tools phase complete` (re-verified against
next; the three severe corruption paths the issue filed are superseded by the
ADR-1769 Transition Module migration + #2012, so this is the confirmed remainder).

1. Milestone-end mislabel (isLastPhase). The milestone-end determination only
   cleared isLastPhase when a HIGHER-numbered phase existed, so completing the
   numerically-highest phase out of order (e.g. Phase 10 before Phase 9) stamped
   STATE.md `Status: Milestone complete` while a lower phase was still outstanding.
   Added a lower-phase check: after the existing higher-phase scans, if any earlier
   phase in the current milestone has an unchecked roadmap checkbox (`[ ]`),
   isLastPhase becomes false AND next_phase/next_phase_name point at the LOWEST
   outstanding lower phase — so STATE.md advances to the real gap instead of
   parking on the just-completed phase. A completed phase always has `[x]`
   (phase.complete sets it), so all-lower-complete still reports milestone-end;
   heading-only roadmaps (no checkboxes) retain prior behavior. The checkbox regex
   mirrors the sibling phasePattern's anchoring (whitespace/bold + required `:`) so
   unrelated checklist lines mentioning "Phase N" don't match.

2. Workstream root-fallback (no guard). cmdPhaseComplete resolves every path via
   planningDir(cwd); with a `workstreams/` dir present but no active workstream and
   no --ws, that returns root `.planning`, so phase.complete wrote STATE.md/
   ROADMAP.md (and the mislabel) into the shared root other workstreams read.
   Added the same #1912 fail-safe guard init.progress got: refuse (asking for
   `--ws`/active workstream) instead of silently writing root. Resolution itself
   was already wired globally (resolveActiveWorkstream: --ws > GSD_WORKSTREAM >
   pointer, set in bin/gsd-tools.cjs), so only the refusal guard was missing.

The workstream-mode detection (`listAvailableWorkstreams`) is extracted into
planning-workspace.cts as the single source of truth and consumed by BOTH
init.progress and phase.complete, so the two fail-safe paths cannot drift.

Tests (tests/phase.test.cjs, new #2028 describe): out-of-order completion becomes
`Ready to plan` with is_last_phase=false, next_phase pointing at the outstanding
phase and Current Phase advancing to it (not the completed phase); all-lower-
complete still reports milestone-end; workstream-mode-no-active refuses with an
`--ws` hint; `--ws` completes in the workstream leaving root untouched; flat mode
unaffected. Fail-first verified locally via direct gsd-tools invocation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 14:40:44 -04:00
Tom Boucher
3742270c87 Merge branch 'next' into fix/2046-config-unset-null 2026-07-07 13:26:16 -04:00
Tom Boucher
9accce2b72 Merge branch 'next' into fix/2043-phase-token-single-digit-slug 2026-07-07 13:13:31 -04:00
Tom Boucher
603593d41d fix(#1857): test gates normalize to one-shot + bounded timeout (no watch-mode hang)
A GSD verification gate resolves a project's test command and runs it. vitest
defaults to WATCH mode in an interactive TTY — exactly where a user runs
`gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never
exited and the orchestrator waited indefinitely. Recovery needed the user to
manually prompt "something blocking?".

Fix — one shared helper + a bounded, surfacing timeout on the test-command gates:
- New pure module src/normalize-test-command.cts + `gsd-tools query
  normalize-test-command` verb: rewrites a resolved command to a best-effort
  one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`;
  a package-manager `test` script whose package.json runner is watch-vitest →
  `CI=true` prefix; handles `--dir`; already-one-shot commands unchanged —
  never double-flagged). Named `normalize-test-command` (not `test-*`) so the
  file does not match node --test's default `test-*` discovery glob.
- The three gates that HUNG or silently-continued route through that ONE helper
  and bound execution with `timeout $(config-get workflow.test_gate_timeout)`
  (new config key, default 600s): the regression gate (extracted to
  execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen —
  it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the
  audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124.
- verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so
  it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint
  on 124, staying under its frozen 40960-byte tier cap.

Security hardening (review): the normalizer only rewrites a runner named as a
standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never
mangled), is length-capped and uses only linear-time split-based scanning (no
super-linear backtracking on an adversarial `workflow.test_command`), and reads
package.json only when it is a regular file (never blocks on a FIFO via `--dir`).

Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the
schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors
workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores,
inventory manifest/index. All 16 golden-install-parity fixtures + workflow size
baseline regenerated for the changed shipped files; bin/lib is excluded from the
parity manifest.

Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security
hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route
through the shared helper + configured timeout + exit-124 watch-mode hint;
verify-phase asserted as normalize-only/already-bounded).
tests/execute-phase-active-flags.test.cjs repointed at the extracted step;
tests/planner-language-regression.test.cjs allowlist comment updated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 12:14:43 -04:00
Tom Boucher
6efd2b4016 fix(#2043): reject single-digit slug word in phase-token extraction (all sites)
A phase whose slug's first word is a bare single digit (dir
`46-6-rs-pipeline-orchestrator`, roadmap phase "6 Rs Pipeline Orchestrator" →
slug `6-rs-…`) had its phase token over-collected as `46-6` instead of `46`, so
`gsd-tools` phase-by-number lookups (init.plan-phase, init.phase-op, and
downstream execute/verify/ship) resolved phase_dir=null / has_context=false.

Root cause: an over-broad "looks numeric" test (`/^\d/`, `\d+(?:-\d+)*`)
classified token segments and could not distinguish a legitimate zero-padded
sub-phase segment (`01`, `02`) from a single-digit slug word (`6`). Zero-padded
phase/sub-phase segments are always ≥2 digits, so requiring ≥2 digits is the
structural distinguisher. Applied consistently across every same-class
implementation the triage identified (fixing extractPhaseToken alone leaves the
health-check and milestone-filter subsystems exposed):

- src/phase-id.cts extractPhaseToken — a pure-numeric leading segment continues
  only with ≥2-digit segments; a letter-prefixed milestone id (`M1`) still
  admits its single-digit sub-phase (`M1-2`).
- src/validate.cts PHASE_TOKEN_FROM_DIR_RE (W005/W006/W007 health checks) and
  canonicalPlanStem (I001 plan/summary pairing).
- src/roadmap-parser.cts isDirInMilestone numericRe (getMilestonePhaseFilter —
  highest blast radius: a false negative silently excludes a phase from the
  milestone).
- src/core-utils.cts + its verbatim duplicate src/phase.cts extractCanonicalPlanId.

Regression tests added across tests/{phase-id,health-validation,core-utils,
roadmap-parser}.test.cjs using the shared `46-6-rs-…` fixture, asserting the
token resolves to `46` (not `46-6`) while legit multi-segment tokens
(`01-02`, `02-03-04`, `M1-2`, `68-01`, `02-01`) are unchanged. The
roadmap-parser test exercises the real getMilestonePhaseFilter end-to-end.

The residual ≥2-digit-leading-slug case (e.g. `NN-2024-roadmap`) stays out of
scope — subsumed by the #612 / #565 phase-ID convention work, per the issue.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 10:15:37 -04:00
Tom Boucher
dd21bb7451 fix(#2046): config-set <key> null unsets (removes) the key instead of writing "null"
`gsd-tools config-set <key> null` — the "Clear" action documented in
settings-integrations.md / settings-advanced.md — previously fell through the
value parser as the literal STRING "null" and persisted it. Consequences:
"cleared" keys stayed set (config-get returned truthy "null"), and for secret
keys (brave_search/firecrawl/exa_search) a masked success line hid a truthy
4-char value on disk that integrations could pass along as a real credential.
There was also no unset/delete verb at all.

Fix: parse a bare `null` to JS null and short-circuit to a real UNSET that
DELETES the key from config.json — the semantic the docs already describe
("Remove the stored key" / "remove the key by setting it to null"). Deleting
(not persisting JSON null) is the correct clear: a persisted null is still a
present value consumers must special-case.

- src/config.cts:
  - parse block: `else if (val === 'null') parsedValue = null;`
  - cmdConfigSet: when parsedValue === null, short-circuit BEFORE the typed
    per-key validator gauntlet (so clearing an enum/boolean/number key removes
    it rather than being rejected) and before the project_code special-case;
    mask the previous value for secret keys in the output.
  - new `unsetConfigValue()` + `_unsetNestedValue()` mirroring setConfigValue/
    _setNestedValue: same prototype-pollution guard, but never creates missing
    intermediates and never prunes empty parents; returns { previousValue,
    existed }. Unsetting a never-set key is an idempotent no-op success.
- tests/config.test.cjs: new suite covering non-secret routing key, secret key,
  typed-enum-key bypass (context), idempotent unset, literal-"null"-on-disk
  guard, the unset-path prototype-pollution guard (alert #26 parity), and a
  4-segment deep-nested unset.
- tests/review-model-config.test.cjs: update the stale round-trip test that
  codified the bug (asserted config-set null → config-get returns "null") to the
  fixed contract — the model key is removed; the review workflow's
  `[ -n "$VAR" ] && [ "$VAR" != "null" ]` guard handles the empty read as
  "no override → reviewer default", same as the old "null" sentinel.

The 4 documented "Clear" flows (settings-integrations.md, settings-advanced.md)
were verified — their prose already describes removal, so the fix makes them
accurate rather than aspirational; no doc wording change required.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 09:41:14 -04:00
Tom Boucher
847de596b8 fix: third-party capability skills surface correctly (#2045)
A skills-only role:feature third-party capability installed 'active' but its skills never reached the runtime surface, capability enable/set rejected it as 'unknown capability', and capability list disagreed with capability state. Three defects, fix shape 1b (teach resolveSurface, no on-disk linking):

D1 (materialization): resolveSurface built the surfaced skills Set only from the on-disk manifest; third-party cap skills live at ~/.gsd/capabilities/<id>/skills/ and never entered the Set -> surfaced:false. Fix: union registry.capabilityClusters values into the Set in the full-profile branch (idempotent for first-party, additive for third-party; prototype-pollution guarded).

D2 (enable/set unknown): setCapabilityState validated the capId against the static first-party registry instead of the composed overlay-aware registry. Fix: validate against loadRegistry({includeInstalled}) like capability-state.cts does.

D3 (list vs state): capability list derived status purely from ledger existence, never surface composition. Fix: add a 'surfaced' field to list rows sourced from the same resolver capability state uses, so list and state agree.

Regression coverage: tests/issue-2045-third-party-skills-surface.test.cjs asserts all five acceptance criteria (resolveSurface union, state surfaced:true, enable/set not-unknown, list/state agreement, first-party + unknown-id regression). gsd-test: 23916/23916 green (linux-node22 + linux-node24).
2026-07-07 08:25:33 -04:00
Tom Boucher
ea4378063f Merge branch 'next' into feat/1820-specless-predicate-rail 2026-07-07 07:49:07 -04:00
Tom Boucher
0d7c15badc Merge branch 'next' into fix/2003-capability-state-runtime-flag 2026-07-07 00:24:27 -04:00
Tom Boucher
8a935a08a3 Merge branch 'next' into fix/2003-capability-state-runtime-flag 2026-07-07 00:19:08 -04:00
Tom Boucher
327b6409e8 fix(#2003): address code+security review findings
- warn (don't silently ignore) when --runtime is an unknown runtime that
  canonicalizeRuntimeName rejects; the warning surfaces via warnings[] so a
  typo like --runtime cluade or a runtime known to runtime-homes but not the
  alias manifest (e.g. grok) no longer silently resolves to the persisted
  runtime's config dir on this diagnostic command [M-1]
- add end-to-end CLI test for loop render-hooks --runtime (the exact command
  the bug report calls out as silently no-op'ing) [L-2]
- add closed-vocabulary rejection test: crafted --runtime values
  (../../etc/passwd, __proto__, --config-dir, garbage) are rejected, warn,
  and fall through to the persisted runtime — pins the security-load-bearing
  contract [NIT-01]
- add boundary tests: --config-dir wins over --runtime (precedence); missing
  --runtime value errors with USAGE [N-1]

Both orthogonal reviews returned APPROVE with no Critical/High findings.
Security review confirmed --runtime cannot coerce getGlobalConfigDir into an
arbitrary path (closed-vocabulary Map lookup + registry hash-key gate) and
does not expand the trust surface beyond the existing operator-controlled
--config-dir flag.
2026-07-06 23:32:06 -04:00
Tom Boucher
ab82e73af3 fix(#2003): add --runtime override to capability state + loop render-hooks
resolveCapabilityRuntimeState derived the config dir from resolveRuntime(cwd)
(GSD_RUNTIME -> config.runtime -> 'claude') when no --config-dir was passed,
so a repo with persisted runtime:'codex' resolved the config dir to ~/.codex
where the Claude skill isn't installed -> surfaced:false / hooks silently
no-op when the operator drove from Claude Code. capability state and loop
render-hooks parsed only --config-dir, never --runtime, so there was no way
to assert the actually-active runtime.

Add a runtimeOverride param to resolveCapabilityRuntimeState (canonicalized
via runtime-name-policy so aliases like codex-app work); when present it
short-circuits the persisted-runtime fallback and resolves getGlobalConfigDir
for the explicit runtime. Thread --runtime through cmdCapabilityState and
cmdLoopRenderHooks, and parse it in gsd-tools.cjs for both commands (dual
--runtime X / --runtime=X form, mirroring --config-dir and the existing
capability-set --runtime precedent). Help text updated.

Without the override, behavior is byte-identical to today (regression-guarded).
2026-07-06 22:57:05 -04:00
Tom Boucher
080bacdb4b Merge branch 'next' into codex/gsd-onboard 2026-07-06 22:43:32 -04:00
Tom Boucher
12b4c625ad Merge branch 'next' into fix/1858-flat-layout-skill-manifest 2026-07-06 22:33:07 -04:00
Tom Boucher
2c853822a1 fix(#1858): address code+security review findings
- restructure _loadFlatCommandsGsdManifest try/catch to wrap read+parse+set
  together (mirrors loadSkillsManifest exactly), so a thrown parser degrades
  both keys to [] — closes the latent catch-scope parity drift [Nit-1]
- add boundary test: gsd-.md (empty stem) skipped, gsd-x.md single-char stem
  kept (slice(4,-3) boundary) [Low-1]
- add unreadable-file test (POSIX-gated): both keys degrade to [] [Low-2]
- strengthen parity test to also compare requires + _calls_agents_ VALUES,
  not just the stem set [Nit-2]

Both orthogonal reviews returned APPROVE with no Critical/High findings.
Security review confirmed no new trust-boundary crossing, prototype-pollution
immune (Map/Set throughout), and symlink/path-traversal surface identical to
the pre-existing nested loader (not a regression).
2026-07-06 22:06:30 -04:00
Tom Boucher
9dcd370332 fix(#1858): detect flat commands/gsd-<stem>.md layout in _resolveManifest
_resolveManifest only recognized the nested source layout (commands/gsd/*.md)
and the installed-runtime skills layout (skills/gsd-<stem>/SKILL.md). A flat
source install (Claude local project shape: commands/gsd-<stem>.md, no
commands/gsd/ subdir) matched neither branch, so the manifest came back empty
and resolveSurface materialized the full profile to an empty Set — silently
reporting every skill-bearing capability as surfaced:false / enabled:false /
active:false. The nyquist/code-review/security/ui verify:post and execute:post
hooks never fired even with their workflow.* toggles on.

Add a third branch: when commandsGsdDir is absent, scan dirname(commandsGsdDir)
for gsd-<stem>.md files, strip the gsd- prefix, and build the same Map shape
the nested loader produces (requires via shared parseRequires, companion
_calls_agents_<stem> via shared parseCallsAgents). Falls through to the
installed-skills branch when the flat dir has no gsd-*.md files (precedence:
nested > flat-source > installed).

Also export parseCallsAgents from install-profiles so capability-state reuses
the SAME parser the nested loader uses (no drift; mirrors the existing
parseRequires export+reuse pattern).
2026-07-06 21:13:42 -04:00
Tom Boucher
e95af39a8c fix(#2041): address code+security review findings
- add typeof guard so a non-string override passes through verbatim instead of
  crashing on .startsWith (preserves pre-fix no-crash behaviour) [LOW-1]
- use Object.hasOwn() for the alias lookup so __proto__/constructor cannot
  return a truthy non-string from the plain object literal [LOW-D3]
- cap the unmappable-override stderr warning at 64 chars so an oversized or
  secret-shaped value cannot leak in full to stderr/logs [LOW-D4]
- remove the unused mapClaudeOverrideForRuntime export (helpers are covered
  behaviourally via resolveModelInternal/resolveModelForTier) [NIT]
- add resolveModelForTier unmappable-override fall-through test (closes the
  mutation-score gap) [MEDIUM-1]
- add case-sensitivity contract test (Claude-Sonnet-5 passes through verbatim) [LOW-2]

Both orthogonal reviews returned APPROVE with no Critical/High findings.
2026-07-06 20:45:40 -04:00
Tom Boucher
f214f1320d fix(#2041): map model_overrides full claude IDs to agent-tool aliases
model_overrides values that are full Claude model IDs (claude-sonnet-5,
claude-opus-4-8, claude-haiku-4-5, claude-fable-5) were returned verbatim on
the claude runtime and handed to the Claude Agent tool, whose typed model
parameter documents only tier aliases (opus/sonnet/haiku/fable). The
model_policy path already mapped full IDs -> aliases via
CLAUDE_POLICY_ID_TO_ALIAS (#1144); model_overrides skipped that mapping, so
the two resolver paths produced different shapes for the same underlying
Claude model. The fix mirrors #1144 on the override path via a shared
mapClaudeOverrideForRuntime helper used by both resolveModelInternal and
resolveModelForTier. Bare aliases pass through verbatim; non-Claude runtimes
and non-Claude custom/vendor values keep full IDs verbatim (parity). An
unmappable Claude ID (e.g. claude-opus-4-5) warns once to stderr and falls
through to tier resolution, exactly as the model_policy path already does.
Alias mapping is also the documented best practice (prevents staleness when
new model versions ship).
2026-07-06 20:06:30 -04:00
Tom Boucher
f433db8b88 fix(#1143): address adversarial review — full semver precedence, docs/reality alignment
- compareSemver: implement full SemVer 2.0.0 §11 pre-release identifier
  comparison (two pre-releases of the same triple now order correctly; was 0).
- capability description + fragment: scope the plan-checker/verifier claim
  (this capability delivers the parallel-execution backend; those gates remain
  inline until separately wired). Correct the 'each wave is one barrier' prose
  (a wave splits into multiple sequential parallel() barriers on files_modified
  overlap). Frame detect-backend CLI as a simulation harness; the pure function
  with the live host descriptor is the real detection seam.
- partitionStages docstring: 'near-minimal via greedy first-fit' (not 'fewest');
  document empty-files_modified behavior.
2026-07-06 15:41:19 -04:00
Tom Boucher
e3262d94d3 feat(capabilities): add claude-orchestration capability (Workflow backend) (#1143)
Default-off, BETA, claude-only capability adopting Claude Code's Workflow tool
(/effort ultracode, Agent SDK >= v0.3.149) as an optional parallel-execution
backend for the GSD loop. Restores the wave parallelism + plan-checker + verifier
that #853 forces inline on Claude Code, and folds gsd-ultraplan-phase under one
runtime gate.

- Pure fail-closed core (src/claude-orchestration.cts): detectWorkflowBackend
  (gate ladder: enabled -> Claude -> backend != inline -> nested+background host
  -> valid Agent SDK -> SDK >= floor; every miss degrades to inline) and
  emitWorkflowScript (waves -> parallel() barriers, plans -> gsd-executor +
  worktree, files_modified overlap -> separate stages, resumeFromRunId, budget).
  All interpolated identifiers validated script-safe; briefs JSON-quoted.
- claude-orchestration command family (gsd-tools claude-orchestration
  detect-backend|emit-workflow) for orchestrator invocation.
- Two gated loop contributions at wired points (execute:wave:post, plan:post);
  federated config keys (enabled/execution_backend/min_agent_sdk_version).
- ADR-1143 implementation amendment; CONTEXT.md glossary entry; explanation doc.

On any runtime lacking the Workflow tool, behaviour is byte-identical to today.

closes #1143
2026-07-06 15:18:23 -04:00
Dave
7ef834cabc feat(#1820): spec-optional predicate rail — author probe predicates into must_haves when SPEC omits them 2026-07-06 14:40:38 -04:00
Tom Boucher
dc746ce328 fix(#1575): address code review M1+M2+L1
M1: gate skills:'*' sentinel on unmodified-full profile (base profile must be
'full' AND no surface mods) so tiered profiles (core/standard) don't over-stage.
M2: thread resolveAttribution through capability-writer materialize opts; add
parity test variant with non-undefined Co-Authored-By attribution.
L1: remove redundant .agent.md filter condition (.endsWith('.md') already covers it).
2026-07-06 11:34:32 -04:00
Tom Boucher
d671171698 feat(#1575): complete agent-converter descriptor cutover for copilot/antigravity + surface path parity
- Teach applySurface to build agentCtx (pathPrefix + attribution) and pass it
  to kind.stage() for agents kind, mirroring createRuntimeArtifactInstallPlan
  (ADR-1235 §1). Surface-path agents now receive path-rewrite + attribution +
  converter + normalize, matching install output byte-for-byte.

- Pass skills:'*' sentinel for agents staging when no surface state modifications
  exist, so ALL agents are staged (not just those referenced by _calls_agents_).

- Declare converted agents kind in copilot and antigravity capability.json;
  add to _DESCRIPTOR_AGENTS_RUNTIMES in bin/install.js.

- Handle copilot .agent.md filename rename in both _copyStaged (install path)
  and _syncGsdDir (surface path).

- Ship golden-parity harness (ADR-1235 §0): tests/issue-1575-agent-descriptor-
  parity.test.cjs asserts applySurface output is byte-identical to
  installRuntimeArtifacts for all 7 descriptor-driven runtimes, plus stale-
  cleanup convergence and prune data-loss coverage.

- Update ADR-1235 with cutover progress.

Cline remains deferred (rules-only local branch + local/global complication).
2026-07-06 11:34:32 -04:00
Tom Boucher
7acbdc71b8 test(#1925): register zcode across installer surfaces + fluidify count pins
Resolve the zcode test cascade exposed by gsd-test:
- model-catalog.json: add zcode runtimeTierDefaults (null tiers, matching the
  other profile-marker-only runtimes) so KNOWN_RUNTIMES stays parity with allRuntimes.
- runtime-name-policy: add zcode to GLOBAL_CONFIG_HOME_FRAGMENTS (~/.zcode) so
  getGlobalConfigHomeFragment stops falling through to the .claude default.
- installer-migration-install: add the zcode fresh-install contract (flat-skills,
  no settings, no package.json — same shape as trae).
- golden-install-parity: capture the zcode fixture via a standalone generator
  script (not node --test — the gate stays gsd-test).
- capability-matrix.md: regenerate so zcode appears as a first-party row.

Fluidify the remaining count-pinned guards so adding a runtime no longer trips a
hand-pinned snapshot: gemini-runtime-removed (flag count derives from the
registry), non-claude-runtimes-registry-derivation (golden list derived from the
registry).
2026-07-06 08:31:52 -04:00
Tom Boucher
69b309e4e0 feat(#1925): add ZCode (Z.ai) as a pluggable runtime descriptor
Add ZCode as a first-party runtime via a declarative capability descriptor
(capabilities/zcode/capability.json) with zero hardcoded runtime === 'zcode'
branches — exercising the de-hardcoded, data-driven runtime path that 1.7.0
(ADR-1016 / ADR-1239) enables.

Descriptor (all axes sourced verbatim from zcode.z.ai docs):
- configHome ~/.zcode; nested skills + flat commands/agents; profile-marker install
- Claude-shaped skill format reuses convertClaudeCommandToClaudeSkill (no new converter)
- hostIntegration: declarative / slash-file / mcp / electron; dispatch background=false
  (foreground-only per docs); nested+maxDepth undocumented; passive model mode

Installer registration (data, not branches): --zcode flag, allRuntimes, runtimeMap,
interactive menu, --all list. getGlobalConfigDir + resolveRuntimeArtifactLayout +
ALLOWED_CONFIG_RUNTIMES + resolveInstallPlan all derive from the descriptor.

Revamped the brittle per-runtime golden-master tests to be count-agnostic,
descriptor-derived property tests (1.7.0 makes runtimes pluggable data, so pinning
frozen '15 runtime' snapshots is the wrong invariant): getdirname, label-policy,
config-adapter-registry (intent + install-plan golden master), capability-registry,
host-integration-descriptors (counts derive from curated maps). Adding a runtime
descriptor now extends coverage with zero edits to those suites.

Welcome banner, --zcode help, supported-runtimes how-to, and the host-integration
capability matrix (every axis cited) updated.
2026-07-06 08:31:51 -04:00