Commit Graph

446 Commits

Author SHA1 Message Date
Tom Boucher
93c5ecd645 feat(#792): add devin-desktop runtime alias for windsurf (#1086)
windsurf now also answers to devin-desktop (CLI --devin-desktop) for the Windsurf→Devin Desktop rebrand; all paths unchanged. The .devin/skills/ workspace migration is split to #1085. Closes #792.
2026-06-11 22:32:49 -04:00
Tom Boucher
e644819d59 feat(#767): inject disallowedTools deny-list into Claude read-only agents (#1081)
Install-time injection of a disallowedTools deny-list into Claude copies of read-only verifier/auditor agents (mirrors the #443 effort injection); source agents stay runtime-neutral so Gemini/Qwen/Hermes are unaffected. Closes #767.
2026-06-11 22:25:43 -04:00
Tom Boucher
58bfae9d6a refactor(#1059): phase 5f-1 — extract standalone hook-surface writers into a module — ADR-857/1016 (#1064)
* refactor(#1059): phase 5f-1 — extract standalone hook-surface writers into a module

Extract the structurally-isolated hook-surface writer functions (cline/cursor/
copilot/codex-hooks-json + buildHookCommand + atomicWriteFileSync + node/bash runner
resolvers) out of bin/install.js into a new src/runtime-hooks-surface.cts module
(-693 LOC from install.js). Behavior-preserving: install.js requires + re-exports
the moved functions (module.exports surface preserved); no descriptor reads, no
behavior change. Prerequisite for the descriptor-drive (5f-2), mirroring ADR-3660's
artifactLayout extract→drive split.

Review caught + fixed 3 coupling issues: (HIGH) the module's atomicWriteFileSync
dropped the shared __atomicWrittenTmps temp-tracking → now ONE shared set (module
owns it, install.js aliases it, both cleanups read it); (drift) buildHookCommand
called resolveNodeRunner(opts) vs the original resolveNodeRunner() → reverted; two
source-grep tests (workflow-guard, sh-hook-paths) that scanned install.js for the
moved functions → made behavioral/non-vacuous; duplicate runner resolvers consolidated.

Settings-json hook block (~648 LOC) deferred to 5f-1b; descriptor-drive to 5f-2.
New-module checklist done. ~62 hook test files green; gsd-test 17592/0.

Closes #1059

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1059): reconcile CLI Modules count after merging next (uat-predicate)

Merging current next (which added uat-predicate.cjs via #247) alongside this
branch's runtime-hooks-surface.cjs put the filesystem at 107 bin/lib modules, but
both sides had independently bumped the INVENTORY headline 105→106 so the merge
under-counted. Set "CLI Modules (107 shipped)" + regenerate INVENTORY-MANIFEST.json.
Both module rows already present. Fixes inventory-counts.test.cjs (the only CI red).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 17:21:04 -04:00
Tom Boucher
9223f2f4c8 feat(#247): runtime-neutral phase uat-passed predicate from HUMAN-UAT results (#1063)
* feat(#247): runtime-neutral phase uat-passed predicate from HUMAN-UAT results

Wire the already-reserved `phase.uat-passed` alias (subcommand `uat-passed`,
mutation:false) into the phase command router with a new markdown-aware
predicate that evaluates HUMAN-UAT results and reports pass only when every
required check passes. Post-SDK-retirement (ADR-0174/#174) successor to the
SDK-framed #70, with no SDK-specific API surface.

New pure module src/uat-predicate.cts:
- stripFalsePositiveContexts: frontmatter -> HTML-comment -> CommonMark-style
  fenced-block state machine (tracks delimiter char+length) -> blockquote,
  each a small composable step, so a `result: passed` inside frontmatter, a
  fenced/~~~ block (incl. ~~~ nested in a ``` fence), a comment, or a
  blockquote is never counted.
- parseUatResultItems: heading-block parser, column-0-anchored same-line
  result; a heading with no result -> `missing` (fail-closed).
- analyzeMarkdown: unterminated fence/comment detection (malformed -> blocker).
- evaluateUatPassed: allowlist pass/verification semantics; passed = no
  blockers && >=1 check && all passing; no_uat_artifacts discriminator (no
  vacuous pass); optional requireVerification policy hook.

Thin cmdPhaseUatPassed handler in phase.cts; router closure rejects unknown
flags via makeInvalidArgs. Hardened across two Codex adversarial passes
(vacuous pass, dropped failing tests, permissive verification status,
nested-fence escape, cross-line result value, masked unterminated comment) —
all fixed fail-closed. New unit + CLI-integration suites incl. a fast-check
property test; docs, CONTEXT glossary, inventory, and changeset updated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#247): backfill changeset PR number (#1063)

* fix(#247): indexOf paired-scan for unterminated-comment detection

CodeQL js/incomplete-multi-character-sanitization (high) flagged the
`raw.replace(/<!--[\s\S]*?-->/g,'')`-then-`.includes('<!--')` detection in
analyzeMarkdown as incomplete sanitization (a single regex pass can leave a
residual `<!--`). Replace it with a paired left-to-right indexOf scan that
contains no `.replace()` of the comment token — CodeQL-clean and strictly
more correct (a closed earlier comment can never mask a later unterminated
one). Behaviour unchanged; 98 predicate tests + scoped docker run green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 16:37:17 -04:00
Tom Boucher
8813ee5f95 feat(#429): HARD GATE on negative-grep literals echoed in plan <action> bodies (#1062)
Convert the planner's soft comment-text guideline into a plan-write-time
HARD GATE. When an acceptance criterion negative-greps for a literal
(`grep -c 'LIT' file == 0`) and that same literal appears verbatim in an
`<action>` body (JSDoc samples, head-comment references, "what NOT to do"
snippets), the executor's commit-time verify gate later fails on the
comment echo rather than a real regression — wasting cycles and training
the executor to distrust the gate.

`verify.plan-structure` (the `validate_plan` step) now scans for this:
- confidently-extracted (quoted) negative-grep literal echoed in an
  <action> → error (valid:false), failing plan creation
- unquoted/ambiguous grep target → warning (fallback policy)
- `<!-- planner-discipline-allow: LIT -->` escape hatch skips a literal
- positive-count gates (`== N`) and `!= 0`/`>= 0` are out of scope

Adds the `<comment_text_discipline>` block to gsd-planner.md, the full
rules + allowlist example to planner-antipatterns.md, and regression
fixtures for downstream incidents 12-04, 11-04, 12-02 (plus a boundary
case proving positive-count gate 11-02 is not flagged).

Closes #429

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 15:46:49 -04:00
Tom Boucher
9e3b056b15 fix(#779): correct stale model-catalog model IDs verified against live providers (#1047)
Verify-first audit: gemini opus gemini-3-pro→gemini-3.1-pro-preview (undefined in gemini-cli source), codex sonnet gpt-5.3-codex→gpt-5.4 (deprecated per OpenAI); qwen3-coder-next verified valid, unchanged. Adds a regression guard + sourcing note. Closes #779.
2026-06-11 13:36:23 -04:00
Tom Boucher
1fab2e10ba fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver (#1045)
* fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver

Source agents/*.md (gsd-planner, gsd-executor, gsd-verifier, gsd-plan-checker,
gsd-intel-updater, gsd-debugger, …) called bare "gsd-tools …" in shell blocks.
On a shim-only install — where gsd-tools.cjs exists under the runtime home but
gsd-tools is NOT on PATH — those calls fail with "command not found" and the
agent silently skips init/state/validate/commit ceremony, deferring to the
orchestrator or bypassing GSD bookkeeping entirely.

were never migrated, so it persisted on Claude Code and every other runtime that
consumes the source agents directly. Only gsd-phase-researcher.md carried a
resolver — and a stale, claude-only truncated one.

Fix (all runtimes):
- Inject the canonical multi-runtime gsd_run preamble (byte-equal to
  _runtime-launcher.snippet.sh — claude/codex/cursor/gemini/copilot/windsurf/
  augment/trae/qwen/cline/opencode/kilo/hermes/antigravity homes) at the top of
  the first gsd_run block of all 12 gsd-tools-calling agents, and rewrite every
  command-position bare gsd-tools to gsd_run.
- Upgrade gsd-phase-researcher.md's stale resolver to the canonical one.
- Extend scripts/sync-runtime-launcher.cjs to maintain agents/ in parity (the
  sync caught and corrected a mis-placed preamble during development).
- Extend the bare-gsd-tools (#2851) and launcher-parity (#373) regression guards
  to agents/ so no runtime can silently regress.

Closes #1041

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1041): backfill changeset PR number to 1045

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:42:59 -04:00
Tom Boucher
e4dfa6b9ea fix(#1012): invoke fallow with its real CLI and wire the report normalizer (#1044)
* fix(#1012): invoke fallow with its real CLI and wire the report normalizer

The /gsd-code-review structural pre-pass invoked fallow with flags no published
fallow version accepts (--json, --profile, --stdin-files), so it failed on every
run and degraded silently per REQ-FALLOW-02 — the feature never delivered on any
fallow version. Three compounding defects:

1. Invalid flags. Real fallow audit uses --format json (not --json), -q/--quiet,
   --changed-since/--base for changed-files scoping (no file-list input), and
   --max-crap for thresholds. There is no --profile or --stdin-files.
2. Exit-code handling. fallow audit exits 1 when it FINDS issues (verdict=fail),
   0 when clean. The pre-pass treated any non-zero exit as a crash and discarded
   the output — i.e. it threw away exactly the findings it exists to surface.
   Success is now decided by whether a valid fallow JSON report was produced,
   not by the exit code.
3. Schema mismatch. normalizeFallowReport parsed a fictional top-level schema
   (unusedExports/duplicates/circularDependencies) fallow never shipped, and was
   dead code (the workflow embedded raw JSON; its tests asserted the fictional
   schema, one even calling a non-existent runFallowAudit and passing vacuously).

Fixes: align the invocation to fallow's documented agent-facing pattern; map the
profile preset (minimal/standard/strict) to --max-crap (50/30/15); scope phase
runs via --changed-since with a repo-scope fallback; rewrite the normalizer to
fallow's real schema (dead_code.unused_exports/unused_files/circular_dependencies
+ duplication.clone_groups) and wire it into the workflow so the reviewer
receives normalized findings; replace the fictional-schema fixtures and tests
with real-schema ones and delete the vacuous runFallowAudit test.

Closes #1012

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1012): backfill changeset PR number to 1044

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:36:50 -04:00
Tom Boucher
ab8b84286e fix(#1013): resolve worktree.baseRef from user/global settings cascade (#1038)
* fix(#1013): resolve worktree.baseRef from user/global settings cascade

cmdWorktreeBaseCheck resolved worktree.baseRef from the project checkout's
.claude/ only (settings.local.json then settings.json). A user/global
worktree.baseRef:"head" — the layer /config writes and the harness honors,
and the only sensible place for a machine-wide preference — was invisible. On a
phase lane (HEAD ahead of origin/HEAD, or no origin/HEAD symref) base-check
returned shouldDegrade:true and execute-phase forced sequential execution,
silently losing the parallel worktree execution the user configured.
CLAUDE_CONFIG_DIR (relocated user config dir) was also ignored.

resolveEffectiveBaseRef now accepts an optional user/global config dir and reads
its settings.json as a third, lowest-precedence layer (project local > project
shared > user/global). cmdWorktreeBaseCheck resolves it via
getGlobalConfigDir('claude'), which honors CLAUDE_CONFIG_DIR. The existing
project-level reads stay as higher-precedence overrides and the injectable
readFile seam is preserved, keeping the unit tests hermetic.

Closes #1013

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1013): backfill changeset PR number to 1038

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:36:30 -04:00
Tom Boucher
fbd62cd84f feat(#1035): phase 5a — author 16 role:runtime capability descriptors (registry-only) (#1039)
Author capabilities/<runtime>/capability.json for all 16 runtimes (role:runtime),
populating the registry's runtimes index ({} -> 16). Each carries the 6 ADR-1016
axes (configHome structured, configFormat, artifactLayout structured, commandStyle,
hooksSurface + hookEvents, sandboxTier, supportTier), extracted from the live
modules (runtime-homes/runtime-slash/runtime-artifact-layout/runtime-config-adapter
+ CODEX_AGENT_SANDBOX). validateRuntimeBody tightened to enforce the closed
vocabularies + structured configHome/artifactLayout (rejects old string configHome;
env required; skillsHome recursively validated).

Registry-only: bin/install.js + src/*.cts untouched, zero behavior change. The 5b-5f
drive steps consume these one axis at a time (ADR-1016 §8).

Codex accuracy review caught + fixed: opencode/kilo register ZERO lifecycle hooks
(install.js skips the whole block) -> hooksSurface:none (configFormat stays
settings-json); kimi probe selects on <root>/skills existence -> probeExists:"skills".
ArtifactKind required-field strictness + ConverterName closing deferred to 5d/5e.
ADR-1016 (Proposed) amended to match (Decisions 1 + 5).

Closes #1035

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 09:34:07 -04:00
github-actions[bot]
b2222fa769 docs(#860): ADR-3660 addendum — Nth-runtime-via-existing-layout is an addendum, not a new ADR
Maintainer governance decision (waiving a standalone ADR for Qoder, PR #1021):
registering a runtime that reuses the existing profile-marker-only install
surface + a single skills kind is an enhancement governed by ADR-3660 via
addendum. Records the addendum-vs-new-ADR qualifying criteria, the normative
agent-frontmatter contract (name+description only — the sibling converter
shape), and Qoder as the first runtime logged under this path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 08:39:22 -04:00
Tom Boucher
5f48a42514 docs(#1022): resolve step-vs-gate model question — steps additive, gates block, mode self-gates (#1025)
Record the resolution of #1022 (surfaced scoping the §5.6 ui-phase cutover):
a step is purely additive and never halts the host; host-blocking preconditions
are gates (blocking/onError:halt) — no hook-model change. Runtime/mode context
(auto/chain vs manual) self-gates in the skill, not via when (config-only).

§5.6 decomposes into the existing plan:pre step (ui-phase, self-gates on
frontend + pipeline) + a new plan:pre gate (frontend-and-no-UI-SPEC → halt,
when: workflow.ui_safety_gate) that blocks planning in manual mode — preserving
the "run /gsd:ui-phase first" UX (maintainer call: pipelines-only auto-fire).
The render-hooks dispatch template grows to handle gates, not just steps.

Recorded in ADR-894 (clarification) + CONTEXT.md
(RULESET.CAPABILITY.step-additive-gate-blocks). Unblocks the §5.6 cutover.

Closes #1022

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 22:51:31 -04:00
Tom Boucher
61b29d2af6 docs(#1016): ADR-1016 runtime capability descriptor (ADR-857 phase 5 design) (#1019)
Realize ADR-857 Branch 8 (host-CLI support as role:runtime Capabilities) +
materialize ADR-58's InstallPlan. Closes the 6-axis descriptor vocabulary,
absorbs the hard-case runtimes as data, and stages the install migration as a
5a→5g ladder (4/6 axes already modularized; InstallPlan is the 5g capstone,
reachable by collection, not a big-bang rewrite).

Amended after a plugin-side design grill (rubber-duck + grill-with-docs vs
#956 MemPalace / #999 Impeccable):
- New term Connected Capability (CONTEXT.md) — a Capability whose integration
  shape brings its own external process/service/state; orthogonal to authorship.
  Named, tracked gap (vehicle #956); current schema does not express it.
- Narrowed the dogfood claim: the descriptor dogfoods the runtime interface
  only, not the feature-plugin/Connected path.
- Structural "off means off" rule (CONTEXT.md RULESET.CAPABILITY.off-means-off):
  the host derives shared outputs from active hooks; a hook adds/is-counted,
  never mutates host source. Ratify in ADR-894; proven by spike #1018.
- Hand-waves resolved by code: sandboxTier real-but-thin; model-catalog
  orthogonal (not a 7th axis); converters closed into a ConverterName enum +
  added the kimi-agents artifact kind; configHome is pure read-only.
- Hook-firing path is unproven (render-hooks built, never consumed) → spike
  #1018 must prove render-hooks→live-workflow execution before phase-5 build.

Design-only; no code. Status: Proposed.

Closes #1016

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 21:00:37 -04:00
Colin Johnson
76f42ddb4b feat(#1014): add Claude Fable 5 model config (#1015) 2026-06-10 20:32:26 -04:00
Jeremy McSpadden
092340d18a fix(#711): wire autonomous convergence flag (#729)
* fix(#711): wire autonomous convergence flag

* Update wise-ibex-tumble.md

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 16:10:07 -04:00
Jeremy McSpadden
f61b97276e fix(#724): block convergence on actionable review findings (#728)
* fix(#724): block convergence on actionable review findings

* merge: integrate clean next (#936 inline) onto author tip + re-apply cursor fixes (Mode field, REVIEWS.md extraction) and review hardening (#724)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 14:54:15 -04:00
Joe
5e8a723089 feat(templates): add optional Business Context section to PROJECT.md template (#756)
* feat(templates): add optional Business Context section to PROJECT.md template

Adds an optional `## Business Context` section (Customer, Revenue model,
Success metric, Strategy notes) between Core Value and Requirements, for
monetized or customer-facing projects. Optional by default — an HTML comment
tells non-business projects to delete it; capped at four one-line fields to
stay a constraint reference, not a business plan. The milestone evolution
review in complete-milestone.md checks it only when the section is present.

Refs #72

* chore(changeset): set pr number for #72 fragment

* test(#72): add source-text-is-the-product exemption marker

Addresses review Minor #1 on PR #756. The contract test reads the
PROJECT.md template and complete-milestone workflow .md files and
asserts on their content (the local/no-source-grep pattern). Those
.md files ARE the product surface, so this is a valid
source-text-is-the-product case. Add the explicit // allow-test-rule
marker per RULESET.TESTS.no-source-grep.exemption so intent is
audit-traceable before the rule promotes to error (#453).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-10 14:13:01 -04:00
Tom Boucher
972a41a528 fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:) (#990)
* fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:)

Closes #967

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#967): backfill changeset pr number (990)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 11:10:40 -04:00
Tom Boucher
2981983bae fix(#973): add Edit to gsd-planner tools and forbid whole-file Write of ROADMAP.md (#989)
* fix(#973): add Edit to gsd-planner tools and forbid whole-file Write of ROADMAP.md

gsd-planner shipped Write but not Edit — the same writer-agent gap fixed for six
agents in #571/#581. Without Edit, an in-place ROADMAP update fell back to a
whole-file Write that truncated committed milestone history (292→16 lines in a
real incident).

Changes:
- agents/gsd-planner.md: add Edit to tools: frontmatter (adjacent to Write)
- agents/gsd-planner.md: update_roadmap step now directs Edit (scoped), with an
  explicit blocking prohibition on whole-file Write of ROADMAP.md or any existing
  curated .planning/ file
- agents/gsd-planner.md: Write contract section clarifies Write is authorized only
  for net-new PLAN.md creation; existing files must use Edit
- tests/agent-frontmatter.test.cjs: extend SECTION_WRITER_AGENTS list (#581 test)
  to cover gsd-planner — fails before fix, passes after
- .changeset/973-gsd-planner-edit-tool.md: Fixed changeset, pr:0

Closes #973

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#973): backfill changeset pr number (989)

* fix(#973): trim gsd-planner.md prose under agent size cap (keep Edit + scoped-Edit-for-ROADMAP rule)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 10:44:20 -04:00
Tom Boucher
caca4d255c feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4) (#988)
* feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4)

Migrate the intel CLI command family from a hardcoded gsd-tools.cjs case arm to
a registry-dispatched Capability (commandFamilies mechanism, #961), mirroring
the graphify (#972) and audit (#984) cutovers. New src/intel-command-router.cts
exports routeIntelCommand reproducing all 9 subcommands verbatim (incl. the
status timeAgo non-raw post-processing), lazily requiring intel.cjs inside the
route fn. capabilities/intel/capability.json declares the intel command family;
commands-only (skills:[]), declares the existing intel.enabled gate (default
false — behavior unchanged).

Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing intel.test.cjs passes unchanged. Completes the first-party command-
family cutover sequence (graphify/audit/intel).

Closes #985

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#985): compute expected planningDir via path.join in intel cutover unit tests (Windows CI)

The intel-command-cutover unit-mock assertions hardcoded a POSIX `/.planning`
expectation while the router builds it with path.join(cwd, '.planning') →
backslashes on Windows, so the planningDir-arg assertions (query/status/diff/
snapshot/validate/update/api-surface) failed only on windows-latest CI. Compute
the expectation with path.join (cross-platform); production router unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 10:00:30 -04:00
Tom Boucher
9a03539c2d feat(#981): audit-uat + audit-open command cutover — commands-only capability (ADR-857 phase 4d-impl-3) (#984)
Migrate the audit-uat + audit-open CLI commands from hardcoded gsd-tools.cjs
case arms to a registry-dispatched Capability (commandFamilies mechanism, #961),
mirroring the graphify cutover (#972). New src/audit-command-router.cts exports
routeAuditUat/routeAuditOpen, each lazily requiring only its backing module
(uat.cjs/audit.cjs) inside the route fn — matching the old per-case lazy loads.
capabilities/audit/capability.json declares the two command families;
commands-only (skills:[]), no config gate, audit_review cluster untouched.

Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing audit regression tests pass unchanged.

Closes #981

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 08:34:49 -04:00
Tom Boucher
77bfd943dd feat(#972): graphify command cutover — first capability owning a command family (ADR-857 phase 4d-impl-2) (#975)
Migrate graphify into a Capability that owns its `graphify` command family,
dispatched via the registry (#961 mechanism) instead of a hardcoded case.
graphify is now an enable/disable plug-in.

- src/graphify-command-router.cts: routeGraphifyCommand (standard route*Command),
  reproduces the removed case EXACTLY (query +--budget, status, diff, build,
  hidden build snapshot, usage/unknown errors); injectable _graphify test seam.
- capabilities/graphify/capability.json: role feature, tier:full, skills:[graphify],
  config:{graphify.enabled default false}, commands:[{family:graphify, module,
  router:routeGraphifyCommand}].
- removed case 'graphify' from gsd-tools.cjs; graphify now flows
  default -> dispatchCapabilityCommand -> commandFamilies.graphify -> router.
- regenerated registry (commandFamilies/bySkill/configSchema/profileMembership/
  capabilityClusters for graphify); tier:full keeps 4c install/surface a no-op.

Equivalence-proven: 8 recording-mock unit tests assert the exact fn+args per
subcommand (budget, snapshot-vs-build); 9 subprocess tests assert distinguishing
output shapes; existing graphify tests pass unchanged through the new path.

Surfaced (not silently accepted): the pre-existing --budget-no-value NaN no-op
quirk, preserved for equivalence, filed separately.

Closes #972

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 07:12:57 -04:00
Colin Johnson
354e0e1b94 fix(ratchet): add --update drift repair + inherited-drift guidance to regression-name lint (#971) 2026-06-10 01:05:24 -04:00
Tom Boucher
1c86368785 fix(#947): restore gsd- prefix on Hermes skills for canonical dispatch (#955)
* test(#947): add regression tests and update stale Hermes assertions

- Add bug-947-hermes-gsd-prefix.test.cjs: 12 TDD tests covering fresh
  install canonical layout, bare-stem migration, manifest key format,
  and non-Hermes runtime isolation
- Update hermes-skills-migration.test.cjs: bare-stem → gsd-prefixed
  path and name assertions (#947 canonical layout)
- Update install-nested-layout.test.cjs: Hermes NEST matrix prefix ''
  → 'gsd-'
- Update install-regressions.test.cjs: Defect #1 now seeds bare-stem
  dirs (help/, quick/) and asserts gsd-help/ canonical output; use
  real GSD stems so readGsdCommandNames() migration finds them
- Update install-runtime-artifacts.test.cjs: Hermes nested layout and
  legacy migration assertions align with gsd- prefix
- Update install.test.cjs: Hermes install test uses gsd- prefixed paths
- Update runtime-artifact-layout.test.cjs: prefix '' → 'gsd-'

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#947): restore gsd- prefix on Hermes skills for canonical dispatch

Hermes skills were installing under bare-stem paths
(skills/gsd/<stem>/SKILL.md, name: <stem>) due to prefix: '' set in
ADR-3660 / #3664. This broke /gsd-<stem> dispatch and forced users to
invoke skills without the gsd- namespace prefix.

- src/runtime-artifact-layout.cts: change Hermes skillsKind prefix
  from '' to 'gsd-'; skills now land at skills/gsd/gsd-<stem>/SKILL.md
  with name: gsd-<stem>
- bin/install.js _runLegacyInstallMigrations: invert the #3664
  migration — remove stale bare-stem dirs (using readGsdCommandNames()
  to distinguish GSD-owned stems from user content), keep gsd-* dirs
  which are now canonical
- bin/install.js _runLegacyUninstallCleanup: also remove bare-stem
  dirs on uninstall for clean teardown
- bin/install.js uninstallRuntimeArtifacts: post-cleanup removes
  DESCRIPTION.md and empty skills/gsd/ category dir on Hermes
- bin/install.js: remove skillListPrefix Hermes exception (now uses
  shared 'gsd-' path)
- docs/adr/3660-runtime-artifact-layout-module.md: document #947
  reversal of the bare-stem sub-decision

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: add changeset for #947 fix (#955)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#947): remove ALL pre-migration bare-stem Hermes skills on reinstall (adversarial review)

Replace readGsdCommandNames()-based bare-stem cleanup (which missed skills
not in the commands source tree, e.g. dev-preferences) with
_removeHermesBareStemDirs(), called AFTER the install loop when the exact
set of installed gsd-<stem>/ dirs is authoritative. For every gsd-<stem>/
written this run, the corresponding bare skills/gsd/<stem>/ is removed.
User-owned bare dirs with no gsd-<stem> counterpart are preserved.

Add two adversarial-review regression tests that FAIL on old code:
- bare skills/gsd/dev-preferences/ removed when gsd-dev-preferences/ installed
- user-owned bare dir with no gsd-<stem> counterpart is preserved (no over-deletion)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 00:22:36 -04:00
Colin
9db7958a70 fix(review): single eslint home, threshold co-location, helper reuse, docs clarity
Review-pass fixes: lint:ci composes npm run lint (one eslint invocation
home); the scripts/ coverage floor moves to package.json
(test:coverage:scripts-floor) so both thresholds live together; the ratchet
test uses helpers.createTempDir; TESTING-SUITES.md clarifies what the
Windows scoped lane runs and why feat-*/enh-* files are exempt from the
bug-* ratchet.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:53:25 -04:00
Colin
cd5db1f8db test(suites): seed security/slow/integration suites via measured retags
Renames (git mv) with all references updated (ci-test-scope RULES,
windows-parity allowlist, test-file-count allowlist, docs in 6 locales):

- 5 scanner tests -> *.security.test.cjs — the 'Run security tests' CI step
  ran zero files since the suite taxonomy landed; it is now honest.
- graphify-auto-update -> *.slow.test.cjs (36s, slowest file in the suite;
  e2e gsd-tools spawns) — runs on full-matrix lanes and push to next.
- installer-migration-install-integration -> *.integration.test.cjs
  (13s; an integration test by its own name).

Coverage gate measured after retags: 88.55% lines (gate 70%).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:50:41 -04:00
Tom Boucher
34435736e4 docs(#959): ADR-959 capability command contribution (ADR-857 phase 4d design) (#960)
Design how a Capability contributes a gsd-tools CLI command family and how the
hardcoded 73-case runCommand switch opens to registry-driven dispatch. Realizes
ADR-857 decision 7's reserved commands/module field (deferred by ADR-894).

Grilled to its leanest form: the registry DISCOVERS a standard route*Command
(no rebuilt handler table, no new arg convention); dispatch sits in the default
case (collision structurally impossible, no shadowing gate needed); graphify is
the first real cutover (lowest blast radius, has skill+cluster+config gate,
full-only so 4c stays no-op), proven equivalent and serving as the phase-6
template.

Design-only; CONTEXT.md gains a "Capability Command Family [Planned]" entry.

Closes #959

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 23:20:51 -04:00
Tom Boucher
e006ff74fd docs(#956): MemPalace capability pre-proposal (PRD/ADR draft) (#957)
Combined PRD + ADR for wiring MemPalace (local-first AI memory) into the
GSD loop as an ADR-857 feature capability. Bidirectional sync, three
selectable memory-relationship modes (augment/kg_backend/replace),
loop-point recall+capture map, opt-in tier:full, MCP-primary/CLI-fallback.

Marked Pre-Proposal: the first-party-plugin proposal standard is not yet
established and ADR-857 phase-6 loop wiring is pending. First of a planned
series; PRD/ADR format is provisional pending PM-method evaluation.

Refs #956

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 22:30:48 -04:00
Tom Boucher
85cfa5dc13 feat(#945): unified capability-state resolver (ADR-857 phase 4b) (#946)
Add a read-side query composing the three toggle systems into one
per-capability view. resolveCapabilityState({registry, installedSkills,
surfacedSkills, config, cwd}) reports installed (skills ⊆ resolved install
profile), surfaced (skills ⊆ resolved surface), and per-hook active (no when →
active; non-empty-string when → resolved via _resolveActivationValue; empty/
non-string → inactive), with no forced composite verdict. cmdCapabilityState
does the I/O (resolveProfile + resolveSurface + loadConfig), resolves the
runtime config dir via the canonical getGlobalConfigDir (--config-dir override),
and surfaces resolution failures as warnings rather than a false installed='*'.
Routed as `gsd-tools capability state`.

Additive: install/surface/workflows untouched; consumed by nothing.

Closes #945

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 16:18:23 -04:00
Colin Johnson
27aca73b6e Merge branch 'next' into kimi-runtime-support 2026-06-09 10:15:41 -04:00
Tom Boucher
86845340dc chore(#930): remove self-masking next dist-tag repoint from release finalize (#931)
* chore(#930): remove self-masking next dist-tag repoint from release finalize

The "Clean up next dist-tag" step silently failed under OIDC trusted
publishing (which can't write dist-tags) while unconditionally reporting
success via || true + an echo. It also violated the release model by
trying to repoint @next→stable; @next is managed exclusively by the rc
job's --tag next publish.

Closes #930

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs: update ADR-660 to reflect removal of next dist-tag repoint

The finalize job no longer runs `npm dist-tag add … next`; update the
ADR-660 description of step 4 to match the new behavior — @next is
managed exclusively by the rc job.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-09 09:17:00 -04:00
Tom Boucher
b866b95296 fix(#921,#922): orchestrators must not fork; plan-phase Agent gate is attempt-based (#926)
`context: fork` strips the `Agent` tool from a subagent's environment.
Spawning orchestrators (`/gsd-autonomous`, `/gsd-execute-phase`,
`/gsd-plan-phase`) depend on `Agent` to dispatch sub-agents; running
them forked silently disables the core capability they exist to provide
(#921). Remove `context: fork` from all three command frontmatter files.
`effort: xhigh` (introduced by #769) is preserved.

The `<runtime_compatibility>` Agent-availability guard added by #913 was
checking whether `Agent` was present *before* attempting the call. On
runtimes where the tool list is dynamically resolved this produced
false-negative aborts in sessions that have the tool (#922). Replace
the introspection-based pattern with an attempt-based gate: always
attempt the `Agent()` call; stop only if a real tool-unavailable error
is returned. This preserves #853's backgrounded-session close-off and
#913's intent of preventing inline role-collapse, while eliminating
false negatives.

Tests updated: enh-769-context-fork-effort.install.test.cjs asserts the
three orchestrators lack `context: fork` and that the converter still
passes the field through for non-orchestrator commands; plan-phase-drift-
guard.test.cjs adds four assertions for the attempt-based gate language;
workflow-size-budget unchanged (budgets not exceeded).

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 08:42:25 -04:00
Viktorplus
c6c3a51277 Merge branch 'next' into kimi-runtime-support 2026-06-09 13:31:21 +02:00
Tom Boucher
5670feaef5 feat(#918): loop.render-hooks resolver — consume the Capability Registry (ADR-857 phase 3c) (#920)
Add the loop.render-hooks resolver: the first registry-consuming query.
`gsd-tools loop render-hooks <point>` validates the point against the
authoritative canonical 12, reads the registry's materialized byLoopPoint
hooks, filters them by activation, and emits a JSON envelope {point,
activeHooks, rendered} with ordered markdown.

Activation resolves each hook's `when` key by precedence: loadConfig value
(post-cutover federated) -> raw config.json workstream/root single-key lookup
(pre-cutover central override) -> registry configSchema default (so a
default:true capability hook is active out-of-the-box) -> inactive. Guarded
single-value reads only (no merged object built from untrusted keys).

Registry-only: no workflow calls the resolver yet (wiring is the phase-6
cutover). Completes the phase-3 trio.

Closes #918

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 00:24:12 -04:00
Tom Boucher
6dbd895028 feat(#910): federated config merge in config-loader (ADR-857 phase 3b) (#914)
Build the federated config merge: each capability owns its config-key slice
(ADR-857 decision 3 / ADR-894), and loadConfig merges them defensively. The
registry now emits a full configSchema index ({key:{owner,type,default,
description}}, generator-validated); a new src/federated-config.cts resolves
federated keys defensively (skip central keys -> pending-migration warning,
skip malformed slices -> warning never throw, else type-checked user override
?? default, with nested dotted-path lookup and enum validation); and loadConfig
applies the overlay on every return path.

Wired as a provably-empty no-op channel: every UI-pilot key is still central,
so validKeys is empty and loadConfig returns byte-identical output on all paths
(identity return when the overlay is empty; shared CONFIG_DEFAULTS never
mutated). Registry-only; no key is cut over; nothing in the live loop changes.

Closes #910

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 23:37:41 -04:00
Colin Johnson
bf8813b7a0 Merge branch 'next' into kimi-runtime-support 2026-06-08 23:04:37 -04:00
Tom Boucher
6edc39c4eb chore(#893): remove dead loadConfig from configuration.cts (#912)
`loadConfig` in configuration.cts was superseded by config-loader.cts
(ADR-857 phase 2e, #885). Exhaustive grep confirms no caller imports
loadConfig from configuration.cjs — all live callers use config-loader.cjs
or the core.cjs back-compat re-export. configuration.cts now provides only
the pure normalization and defaults primitives (normalizeLegacyKeys,
mergeDefaults, migrateOnDisk, CONFIG_DEFAULTS) that config-loader.cts
depends on. Updated CONTEXT.md and docs/INVENTORY.md to reflect the
narrowed module surface.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 22:51:43 -04:00
Tom Boucher
48cc27bd84 feat(#903): generate Loop Host Contract from workflow markers (ADR-857 phase 3a-impl-2) (#906)
Replace the inline LOOP_HOST_CONTRACT constant in the Capability Registry
generator with a generated-from-workflows contract (ADR-894 §3). The contract
is now derived from inert `<!-- gsd:loop-host ... -->` marker blocks in the five
step workflows, emitted as the committed gsd-core/bin/lib/loop-host-contract.cjs,
and required by gen-capability-registry.cjs — one source of truth, no drift.

Drift guards in gen-loop-host-contract.cjs: per-step point ownership (each step
must declare exactly its canonical loop points), multiple-block + duplicate-key
hard errors, and a word-boundary agent-role cross-check. Contract content is
byte-identical to the former constant; registry-only, nothing wired into the
live loop.

Closes #903

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 22:16:58 -04:00
Tom Boucher
ad754ca6cd feat(#896): Capability Registry generator + UI pilot (ADR-857 phase 3a-impl) (#902)
* feat(#896): Capability Registry generator + UI pilot (ADR-857 phase 3a-impl)

First phase-3 code: the Capability Registry generation pipeline, built against
the ADR-894 contract and NOT wired into the live loop (registry-only, per the
staged-cutover design).

- capabilities/ui/capability.json — the UI pilot (ADR-894 worked example): 2
  skills, 2 agents, 3 config keys, 2 steps + 1 gate, with `when` activation.
- scripts/gen-capability-registry.cjs — --write/--check generator. Hand-rolled
  schema validation (envelope + role-typed feature/runtime bodies + typed
  steps/contributions/gates + when + gate-check variants); cross-capability
  invariants (single ownership; requires exist+acyclic+tier-monotone; config-key
  ownership exclusive, collision-vs-central as a pending-migration warning);
  hooks validated against an inline LOOP_HOST_CONTRACT (3a-impl-2 swaps its
  source to the generated-from-workflows contract); GLOBAL point-ordered
  consumes-satisfiability; materialized byLoopPoint ordering (produces/consumes
  topo-sort); emits gsd-core/bin/lib/capability-registry.cjs (role-partitioned
  indexes + requiresClosure). Prototype-pollution guards (Object.create(null) +
  inline literal key checks) + fragment.path traversal guard.
- gsd-core/bin/lib/capability-registry.cjs — committed generated artifact
  (mirrors package-identity.cjs: script-generated, tracked, linted, regenerated
  on build, drift-tested), wired via the new `gen:capability-registry` build step.
- tests/capability-registry.test.cjs — 72 tests: schema + invariant + hook +
  ordering + adversarial (path-traversal, proto-pollution, runtime body,
  self-consume, cycles, collisions) + committed-file staleness guard.

New-CLI-module checklist (INVENTORY 97->98, MANIFEST, ARCHITECTURE), CONTEXT.md
"Capability Registry" un-[Planned]'d. Nothing wired into install/surface/loop.

Gates: lint, code-review (4 bugs fixed), security-review (path-traversal +
prototype-pollution fixed), codex adversarial-review ×3 (8+ findings fixed,
confirmed sound), clean-build docker 13190 pass / 0 fail.

Closes #896

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#896): CRLF-agnostic --check for capability-registry staleness (Windows)

The committed capability-registry.cjs staleness guard failed on Windows CI only:
git checks out the committed .cjs as CRLF (autocrlf, no .gitattributes) while the
generator emits LF, so the byte-for-byte --check comparison mismatched. Normalize
line endings on both sides of the --check comparison (no .gitattributes change,
no change to the LF the generator writes). Adds a regression test simulating the
Windows CRLF checkout.

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 21:15:41 -04:00
Tom Boucher
197d6bffe2 docs(#894): ADR-894 Capability declaration format + registry generation (#895)
* docs(#894): ADR-894 Capability declaration format + registry generation

ADR-857 rollout phase 3a (design-only). Resolve ADR-857's deferred open
question — the on-disk Capability declaration format — as a reviewable design
ADR before any generator code.

Specifies: the capabilities/<id>/capability.json folder layout (migration-staged
ownership — declarations reference existing stems until the phase-6 move); the
capability.json schema for role:feature (skills/agents/hooks/federated config/
loopHooks) and role:runtime (the six closed projection-primitive axes); the 12
named Loop Extension Points; the gen-capability-registry.cjs generator design
(validation + cross-capability invariants + --write/--check drift gate, mirroring
gen-inventory-manifest); the generated capability-registry.cjs shape (by-id /
by-skill / by-loop-point indexes + requires-closure); and a full worked example
(the UI capability: ui-phase + ui-review + agents + config + two loop hooks).

No code — design artifact only; the generator build, federated config loader
(3b), and loop seam (3c) implement against this contract.

Closes #894

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#894): amend ADR-894 with grilled capability declaration format

Stress-tested the declaration format before merge; the format changed
materially. Amendments:
- loopHooks[] -> three typed arrays (steps/contributions/gates), each with its
  own shape (step: ref+produces/consumes; contribution: fragment+into agent-role;
  gate: check+blocking).
- Add the Loop Host Contract (§3): each step publishes its points, agent roles,
  and core artifacts so the generator validates hooks against reality, not
  trusted strings.
- requires = capability ids only (host implicit); add tier-monotone invariant;
  drop the requires:["plan"] error from the example.
- Config federation = atomic move: a migrated key leaves the central schema in
  the same PR; presence in both is a collision (invariant stays).
- One registry, role-partitioned indexes (feature indexes vs runtimes index).
- Rework the UI worked example to the split-array shape (2 steps + 1 gate) +
  a contribution illustration.

Adds a "Grilling amendments" section recording the six changes.

* docs(#894): amend ADR-894 with round-2 grilling (operational reality)

Second design-grill round, folded in before merge:
- Loop Host Contract is GENERATED from structured workflow markers
  (<loop-point>/<agent-role>/<loop-artifact>) via gen-loop-host-contract.cjs —
  it can't drift from the real workflows.
- Hook activation `when`: cheap deterministic config-level gating evaluated by
  loop.render-hooks; deeper phase-context applicability self-gates inside the
  dispatched skill (no phase-context vocabulary to keep honest).
- `tier` is the source of install-profile + cluster membership; profiles and
  clusters are generated from tier + requires-closure (/gsd:surface operates on
  capabilities) — collapses ADR-857's multiple toggle systems.
- Gate `check` = query | declarative-predicate | agentVerdict; agentVerdict is
  forced advisory; only deterministic checks may block.
- byLoopPoint ordering is materialized in the registry; render-hooks filters the
  active set + renders. Same-capability hooks degrade gracefully when an entry
  step self-gates.
- Rollout: registry-only until atomic per-feature cutover (no double-execution
  with still-inlined workflow features).

Updates the Grilling amendments / Consequences / Alternatives / Open questions
sections; reworks the UI example with `when`.

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 18:51:20 -04:00
Tom Boucher
185935379a refactor(#888): extract model+effort resolution into model-resolver.cts (#890)
ADR-857 rollout phase 2f — the FINAL core.cts decomposition. Move the model and
effort resolution cluster (resolveModelInternal, resolveModelPolicy,
resolveTierEntry, _resolveRuntimeTier, resolveModelForTier,
resolveGranularityInternal, assertValidGranularityOverride, resolveEffortInternal,
resolveFastModeInternal, resolveEffortForTier, nextEffort + VALID_GRANULARITIES/
VALID_EFFORTS/EFFORT_SET + interfaces) out of core.cts into a new leaf module
src/model-resolver.cts. core.cts re-exports the 13 public symbols (callers in
init/docs/commands unchanged; export= set byte-identical).

Cycle-free: model-resolver imports only leaves (config-loader for loadConfig,
configuration for defaults, model-profiles + model-catalog for the static
tables). Removed 6 now-unused imports from core (verified zero remaining
references, none re-exported).

This completes the god-module decomposition: core.cts 2271 -> 389 lines (~83%),
now a thin re-export spine over seven clean leaves (io, phase-id, roadmap-parser,
core-utils, phase-locator, config-loader, model-resolver).

New-CLI-module checklist done (.gitignore, eslint, INVENTORY 96->97 + row,
manifest, ARCHITECTURE, CONTEXT.md "Model Resolver Module"). Adds
tests/model-resolver.test.cjs (81 tests: behavioral + shim-identity + adversarial).

Gates: lint, code-review (export set byte-identical; import-removal verified),
security-review, codex adversarial-review (all 0 findings; verbatim move). Mac
4303 pass; clean-build docker 13117 pass, 0 fail.

Closes #888

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 17:26:40 -04:00
Viktorplus
69b7bbd062 Merge branch 'next' into kimi-runtime-support 2026-06-08 22:00:34 +02:00
Tom Boucher
a74e71b049 refactor(#885): extract loadConfig cluster into config-loader.cts (#886)
ADR-857 rollout phase 2e — the largest core.cts extraction. Move the
configuration-loading subsystem (loadConfig + _getConfigDefault/
_getNestedConfigDefault/CONFIG_DEFAULTS/_deepMergeConfig, isGitIgnored +
_gitIgnoredCache, _warnUnknownProfileOverrides + RUNTIME_OVERRIDE_TIERS + the
dedup Sets, _resetRuntimeWarningCacheForTests) out of core.cts into a new leaf
module src/config-loader.cts. core.cts re-exports the public surface
(loadConfig, isGitIgnored, CONFIG_DEFAULTS, RUNTIME_OVERRIDE_TIERS,
_resetRuntimeWarningCacheForTests); 12+ callers unchanged.

Cycle-free: config-loader imports only leaves (configuration, config-schema,
planning-workspace, shell-command-projection, core-utils, model-catalog). All
core-internal helpers loadConfig touches moved with it to avoid a cycle. core
keeps CANONICAL_CONFIG_DEFAULTS for its model-resolver functions, which now
resolve loadConfig via the binding — this unblocks the final model-resolver
extraction (2f). core.cts: 1275 -> 792 lines.

Repointed tests/config-field-docs.test.cjs (a docs-parity source check) to read
the CONFIG_DEFAULTS literal from its new home (config-loader.cjs). New-CLI-module
checklist done (.gitignore, eslint, INVENTORY 95->96 + row, manifest,
ARCHITECTURE, CONTEXT.md "Config Loader Module"). Adds tests/config-loader.test.cjs
(27 tests: behavioral + shim-identity + adversarial config fixtures).

Gates: lint, code-review, security-review (prototype-pollution guard confirmed
intact), codex adversarial-review (0 findings; byte-identical move). Mac 4115
pass; clean-build docker: full-suite hit the local mirror's known incremental-tsc
non-determinism on an unrelated re-exported symbol (findPhaseInternal, from
already-merged 2d), but a clean targeted rebuild of the affected file passed
163/0 — CI's clean full matrix is the authoritative gate.

Closes #885

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 15:36:38 -04:00
Viktorplus
8050516094 Merge remote-tracking branch 'origin/next' into kimi-runtime-support
# Conflicts:
#	docs/ARCHITECTURE.md
2026-06-08 21:27:05 +02:00
Tom Boucher
0a11d361ca feat(#69): nest concrete skills under namespace routers at install (#883)
Emit the 6 gsd-ns-* routers as the only top-level skill bundles and nest
the ~61 concrete skills under <router>/skills/<name>/SKILL.md on runtimes
with confirmed non-recursive skill loaders (claude global, cline, qwen,
hermes, augment, trae, antigravity). Router bodies rewrite their routing
tables from Skill-tool dispatch to a Read skills/<name>/SKILL.md pattern.
Recursive/unconfirmed loaders (cursor, codex, copilot, windsurf, codebuddy,
opencode, kilo) keep the flat layout. Completes the v1.40 namespace
architecture (#2792) so the eager skill listing drops to ~6 entries.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 15:09:34 -04:00
Tom Boucher
dd81e3d120 refactor(#881): extract phase-locator fs-search into phase-locator.cts (#882)
ADR-857 rollout phase 2d. Move the phase-directory search/location functions
(searchPhaseInDir, findPhaseInternal, getArchivedPhaseDirs) + their interfaces
(PhaseSearchResult, ArchivedPhaseDir) out of core.cts into a new module
src/phase-locator.cts. core.cts re-exports all three (callers unchanged).

Cycle-free: phase-locator depends only on leaves (phase-id for token/name
matching, core-utils for fs-scan/path helpers, planning-workspace for
planningDir) — unblocked by the core-utils leaf (2c). This completes the
phase-search split: parsing in phase-id (2a), fs-search in phase-locator.

New-CLI-module checklist done (.gitignore, eslint, INVENTORY 94->95 + row,
manifest, ARCHITECTURE, CONTEXT.md "Phase Locator Module"). Adds
tests/phase-locator.test.cjs (37 tests: behavioral + shim-identity +
adversarial phase-dir fixtures).

Gates: lint, code-review, security-review, codex adversarial-review (0
findings; verbatim move checksum-verified). Mac 4078 pass; clean-build docker
12972 pass, 0 fail.

Closes #881

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 14:38:26 -04:00
Viktorplus
87b7ab2c9f Merge branch 'next' into kimi-runtime-support 2026-06-08 20:02:57 +02:00
Viktorplus
69f6efebf1 docs(#743): address Kimi review — runtime count + support boundary
Address remaining review feedback on PR #743:

- Bump supported-runtime count 15 -> 16 across user-facing READMEs
  (docs/README.md + ja-JP/zh-CN translations) now that Kimi is added.
- Add an explicit support-boundary note in the Kimi CLI install section:
  the kimi --agent-file custom-agent contract targets legacy/Python
  kimi-cli; newer npm Kimi Code (@moonshot-ai/kimi-code) rejects
  --agent-file and consumes the same /skill:gsd-* skills via --skills-dir.
- Mirror the boundary in the changeset so release notes are unambiguous.

Historical ADR/changeset/INVENTORY references to '15 runtimes' are left
intact as point-in-time records.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 19:47:41 +02:00
Tom Boucher
282f145745 refactor(#877): extract shared low-level utilities into core-utils.cts (#878)
ADR-857 rollout phase 2c. Move 11 shared low-level utilities out of core.cts
into a new leaf module src/core-utils.cts: POSIX path normalization (toPosixPath),
filesystem scanning (detectSubRepos, readSubdirectories, getPhaseFileStats,
pathExistsInternal), and small pure helpers (generateSlugInternal,
extractOneLinerFromBody, filterPlanFiles, filterSummaryFiles, timeAgo, and the
private extractCanonicalPlanId). core.cts re-exports the 10 public ones
(callers unchanged); extractCanonicalPlanId stays private (exported from the
leaf for core's fs-search functions).

Cycle-free: core-utils depends only on Node built-ins + already-leafed modules
(phase-id for comparePhaseNum, planning-workspace for findContextMdIn). This is
the shared leaf that unblocks the phase-locator fs-search extraction (2d) —
searchPhaseInDir/findPhaseInternal/getArchivedPhaseDirs can now take their
utilities from a leaf instead of from core.

New-CLI-module checklist done (.gitignore, eslint, INVENTORY 93->94 + row,
manifest, ARCHITECTURE, CONTEXT.md "Core Utilities Module"). Adds
tests/core-utils.test.cjs (88 tests: behavioral + shim-identity + adversarial).

Gates: lint, code-review, security-review, codex adversarial-review (0
findings). Mac 4041 pass; clean-build docker 12935 pass, 0 fail.

Closes #877

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 13:29:13 -04:00
Viktorplus
4b0fcd83a0 Merge branch 'next' into kimi-runtime-support 2026-06-08 19:05:16 +02:00