* feat(#3661): make the code-review hook point configurable
Add `workflow.code_review_point` (`execute:post` default, or
`execute:wave:post`) so a multi-wave phase can run code review once per
wave instead of once at the end, scoped to what changed since the phase's
prior review.
The code-review capability now declares its step at both loop points via a
new generic `pointFrom` step field: `pointFrom` names an enum config key,
and the step is only active at its own `point` when that key resolves to
a matching value. `_resolvePointGate` (capability-activation.cts) is the
single shared implementation consumed identically by loop-resolver.cts and
capability-state.cts, and capability-validator.cjs enforces that `pointFrom`
references an enum key whose values cover the declaring step's own point.
code-review.md's manual-invocation gate now reads `workflow.code_review`
directly instead of probing registry presence at the hardcoded execute:post
point (so manual `/gsd-code-review` keeps working regardless of which
automatic point is configured), and its file-scope tiers narrow to what
changed since the phase's last review commit when one exists.
execute-phase.md's wave-post step dispatch gets a small, precedented
carve-out so the code-review skill still receives its required phase
argument when dispatched generically (caught by the isolated spec review).
Closes#3661
Emitted-Drift-Ack-Growth: code-review.md — #3661 adds a point-aware config gate check and LAST_REVIEW_COMMIT-based incremental scoping to the file-scope tiers.
Emitted-Drift-Ack-Growth: execute-phase.md — #3661 adds one carve-out sentence so the wave-post generic step dispatch passes PHASE_NUMBER to the code-review skill.
* docs: backfill changeset PR number for #3661 (#4159)
* fix: scope tests/io.test.cjs's fs.writeSync fault-injection mocks by fd
Five fault-injection mocks in the "bug #1008" describe blocks intercepted
every fs.writeSync call regardless of file descriptor, and several threw or
truncated unconditionally on the first call. This surfaced as an
intermittent macOS CI failure: node:test's own IPC channel back to the
parent process (which also goes through fs.writeSync internally) could get
a bogus injected error or truncated write if node's internal machinery
called it while one of these mocks was active, corrupting the message
frame the parent tried to deserialize ("Unable to deserialize cloned
data.", location tests/io.test.cjs:1:1, uncaughtException — a whole-file
IPC crash, not a test assertion failure).
Root cause confirmed by a working counter-example already in the same
file: the "#3912 A6" mocks gate on `fd !== 2` before any fault injection
and were never implicated. Applied the same fd-scoped pattern to the five
unscoped mocks (four output()-targeting tests gate on fd 1, one
error()-targeting test gates on fd 2), and added a regression test proving
an unrelated fd passes through untouched while the fault-injection mock is
active.
Found while verifying #3661; unrelated to that change's own diff.
---------
Co-authored-by: sim <sim@local>
* test(#3798): tiered profiles must install the agents their workflows spawn
* fix(#3798): the profile closure follows command references into workflow spawn surfaces
* chore(#3798): changeset fragment (pr number backfilled after PR creation)
* chore(#3798): backfill changeset PR number (4009)
---------
Co-authored-by: sim <sim@local>
- warn (don't silently ignore) when --runtime is an unknown runtime that
canonicalizeRuntimeName rejects; the warning surfaces via warnings[] so a
typo like --runtime cluade or a runtime known to runtime-homes but not the
alias manifest (e.g. grok) no longer silently resolves to the persisted
runtime's config dir on this diagnostic command [M-1]
- add end-to-end CLI test for loop render-hooks --runtime (the exact command
the bug report calls out as silently no-op'ing) [L-2]
- add closed-vocabulary rejection test: crafted --runtime values
(../../etc/passwd, __proto__, --config-dir, garbage) are rejected, warn,
and fall through to the persisted runtime — pins the security-load-bearing
contract [NIT-01]
- add boundary tests: --config-dir wins over --runtime (precedence); missing
--runtime value errors with USAGE [N-1]
Both orthogonal reviews returned APPROVE with no Critical/High findings.
Security review confirmed --runtime cannot coerce getGlobalConfigDir into an
arbitrary path (closed-vocabulary Map lookup + registry hash-key gate) and
does not expand the trust surface beyond the existing operator-controlled
--config-dir flag.
resolveCapabilityRuntimeState derived the config dir from resolveRuntime(cwd)
(GSD_RUNTIME -> config.runtime -> 'claude') when no --config-dir was passed,
so a repo with persisted runtime:'codex' resolved the config dir to ~/.codex
where the Claude skill isn't installed -> surfaced:false / hooks silently
no-op when the operator drove from Claude Code. capability state and loop
render-hooks parsed only --config-dir, never --runtime, so there was no way
to assert the actually-active runtime.
Add a runtimeOverride param to resolveCapabilityRuntimeState (canonicalized
via runtime-name-policy so aliases like codex-app work); when present it
short-circuits the persisted-runtime fallback and resolves getGlobalConfigDir
for the explicit runtime. Thread --runtime through cmdCapabilityState and
cmdLoopRenderHooks, and parse it in gsd-tools.cjs for both commands (dual
--runtime X / --runtime=X form, mirroring --config-dir and the existing
capability-set --runtime precedent). Help text updated.
Without the override, behavior is byte-identical to today (regression-guarded).
- restructure _loadFlatCommandsGsdManifest try/catch to wrap read+parse+set
together (mirrors loadSkillsManifest exactly), so a thrown parser degrades
both keys to [] — closes the latent catch-scope parity drift [Nit-1]
- add boundary test: gsd-.md (empty stem) skipped, gsd-x.md single-char stem
kept (slice(4,-3) boundary) [Low-1]
- add unreadable-file test (POSIX-gated): both keys degrade to [] [Low-2]
- strengthen parity test to also compare requires + _calls_agents_ VALUES,
not just the stem set [Nit-2]
Both orthogonal reviews returned APPROVE with no Critical/High findings.
Security review confirmed no new trust-boundary crossing, prototype-pollution
immune (Map/Set throughout), and symlink/path-traversal surface identical to
the pre-existing nested loader (not a regression).
_resolveManifest only recognized the nested source layout (commands/gsd/*.md)
and the installed-runtime skills layout (skills/gsd-<stem>/SKILL.md). A flat
source install (Claude local project shape: commands/gsd-<stem>.md, no
commands/gsd/ subdir) matched neither branch, so the manifest came back empty
and resolveSurface materialized the full profile to an empty Set — silently
reporting every skill-bearing capability as surfaced:false / enabled:false /
active:false. The nyquist/code-review/security/ui verify:post and execute:post
hooks never fired even with their workflow.* toggles on.
Add a third branch: when commandsGsdDir is absent, scan dirname(commandsGsdDir)
for gsd-<stem>.md files, strip the gsd- prefix, and build the same Map shape
the nested loader produces (requires via shared parseRequires, companion
_calls_agents_<stem> via shared parseCallsAgents). Falls through to the
installed-skills branch when the flat dir has no gsd-*.md files (precedence:
nested > flat-source > installed).
Also export parseCallsAgents from install-profiles so capability-state reuses
the SAME parser the nested loader uses (no drift; mirrors the existing
parseRequires export+reuse pattern).
* feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2)
Promote the registry from a frozen data file to loadRegistry({includeInstalled}),
composing the first-party registry with a validated installed overlay (ADR-1244 D2):
- Extract the conformance validator to a shared runtime-callable module
(gsd-core/bin/lib/capability-validator.cjs); the generator re-exports it
verbatim, guarded by a generative-parity test (no build-time/runtime drift).
- capability-loader.cts: loadRegistry({includeInstalled}) composes first-party
∪ validated overlay from $GSD_HOME/.gsd/capabilities (global) and
<root>/.gsd/capabilities (project) via the canonical buildRegistry. First-party
always wins (id/skill/agent/config/command-family + reserved gsd-/anthropic-
prefixes); full merged-set cross-capability validation; engines.gsd load-time
re-gate (skip-with-warning); gate-kind capabilities FAIL CLOSED; fragment-path
escapes rejected.
- semverSatisfies (hand-written, no dep) for the engines.gsd gate, fail-closed.
- Wire surface/state + loop to the overlay; loop injects a blocking gate for each
skipped gate-kind overlay (fail-closed).
- cwd-aware overlay config-key federation: config-loader _federatedConfigSchema(cwd)
+ config-schema isValidConfigKey(key, cwd) compose the overlay per loadConfig/
config-set call (never eager at module load, never wrong-cwd); first-party path
unchanged with no cwd.
- run-tests.cjs sandboxes GSD_HOME (idempotent — nested spawns reuse it) for test
hermeticity; capability-loader.cjs git+eslint-ignored (tsc artifact);
capability-validator.cjs stays linted (#551 migration coverage).
Closes#1431
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1431): add changeset for runtime capability registry overlay
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1431): kill config-schema cwd-aware federation mutants (Stryker ≥52)
The cwd-aware overlay config-key federation added to config-schema.cts
(_capabilityConfigSchema(cwd) + isCapabilityConfigKey/isValidConfigKey cwd
threading) introduced mutable surface uncovered by config-schema's mutation
test set, dropping its score to 39.58% (below the 52 break threshold). Add a
real-overlay-fixture describe block exercising every branch (cwd guard, overlay
loadRegistry, found-branch, first-party fallback, cwd threading); local Stryker
score 39.58% -> 77.08%.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Narrows P3 of ADR-1411 (Resolution Provenance, epic #1411) based on an
adversarial fit-analysis that showed a single Resolution<T> envelope adopted
by agent-skills, capability-state, and capability-writer fails the deletion
test: configured/reason are meaningless for capability verbs, and
capability-writer's errors[] (operation-not-applied) cannot fold into
warnings[]. The only genuinely shared seam is warnings: string[].
Changes:
- src/resolution.cts: new pure types+builder leaf — exports Resolution<T>
{value, configured, reason, warnings}, makeResolution<T>() builder, and
AgentSkillsValue {block, skills_count}. No other src/ imports.
- src/init.cts: cmdAgentSkills --json IR gains additive value:{block,
skills_count} field (built via makeResolution). All existing flat fields
(agent_type, block, skills_count, warnings, configured, reason, source,
degraded) are retained unchanged for back-compat.
- src/capability-state.cts: doc comment on ResolveCapabilityRuntimeStateResult
naming it the canonical read-verb envelope. No JSON change.
- src/capability-writer.cts: doc comment on SetCapabilityStateResult naming it
the canonical mutation-verb result (warnings=advisory, errors=operation-
not-applied). No JSON change.
- CONTEXT.md: new ### Resolution Convention glossary entry after
### Resolution Provenance.
- docs/adr/1411-resolution-provenance.md: P3 narrowing amendment appended.
- tests/resolution.test.cjs: 9 unit tests for makeResolution (new).
- tests/agent-skills.test.cjs: 2 P3 tests for value.block/value.skills_count
and back-compat of all flat fields.
All 277 tests pass (5 suites). npm run lint clean. All lint checks pass.
Part of #1411
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Make src/capability-activation.cts the sole owner of the four-level config-key
precedence walk via a new raw-value primitive resolveConfigKey(dotKey,{config,
cwd,registry}); the boolean wrapper _resolveActivationValue and loop-resolver's
resolveConfigValues are both rebuilt on it. loop-resolver deletes its
byte-identical copy and inline re-walk and imports the engine.
resolveCapabilityRuntimeState no longer returns registry/config (leaked internal
detail); the caller loads one fail-closed config snapshot and threads it in via a
new optional configOverride param, so capability `active` and hook when/configValues
resolve against the same object. capability-writer requires the registry module
directly.
Adds a DEFECT.GENERATIVE-FIX parity gate (identity + behavioral matrix +
end-to-end resolveLoopHooks configValues) that fails if the two precedence
surfaces ever diverge. Pure internal refactor, no user-facing change.
Closes#1308
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(#1306): gate graphify on isCapabilityActive (tri-state), not config-only
graphify's command gate moves from the config-only isGraphifyEnabled to the
shared isCapabilityActive('graphify', cwd) — so graphify is off unless installed
AND surfaced AND graphify.enabled. Fixes a latent claude-hardcoding in the
resolver: resolveCapabilityRuntimeState now detects the active runtime via
resolveRuntime(cwd) (GSD_RUNTIME -> config.runtime -> 'claude') so non-Claude
runtimes (Codex/Cursor) read their own surface, not ~/.claude. Hermetic
regression test proves config-on+unsurfaced -> disabled; cross-runtime test
proves GSD_RUNTIME=codex honors CODEX_HOME. Gate fails closed on every error
path. Part of #1302.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1306): add changeset for graphify tri-state gate
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#1305): add per-capability active tri-state + isCapabilityActive to the Capability State Resolver
CapabilityStateEntry gains active = enabled && configActivation, where
configActivation resolves the capability's optional activationKey via the
shared _resolveActivationValue (absent activationKey -> true). enabled stays
installed && surfaced (unchanged). Each hook's active now also cascades the
capability config gate (active && configured). Adds isCapabilityActive(capId,
cwd) — a thin convenience over resolveCapabilityRuntimeState. cmdCapabilityState
emits active per capability. No consumer cutover yet (graphify/intel: #1306/#1307;
loop-resolver: #1310). Part of #1302.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1305): add changeset for capability active tri-state
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Per the T1 design rubber-duck, batch by FILE so each tranche drops
convergence-lint allowlist entries. Migrate 12 files' core imports to the
leaf modules directly (behaviour-identical — leaves are the objects core
re-exports by reference):
- io (output/error/ERROR_REASON): agent-command-router, capability-state,
capability-writer, frontmatter, gsd2-import, learnings, loop-resolver,
task-command-router
- roadmap-command-router -> config-loader; workstream-inventory -> core-utils
- milestone, verify -> their full leaf sets (both were multi-leaf, not
single-leaf as first scoped; migrated completely)
All 12 files now import zero core symbols and are removed from the
allowlist (30 -> 18). core.cts re-exports untouched (still serve the
remaining 18 files); teardown is T-final. Stale core.cjs docstrings in the
migrated files corrected to reference io.cjs. No behaviour change.
Closes#1281
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1160): resolve capability surface from installed skill layouts
In a global skills-runtime install (e.g. Codex at ~/.codex), gsd-tools.cjs
runs from <configDir>/gsd-core/bin/ and the commands/gsd source tree is
absent — only <configDir>/skills/gsd-<stem>/SKILL.md files exist.
_resolveCommandsGsdDir() returned a path that does not exist there, so
loadSkillsManifest returned an empty Map. resolveSurface then materialised
the '*' (full) profile sentinel by enumerating that empty manifest → empty
surfaced Set → every capability reported surfaced=false/enabled=false
regardless of project config. As a result `loop render-hooks verify:post`
returned activeHooks:[] even with workflow.security_enforcement and
workflow.nyquist_validation enabled, silently disabling the security and
Nyquist gates.
Fix: add _loadInstalledSkillsManifest(configDir) that scans configDir/skills/
for gsd-<stem>/SKILL.md dirs and builds the same Map shape, and
_resolveManifest(commandsGsdDir, configDir) that prefers the source tree when
present (preserving repo-checkout behaviour) and falls back to the installed
skills layout otherwise. Both resolveCapabilityRuntimeState call sites use
_resolveManifest. Both helpers are exported for direct unit-testing.
Tests: capability-state.test.cjs gains a faithful installed-runtime e2e block
that copies gsd-core/bin + scripts + package.json into a temp install root
with no reachable commands/gsd, then runs the real gsd-tools.cjs against an
installed skills/ layout. It asserts capability state reports security &
nyquist enabled and verify:post includes security->secure-phase and
nyquist->validate-phase; a disabled-config negative confirms no
over-activation. This block FAILS before the fix (activeHooks:[]) and PASSES
after. Plus unit coverage for the two new helpers and the empty-surface
pre-fix scenario.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(changeset): backfill PR number
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The render-hooks resolver and capability-state resolver headers (and their CONTEXT.md glossary entries) still claimed "no workflow calls this yet" / "consumed by nothing yet (phase-6 wiring is out of scope)". Both are now false: loop.render-hooks is consumed live by plan-phase.md/autonomous.md (plan:pre ui-phase, verify:post ui-review), and capability-state is the `gsd-tools capability state` CLI diagnostic.
Comment/prose only — no behavior change. Surfaced by the ADR-857 phase 1-5 completeness audit.
Closes#1125
Refs #857
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Make install + surface read the registry's derived profileMembership/
capabilityClusters so a capability's tier drives what installs + surfaces.
resolveProfile (when given the registry) unions capability skills for the
profiles its tier implies before the requires: closure; resolveSurface merges
capabilityClusters into the cluster map. bin/install.js, /gsd:surface, and the
capability-state resolver all thread the registry.
Shipped as a proven no-op: the UI capability is reconciled to tier:full (its
skills were full-only in the hand-authored profiles), so it contributes only to
the full profile (already the '*' sentinel) and core/standard are unchanged.
Equivalence tests prove resolveProfile/resolveSurface/listSurface/staging/
capability-state are identical with vs without the registry; the core-alias
staging path is verified equivalent (empty manifest → raw PROFILES.core).
Closes#949
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Add a read-side query composing the three toggle systems into one
per-capability view. resolveCapabilityState({registry, installedSkills,
surfacedSkills, config, cwd}) reports installed (skills ⊆ resolved install
profile), surfaced (skills ⊆ resolved surface), and per-hook active (no when →
active; non-empty-string when → resolved via _resolveActivationValue; empty/
non-string → inactive), with no forced composite verdict. cmdCapabilityState
does the I/O (resolveProfile + resolveSurface + loadConfig), resolves the
runtime config dir via the canonical getGlobalConfigDir (--config-dir override),
and surfaces resolution failures as warnings rather than a false installed='*'.
Routed as `gsd-tools capability state`.
Additive: install/surface/workflows untouched; consumed by nothing.
Closes#945
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>