diff --git a/.changeset/1278-prohibition-check-descriptor.md b/.changeset/1278-prohibition-check-descriptor.md new file mode 100644 index 000000000..d380cb9b7 --- /dev/null +++ b/.changeset/1278-prohibition-check-descriptor.md @@ -0,0 +1,5 @@ +--- +type: Changed +pr: 1301 +--- +**The test-tier prohibition gate now has a deterministic SOURCE for its wired check** — a resolved `test`-tier `must_haves.prohibitions` item MAY carry an optional `check` descriptor authored at spec-phase: the flat-scalar keys `check_kind` (`node-test` | `lint-rule`), `check_target`, and `check_rule` (lint-rule only). `projectProhibitions` projects these scalars deterministically and verify-phase reads them back (via `descriptorFromProjection`) to locate the check handed to `check prohibition-enforcement` — so a wired, passing test closes the gap with **zero manual descriptor authoring** (previously the verify-phase LLM had to invent `{kind, target, rule}` each run, #1259). This extends the ADR-550 Decision 3 prohibition-item shape (ratified in a dated 2026-06-15 ADR-550 addendum). The descriptor is **optional and fully backward-compatible** — a prohibition with no descriptor parses and disposes byte-identically to today — and **fail-closed**: a partial, invalid, or absent descriptor falls through to the producer's existing fail-closed locate, never a silent green. The descriptor is represented as flat scalars (not a nested `check:{}` object) to keep the shared `parseMustHavesBlock` round-trip regression-free. Out of scope: machine-proven fail-first (#1279) and the `dispositionForProhibition` policy stay unchanged. (#1278) diff --git a/.changeset/nimble-newts-rally.md b/.changeset/nimble-newts-rally.md new file mode 100644 index 000000000..ecca6fc3e --- /dev/null +++ b/.changeset/nimble-newts-rally.md @@ -0,0 +1,5 @@ +--- +type: Changed +pr: 1313 +--- +**Graphify now respects surface/profile state, not just `graphify.enabled`** — `gsd-tools graphify` is off unless graphify is installed AND surfaced AND `graphify.enabled` is true (previously only the config key was checked). The gate is now runtime-aware: Codex/Cursor/etc. read their own runtime's surface instead of `~/.claude`. (#1313) diff --git a/.changeset/quick-yaks-march.md b/.changeset/quick-yaks-march.md new file mode 100644 index 000000000..7d66f985e --- /dev/null +++ b/.changeset/quick-yaks-march.md @@ -0,0 +1,5 @@ +--- +type: Changed +pr: 1315 +--- +**Intel and loop-hook rendering now honor the single capability `active` state** — `gsd-tools intel` gates through the shared resolver (consistency; intel stays governed by `intel.enabled`), and loop-hook rendering now suppresses a config-disabled capability's hooks via the capability-level `active` gate (fail-closed), not just per-hook `when`. (#1315) diff --git a/.changeset/sharp-geese-glide.md b/.changeset/sharp-geese-glide.md new file mode 100644 index 000000000..5d207cb60 --- /dev/null +++ b/.changeset/sharp-geese-glide.md @@ -0,0 +1,5 @@ +--- +type: Changed +pr: 1311 +--- +**Capability state now reports a tri-state `active`** — `gsd-tools capability state` adds an `active` field per capability (installed && surfaced && config-enabled), alongside the existing `enabled` (installed && surfaced). Internal `isCapabilityActive(capId, cwd)` lets consumers honor the single resolved on/off answer. (#1311) diff --git a/CONTEXT.md b/CONTEXT.md index de55b1fab..4078f1bb2 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -170,7 +170,7 @@ ADR-857 phase 3b seam that merges capability-declared config slices into the `lo A named, stable site on a host loop step (per-step `pre`/`post` plus per-wave in Execute; 12 total) where Capabilities register hooks. Three hook kinds: `step` (runs as its own sequenced unit), `contribution` (injects into the core step's prompt/context), and `gate` (checks and optionally blocks via a declared `blocking` flag). Each hook declares the artifacts it produces and consumes; hook order is derived by topological sort of that produces/consumes graph (capability-id tiebreak), which also defines data flow — file-artifact based, surviving `/clear` and fresh executor contexts. Hooks are surfaced by runtime resolution with concrete projection: the workflow calls a query that resolves the active hooks and returns fully-rendered, ordered markdown for the executor. Failure is default-resilient — a non-gate hook that errors is skipped with a warning; a hook may opt into `onError: halt`. Part of the Capability system. ADR-857 phase 3c ships the registry-consuming query layer: `gsd-core/bin/lib/loop-resolver.cjs` exposes `resolveLoopHooks({ point, registry, config })` (pure, no I/O), `renderLoopHooks(resolved)` (pure markdown renderer), and `cmdLoopRenderHooks(cwd, point, raw, opts)` (I/O entry point); activated via `gsd-tools loop render-hooks ` which emits `{ point, activeHooks[], rendered }`. Activation is driven by `when` (dotted config key resolved against `loadConfig`), with inline literal `__proto__`/`constructor`/`prototype` prototype-pollution guard. The first phase-6 cutovers wiring workflows to this query have landed — ui-phase at `plan:pre` and ui-review at `verify:post` (in `plan-phase.md`/`autonomous.md`); further per-feature cutovers are ongoing. ### Capability State Resolver -ADR-857 phase 4b/6 unified resolver that composes the three toggle systems (install profile, runtime surface, config activation) into one per-capability view consumed by workflow hook rendering. Source of truth: `gsd-core/bin/lib/capability-state.cjs` (generated from `src/capability-state.cts`). Interface: `resolveCapabilityState({ registry, installedSkills, surfacedSkills, config, cwd? }) → { capabilities: CapabilityStateEntry[] }` (pure, no I/O); `resolveCapabilityRuntimeState(cwd, runtimeConfigDir)` (I/O resolver shared by workflow dispatch and diagnostics); `cmdCapabilityState(cwd, runtimeConfigDir, raw, opts)` (I/O output entry point). CLI surface: `gsd-tools capability state [--config-dir ]` — emits `{ runtimeConfigDir, capabilities[] }`. Per-capability output: `{ id, tier, skills[], installed, surfaced, enabled, hooks[] }` where `installed` = every owned skill ∈ installedSkills (or `installedSkills==='*'`; vacuously true for empty-skills caps), `surfaced` = every owned skill ∈ surfacedSkills (vacuously true for empty-skills caps), `enabled = installed && surfaced`, and `hooks` = `[{ point, kind: 'step'|'gate'|'contribution', when, configured, active }]` derived from the cap's `steps`, `gates`, `contributions` arrays (`configured` resolves `when`; `active = enabled && configured`). Capabilities sorted by `id` for determinism. Defensive: malformed registry → `{ capabilities: [] }`, never throws; inline literal `__proto__`/`constructor`/`prototype` prototype-pollution guard on capability id keys. +ADR-857 phase 4b/6 unified resolver that composes the three toggle systems (install profile, runtime surface, config activation) into one per-capability view consumed by workflow hook rendering. Source of truth: `gsd-core/bin/lib/capability-state.cjs` (generated from `src/capability-state.cts`). Interface: `resolveCapabilityState({ registry, installedSkills, surfacedSkills, config, cwd? }) → { capabilities: CapabilityStateEntry[] }` (pure, no I/O); `resolveCapabilityRuntimeState(cwd, runtimeConfigDir)` (I/O resolver shared by workflow dispatch and diagnostics); `cmdCapabilityState(cwd, runtimeConfigDir, raw, opts)` (I/O output entry point); `isCapabilityActive(capId, cwd): boolean` (convenience predicate — calls `resolveCapabilityRuntimeState`, finds the entry, returns `entry.active`; `false` when capability not found). CLI surface: `gsd-tools capability state [--config-dir ]` — emits `{ runtimeConfigDir, capabilities[] }`. Per-capability output: `{ id, tier, skills[], installed, surfaced, enabled, active, hooks[] }` where `installed` = every owned skill ∈ installedSkills (or `installedSkills==='*'`; vacuously true for empty-skills caps), `surfaced` = every owned skill ∈ surfacedSkills (vacuously true for empty-skills caps), `enabled = installed && surfaced` (unchanged — install+surface toggle only), `active = enabled && configActivation` (tri-state deepening: configActivation resolves the capability's `activationKey` via `_resolveActivationValue`; absent `activationKey` → configActivation=true, so `active===enabled` for ungated capabilities), and `hooks` = `[{ point, kind: 'step'|'gate'|'contribution', when, configured, active }]` derived from the cap's `steps`, `gates`, `contributions` arrays (`configured` resolves `when`; `active = enabled && configured`). Capabilities sorted by `id` for determinism. Defensive: malformed registry → `{ capabilities: [] }`, never throws; inline literal `__proto__`/`constructor`/`prototype` prototype-pollution guard on capability id keys. ### Capability State Writer The write-side mirror of the Capability State Resolver. Takes a desired capability state — per-capability `enabled` plus per-hook `gates` — and projects it onto the substrates: `enabled` drives the runtime surface (`.gsd-surface.json`) as the capability on/off switch; `gates` drive the federated config keys (`config.json` `workflow.*`) for hook-level granularity; the install profile (`.gsd-profile`) is a read-only floor it never writes. Writes the surface once and config once (atomic per substrate), then re-runs the resolver and reports divergence (assert-and-report) — so 'off means off' holds as a write-time invariant rather than by caller discipline. Source of truth: `src/capability-writer.cts`; the surface and config writers become its internal adapters. @@ -221,7 +221,10 @@ Module owning the projection of gsd-core's artifact surfaces (`commands`, `agent The repo-root `gemini-extension.json` + `GEMINI.md` pair that projects gsd-core onto the Gemini CLI extension contract, enabling one-step lifecycle management via `gemini extensions install ` / `update` / `remove` (and `gemini extensions link ` for dev). The Gemini-CLI sibling of the Claude Code Plugin Manifest Module — same additive idea, different runtime package format. Defined mapping: `name`=`binName` (`gsd-core`; lowercase-dashes per Gemini's extension naming rule), `version` tracks `package.json` (Gemini's `gemini extensions update` keys off the manifest `version` field), `description` (required by the manifest schema), `contextFileName`=`GEMINI.md` (the extension's context payload, loaded into every Gemini session). Intentionally minimal: no `mcpServers` (gsd-core ships no MCP server). Slash-command / agent / hook projection into the extension (which would require committing the Gemini-format TOML/agent conversions the Installer Module produces at `--gemini` install time) is deferred — the manual `npx gsd-core --gemini` path remains the way to install the `/gsd:*` commands, and is unchanged (additive, no breaking change). Conformance is guarded by the in-repo drift test `tests/issue-775-gemini-extension.test.cjs` (manifest validity, `version`↔`package.json` parity, `contextFileName` existence, `files[]` publication). _Avoid_: "the Gemini plugin" (Gemini calls them extensions, not plugins). See #775, ADR-766, Claude Code Plugin Manifest Module, and Runtime Artifact Layout Module. ### Knowledge Graph Module -Module owning the graphify integration: config gate (`isGraphifyEnabled`), disabled response (`disabledResponse`), subprocess helper (`execGraphify`, typed `GRAPHIFY_REASON` enum), presence detection (`checkGraphifyInstalled`), version checking (`checkGraphifyVersion`), query surface (`graphifyQuery` — BFS seed-expand + budget trim), status surface (`graphifyStatus` — node/edge counts, mtime staleness, commit-staleness tri-state via `built_at_commit`/`commits_behind`/`commit_stale`), diff surface (`graphifyDiff` — added/removed/changed nodes+edges), build pre-flight (`graphifyBuild`), snapshot management (`writeSnapshot`). Reads `.planning/config.json:graphify.enabled` as config gate; writes to `.planning/graphs/`. Auto-update hook (`hooks/gsd-graphify-update.sh`) triggers a detached background rebuild after HEAD-advancing git operations on the default branch when `graphify.auto_update=true`. Status file `.planning/graphs/.last-build-status.json` carries `{ ts, status, exit_code, duration_ms, head_at_build, graphify_version }`. Graph IR uses `nodes[]`, `edges[]` (or `links[]` for graphify ≥0.7 compat), `hyperedges[]`, `built_at_commit`. `commit_stale` is tri-state: `false` (known fresh), `true` (stale), `null` (unknown — no git or pre-v0.7 graph). Source: `gsd-core/bin/lib/graphify.cjs`. Skill: `commands/gsd/graphify.md`. +Module owning the graphify integration: tri-state capability gate (`isCapabilityActive('graphify', cwd)` from capability-state.cjs — requires installed AND surfaced AND config-enabled; replaces the former config-only `isGraphifyEnabled` gate, cutover in #1306), disabled response (`disabledResponse`), subprocess helper (`execGraphify`, typed `GRAPHIFY_REASON` enum), presence detection (`checkGraphifyInstalled`), version checking (`checkGraphifyVersion`), query surface (`graphifyQuery` — BFS seed-expand + budget trim), status surface (`graphifyStatus` — node/edge counts, mtime staleness, commit-staleness tri-state via `built_at_commit`/`commits_behind`/`commit_stale`), diff surface (`graphifyDiff` — added/removed/changed nodes+edges), build pre-flight (`graphifyBuild`), snapshot management (`writeSnapshot`). Config leg reads `.planning/config.json:graphify.enabled`; all three legs (install, surface, config) must be active; writes to `.planning/graphs/`. Auto-update hook (`hooks/gsd-graphify-update.sh`) triggers a detached background rebuild after HEAD-advancing git operations on the default branch when `graphify.auto_update=true`. Status file `.planning/graphs/.last-build-status.json` carries `{ ts, status, exit_code, duration_ms, head_at_build, graphify_version }`. Graph IR uses `nodes[]`, `edges[]` (or `links[]` for graphify ≥0.7 compat), `hyperedges[]`, `built_at_commit`. `commit_stale` is tri-state: `false` (known fresh), `true` (stale), `null` (unknown — no git or pre-v0.7 graph). Source: `gsd-core/bin/lib/graphify.cjs`. Skill: `commands/gsd/graphify.md`. + +### Intel Module +Module owning the code-intelligence store: tri-state capability gate (`isCapabilityActive('intel', cwd)` from capability-state.cjs — honours installed+surfaced+config-enabled; replaces the former config-only `isIntelEnabled` gate, cutover in #1307; intel has `skills:[]` so installed/surfaced are vacuously true and the effective gate is `intel.enabled` in config), disabled response, query surface (`intelQuery` — full-text search across all intel JSON files), status surface (`intelStatus` — per-file freshness, 24-hour staleness threshold), diff surface (`intelDiff` — added/changed/removed files vs last-refresh snapshot), snapshot management (`saveRefreshSnapshot`/`intelSnapshot`), validation (`intelValidate` — existence, JSON validity, _meta.updated_at recency), api-surface render (`intelApiSurface` — generates `.planning/intel/API-SURFACE.md` from `api-map.json`), plus ungated utilities (`intelPatchMeta` — patches `_meta.updated_at` in any JSON file; `intelExtractExports` — extracts CJS/ESM exports from any JS file). Loop hook rendering gates on `state.active` (not `state.enabled`) so the `activationKey` config gate is honoured even without a per-hook `when` guard (Phase 4 tri-state alignment, #1307). Source: `gsd-core/bin/lib/intel.cjs` (generated from `src/intel.cts`). Router: `gsd-core/bin/lib/intel-command-router.cjs`. See Capability Command Family Module (ADR-959 4d-impl-4) and Loop Extension Point. ### Research Module The GSD-RESEARCH capability behind an L2-hybrid seam: code owns cache + provider policy + package legitimacy; MCP owns the actual fetch. Reachable via `gsd-tools query research-plan|research-store|package-legitimacy`. Source: `src/research-{store,provider}.cts` + `src/package-legitimacy.cts` (generated to `gsd-core/bin/lib/*.cjs` per ADR-457). Replaces the prose provider-waterfall duplicated across the researcher agents and the pip-install `slopcheck` bolt-on. @@ -256,7 +259,7 @@ The orthogonal `verification` dimension a resolved probe item carries alongside The ownership seam between the prohibition probe and security/compliance tooling (ADR-550 D6). The probe owns **bespoke** product/values prohibitions — the unwritten must-NOTs specific to this feature's intent (e.g. "the streak reminder must not manipulate the user into returning"). When precision classifies an item as a **canon** security/compliance concern (OWASP / GDPR / fairness / prototype-pollution / path-traversal — the codified, cross-project rule sets), the probe does **not** mint a SPEC prohibition: it emits a one-line breadcrumb (*"possible canon-security concern X — owned by `/gsd:secure-phase` / eslint"*) and stops. Canon checks are **referred, not duplicated** — keeping the surfaced list short (#644's ~2–3-item precision goal) and the secure-phase boundary explicit. See Prohibition Probe Module. ### Prohibition Probe Module -Second adapter of the Probe Core Module (ADR-550 Decision 7): the spec-phase prohibition-completeness probe wired into spec-phase Step 5.6, surfacing the unwritten *must-NOT* constraints (values/safety/ethics) the spec never forbids. Unlike the Edge Probe, recall is **prose-orchestrated, not a compiled engine** (ADR-550 D7b) — a two-stage pass per requirement: Stage 1 an adversarial recall question, Stage 2 a one-pass precision classifier (drop routine engineering, keep genuine prohibitions). The code surface is schema/projection only: `projectProhibitions()` (deterministic SPEC↔`must_haves.prohibitions` projection backing the `DEFECT.GENERATIVE-FIX` parity assertion), the `{test, judgment}` `PROHIBITION_VALIDATORS`, `validateProhibitionResolution`, and `dispositionForProhibition()` (the fail-closed default — an unwired `test`-tier item resolves to `unverified`/flagged, never green). No `proposeProhibitions()` — recall is LLM prose. `plan-phase` lifts every resolved prohibition from the SPEC `## Prohibitions (must-NOT)` section into the `must_haves.prohibitions` sibling block (never `truths`). Exports (locked surface): `projectProhibitions`, `PROHIBITION_VALIDATORS` (the `{test, judgment}` validators bundle injected into `probe-core`'s generic engine), `validateProhibitionResolution`, and `dispositionForProhibition` (the fail-closed disposition) — the prohibition adapter surface, shipped from `probe-core` alongside the generic engine. Source of truth: `gsd-core/bin/lib/probe-core.cjs` (the prohibition exports live in `src/probe-core.cts`, gitignored per ADR-457) + `gsd-core/references/prohibition-probe.md`. Tests: `tests/prohibition-probe.*.test.cjs`. See ADR-550, Probe Core Module, Edge Probe Module, Verification Tier, Bespoke vs Canon Prohibition. +Second adapter of the Probe Core Module (ADR-550 Decision 7): the spec-phase prohibition-completeness probe wired into spec-phase Step 5.6, surfacing the unwritten *must-NOT* constraints (values/safety/ethics) the spec never forbids. Unlike the Edge Probe, recall is **prose-orchestrated, not a compiled engine** (ADR-550 D7b) — a two-stage pass per requirement: Stage 1 an adversarial recall question, Stage 2 a one-pass precision classifier (drop routine engineering, keep genuine prohibitions). The code surface is schema/projection only: `projectProhibitions()` (deterministic SPEC↔`must_haves.prohibitions` projection backing the `DEFECT.GENERATIVE-FIX` parity assertion), the `{test, judgment}` `PROHIBITION_VALIDATORS`, `validateProhibitionResolution`, and `dispositionForProhibition()` (the fail-closed default — an unwired `test`-tier item resolves to `unverified`/flagged, never green). **Deterministic test-tier locate (#1278, ADR-550 D3 addendum):** a resolved `test`-tier prohibition MAY carry an optional flat-scalar `check` descriptor — `check_kind` (`node-test` | `lint-rule`), `check_target`, and `check_rule` (lint-rule only) — that `projectProhibitions` emits into `must_haves.prohibitions` when well-formed, and `descriptorFromProjection()` (the read-back seam in the #1259 enforcement producer, `src/prohibition-enforcement.cts`) reconstructs into a `{kind, target, rule?}` `CheckDescriptor`, so verify-phase locates the wired check with zero LLM/author authoring. **Flat scalars, never a nested `check:{}` object** — so the round-trip rides the *unchanged* shared `parseMustHavesBlock` (the #644 no-parser-rewrite precedent); an absent/partial descriptor falls through to the producer's existing fail-closed locate, and `failFirst` stays caller-attested (machine-proof is #1279). No `proposeProhibitions()` — recall is LLM prose. `plan-phase` lifts every resolved prohibition from the SPEC `## Prohibitions (must-NOT)` section into the `must_haves.prohibitions` sibling block (never `truths`). Exports (locked surface): `projectProhibitions`, `PROHIBITION_VALIDATORS` (the `{test, judgment}` validators bundle injected into `probe-core`'s generic engine), `validateProhibitionResolution`, and `dispositionForProhibition` (the fail-closed disposition) — the prohibition adapter surface, shipped from `probe-core` alongside the generic engine. Source of truth: `gsd-core/bin/lib/probe-core.cjs` (the prohibition exports live in `src/probe-core.cts`, gitignored per ADR-457) + `gsd-core/references/prohibition-probe.md`. Tests: `tests/prohibition-probe.*.test.cjs`. See ADR-550, Probe Core Module, Edge Probe Module, Verification Tier, Bespoke vs Canon Prohibition. ### MVP Mode Phase-level planning mode that frames work as a vertical slice (UI → API → DB) of one user-visible capability instead of horizontal layers. Resolved at workflow init via the precedence chain: `--mvp` CLI flag → ROADMAP.md `**Mode:** mvp` field → `workflow.mvp_mode` config → false. All-or-nothing per phase (PRD #2826 Q1). Surfaced as `MVP_MODE=true|false` to the planner, executor, verifier, and discovery surfaces (progress, stats, graphify). Canonical parser: `roadmap.cjs` `**Mode:**` field; canonical resolution chain documented in `workflows/plan-phase.md`. Concept index: `references/mvp-concepts.md`. diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index d9607c92a..8282310de 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -375,7 +375,7 @@ Node.js CLI utility (`gsd-tools.cjs`) with domain modules split across `gsd-core | `loop-host-contract.cjs` | Generated Loop Host Contract — 12 loop points, per-step agent roles, and core artifacts; emitted by `scripts/gen-loop-host-contract.cjs` from workflow markers (ADR-894 §3); consumed by `gen-capability-registry.cjs` | | `capability-registry.cjs` | Generated central Capability Registry — role-partitioned index of all co-located capability declarations; emitted by `scripts/gen-capability-registry.cjs` (ADR-894 §5) | | `loop-resolver.cjs` | Loop Extension Point resolver — ADR-857 phase 3c registry-consuming query; consumes resolved Capability State, filters `byLoopPoint` by capability enablement plus config activation, renders active hooks as markdown, emits `{ point, activeHooks, rendered }` envelope; `gsd-tools loop render-hooks [--config-dir ]` | -| `capability-state.cjs` | Unified capability-state resolver — ADR-857 phase 4b/6; composes install profile, runtime surface, and config activation into one per-capability view consumed by workflow hook rendering; pure `resolveCapabilityState`, reusable `resolveCapabilityRuntimeState`, and I/O `cmdCapabilityState`; `gsd-tools capability state [--config-dir ]` | +| `capability-state.cjs` | Unified capability-state resolver — ADR-857 phase 4b/6; composes install profile, runtime surface, and config activation into one per-capability view consumed by workflow hook rendering; pure `resolveCapabilityState`, reusable `resolveCapabilityRuntimeState`, I/O `cmdCapabilityState`, and convenience predicate `isCapabilityActive(capId, cwd)`; `gsd-tools capability state [--config-dir ]` emits `{ runtimeConfigDir, capabilities[] }` where each entry carries `enabled` (installed && surfaced) and `active` (enabled && configActivation via the capability's `activationKey`; absent key → active===enabled) | | `graphify-command-router.cjs` | ADR-959 capability command router — first real capability command cutover (phase 4d-impl-2); extracted from the `case 'graphify':` arm in `gsd-tools.cjs`; dispatches build/query/status/diff subcommands; discovered via `commandFamilies` in the capability registry | | `audit-command-router.cjs` | ADR-959 capability command router (phase 4d-impl-3); extracted from the `case 'audit-uat':` and `case 'audit-open':` arms in `gsd-tools.cjs`; `routeAuditUat` → `uat.cjs:cmdAuditUat`, `routeAuditOpen` → `audit.cjs:{auditOpenArtifacts,formatAuditReport}`; discovered via `commandFamilies` in the capability registry | | `intel-command-router.cjs` | ADR-959 capability command router (phase 4d-impl-4, last first-party cutover); extracted from the `case 'intel':` arm in `gsd-tools.cjs`; `routeIntelCommand` → all 9 intel subcommands via lazy `require('./intel.cjs')`; preserves non-raw `timeAgo` transform on `status.files[*].updated_at`; discovered via `commandFamilies` in the capability registry | diff --git a/docs/FEATURES.md b/docs/FEATURES.md index 03ce0b358..808a67b68 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -3173,6 +3173,8 @@ Each resolved prohibition carries a `verification` tier — `test` (a negative t The load-bearing wire is the `plan-phase` lift into `must_haves.prohibitions`, so the section is not merely documentation. +**Deterministic prohibition-check descriptor source (#1278).** A resolved `test`-tier prohibition MAY carry an optional **`check` descriptor** — the flat-scalar keys `check_kind` (`node-test` | `lint-rule`), `check_target`, and `check_rule` (lint-rule only) — authored at spec-phase. `projectProhibitions` projects these scalars deterministically and verify-phase reads them back to locate the check handed to `check prohibition-enforcement`, so a wired, passing test closes the gap with **zero manual descriptor authoring** (previously the verify-phase LLM had to invent `{kind, target, rule}` each run, #1259). The descriptor is **optional and backward-compatible** — a descriptor-less prohibition parses and disposes byte-identically to today — and **fail-closed**: a partial, invalid, or absent descriptor falls through to the producer's existing fail-closed locate, never a silent green. `failFirst` stays a verify-time caller attestation (machine-proven fail-first is tracked in #1279). + **Requirements:** - REQ-PROHIB-01: The prohibition pass MUST run after the edge probe and emit a `## Prohibitions (must-NOT)` SPEC section. - REQ-PROHIB-02: Stage 1 MUST ask the adversarial recall question; Stage 2 MUST drop routine-engineering items and keep values/safety/ethics prohibitions. diff --git a/docs/USER-GUIDE.md b/docs/USER-GUIDE.md index 0efd6ab1e..b8ac37ef9 100644 --- a/docs/USER-GUIDE.md +++ b/docs/USER-GUIDE.md @@ -459,6 +459,26 @@ The review step slots in after execution and before UAT: - **Configuration Reference:** see [`docs/CONFIGURATION.md`](CONFIGURATION.md) for the full `config.json` schema, model-profile table, git branching strategies, and security settings. - **Discuss Mode:** see [`docs/workflow-discuss-mode.md`](workflow-discuss-mode.md) for interview vs assumptions mode. +### Graphify capability gate (tri-state, v1.43+) + +Graphify commands (`graphify status`, `graphify build`, `graphify query`, `graphify diff`) now respect the **full tri-state capability gate**: + +1. **Installed** — the `gsd-graphify-*` skills are present in the active install profile. +2. **Surfaced** — those skills appear on the current runtime surface (e.g., in `~/.claude/commands/gsd/`). +3. **Config-enabled** — `graphify.enabled: true` is set in `.planning/config.json`. + +All three conditions must be true. Setting `graphify.enabled: true` alone is no longer sufficient if graphify has not been installed and surfaced. If graphify commands return `{ disabled: true }` after upgrading, verify that the install profile includes graphify skills (`gsd-tools capability state`) and re-run the installer to surface them. + +### Intel capability gate (tri-state, v1.44+) + +Intel commands (`intel status`, `intel query`, `intel diff`, `intel snapshot`, `intel validate`, `intel api-surface`) now respect the **full tri-state capability gate** (same resolver as graphify above): + +1. **Installed** — the intel capability is present in the active install profile (intel has no skill files, so this is vacuously true for all profiles). +2. **Surfaced** — the intel capability is on the current runtime surface (vacuously true for all surfaces since intel registers no skill stems). +3. **Config-enabled** — `intel.enabled: true` is set in `.planning/config.json`. + +For intel, conditions 1 and 2 are always satisfied (intel has no skill files). The effective gate is `intel.enabled` in config — the same behaviour as before, but now enforced through the shared `isCapabilityActive('intel', cwd)` resolver rather than a direct config read. This means intel honours the full capability-state pipeline, including any future install-profile or surface restrictions. If intel commands return `{ disabled: true }`, ensure `intel.enabled: true` is set in `.planning/config.json` and verify `gsd-tools capability state` shows intel as active. + --- ## Usage Examples diff --git a/docs/adr/550-spec-phase-probe-contract.md b/docs/adr/550-spec-phase-probe-contract.md index 4a8cdd84c..2eb81c7d7 100644 --- a/docs/adr/550-spec-phase-probe-contract.md +++ b/docs/adr/550-spec-phase-probe-contract.md @@ -112,3 +112,19 @@ This addendum ratifies three contract points: > **PR-review flag — PROPOSED, renamable conventions (zero live consumers).** Both **`GSD_PROHIB_SUBJECT`** and **`CheckDescriptor.violationFixture`** are net-new surface introduced by this PR with **ZERO live in-tree consumers** — there is no in-tree `node --test` prohibition yet (the #1259 dogfood anchor and the #1279 lint-rule dogfood are both the LINT-rule `local/no-source-grep`; node-test fail-first is exercised only by SYNTHETIC temp fixtures in tests). They are therefore forward-looking scaffolding, and a later rename (or replacing the env var with an argv) is a **mechanical, zero-migration find/replace**. They are surfaced here explicitly so the maintainer can **rename or replace them at PR review** — the natural ratification point, exactly as #1278's ADR addendum was reviewed at PR time — without any migration cost. The `failFirst` DEMOTION is likewise open to the reviewer weighing outright removal; the rationale for keeping it as a hint is recorded above. Net effect on D4: the *guarantee* ("a `test`-tier prohibition is never a silent pass") was preserved at every step — fail-closed-now (#644), genuine-execution (#1259), and now **machine-proven fail-first (#1279)**. A `test`-tier prohibition reaches `green`/`passed` ONLY when the wired check both genuinely, non-vacuously passes AND is independently proven to fail on a violation; every miss/fail/un-provable hard-gates. The decision also lives in `src/prohibition-enforcement.cts` comments, `gsd-core/references/prohibition-probe.md`, `gsd-core/workflows/verify-phase.md`, and the #1279 changeset. + +## Addendum (2026-06-15): optional `check` descriptor on the prohibition item — D3 shape extension (#1278) + +This ratifies the **deterministic SOURCE** for the test-tier `CheckDescriptor` that #1259 (PR #1273) left caller/verifier-supplied. #1259 shipped the PRODUCER (`check prohibition-enforcement`) that *runs* a wired check given a `{kind, target, rule?}` descriptor, but the descriptor itself was invented by the verify-phase LLM each run (the "locate" half). #1278 makes that locate half **deterministic**: an optional `check` descriptor is authored at spec-phase on the resolved `test`-tier prohibition, projected by `projectProhibitions`, and read back by verify-phase — so a wired, passing test closes the gap with **zero manual authoring**. This extends the **Decision 3 prohibition-item shape** (it adds optional keys to that item), so it is ratified here rather than rewriting D3 in place. + +1. **The D3 item shape gains OPTIONAL flat-scalar keys.** Alongside `statement`, `status`+`verification`, and the dismissed-only `reason`, a resolved `test`-tier prohibition MAY carry `check_kind` (`node-test` | `lint-rule`), `check_target`, and `check_rule` (lint-rule only). All three are **optional**; a descriptor-less prohibition parses and disposes byte-identically to today (backward compatibility, CHK-07). + +2. **Flat scalars — NOT a nested `check:{}` object (load-bearing).** The representation is three flat scalar keys, never a nested object. The shared `parseMustHavesBlock` (`src/frontmatter.cts:252`) is a **flat parser**: continuation lines under a list item are handled only as nested ARRAYS or scalar `key: value` pairs, and the `reconstructFrontmatter` serializer is deliberately lossy for nested object-lists. A nested `check:{}` would flatten its keys into the parent item and mangle the round-trip. We follow the #644 precedent — **no parser rewrite, no `parseMustHavesBlock` change** — so the `truths`/`artifacts`/`key_links` shared-parser readers stay regression-free (the untouched frontmatter suite is the proof). + +3. **Deterministic projection + read-back.** `projectProhibitions` (`src/probe-core.cts`) emits the scalar keys ONLY for a well-formed descriptor (valid `check_kind` + non-empty `check_target`; `check_rule` only on the lint-rule path). `descriptorFromProjection` (`src/prohibition-enforcement.cts`) reads them back into a `{ kind: check_kind, target: check_target, rule?: check_rule }` `CheckDescriptor` to build the producer request. The `CheckDescriptor` type itself is unchanged; `failFirst` is **NOT** sourced from the projection — it stays a verify-time caller attestation. verify-phase locates from the projection, so a wired passing test needs no hand-authored descriptor. + +4. **Fail-closed on partial / invalid / absent descriptor.** A `lint-rule` descriptor missing `rule`, an unknown `kind`, or an absent descriptor on a test-tier prohibition MUST fall through to the producer's existing fail-closed locate ("no well-formed check descriptor is locatable → fail-closed") — never a silent green. The producer's locate semantics from #1259 are unchanged. + +5. **Out of scope (unchanged boundaries).** Machine-proven fail-first (a violation-fixture / RuleTester-invalid proof replacing the `failFirst` caller attestation) stays tracked as **#1279**. The `dispositionForProhibition` green/fail-closed **policy** is untouched. No new check kinds are added. + +Net effect on D3: the prohibition-item shape is extended with three optional, backward-compatible flat-scalar keys that give the test-tier locate a deterministic spec-phase source; the contract's CI-testable surface (D5) gains the projection round-trip parity (CHK-03), the fail-closed guard (CHK-06), and the byte-stable backward-compat fixture (CHK-07). The decision also lives in `src/probe-core.cts` / `src/prohibition-enforcement.cts` comments, the `verify-phase.md` / `spec-phase.md` prose, and the #1278 changeset. diff --git a/gsd-core/references/prohibition-probe.md b/gsd-core/references/prohibition-probe.md index 605027b8d..f9b89a05b 100644 --- a/gsd-core/references/prohibition-probe.md +++ b/gsd-core/references/prohibition-probe.md @@ -151,16 +151,51 @@ Splitting these axes keeps the lifecycle enum free of a verification fact and le prohibition adapter declare `test | judgment` without forking the shared lifecycle enum that the edge-probe's `explicit | backstop` also uses. +## Optional wired-check descriptor (deterministic locate, #1278) + +A `resolved`/`test`-tier prohibition MAY carry an **optional `check` descriptor** that names +the wired mechanical check, so verify-phase locates it deterministically instead of inventing +`{kind, target, rule}` each run. The descriptor is captured at spec-phase (soft / optional — +the author wires it when the negative test or lint rule already exists) and is represented as +**three flat scalar keys** on the `must_haves.prohibitions` item — never a nested `check: {}` +object: + +- `check_kind` — `node-test` | `lint-rule` (which producer mechanism runs the check). +- `check_target` — the test file (`node-test`) or the file the rule runs against (`lint-rule`). +- `check_rule` — the `ruleId` to filter on, **lint-rule only** (absent for `node-test`). + +The flat-scalar shape is load-bearing: the shared `parseMustHavesBlock` is a flat parser and a +nested object would flatten/mangle the round-trip (ADR-550 2026-06-15 addendum; #644 "no parser +rewrite" precedent). `projectProhibitions` emits these keys **only for a well-formed descriptor** +(valid `check_kind` + non-empty `check_target`; `check_rule` only on the lint-rule path), and +verify-phase reads them back via `descriptorFromProjection` into the `CheckDescriptor` handed to +`check prohibition-enforcement`. A wired, passing test then closes the gap with **zero manual +descriptor authoring**. + +**Fail-closed + backward-compat.** A partial descriptor (`lint-rule` missing `check_rule`), an +unknown `check_kind`, or an **absent** descriptor on a test-tier prohibition falls through to the +producer's existing fail-closed locate — never a silent green. A prohibition with no descriptor +parses and disposes byte-identically to today. `failFirst` is **not** sourced from the +descriptor — it stays a verify-time caller attestation (machine-proven fail-first is tracked in +#1279; the `dispositionForProhibition` policy is unchanged). + ## Output schema The probe emits, per kept prohibition, an item of the form: ``` -{ requirement_id, category, status, verification, resolution, reason, statement } +{ requirement_id, category, status, verification, resolution, reason, statement, + check_kind?, check_target?, check_rule? } ``` where `statement` is the must-NOT sentence and `category` is the values/safety/ethics class -(`values`, `fairness`, `privacy`, `transparency`, `safety`, …), plus a coverage summary: +(`values`, `fairness`, `privacy`, `transparency`, `safety`, …). The optional **flat-scalar +`check_*` descriptor** (#1278) is present only on a resolved `test`-tier prohibition carrying a +wired check: `check_kind` (`node-test` | `lint-rule`), `check_target`, and `check_rule` (lint-rule +only). `projectProhibitions` emits these into `must_haves.prohibitions` and `descriptorFromProjection` +reads them back into a `{ kind, target, rule? }` `CheckDescriptor`; they are flat scalars (never a +nested `check:{}` object) so they round-trip through the unchanged `parseMustHavesBlock`. Plus a +coverage summary: ``` coverage: { applicable, resolved, unresolved, byVerification: { test, judgment } } diff --git a/gsd-core/workflows/spec-phase.md b/gsd-core/workflows/spec-phase.md index 9a93c6758..7a6d49bde 100644 --- a/gsd-core/workflows/spec-phase.md +++ b/gsd-core/workflows/spec-phase.md @@ -356,6 +356,21 @@ For each Requirement gathered so far, run the two-stage recall→precision pass: Criteria AND mark the prohibition `resolved` with a verification tier: `test` (a mechanical negative test/lint/assertion exists) or `judgment` (real but not mechanically checkable — routes to judgment review). + - **Capture the wired-check descriptor on `test`-tier (#1278, SOFT).** When a prohibition is + resolved `verification: test`, ALSO capture the descriptor of the wired check so + `verify-phase` can LOCATE it deterministically (no verifier invention at verify time). + Capture the flat scalars — persisted into SPEC and projected onto the + `must_haves.prohibitions` item by `projectProhibitions`: + - `check_kind` — `node-test` | `lint-rule`. + - `check_target` — the negative-test file path (for `node-test`), or the path to lint + (for `lint-rule`). + - `check_rule` — the eslint rule id (e.g. `local/no-source-grep`); `lint-rule` only. + This is a **SOFT capture (CHK-04): a `test`-tier prohibition WITHOUT a descriptor is still + allowed** — if the author cannot yet name the wired check, leave the descriptor empty and + proceed. It is NOT a hard authoring block; the item simply stays fail-closed/flagged + downstream (an absent/partial descriptor → `descriptorFromProjection` null/under-specified + → producer fail-closed locate, never green). Do NOT capture `failFirst` here — it is a + verify-time caller attestation, not a spec-authored field (#1279). - **Dismiss (reason)** → mark `dismissed` with a REQUIRED non-empty reason (PROB-05). The reason string is the audit trail; silence is not a valid dismissal. - **Defer** → leave `unresolved`. @@ -374,7 +389,11 @@ For each Requirement gathered so far, run the two-stage recall→precision pass: **`--auto` mode:** auto-`resolved` where a defensible negative acceptance criterion can be written (test or judgment tier); otherwise leave `unresolved`. **`--auto` NEVER auto-dismisses a prohibition** — a wrong dismissal is the exact silent failure this probe eliminates (PROB-06, -the load-bearing safety property). Log: `[auto] prohibitions: R resolved, U unresolved`. +the load-bearing safety property). On a `test`-tier auto-resolution, capture the `check_kind` / +`check_target` / `check_rule` descriptor **only when a wired check is unambiguous**; otherwise +leave it empty — `--auto` NEVER fabricates a check path (a wrong locate is re-validated and +fails closed at the producer, but a fabricated path is still noise to avoid). Log: +`[auto] prohibitions: R resolved, U unresolved`. **Text mode (PROB-09):** per Step 5's text-mode rule, replace the AskUserQuestion menus above with plain-text numbered lists — there is NO hard AskUserQuestion dependency, so the probe @@ -382,7 +401,11 @@ runs identically for non-Claude / text-mode hosts. Populate the `## Prohibitions` section of SPEC.md from the resolved prohibitions (each `resolved`/`test` row is a checkable negative acceptance criterion; `resolved`/`judgment` -rows route to judgment review; `⚠ UNRESOLVED` rows are flagged as assumptions). +rows route to judgment review; `⚠ UNRESOLVED` rows are flagged as assumptions). A +`resolved`/`test` row ALSO carries its captured `check_kind` / `check_target` / `check_rule` +descriptor when present (so the projection feeds `verify-phase`'s deterministic locate, #1278); +a `test` row with no captured descriptor is still valid — it stays fail-closed/flagged +downstream rather than blocking authoring. ## Step 6: Generate SPEC.md diff --git a/gsd-core/workflows/verify-phase.md b/gsd-core/workflows/verify-phase.md index c77fcb1ce..54695c22e 100644 --- a/gsd-core/workflows/verify-phase.md +++ b/gsd-core/workflows/verify-phase.md @@ -70,17 +70,17 @@ Aggregate all must_haves across plans for phase-level verification. **Prohibitions (`must_haves.prohibitions`, ADR-550 D3 — the must-NOT sibling block):** When a plan carries `must_haves.prohibitions`, extract each `{ statement, status, verification }` item and route it by `verification` tier in verdict assembly (ADR-550 D4, "B-with-guard", 2026-06-12 maintainer decision). These are NEGATIVE checks (the must-NOT must NOT have happened), distinct from positive `truths`: - **judgment-tier → mode-dependent soft-gate.** Interactive verify defers each item to the end-of-phase human checkpoint (`human_verify_mode: end-of-phase`). Autonomous verify records a NON-AUTHORITATIVE LLM-judge verdict + a prominent `unverified-prohibition — human review recommended` flag (autonomous completion reads "complete with N flagged prohibitions"). NEVER a silent pass; NEVER a hard halt of an AFK run. -- **test-tier → ENFORCED via `check prohibition-enforcement` (green on pass, hard-gate on miss/fail).** Accept the `verification: test` value (the SPEC↔must_haves.prohibitions projection contract holds — no forced schema change later). For each test-tier item, the verifier invokes the deterministic producer: +- **test-tier → ENFORCED via `check prohibition-enforcement` (green on pass, hard-gate on miss/fail).** Accept the `verification: test` value (the SPEC↔must_haves.prohibitions projection contract holds — no forced schema change later). For each test-tier item, the verifier builds `request.check` **DETERMINISTICALLY from the projected descriptor** — it does NOT invent `{ kind, target, rule }`. Read the flat scalar keys `check_kind` / `check_target` / `check_rule` off the `must_haves.prohibitions` item and reconstruct the `CheckDescriptor` via the `descriptorFromProjection` adapter in `prohibition-enforcement` (`descriptorFromProjection(projectedItem)` → `{ kind: check_kind, target: check_target, rule?: check_rule }`). Then attest `failFirst: true` in the request — `failFirst` is the ONE field NOT sourced from the projection; it stays a verify-time caller attestation (#1279 machine-proves it against a violation fixture). Invoke the producer (CLI surface unchanged): ```bash gsd_run check prohibition-enforcement ``` - where `` carries `{ prohibition, check, mode }` — `check` being the wired mechanical-check descriptor `{ kind: 'node-test' | 'lint-rule', target, rule?, violationFixture, failFirst? }`. For `node-test`, `target` is the negative-test file path; for `lint-rule`, `target` is the PATH to lint and `rule` is the eslint rule id (e.g. `local/no-source-grep`) — both required (a lint-rule without `rule` is not a valid wired check). `violationFixture` is the author-supplied path to a KNOWN-BAD subject the producer runs the check against to **machine-prove fail-first** (for `node-test`, injected via the `GSD_PROHIB_SUBJECT` env convention); `failFirst` is a DEMOTED, non-authoritative hint kept only for backward route-JSON shape (no path greens on it alone — FF-08). The producer LOCATES the wired check, **machine-proves it is fail-first** by running it against the violation and confirming it goes RED, RUNS it for a genuine non-vacuous pass, builds `enforcementEvidence`, and emits the `dispositionForProhibition()` verdict (#1259 + #1279, ADR-550 D5d). Fail-first is **machine-proven, not caller-attested** — absent a provable violation the producer fails closed, never falling back to attestation. Route the result by its typed fields: + where `` carries `{ prohibition, check, mode }` — `check` being the wired mechanical-check descriptor `{ kind: 'node-test' | 'lint-rule', target, rule?, violationFixture, failFirst? }`, with `kind`/`target`/`rule` now sourced from the projected `check_*` scalars (not author/verifier invention — #1278). For `node-test`, `target` (from `check_target`) is the negative-test file path; for `lint-rule`, `target` is the PATH to lint and `rule` (from `check_rule`) is the eslint rule id (e.g. `local/no-source-grep`) — both required (a lint-rule without `rule` is not a valid wired check). `violationFixture` is the author-supplied path to a KNOWN-BAD subject the producer runs the check against to **machine-prove fail-first** (for `node-test`, injected via the `GSD_PROHIB_SUBJECT` env convention — #1279); `failFirst` is a DEMOTED, non-authoritative hint kept only for backward route-JSON shape (no path greens on it alone — FF-08). The producer LOCATES the wired check from the projection, **machine-proves it is fail-first** by running it against the violation and confirming it goes RED, RUNS it for a genuine non-vacuous pass, builds `enforcementEvidence`, and emits the `dispositionForProhibition()` verdict (#1259 + #1278 + #1279, ADR-550 D5d). Fail-first is **machine-proven, not caller-attested** — absent a provable violation the producer fails closed, never falling back to attestation. Route the result by its typed fields: - **`status: 'green'`, `flagged: false`** (a genuinely-passing wired negative test / lint rule, `located: true`, non-empty `evidence`) → the item is satisfiable → it can reach **passed**. - **missing, non-attested, or genuinely-non-passing check** (`located: false` OR `status: 'unverified'`, `flagged: true`) → **hard-gate**: disposes flagged-unverified, NEVER green, routing to `gaps_found` in BOTH interactive and autonomous modes (a failing mechanical check blocks even AFK; ADR-550 D4 / D3). The deterministic fail-closed default backing every miss/fail is `dispositionForProhibition()` in probe-core (`status: 'unverified'`, `flagged: true` on empty `enforcementEvidence`). - > **Descriptor authoring — current scope (#1259).** The `check` descriptor is **supplied by the phase author / verifier**; there is no projection field yet that deterministically derives `{ kind, target, rule, failFirst }` from a prohibition in `must_haves.prohibitions` (which carries only `{ statement, status, verification }`). So #1259 lands the **deterministic run+verdict half** (locate→run→evidence→disposition, all CI-testable) while the **locate→descriptor half is author-provided** for now. Deterministic auto-locate — so a wired passing test closes the gap with zero manual descriptor authoring — is a **tracked follow-up: #1278**. Until then, the green path requires the author to wire the descriptor explicitly. + > **Descriptor source — deterministic locate (#1278, DELIVERED).** The `check` descriptor's `{ kind, target, rule }` is now sourced **deterministically from the projected `check_kind` / `check_target` / `check_rule` scalars** on the `must_haves.prohibitions` item (authored at `/gsd:spec-phase`, projected by `projectProhibitions`, read back via the `descriptorFromProjection` adapter). So a wired passing test closes the gap with **zero manual descriptor authoring** — the verifier no longer invents the locate (removing the spoofable invent-at-verify-time surface; ADR-857 §147 exogenous grading). **Fail-closed is preserved:** an item with NO projected descriptor — or a PARTIAL one (e.g. a `lint-rule` missing `check_rule`) — makes `descriptorFromProjection` return `null` / an under-specified descriptor, which falls through to the producer's existing fail-closed LOCATE (`located: false`) → flagged-unverified, NEVER green, in BOTH modes. The only field still attested at verify time (not projected) is `failFirst`; machine-proving it against a violation fixture is the remaining **tracked follow-up: #1279**. **Option B: Use Success Criteria from ROADMAP.md** diff --git a/src/capability-state.cts b/src/capability-state.cts index dc0c58475..c981af516 100644 --- a/src/capability-state.cts +++ b/src/capability-state.cts @@ -25,6 +25,7 @@ * - ./surface.cjs (resolveSurface) * - ./config-loader.cjs (loadConfig) * - ./runtime-homes.cjs (getGlobalConfigDir — for runtimeConfigDir auto-detection) + * - ./runtime-slash.cjs (resolveRuntime — GSD_RUNTIME > config.runtime > 'claude' precedence) * - capability-registry.cjs (loaded at call time) */ @@ -89,6 +90,18 @@ interface CapabilityStateEntry { surfaced: boolean; /** True when the capability is both installed and surfaced. */ enabled: boolean; + /** + * True when the capability is enabled AND its config activation resolves to + * true. Config activation is determined by resolving the capability's + * `activationKey` (a dotted config key, e.g. `graphify.enabled`) via + * `_resolveActivationValue`. When `activationKey` is absent, configActivation + * defaults to `true` — the capability has no config gate. + * + * active = enabled && configActivation + * + * Note: `enabled` stays exactly `installed && surfaced` (unchanged). + */ + active: boolean; /** Resolved hook activation state across steps, gates, and contributions */ hooks: HookEntry[]; } @@ -211,6 +224,20 @@ function resolveCapabilityState(input: ResolveCapabilityStateInput): ResolveCapa const enabled = installed && surfaced; + // ── per-capability config activation ────────────────────────────────────── + // Resolve the capability's own activationKey (if present). This is the + // config-level toggle that gates the whole capability — separate from the + // per-hook `when` keys that gate individual hooks. When activationKey is + // absent, configActivation defaults to true (no config gate on the cap). + // active = enabled && configActivation (enabled unchanged: installed && surfaced) + const activationKey = typeof capObj['activationKey'] === 'string' && capObj['activationKey'].length > 0 + ? capObj['activationKey'] + : undefined; + const configActivation: boolean = activationKey !== undefined + ? _resolveActivationValue(activationKey, config, cwd, registry) + : true; + const active = enabled && configActivation; + // ── hooks ────────────────────────────────────────────────────────────────── // Collect from steps, gates, contributions. Each may have a `when` key. // Activation semantics (mirrors loop-resolver.isActive exactly): @@ -242,7 +269,12 @@ function resolveCapabilityState(input: ResolveCapabilityStateInput): ResolveCapa // (mirrors loop-resolver.isActive: `typeof when !== 'string' || when.length === 0` → false) configured = false; } - hooks.push({ point, kind, when: whenRaw, configured, active: enabled && configured }); + // Hook active = capability-level active AND hook's own config gate. + // The capability's `active` constant (= enabled && configActivation) is + // used here so that a config-disabled capability (active=false) cannot + // produce active hooks even when the hook's own `when` is unconditional + // (configured=true). The capability gate cascades to all its hooks. + hooks.push({ point, kind, when: whenRaw, configured, active: active && configured }); } } @@ -254,7 +286,7 @@ function resolveCapabilityState(input: ResolveCapabilityStateInput): ResolveCapa processHooks(Array.isArray(gatesRaw) ? gatesRaw : [], 'gate'); processHooks(Array.isArray(contributionsRaw) ? contributionsRaw : [], 'contribution'); - results.push({ id: capId, tier, skills, installed, surfaced, enabled, hooks }); + results.push({ id: capId, tier, skills, installed, surfaced, enabled, active, hooks }); } // Deterministic sort by id for stable output across calls @@ -368,11 +400,15 @@ function _resolveManifest(commandsGsdDir: string, configDir: string): Map string; }; - // Delegate runtime detection entirely to getGlobalConfigDir: calling it - // with 'claude' causes it to check CLAUDE_CONFIG_DIR first, falling back - // to ~/.claude. The canonical resolver already encodes the correct env-var - // precedence for each runtime — we do not re-implement that logic here. - // For non-claude runtimes, the caller should pass --config-dir explicitly - // (or set the runtime-specific env var, which getGlobalConfigDir honors). - resolvedConfigDir = runtimeHomes.getGlobalConfigDir('claude'); + // eslint-disable-next-line @typescript-eslint/no-require-imports + const runtimeSlash = require('./runtime-slash.cjs') as { + resolveRuntime: (projectDir: string | null | undefined) => string; + }; + // Detect the active runtime via GSD_RUNTIME → config.runtime → 'claude'. + // resolveRuntime reads config.json directly (no side effects) and returns + // a lowercased canonical runtime name. + const detectedRuntime = runtimeSlash.resolveRuntime(cwd); + resolvedConfigDir = runtimeHomes.getGlobalConfigDir(detectedRuntime); } catch { // Defensive fallback: use ~/.claude if the canonical resolver throws. // eslint-disable-next-line @typescript-eslint/no-require-imports @@ -533,9 +574,28 @@ function cmdCapabilityState( coreOutput(envelope, raw); } +/** + * Convenience predicate: returns true if the capability identified by `capId` + * is active (installed && surfaced && config-enabled) in the current runtime + * environment at `cwd`. + * + * Internally calls `resolveCapabilityRuntimeState(cwd, undefined)` and returns + * the `active` field of the matching CapabilityStateEntry. + * Returns `false` when the capability is not found in the registry. + * + * @param capId Capability identifier (e.g. 'graphify', 'intel') + * @param cwd Project root directory for config resolution + */ +function isCapabilityActive(capId: string, cwd: string): boolean { + const result = resolveCapabilityRuntimeState(cwd, undefined); + const entry = result.capabilities.find((c) => c.id === capId); + return entry !== undefined ? entry.active : false; +} + export = { resolveCapabilityState, resolveCapabilityRuntimeState, + isCapabilityActive, cmdCapabilityState, // Exported for tests _resolveCommandsGsdDir, diff --git a/src/graphify.cts b/src/graphify.cts index d8432cfb6..bd095ecca 100644 --- a/src/graphify.cts +++ b/src/graphify.cts @@ -11,32 +11,11 @@ import fs from 'node:fs'; import path from 'node:path'; import { execTool, execGit, platformWriteSync } from './shell-command-projection.cjs'; -// ─── Config Gate ───────────────────────────────────────────────────────────── +// eslint-disable-next-line @typescript-eslint/no-require-imports +import capabilityStateMod = require('./capability-state.cjs'); +const { isCapabilityActive } = capabilityStateMod; -/** - * Check whether graphify is enabled in the project config. - * Reads config.json directly via fs. Returns false by default - * (when no config, no graphify key, or on error). - */ -function isGraphifyEnabled(planningDir: string): boolean { - try { - const configPath = path.join(planningDir, 'config.json'); - if (!fs.existsSync(configPath)) return false; - const config: unknown = JSON.parse(fs.readFileSync(configPath, 'utf8')); - if ( - config && - typeof config === 'object' && - 'graphify' in config && - config.graphify && - typeof config.graphify === 'object' && - 'enabled' in config.graphify && - (config.graphify as Record).enabled === true - ) return true; - return false; - } catch { - return false; - } -} +// ─── Config Gate ───────────────────────────────────────────────────────────── interface DisabledResponse { disabled: true; @@ -396,7 +375,7 @@ function countCommitsBetween(cwd: string, from: string, to: string): number | nu */ function graphifyQuery(cwd: string, term: string, options: { budget?: number | null } = {}): unknown { const planningDir = path.join(cwd, '.planning'); - if (!isGraphifyEnabled(planningDir)) return disabledResponse(); + if (!isCapabilityActive('graphify', cwd)) return disabledResponse(); const graphPath = path.join(planningDir, 'graphs', 'graph.json'); if (!fs.existsSync(graphPath)) { @@ -435,7 +414,7 @@ function graphifyQuery(cwd: string, term: string, options: { budget?: number | n */ function graphifyStatus(cwd: string): unknown { const planningDir = path.join(cwd, '.planning'); - if (!isGraphifyEnabled(planningDir)) return disabledResponse(); + if (!isCapabilityActive('graphify', cwd)) return disabledResponse(); const graphPath = path.join(planningDir, 'graphs', 'graph.json'); if (!fs.existsSync(graphPath)) { @@ -498,7 +477,7 @@ function graphifyStatus(cwd: string): unknown { */ function graphifyDiff(cwd: string): unknown { const planningDir = path.join(cwd, '.planning'); - if (!isGraphifyEnabled(planningDir)) return disabledResponse(); + if (!isCapabilityActive('graphify', cwd)) return disabledResponse(); const snapshotPath = path.join(planningDir, 'graphs', '.last-build-snapshot.json'); const graphPath = path.join(planningDir, 'graphs', 'graph.json'); @@ -554,7 +533,7 @@ function graphifyDiff(cwd: string): unknown { */ function graphifyBuild(cwd: string): unknown { const planningDir = path.join(cwd, '.planning'); - if (!isGraphifyEnabled(planningDir)) return disabledResponse(); + if (!isCapabilityActive('graphify', cwd)) return disabledResponse(); const installed = checkGraphifyInstalled(); if (!installed.installed) return { error: installed.message }; @@ -619,7 +598,6 @@ function writeSnapshot(cwd: string): SnapshotResult | { error: string } { export = { // Config gate - isGraphifyEnabled, disabledResponse, // Subprocess execGraphify, diff --git a/src/intel.cts b/src/intel.cts index 8b1925da1..e228fcbe1 100644 --- a/src/intel.cts +++ b/src/intel.cts @@ -5,7 +5,8 @@ * Intel files live in .planning/intel/ and store structured data about * the project's files, APIs, dependencies, architecture, and tech stack. * - * All public functions gate on intel.enabled config (no-op when false). + * All public functions gate on isCapabilityActive('intel', cwd) — the shared + * tri-state resolver (installed + surfaced + intel.enabled config key). * * ADR-457 build-at-publish: the hand-written bin/lib/intel.cjs collapsed * to a TypeScript source of truth. Behaviour is preserved byte-for-behaviour @@ -17,6 +18,10 @@ import path from 'node:path'; import crypto from 'node:crypto'; import { platformWriteSync, platformReadSync, platformEnsureDir } from './shell-command-projection.cjs'; +// eslint-disable-next-line @typescript-eslint/no-require-imports +import capabilityStateMod = require('./capability-state.cjs'); +const { isCapabilityActive } = capabilityStateMod; + // ─── Constants ─────────────────────────────────────────────────────────────── const INTEL_DIR = '.planning/intel'; @@ -41,29 +46,20 @@ function ensureIntelDir(planningDir: string): string { } /** - * Check whether intel is enabled in the project config. - * Reads config.json directly via fs. Returns false by default - * (when no config, no intel key, or on error). + * Check whether intel is active (installed, surfaced, and config-enabled) for the project at cwd. + * Delegates to the shared tri-state capability resolver (isCapabilityActive) which honours the + * install profile, runtime surface, and activationKey (intel.enabled config gate). + * + * NOTE: planningDir is the legacy entry-point; cwd is derived as path.dirname(planningDir). + * Callers that have cwd directly may call isCapabilityActive('intel', cwd) themselves. + * + * INVARIANT: planningDir is always `/.planning` (i.e. path.join(cwd, '.planning')). + * The intel-command-router always constructs planningDir as path.join(cwd, '.planning'), + * so path.dirname(planningDir) === cwd is guaranteed. If a workstream-aware planningDir + * were ever passed here, the dirname would be wrong — but no caller does that. */ -function isIntelEnabled(planningDir: string): boolean { - try { - const configPath = path.join(planningDir, 'config.json'); - const raw = platformReadSync(configPath); - if (raw === null) return false; - const config: unknown = JSON.parse(raw); - if ( - config && - typeof config === 'object' && - 'intel' in config && - config.intel && - typeof config.intel === 'object' && - 'enabled' in config.intel && - (config.intel as Record).enabled === true - ) return true; - return false; - } catch { - return false; - } +function isIntelCapabilityActive(planningDir: string): boolean { + return isCapabilityActive('intel', path.dirname(planningDir)); } interface DisabledResponse { @@ -188,7 +184,7 @@ interface IntelQueryResult { * Searches across all JSON intel files in INTEL_FILES (keys and values), including arch-decisions.json (parsed as JSON, not as text). */ function intelQuery(term: string, planningDir: string): IntelQueryResult | DisabledResponse { - if (!isIntelEnabled(planningDir)) return disabledResponse(); + if (!isIntelCapabilityActive(planningDir)) return disabledResponse(); const matches: Array<{ source: string; entries: SearchMatch[] }> = []; let total = 0; @@ -225,7 +221,7 @@ interface IntelStatusResult { * A file is considered stale if its updated_at is older than 24 hours. */ function intelStatus(planningDir: string): IntelStatusResult | DisabledResponse { - if (!isIntelEnabled(planningDir)) return disabledResponse(); + if (!isIntelCapabilityActive(planningDir)) return disabledResponse(); const STALE_MS = 24 * 60 * 60 * 1000; // 24 hours const now = Date.now(); @@ -273,7 +269,7 @@ interface IntelDiffResult { * Show changes since the last full refresh by comparing file hashes. */ function intelDiff(planningDir: string): IntelDiffResult | { no_baseline: true } | DisabledResponse { - if (!isIntelEnabled(planningDir)) return disabledResponse(); + if (!isIntelCapabilityActive(planningDir)) return disabledResponse(); const snapshotPath = intelFilePath(planningDir, '.last-refresh.json'); const snapshot = safeReadJson(snapshotPath); @@ -309,7 +305,7 @@ function intelDiff(planningDir: string): IntelDiffResult | { no_baseline: true } * The actual update is performed by the intel-updater agent (PLAN-02). */ function intelUpdate(planningDir: string): { action: string; message: string } | DisabledResponse { - if (!isIntelEnabled(planningDir)) return disabledResponse(); + if (!isIntelCapabilityActive(planningDir)) return disabledResponse(); return { action: 'spawn_agent', @@ -359,7 +355,7 @@ function saveRefreshSnapshot(planningDir: string): SaveRefreshResult { * Writes .last-refresh.json with accurate timestamps and hashes. */ function intelSnapshot(planningDir: string): SaveRefreshResult | DisabledResponse { - if (!isIntelEnabled(planningDir)) return disabledResponse(); + if (!isIntelCapabilityActive(planningDir)) return disabledResponse(); return saveRefreshSnapshot(planningDir); } @@ -373,7 +369,7 @@ interface IntelValidateResult { * Validate all intel files for correctness and freshness. */ function intelValidate(planningDir: string): IntelValidateResult | DisabledResponse { - if (!isIntelEnabled(planningDir)) return disabledResponse(); + if (!isIntelCapabilityActive(planningDir)) return disabledResponse(); const errors: string[] = []; const warnings: string[] = []; @@ -470,7 +466,7 @@ interface IntelApiSurfaceResult { * mistake silence for "nothing exists". */ function intelApiSurface(planningDir: string): IntelApiSurfaceResult | DisabledResponse { - if (!isIntelEnabled(planningDir)) return disabledResponse(); + if (!isIntelCapabilityActive(planningDir)) return disabledResponse(); const intelPath = ensureIntelDir(planningDir); const apiMapPath = path.join(intelPath, INTEL_FILES.apis); @@ -535,7 +531,7 @@ interface IntelPatchMetaResult { * Patch _meta.updated_at in a JSON intel file to the current timestamp. * Reads the file, updates _meta.updated_at, increments version, writes back. * - * NOTE: Does not gate on isIntelEnabled — operates on arbitrary file paths + * NOTE: Does not gate on isCapabilityActive — operates on arbitrary file paths * for use by agents patching individual files outside the intel store. */ function intelPatchMeta(filePath: string): IntelPatchMetaResult { @@ -576,7 +572,7 @@ interface IntelExtractExportsResult { /** * Extract exports from a JS/CJS file by parsing module.exports or exports.X patterns. * - * NOTE: Does not gate on isIntelEnabled — operates on arbitrary source files + * NOTE: Does not gate on isCapabilityActive — operates on arbitrary source files * for use by agents building intel data from project files. */ function intelExtractExports(filePath: string): IntelExtractExportsResult { @@ -711,7 +707,7 @@ export = { // Utilities ensureIntelDir, - isIntelEnabled, + isIntelCapabilityActive, // Constants INTEL_FILES, diff --git a/src/loop-resolver.cts b/src/loop-resolver.cts index c18edcc26..a566f6c38 100644 --- a/src/loop-resolver.cts +++ b/src/loop-resolver.cts @@ -269,8 +269,16 @@ interface ResolveLoopHooksInput { config: Record; /** Optional cwd — enables raw config.json fallback reads (FIX 1 precedence level 2). */ cwd?: string; - /** Optional capability-state map; when present, disabled capabilities do not render hooks. */ - capabilityStatesById?: Map | Record; + /** + * Optional capability-state map; when present, inactive capabilities do not render hooks. + * Each entry carries both `enabled` (installed+surfaced) and `active` (enabled+configActivation). + * The resolver gates on `active` so that the config activation key (activationKey) is + * honoured even when no per-hook `when` guard is present (Phase 4 tri-state alignment). + * + * `active` is REQUIRED (not optional) so the gate is fail-closed: a missing or undefined + * `active` field is a compile error, never silently treated as truthy. + */ + capabilityStatesById?: Map | Record; } interface ResolveLoopHooksResult { @@ -330,13 +338,18 @@ function resolveLoopHooks(input: ResolveLoopHooksInput): ResolveLoopHooksResult return _resolveActivationValue(when, config, cwd, registry); } - function isCapabilityEnabled(capId: string): boolean { + function isCapabilityActive(capId: string): boolean { if (!capabilityStatesById) return true; const state = capabilityStatesById instanceof Map ? capabilityStatesById.get(capId) : capabilityStatesById[capId]; if (!state) return false; - return state.enabled !== false; + // Fail-closed gate: only render the hook when active is explicitly true. + // A capability can be installed and surfaced (enabled=true) but config-disabled + // (active=false); in that case the hook must not render. + // Phase 4 tri-state alignment: `active` is now required (not optional), so + // `=== true` is the correct fail-closed check (not `!== false`). + return state.active === true; } // Helper: safe string array @@ -401,7 +414,7 @@ function resolveLoopHooks(input: ResolveLoopHooksInput): ResolveLoopHooksResult for (const hook of steps) { if (!hook || typeof hook !== 'object') continue; const capId = typeof hook['capId'] === 'string' ? hook['capId'] : ''; - if (!isCapabilityEnabled(capId)) continue; + if (!isCapabilityActive(capId)) continue; if (!isActive(hook)) continue; const ref = (typeof hook['ref'] === 'object' && hook['ref'] !== null) ? (hook['ref'] as HookRef) @@ -427,7 +440,7 @@ function resolveLoopHooks(input: ResolveLoopHooksInput): ResolveLoopHooksResult for (const hook of contributions) { if (!hook || typeof hook !== 'object') continue; const capId = typeof hook['capId'] === 'string' ? hook['capId'] : ''; - if (!isCapabilityEnabled(capId)) continue; + if (!isCapabilityActive(capId)) continue; if (!isActive(hook)) continue; const into = typeof hook['into'] === 'string' ? hook['into'] : undefined; const fragment = toFragment(hook['fragment']); @@ -453,7 +466,7 @@ function resolveLoopHooks(input: ResolveLoopHooksInput): ResolveLoopHooksResult for (const hook of gates) { if (!hook || typeof hook !== 'object') continue; const capId = typeof hook['capId'] === 'string' ? hook['capId'] : ''; - if (!isCapabilityEnabled(capId)) continue; + if (!isCapabilityActive(capId)) continue; if (!isActive(hook)) continue; const when = typeof hook['when'] === 'string' ? hook['when'] : undefined; const check = hook['check'] !== undefined ? hook['check'] : undefined; @@ -613,11 +626,11 @@ function cmdLoopRenderHooks( warnings?: string[]; registry: Record; config: Record; - capabilities: Array<{ id: string; enabled?: boolean }>; + capabilities: Array<{ id: string; enabled?: boolean; active: boolean }>; }; const registry = state.registry; const config = state.config || loadConfig(cwd); - const capabilityStatesById = new Map(); + const capabilityStatesById = new Map(); for (const cap of state.capabilities || []) { capabilityStatesById.set(cap.id, cap); } diff --git a/src/probe-core.cts b/src/probe-core.cts index 85a39f395..7a924fe94 100644 --- a/src/probe-core.cts +++ b/src/probe-core.cts @@ -290,6 +290,15 @@ export interface Prohibition { resolution: string | null; reason: string | null; statement: string; + // Optional flat-scalar wired-check descriptor (#1278). A resolved test-tier prohibition may carry + // these before projection; `projectProhibitions` emits them as the LOCKED flat scalar keys + // `check_kind`/`check_target`/`check_rule` that round-trip the EXISTING flat `parseMustHavesBlock` + // (a nested `check:{}` object is rejected per IMPL-SCOPING §3 — it flattens through the shared + // parser). These mirror `CheckDescriptor.kind/target/rule` (prohibition-enforcement.cts:62) MINUS + // the caller-attested `failFirst`, which is deliberately NOT a Prohibition field (#1279). + check_kind?: 'node-test' | 'lint-rule'; + check_target?: string; + check_rule?: string; } /** @@ -329,6 +338,15 @@ export function validateProhibitionResolution(resolution: Resolution `null`. + * - `check_kind` ABSENT -> `null` (no descriptor -> producer locates nothing -> fail-closed). + * - `check_kind` present -> `{ kind: check_kind, target: check_target }`, adding `rule: check_rule` + * ONLY when `check_rule` is a non-empty string. + * - `failFirst` is NEVER sourced from the projection — it stays a verify-time caller attestation + * (#1279 machine-proves it; out of scope here). The returned descriptor carries no `failFirst`. + * - The adapter does NOT strictly validate kind/target/rule: it faithfully reconstructs whatever + * scalars are present (e.g. `{check_kind:'lint-rule', check_target:'src/'}` with no `check_rule` + * reconstructs to `{kind:'lint-rule', target:'src/'}`), letting the existing LOCATE guard reject + * an under-specified descriptor (located:false, never green). It does NOT re-implement that guard. + * - Pure, deterministic, no-throw. + */ +export function descriptorFromProjection( + projected: Record | null | undefined, +): CheckDescriptor | null { + if (!projected || typeof projected !== 'object') return null; + if (!('check_kind' in projected)) return null; + // The shared `parseMustHavesBlock` (src/frontmatter.cts) coerces /^\d+$/ scalar values to NUMBERS on + // round-trip, so a numeric-looking check_kind/check_target/check_rule arrives here as a number. Normalize + // ONLY string|number scalars back to string — a non-scalar (object/array/bool) or absent value yields '' + // (never an `[object Object]` stringification, no `as string` lie over a number). This keeps the + // descriptor honestly typed and the round-trip lossless across the full string domain; an under-specified + // '' target/kind is rejected by the producer's locate guard (fail-closed; never green). + const scalar = (v: unknown): string => + typeof v === 'string' ? v : typeof v === 'number' ? String(v) : ''; + const kind = scalar(projected.check_kind) as CheckKind; + const target = scalar(projected.check_target); + const descriptor: CheckDescriptor = { kind, target }; + // `rule` belongs only to the lint-rule kind; a stray check_rule on a node-test descriptor is dropped + // (the projector never emits one there — defense in depth). failFirst is NOT sourced here (#1279). + if (kind === 'lint-rule') { + const rule = scalar(projected.check_rule); + if (rule.trim().length > 0) descriptor.rule = rule; + } + return descriptor; +} + /** * The result a check-runner returns: whether the check genuinely, non-vacuously PASSED. The runner * reports only what it can OBSERVE (a real clean pass) — it does NOT determine fail-first. Whether the diff --git a/tests/capability-state.test.cjs b/tests/capability-state.test.cjs index acde0af36..21d2feb87 100644 --- a/tests/capability-state.test.cjs +++ b/tests/capability-state.test.cjs @@ -19,6 +19,7 @@ const { cleanup } = require('./helpers.cjs'); const { resolveCapabilityState, + isCapabilityActive, _isSafePropKey, _loadInstalledSkillsManifest, _resolveManifest, @@ -31,7 +32,7 @@ const realRegistry = require('../gsd-core/bin/lib/capability-registry.cjs'); /** * Build a minimal synthetic registry for a single capability with the given - * skills, steps, gates, contributions, and configSchema entries. + * skills, steps, gates, contributions, configSchema, and optional activationKey. */ function makeRegistry({ id = 'test-cap', @@ -41,18 +42,23 @@ function makeRegistry({ gates = [], contributions = [], configSchema = {}, + activationKey = undefined, } = {}) { + const capEntry = { + id, + tier, + skills, + steps, + gates, + contributions, + config: {}, + }; + if (activationKey !== undefined) { + capEntry.activationKey = activationKey; + } return { capabilities: { - [id]: { - id, - tier, - skills, - steps, - gates, - contributions, - config: {}, - }, + [id]: capEntry, }, configSchema, }; @@ -535,6 +541,64 @@ describe('resolveCapabilityState — hook activation details', () => { assert.strictEqual(hook.when, 42, 'original non-string when value must be preserved'); assert.strictEqual(hook.active, false, 'non-string when → inactive'); }); + + // ── configActivation cascade to hooks ──────────────────────────────────────── + // Bug: a config-disabled capability (active=false) with an unconditional hook + // (no `when`, configured=true) was wrongly yielding hook.active=true because + // the hook loop used `enabled && configured` instead of `active && configured`. + // Fix: hook.active = active(capability) && configured. + + test('config-disabled capability with unconditional hook → hook.active=false, configured=true', () => { + // The capability activationKey resolves false → capability active=false. + // The hook has no `when` → configured=true (unconditional). + // Before the fix: hook.active = enabled(true) && configured(true) = true ← BUG + // After the fix: hook.active = active(false) && configured(true) = false ← correct + const registry = makeRegistry({ + id: 'test-cap', + skills: ['my-skill'], + activationKey: 'myfeature.enabled', + steps: [{ point: 'plan:pre', ref: { skill: 'my-skill' } }], // no `when` → unconditional + }); + const result = resolveCapabilityState({ + registry, + installedSkills: new Set(['my-skill']), + surfacedSkills: new Set(['my-skill']), + config: { myfeature: { enabled: false } }, // capability's configActivation = false + }); + const cap = result.capabilities[0]; + assert.ok(cap, 'capability must be present'); + assert.strictEqual(cap.enabled, true, 'enabled stays true: installed && surfaced'); + assert.strictEqual(cap.active, false, 'capability active=false: configActivation off'); + const hook = cap.hooks.find((h) => h.kind === 'step'); + assert.ok(hook, 'unconditional step hook must be present'); + assert.strictEqual(hook.when, undefined, 'hook has no when field (unconditional)'); + assert.strictEqual(hook.configured, true, 'hook configured=true: hook own gate is open'); + assert.strictEqual(hook.active, false, 'hook active=false: capability configActivation cascades to hook'); + }); + + test('config-ENABLED capability with unconditional hook → hook.active=true (control)', () => { + // Same setup as above but activationKey resolves true → capability active=true. + // hook.active should follow configured as before. + const registry = makeRegistry({ + id: 'test-cap', + skills: ['my-skill'], + activationKey: 'myfeature.enabled', + steps: [{ point: 'plan:pre', ref: { skill: 'my-skill' } }], // no `when` + }); + const result = resolveCapabilityState({ + registry, + installedSkills: new Set(['my-skill']), + surfacedSkills: new Set(['my-skill']), + config: { myfeature: { enabled: true } }, // capability's configActivation = true + }); + const cap = result.capabilities[0]; + assert.ok(cap, 'capability must be present'); + assert.strictEqual(cap.active, true, 'capability active=true: configActivation on'); + const hook = cap.hooks.find((h) => h.kind === 'step'); + assert.ok(hook, 'unconditional step hook must be present'); + assert.strictEqual(hook.configured, true, 'hook configured=true'); + assert.strictEqual(hook.active, true, 'hook active=true: both capability and hook gate open'); + }); }); // ─── resolveCapabilityState — determinism ───────────────────────────────────── @@ -1177,3 +1241,351 @@ describe('regressions: installed-runtime capability surface (#1160)', () => { }); }); + +// ─── resolveCapabilityState — per-capability active field ──────────────────── + +describe('resolveCapabilityState — per-capability active (Phase 2)', () => { + // Boundary: installed && surfaced && config-enabled → active=true + test('active=true when installed && surfaced && activationKey resolves true', () => { + const registry = makeRegistry({ + skills: ['my-skill'], + activationKey: 'myfeature.enabled', + configSchema: { 'myfeature.enabled': { default: false } }, + }); + const result = resolveCapabilityState({ + registry, + installedSkills: new Set(['my-skill']), + surfacedSkills: new Set(['my-skill']), + config: { myfeature: { enabled: true } }, + }); + const cap = result.capabilities[0]; + assert.ok(cap, 'capability must be present'); + assert.strictEqual(cap.installed, true); + assert.strictEqual(cap.surfaced, true); + assert.strictEqual(cap.enabled, true, 'enabled = installed && surfaced (unchanged)'); + assert.strictEqual(cap.active, true, 'active = enabled && config-enabled'); + }); + + // Boundary: installed && surfaced && config-DISABLED → active=false (the key case) + test('active=false when installed && surfaced but activationKey resolves false', () => { + const registry = makeRegistry({ + skills: ['my-skill'], + activationKey: 'myfeature.enabled', + configSchema: { 'myfeature.enabled': { default: false } }, + }); + const result = resolveCapabilityState({ + registry, + installedSkills: new Set(['my-skill']), + surfacedSkills: new Set(['my-skill']), + config: { myfeature: { enabled: false } }, + }); + const cap = result.capabilities[0]; + assert.ok(cap); + assert.strictEqual(cap.enabled, true, 'enabled still true: installed && surfaced unchanged'); + assert.strictEqual(cap.active, false, 'active=false: config gate is off'); + }); + + // Boundary: config-enabled but NOT surfaced → active=false + test('active=false when config-enabled but not surfaced (enabled=false)', () => { + const registry = makeRegistry({ + skills: ['my-skill'], + activationKey: 'myfeature.enabled', + }); + const result = resolveCapabilityState({ + registry, + installedSkills: new Set(['my-skill']), + surfacedSkills: new Set(), // NOT surfaced + config: { myfeature: { enabled: true } }, + }); + const cap = result.capabilities[0]; + assert.ok(cap); + assert.strictEqual(cap.surfaced, false); + assert.strictEqual(cap.enabled, false, 'enabled=false: not surfaced'); + assert.strictEqual(cap.active, false, 'active=false: enabled is false'); + }); + + // Boundary: NOT installed → active=false regardless of config + test('active=false when not installed (installed=false)', () => { + const registry = makeRegistry({ + skills: ['my-skill'], + activationKey: 'myfeature.enabled', + }); + const result = resolveCapabilityState({ + registry, + installedSkills: new Set(), // NOT installed + surfacedSkills: new Set(['my-skill']), + config: { myfeature: { enabled: true } }, + }); + const cap = result.capabilities[0]; + assert.ok(cap); + assert.strictEqual(cap.installed, false); + assert.strictEqual(cap.enabled, false); + assert.strictEqual(cap.active, false, 'active=false: not installed'); + }); + + // Boundary: no activationKey → active === enabled (no config gate) + test('active=enabled when capability has no activationKey', () => { + // No activationKey: configActivation defaults to true so active = enabled + const registry = makeRegistry({ + skills: ['my-skill'], + // No activationKey + }); + // installed && surfaced → enabled=true → active=true + const resultEnabled = resolveCapabilityState({ + registry, + installedSkills: new Set(['my-skill']), + surfacedSkills: new Set(['my-skill']), + config: {}, + }); + const capEnabled = resultEnabled.capabilities[0]; + assert.ok(capEnabled); + assert.strictEqual(capEnabled.enabled, true); + assert.strictEqual(capEnabled.active, true, 'active=true when no activationKey and enabled'); + + // NOT surfaced → enabled=false → active=false + const resultDisabled = resolveCapabilityState({ + registry, + installedSkills: new Set(['my-skill']), + surfacedSkills: new Set(), // not surfaced + config: {}, + }); + const capDisabled = resultDisabled.capabilities[0]; + assert.ok(capDisabled); + assert.strictEqual(capDisabled.enabled, false); + assert.strictEqual(capDisabled.active, false, 'active=false when no activationKey and not enabled'); + }); + + // enabled field semantics unchanged: still installed && surfaced only + test('enabled stays installed && surfaced regardless of activationKey', () => { + const registry = makeRegistry({ + skills: ['my-skill'], + activationKey: 'myfeature.enabled', + }); + // installed && surfaced but config disables it + const result = resolveCapabilityState({ + registry, + installedSkills: new Set(['my-skill']), + surfacedSkills: new Set(['my-skill']), + config: { myfeature: { enabled: false } }, + }); + const cap = result.capabilities[0]; + assert.ok(cap); + // enabled must still be true (installed && surfaced — config doesn't affect it) + assert.strictEqual(cap.enabled, true, 'enabled = installed && surfaced (config does not affect enabled)'); + assert.strictEqual(cap.active, false, 'active reflects config gate'); + }); + + // activationKey defaults to schema default when key absent from config + test('active uses schema default when activationKey not in config', () => { + // schema default = true → active=true when enabled + const registryDefaultTrue = makeRegistry({ + skills: ['my-skill'], + activationKey: 'myfeature.enabled', + configSchema: { 'myfeature.enabled': { default: true } }, + }); + const resultTrue = resolveCapabilityState({ + registry: registryDefaultTrue, + installedSkills: new Set(['my-skill']), + surfacedSkills: new Set(['my-skill']), + config: {}, // no explicit key + }); + assert.strictEqual(resultTrue.capabilities[0].active, true, 'active=true when schema default=true'); + + // schema default = false → active=false when enabled + const registryDefaultFalse = makeRegistry({ + skills: ['my-skill'], + activationKey: 'myfeature.enabled', + configSchema: { 'myfeature.enabled': { default: false } }, + }); + const resultFalse = resolveCapabilityState({ + registry: registryDefaultFalse, + installedSkills: new Set(['my-skill']), + surfacedSkills: new Set(['my-skill']), + config: {}, // no explicit key + }); + assert.strictEqual(resultFalse.capabilities[0].active, false, 'active=false when schema default=false'); + }); +}); + +// ─── isCapabilityActive convenience predicate ───────────────────────────────── + +describe('isCapabilityActive — convenience predicate (Phase 2)', () => { + // isCapabilityActive delegates to resolveCapabilityRuntimeState which does I/O. + // We test it via the CLI+tmpDir path used by the existing e2e tests above, and + // a pure-function boundary test that forces a known-missing capability id. + + test('returns false for unknown capId (not in registry)', () => { + // We can't easily inject the registry into resolveCapabilityRuntimeState + // without a cwd that has been set up. Instead, probe a capability id that + // is guaranteed to never appear in the real registry. + // We use a non-existent cwd so resolveCapabilityRuntimeState degrades + // gracefully (installedSkills/surfacedSkills → empty → all disabled) but + // still returns a capabilities array from the real registry. + // Any truly-unknown capId must return false regardless. + const nonExistentCwd = path.join(os.tmpdir(), 'cap-active-nonexistent-' + Date.now()); + const result = isCapabilityActive('__definitely_not_a_capability__', nonExistentCwd); + assert.strictEqual(result, false, 'unknown capId must return false'); + }); + + test('isCapabilityActive returns same value as the entry active field (e2e)', () => { + // Use a non-existent cwd so resolveCapabilityRuntimeState degrades gracefully. + // resolveCapabilityRuntimeState and isCapabilityActive both resolve against + // the SAME runtime environment: we compare isCapabilityActive(capId, cwd) + // to the matching entry's .active from resolveCapabilityRuntimeState(cwd, undefined) + // called on the identical cwd. This guarantees a real equality check — not typeof. + // + // We need resolveCapabilityRuntimeState for the comparison; require it here + // since it is exported from the same module. + const { resolveCapabilityRuntimeState } = require('../gsd-core/bin/lib/capability-state.cjs'); + + const nonExistentCwd = path.join(os.tmpdir(), 'cap-active-eq-' + Date.now()); + + // Pick a capability that exists in the real registry ('ui' is always present). + const capId = 'ui'; + + // resolveCapabilityRuntimeState resolves state from the real environment. + const runtimeResult = resolveCapabilityRuntimeState(nonExistentCwd, undefined); + const entry = runtimeResult.capabilities.find((c) => c.id === capId); + assert.ok(entry, `'${capId}' must be present in the real registry`); + + // isCapabilityActive must return exactly the same boolean as the resolved entry. + const actual = isCapabilityActive(capId, nonExistentCwd); + assert.strictEqual( + actual, + entry.active, + `isCapabilityActive('${capId}') must equal entry.active=${entry.active} from resolveCapabilityRuntimeState`, + ); + + // Also verify the already-covered false case for unknown capId: + assert.strictEqual(isCapabilityActive('__no_such_cap__', nonExistentCwd), false); + }); +}); + +// ─── Cross-runtime runtime detection (HIGH fix — GSD_RUNTIME → config.runtime → 'claude') ────── + +describe('isCapabilityActive cross-runtime detection (GSD_RUNTIME → config.runtime → claude)', () => { + // Verifies the HIGH bug fix: when GSD_RUNTIME='codex', resolveCapabilityRuntimeState + // must consult the CODEX config dir (via CODEX_HOME), not ~/.claude. + // + // Fixture layout: + // CODEX_HOME → tmpCodexDir/ ← .gsd-surface.json: full profile, graphify SURFACED + // CLAUDE_CONFIG_DIR → tmpClaudeDir/ ← .gsd-surface.json: full profile, graphify NOT surfaced + // GSD_RUNTIME=codex + // project cwd → tmpProjectDir/ ← config.json: graphify.enabled=true + // (config leg must pass; SURFACE leg is what we are isolating) + // + // Expected: isCapabilityActive('graphify', cwd) === true + // (codex dir has graphify surfaced, so the codex surface should win) + // + // Pre-fix (hardcoded 'claude'): would read tmpClaudeDir → graphify NOT surfaced → false (BUG). + // Post-fix (detects 'codex' via GSD_RUNTIME): reads tmpCodexDir → surfaced → true (CORRECT). + + let tmpCodexDir; + let tmpClaudeDir; + let tmpProjectDir; + let prevGsdRuntime; + let prevCodexHome; + let prevClaudeConfigDir; + let prevGsdWorkstream; + let prevGsdProject; + + before(() => { + tmpCodexDir = fs.mkdtempSync(path.join(os.tmpdir(), 'cap-state-codex-cfg-')); + tmpClaudeDir = fs.mkdtempSync(path.join(os.tmpdir(), 'cap-state-claude-cfg-')); + tmpProjectDir = fs.mkdtempSync(path.join(os.tmpdir(), 'cap-state-xrt-project-')); + + // Codex config dir: full profile + graphify SURFACED (disabledClusters empty) + fs.writeFileSync( + path.join(tmpCodexDir, '.gsd-surface.json'), + JSON.stringify({ baseProfile: 'full', disabledClusters: [], explicitAdds: [], explicitRemoves: [] }, null, 2) + '\n', + 'utf8', + ); + + // Claude config dir: full profile + graphify NOT surfaced + fs.writeFileSync( + path.join(tmpClaudeDir, '.gsd-surface.json'), + JSON.stringify({ baseProfile: 'full', disabledClusters: ['graphify'], explicitAdds: [], explicitRemoves: [] }, null, 2) + '\n', + 'utf8', + ); + + // Project: config with graphify.enabled=true so the config leg passes. + // The SURFACE leg (install+surfaced) is what we are testing — it must read + // the CODEX config dir (via CODEX_HOME), not the CLAUDE config dir. + fs.mkdirSync(path.join(tmpProjectDir, '.planning'), { recursive: true }); + fs.writeFileSync( + path.join(tmpProjectDir, '.planning', 'config.json'), + JSON.stringify({ graphify: { enabled: true } }), + 'utf8', + ); + }); + + after(() => { + cleanup(tmpCodexDir); + cleanup(tmpClaudeDir); + cleanup(tmpProjectDir); + }); + + test('GSD_RUNTIME=codex → resolver consults CODEX_HOME, not CLAUDE_CONFIG_DIR (HIGH cross-runtime fix)', () => { + prevGsdRuntime = process.env.GSD_RUNTIME; + prevCodexHome = process.env.CODEX_HOME; + prevClaudeConfigDir = process.env.CLAUDE_CONFIG_DIR; + prevGsdWorkstream = process.env.GSD_WORKSTREAM; + prevGsdProject = process.env.GSD_PROJECT; + + try { + process.env.GSD_RUNTIME = 'codex'; + process.env.CODEX_HOME = tmpCodexDir; // codex dir: graphify SURFACED + process.env.CLAUDE_CONFIG_DIR = tmpClaudeDir; // claude dir: graphify NOT surfaced + delete process.env.GSD_WORKSTREAM; + delete process.env.GSD_PROJECT; + + // With GSD_RUNTIME=codex, the resolver must look at CODEX_HOME (graphify surfaced) + // and return true. Pre-fix it would look at CLAUDE_CONFIG_DIR (not surfaced) → false. + const active = isCapabilityActive('graphify', tmpProjectDir); + assert.strictEqual( + active, + true, + 'isCapabilityActive must return true when GSD_RUNTIME=codex and graphify is surfaced in CODEX_HOME — ' + + 'the pre-fix code hardcoded getGlobalConfigDir("claude") regardless of GSD_RUNTIME, returning false (BUG)', + ); + + // Also verify: if we swap surfaces (codex NOT surfaced, claude surfaced), still uses codex dir → false + fs.writeFileSync( + path.join(tmpCodexDir, '.gsd-surface.json'), + JSON.stringify({ baseProfile: 'full', disabledClusters: ['graphify'], explicitAdds: [], explicitRemoves: [] }, null, 2) + '\n', + 'utf8', + ); + fs.writeFileSync( + path.join(tmpClaudeDir, '.gsd-surface.json'), + JSON.stringify({ baseProfile: 'full', disabledClusters: [], explicitAdds: [], explicitRemoves: [] }, null, 2) + '\n', + 'utf8', + ); + const activeSwapped = isCapabilityActive('graphify', tmpProjectDir); + assert.strictEqual( + activeSwapped, + false, + 'isCapabilityActive must return false when GSD_RUNTIME=codex and graphify is NOT surfaced in CODEX_HOME — ' + + 'even though claude dir has graphify surfaced, the codex dir is the authoritative surface', + ); + } finally { + // Restore env vars + if (prevGsdRuntime === undefined) delete process.env.GSD_RUNTIME; + else process.env.GSD_RUNTIME = prevGsdRuntime; + if (prevCodexHome === undefined) delete process.env.CODEX_HOME; + else process.env.CODEX_HOME = prevCodexHome; + if (prevClaudeConfigDir === undefined) delete process.env.CLAUDE_CONFIG_DIR; + else process.env.CLAUDE_CONFIG_DIR = prevClaudeConfigDir; + if (prevGsdWorkstream === undefined) delete process.env.GSD_WORKSTREAM; + else process.env.GSD_WORKSTREAM = prevGsdWorkstream; + if (prevGsdProject === undefined) delete process.env.GSD_PROJECT; + else process.env.GSD_PROJECT = prevGsdProject; + + // Restore codex dir surface fixture for cleanup consistency + fs.writeFileSync( + path.join(tmpCodexDir, '.gsd-surface.json'), + JSON.stringify({ baseProfile: 'full', disabledClusters: [], explicitAdds: [], explicitRemoves: [] }, null, 2) + '\n', + 'utf8', + ); + } + }); +}); diff --git a/tests/graphify-query.test.cjs b/tests/graphify-query.test.cjs index 6de2a32bc..db5ea73e4 100644 --- a/tests/graphify-query.test.cjs +++ b/tests/graphify-query.test.cjs @@ -6,6 +6,7 @@ const { describe, test, beforeEach, afterEach } = require('node:test'); const assert = require('node:assert/strict'); const fs = require('fs'); +const os = require('os'); const path = require('path'); const { createTempProject, cleanup } = require('./helpers.cjs'); @@ -26,6 +27,55 @@ const { SAMPLE_GRAPH, } = require('./helpers/graphify.cjs'); +// ─── Shared fixture: surfaced-config-dir ───────────────────────────────────── +// +// Positive-path tests (graphifyQuery, graphifyDiff, graceful-degradation) call +// enableGraphify() (config leg only) and assert non-disabled outcomes. With +// the tri-state gate (isCapabilityActive), those outcomes ALSO require graphify +// to be installed+surfaced in the runtime config dir. Without this fixture the +// tests are ambient-dependent: they pass only on machines where the ambient +// ~/.claude has graphify surfaced. +// +// Fix: before each positive-path test, point CLAUDE_CONFIG_DIR at a tmp dir +// with a full-profile .gsd-surface.json (graphify surfaced), and clear +// GSD_RUNTIME / GSD_WORKSTREAM / GSD_PROJECT for hermeticity. + +/** Create a tmp config dir with graphify surfaced (full profile, no disabled clusters). */ +function makeSurfacedConfigDir() { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-graphify-qry-cfg-')); + fs.writeFileSync( + path.join(dir, '.gsd-surface.json'), + JSON.stringify({ baseProfile: 'full', disabledClusters: [], explicitAdds: [], explicitRemoves: [] }, null, 2) + '\n', + 'utf8', + ); + return dir; +} + +/** + * Save the env vars the surfaced-config fixture overrides. + * Returns an object whose .restore() returns env to its original state. + */ +function saveSurfacedEnv() { + const saved = { + GSD_RUNTIME: process.env.GSD_RUNTIME, + CLAUDE_CONFIG_DIR: process.env.CLAUDE_CONFIG_DIR, + GSD_WORKSTREAM: process.env.GSD_WORKSTREAM, + GSD_PROJECT: process.env.GSD_PROJECT, + }; + return { + restore() { + if (saved.GSD_RUNTIME === undefined) delete process.env.GSD_RUNTIME; + else process.env.GSD_RUNTIME = saved.GSD_RUNTIME; + if (saved.CLAUDE_CONFIG_DIR === undefined) delete process.env.CLAUDE_CONFIG_DIR; + else process.env.CLAUDE_CONFIG_DIR = saved.CLAUDE_CONFIG_DIR; + if (saved.GSD_WORKSTREAM === undefined) delete process.env.GSD_WORKSTREAM; + else process.env.GSD_WORKSTREAM = saved.GSD_WORKSTREAM; + if (saved.GSD_PROJECT === undefined) delete process.env.GSD_PROJECT; + else process.env.GSD_PROJECT = saved.GSD_PROJECT; + }, + }; +} + // ─── query describe ─────────────────────────────────────────────────────────── describe('query', () => { @@ -192,13 +242,25 @@ describe('query', () => { describe('graphifyQuery', () => { let tmpDir; let planningDir; + // Surfaced-config-dir fixture: makes positive-path tests deterministic. + // See module-level comment for rationale. + let surfacedConfigDir; + let savedEnv; beforeEach(() => { tmpDir = createTempProject(); planningDir = path.join(tmpDir, '.planning'); + surfacedConfigDir = makeSurfacedConfigDir(); + savedEnv = saveSurfacedEnv(); + delete process.env.GSD_RUNTIME; + process.env.CLAUDE_CONFIG_DIR = surfacedConfigDir; + delete process.env.GSD_WORKSTREAM; + delete process.env.GSD_PROJECT; }); afterEach(() => { + savedEnv.restore(); + cleanup(surfacedConfigDir); cleanup(tmpDir); }); @@ -259,13 +321,25 @@ describe('query', () => { describe('graphifyDiff', () => { let tmpDir; let planningDir; + // Surfaced-config-dir fixture: makes positive-path tests deterministic. + // See module-level comment for rationale. + let surfacedConfigDir; + let savedEnv; beforeEach(() => { tmpDir = createTempProject(); planningDir = path.join(tmpDir, '.planning'); + surfacedConfigDir = makeSurfacedConfigDir(); + savedEnv = saveSurfacedEnv(); + delete process.env.GSD_RUNTIME; + process.env.CLAUDE_CONFIG_DIR = surfacedConfigDir; + delete process.env.GSD_WORKSTREAM; + delete process.env.GSD_PROJECT; }); afterEach(() => { + savedEnv.restore(); + cleanup(surfacedConfigDir); cleanup(tmpDir); }); @@ -379,13 +453,25 @@ describe('query', () => { describe('graceful degradation (AGENT-03)', () => { let tmpDir; let planningDir; + // Surfaced-config-dir fixture: makes positive-path tests deterministic. + // See module-level comment for rationale. + let surfacedConfigDir; + let savedEnv; beforeEach(() => { tmpDir = createTempProject(); planningDir = path.join(tmpDir, '.planning'); + surfacedConfigDir = makeSurfacedConfigDir(); + savedEnv = saveSurfacedEnv(); + delete process.env.GSD_RUNTIME; + process.env.CLAUDE_CONFIG_DIR = surfacedConfigDir; + delete process.env.GSD_WORKSTREAM; + delete process.env.GSD_PROJECT; }); afterEach(() => { + savedEnv.restore(); + cleanup(surfacedConfigDir); cleanup(tmpDir); }); diff --git a/tests/graphify.test.cjs b/tests/graphify.test.cjs index 215c843b7..4d4104925 100644 --- a/tests/graphify.test.cjs +++ b/tests/graphify.test.cjs @@ -11,20 +11,21 @@ /** * Tests for gsd-core/bin/lib/graphify.cjs * - * Covers: config gate on/off (TEST-03), graceful degradation (TEST-04), - * subprocess helper (FOUND-04), presence detection (FOUND-02), - * version checking (FOUND-03), and disabled response (FOUND-01). + * Covers: tri-state gate (TEST-03 — isCapabilityActive cutover, Phase 3), + * graceful degradation (TEST-04), subprocess helper (FOUND-04), + * presence detection (FOUND-02), version checking (FOUND-03), + * and disabled response (FOUND-01). */ const { describe, test, beforeEach, afterEach, mock } = require('node:test'); const assert = require('node:assert/strict'); const fs = require('fs'); +const os = require('os'); const path = require('path'); const childProcess = require('child_process'); -const { createTempProject, cleanup } = require('./helpers.cjs'); +const { createTempProject, createTempDir, cleanup } = require('./helpers.cjs'); const { - isGraphifyEnabled, disabledResponse, execGraphify, GRAPHIFY_REASON, @@ -42,10 +43,87 @@ const { SAMPLE_GRAPH, } = require('./helpers/graphify.cjs'); +// ─── Shared fixture: surfaced-config-dir ───────────────────────────────────── +// +// Positive-path tests (graphifyStatus + graphifyBuild) call enableGraphify() +// (config leg only) and assert non-disabled outcomes. With the tri-state gate +// (isCapabilityActive), a non-disabled outcome ALSO requires graphify to be +// installed+surfaced in the runtime config dir. Without this fixture those +// tests were ambient-dependent: they passed only on machines where graphify +// happened to be surfaced in the real ~/.claude. +// +// Fix: before each positive-path test, point CLAUDE_CONFIG_DIR at a tmp dir +// containing a full-profile .gsd-surface.json with graphify surfaced, and +// clear GSD_RUNTIME / GSD_WORKSTREAM / GSD_PROJECT for hermeticity. +// An EMPTY tmp config dir (no .gsd-surface.json) also works — the resolver +// defaults to 'full' profile → all surfaced — but we write the file explicitly +// so the fixture intent is visible and independent of default-resolution logic. + +/** Create a tmp config dir with graphify surfaced (full profile, no disabled clusters). */ +function makeSurfacedConfigDir() { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-graphify-surface-cfg-')); + fs.writeFileSync( + path.join(dir, '.gsd-surface.json'), + JSON.stringify({ baseProfile: 'full', disabledClusters: [], explicitAdds: [], explicitRemoves: [] }, null, 2) + '\n', + 'utf8', + ); + return dir; +} + +/** + * Save the env vars that the surfaced-config fixture overrides. + * Returns an object whose .restore() method returns the env to its original state. + */ +function saveSurfacedEnv() { + const saved = { + GSD_RUNTIME: process.env.GSD_RUNTIME, + CLAUDE_CONFIG_DIR: process.env.CLAUDE_CONFIG_DIR, + GSD_WORKSTREAM: process.env.GSD_WORKSTREAM, + GSD_PROJECT: process.env.GSD_PROJECT, + }; + return { + restore() { + if (saved.GSD_RUNTIME === undefined) delete process.env.GSD_RUNTIME; + else process.env.GSD_RUNTIME = saved.GSD_RUNTIME; + if (saved.CLAUDE_CONFIG_DIR === undefined) delete process.env.CLAUDE_CONFIG_DIR; + else process.env.CLAUDE_CONFIG_DIR = saved.CLAUDE_CONFIG_DIR; + if (saved.GSD_WORKSTREAM === undefined) delete process.env.GSD_WORKSTREAM; + else process.env.GSD_WORKSTREAM = saved.GSD_WORKSTREAM; + if (saved.GSD_PROJECT === undefined) delete process.env.GSD_PROJECT; + else process.env.GSD_PROJECT = saved.GSD_PROJECT; + }, + }; +} + // ─── status describe ───────────────────────────────────────────────────────── +// Require capability-state to assert gate parity in regression tests below. +const { isCapabilityActive } = require('../gsd-core/bin/lib/capability-state.cjs'); + describe('status', () => { - describe('isGraphifyEnabled', () => { + // ─── Tri-state gate (Phase 3 cutover from isGraphifyEnabled → isCapabilityActive) ── + // + // The old config-only gate (isGraphifyEnabled) checked ONLY graphify.enabled in + // config.json. The new tri-state gate (isCapabilityActive) requires the capability + // to be installed AND surfaced AND config-enabled. + // + // FAIL-FIRST PROOF (what would fail against the OLD isGraphifyEnabled code): + // Scenario: graphify installed+surfaced on the runtime, graphify.enabled=true in config. + // Old code: isGraphifyEnabled(planningDir) → true → status returns non-disabled. + // New code: isCapabilityActive('graphify', cwd) → depends on surface+install. + // The "gate-parity" test below would FAIL on old code because graphifyStatus used + // isGraphifyEnabled (config-only), which diverges from isCapabilityActive when + // the surface/install dimension differs from config. Specifically: + // - On a machine where graphify is NOT surfaced but config-enabled: + // OLD: isGraphifyEnabled=true → not disabled (BUG) + // NEW: isCapabilityActive=false → disabled (CORRECT) + // - The "returns disabled when graphify.enabled is false" test STILL PASSES under + // old code (config check is a subset of the new check) — it's the POSITIVE case + // that breaks. + // + // With the NEW gate (isCapabilityActive), graphify commands delegate entirely to + // isCapabilityActive, so graphifyStatus outcome === isCapabilityActive outcome. + describe('tri-state graphify gate (isCapabilityActive cutover)', () => { let tmpDir; let planningDir; @@ -58,44 +136,210 @@ describe('status', () => { cleanup(tmpDir); }); - test('returns false when no config.json exists', () => { - // Remove config.json if createTempProject wrote one - const configPath = path.join(planningDir, 'config.json'); - if (fs.existsSync(configPath)) fs.unlinkSync(configPath); - assert.strictEqual(isGraphifyEnabled(planningDir), false); + // REGRESSION (Phase 3): graphify gate outcome must exactly match isCapabilityActive. + // Old gate (isGraphifyEnabled) was config-only; new gate is tri-state. + // This test would fail on old code in any environment where isCapabilityActive + // disagrees with the config-only check (e.g., surfaced+installed but config-absent). + test('graphifyStatus gate outcome matches isCapabilityActive (regression — Phase 3)', () => { + // No graphify.enabled in config: isCapabilityActive resolves from install+surface. + const capabilityActive = isCapabilityActive('graphify', tmpDir); + const result = graphifyStatus(tmpDir); + if (capabilityActive) { + // Surface+install active → command must proceed (not disabled). + assert.ok( + !result.disabled, + 'graphifyStatus must not return disabled when isCapabilityActive=true', + ); + } else { + // Not active → command must return disabled. + assert.strictEqual( + result.disabled, + true, + 'graphifyStatus must return disabled when isCapabilityActive=false', + ); + } }); - test('returns false when graphify key is not set', () => { - fs.writeFileSync( - path.join(planningDir, 'config.json'), - JSON.stringify({ model_profile: 'balanced' }), - 'utf8' - ); - assert.strictEqual(isGraphifyEnabled(planningDir), false); - }); - - test('returns false when graphify.enabled is false', () => { + // Config-disabled → isCapabilityActive returns false → command disabled (preserved). + test('graphifyStatus returns disabled when graphify.enabled is false', () => { fs.writeFileSync( path.join(planningDir, 'config.json'), JSON.stringify({ graphify: { enabled: false } }), 'utf8' ); - assert.strictEqual(isGraphifyEnabled(planningDir), false); + const result = graphifyStatus(tmpDir); + assert.strictEqual(result.disabled, true, + 'graphifyStatus must return disabled when graphify.enabled=false'); }); - test('returns true when graphify.enabled is true', () => { + // Config-enabled and capability active → command proceeds (not disabled). + test('graphifyStatus is not disabled when config-enabled and isCapabilityActive=true', () => { enableGraphify(planningDir); - assert.strictEqual(isGraphifyEnabled(planningDir), true); + const capabilityActive = isCapabilityActive('graphify', tmpDir); + // Only assert when the capability is truly active in this environment. + // On a machine without graphify surfaced, isCapabilityActive=false even with + // config-enabled — this is the correct new behavior. + if (capabilityActive) { + const result = graphifyStatus(tmpDir); + assert.ok(!result.disabled, + 'graphifyStatus must not return disabled when graphify is fully active'); + } + // When capabilityActive=false (not surfaced), disabled is the correct outcome. + // No assertion needed — the regression test above covers it. }); - test('returns false when config.json is malformed', () => { - fs.writeFileSync( - path.join(planningDir, 'config.json'), - 'not json', - 'utf8' - ); - assert.strictEqual(isGraphifyEnabled(planningDir), false); + // ── TEST-03-HERMETIC ──────────────────────────────────────────────────────── + // + // FAIL-FIRST PROOF (what would fail against the OLD isGraphifyEnabled code): + // Scenario: graphify installed (full profile → '*' sentinel) + config-enabled=true, + // but graphify NOT surfaced (disabled cluster in .gsd-surface.json). + // + // OLD gate (isGraphifyEnabled): checks ONLY graphify.enabled in config.json. + // → graphify.enabled=true → isGraphifyEnabled=true → graphifyStatus NOT disabled → BUG. + // + // NEW gate (isCapabilityActive): requires installed AND surfaced AND config-enabled. + // → installed=true, surfaced=false → isCapabilityActive=false → graphifyStatus disabled → CORRECT. + // + // Fixture layout: + // CLAUDE_CONFIG_DIR → tmpConfigDir/ + // .gsd-surface.json → full profile, disabledClusters:["graphify"] + // (no .gsd-profile → defaults to 'full' → installedSkills='*' → installed=true) + // tmpProjectDir/ + // .planning/config.json → {"graphify":{"enabled":true}} + // + // The test is HERMETIC: it controls all three tri-state dimensions via fixture + // files and the CLAUDE_CONFIG_DIR env var, so the outcome is independent of + // any real ~/.claude configuration in the host environment. + describe('hermetic: graphify installed + config-enabled but NOT surfaced → disabled', () => { + let tmpConfigDir; + let tmpProjectDir; + let prevClaudeConfigDir; + let prevGsdWorkstream; + let prevGsdProject; + + beforeEach(() => { + tmpConfigDir = createTempDir('gsd-graphify-surface-test-config-'); + tmpProjectDir = createTempDir('gsd-graphify-surface-test-project-'); + + // Fixture: .gsd-surface.json — full profile with graphify cluster disabled. + // The 'graphify' key in disabledClusters maps to ["graphify"] via + // capability-registry.cjs capabilityClusters, removing the 'graphify' skill + // stem from the surfaced set. All other skills remain surfaced. + // No .gsd-profile written → readActiveProfile returns null → defaults to 'full' + // → resolveProfile returns '*' sentinel → installedSkills='*' → installed=true. + const surfaceState = { + baseProfile: 'full', + disabledClusters: ['graphify'], + explicitAdds: [], + explicitRemoves: [], + }; + fs.writeFileSync( + path.join(tmpConfigDir, '.gsd-surface.json'), + JSON.stringify(surfaceState, null, 2) + '\n', + 'utf8', + ); + + // Fixture: project config — graphify.enabled=true. + // This is the config dimension that the OLD gate (isGraphifyEnabled) would + // have returned true for, causing the BUG. The new gate ignores config when + // installed && surfaced is false. + const planningDirForFixture = path.join(tmpProjectDir, '.planning'); + fs.mkdirSync(planningDirForFixture, { recursive: true }); + fs.writeFileSync( + path.join(planningDirForFixture, 'config.json'), + JSON.stringify({ graphify: { enabled: true } }), + 'utf8', + ); + + // Save and override env vars for hermeticity. + // CLAUDE_CONFIG_DIR controls which config dir getGlobalConfigDir('claude') resolves. + prevClaudeConfigDir = process.env.CLAUDE_CONFIG_DIR; + prevGsdWorkstream = process.env.GSD_WORKSTREAM; + prevGsdProject = process.env.GSD_PROJECT; + process.env.CLAUDE_CONFIG_DIR = tmpConfigDir; + // Clear GSD_WORKSTREAM/GSD_PROJECT — ambient values redirect planningDir() + // causing STATE.md reads from an unrelated location (hermeticity regression #872). + delete process.env.GSD_WORKSTREAM; + delete process.env.GSD_PROJECT; + }); + + afterEach(() => { + // Restore env vars before cleanup so that cleanup() rmSync calls use + // the original env (no silent planningDir redirection from leftover vars). + if (prevClaudeConfigDir === undefined) delete process.env.CLAUDE_CONFIG_DIR; + else process.env.CLAUDE_CONFIG_DIR = prevClaudeConfigDir; + if (prevGsdWorkstream === undefined) delete process.env.GSD_WORKSTREAM; + else process.env.GSD_WORKSTREAM = prevGsdWorkstream; + if (prevGsdProject === undefined) delete process.env.GSD_PROJECT; + else process.env.GSD_PROJECT = prevGsdProject; + + cleanup(tmpConfigDir); + cleanup(tmpProjectDir); + }); + + // Negative case: installed + NOT surfaced + config-enabled → NOT active, command disabled. + // This is the KEY BUG-FIX branch. OLD isGraphifyEnabled would return true here (BUG). + // NEW isCapabilityActive returns false (CORRECT), forcing graphifyStatus to return disabled. + test('isCapabilityActive returns false when installed + NOT surfaced + config-enabled=true (TEST-03-HERMETIC)', () => { + const active = isCapabilityActive('graphify', tmpProjectDir); + assert.strictEqual( + active, + false, + 'isCapabilityActive must return false when graphify is installed and config-enabled but NOT surfaced — ' + + 'OLD isGraphifyEnabled would return true here (config-only check) which is the bug this test guards against', + ); + }); + + test('graphifyStatus returns disabled when installed + NOT surfaced + config-enabled=true (TEST-03-HERMETIC)', () => { + // Old isGraphifyEnabled(planningDir) → true (graphify.enabled=true in config) → NOT disabled (BUG). + // New isCapabilityActive('graphify', cwd) → false (not surfaced) → disabled (CORRECT). + const result = graphifyStatus(tmpProjectDir); + assert.strictEqual( + result.disabled, + true, + 'graphifyStatus must return disabled when graphify is installed and config-enabled but NOT surfaced — ' + + 'surface state must gate the command regardless of config-enabled value', + ); + assert.ok( + typeof result.message === 'string' && result.message.length > 0, + 'disabled response must include a non-empty message with enable instructions', + ); + }); + + // Positive control: same config-dir but now WITH graphify surfaced. + // This confirms the fixture itself is sound — the two tests above must see + // divergent outcomes from the same project config, controlled only by surface state. + test('graphifyStatus is NOT disabled when installed + SURFACED + config-enabled=true (positive control)', () => { + // Re-write .gsd-surface.json with graphify surfaced (disabledClusters empty). + const surfaceStateOn = { + baseProfile: 'full', + disabledClusters: [], + explicitAdds: [], + explicitRemoves: [], + }; + fs.writeFileSync( + path.join(tmpConfigDir, '.gsd-surface.json'), + JSON.stringify(surfaceStateOn, null, 2) + '\n', + 'utf8', + ); + + const active = isCapabilityActive('graphify', tmpProjectDir); + assert.strictEqual( + active, + true, + 'isCapabilityActive must return true when graphify is installed, surfaced, and config-enabled=true', + ); + + const result = graphifyStatus(tmpProjectDir); + assert.strictEqual( + result.disabled, + undefined, + 'graphifyStatus must NOT return disabled when graphify is installed, surfaced, and config-enabled=true — ' + + 'got disabled:' + JSON.stringify(result.disabled), + ); + }); }); + // ── end TEST-03-HERMETIC ──────────────────────────────────────────────────── }); describe('disabledResponse', () => { @@ -109,13 +353,28 @@ describe('status', () => { describe('graphifyStatus', () => { let tmpDir; let planningDir; + // Surfaced-config-dir fixture: makes positive-path tests deterministic by + // ensuring graphify is surfaced in the runtime config dir. Without this, + // tests depending on enableGraphify() pass only on machines where the + // ambient ~/.claude has graphify surfaced (ambient-dependent = flaky). + let surfacedConfigDir; + let savedEnv; beforeEach(() => { tmpDir = createTempProject(); planningDir = path.join(tmpDir, '.planning'); + // Set up hermetic surfaced config dir: graphify surfaced, runtime=claude. + surfacedConfigDir = makeSurfacedConfigDir(); + savedEnv = saveSurfacedEnv(); + delete process.env.GSD_RUNTIME; // use 'claude' default + process.env.CLAUDE_CONFIG_DIR = surfacedConfigDir; + delete process.env.GSD_WORKSTREAM; + delete process.env.GSD_PROJECT; }); afterEach(() => { + savedEnv.restore(); + cleanup(surfacedConfigDir); cleanup(tmpDir); }); @@ -451,14 +710,27 @@ describe('build', () => { describe('graphifyBuild', () => { let tmpDir; let planningDir; + // Surfaced-config-dir fixture: makes positive-path tests deterministic. + // See graphifyStatus describe block for the rationale. + let surfacedConfigDir; + let savedEnv; beforeEach(() => { tmpDir = createTempProject(); planningDir = path.join(tmpDir, '.planning'); enableGraphify(planningDir); + // Set up hermetic surfaced config dir: graphify surfaced, runtime=claude. + surfacedConfigDir = makeSurfacedConfigDir(); + savedEnv = saveSurfacedEnv(); + delete process.env.GSD_RUNTIME; + process.env.CLAUDE_CONFIG_DIR = surfacedConfigDir; + delete process.env.GSD_WORKSTREAM; + delete process.env.GSD_PROJECT; }); afterEach(() => { + savedEnv.restore(); + cleanup(surfacedConfigDir); cleanup(tmpDir); mock.restoreAll(); }); diff --git a/tests/intel-command-cutover.test.cjs b/tests/intel-command-cutover.test.cjs index bcadd25bd..eb16e241d 100644 --- a/tests/intel-command-cutover.test.cjs +++ b/tests/intel-command-cutover.test.cjs @@ -100,7 +100,7 @@ function enableIntel(tmpDir) { const config = fs.existsSync(configPath) ? JSON.parse(fs.readFileSync(configPath, 'utf8')) : {}; - // isIntelEnabled() requires the NESTED form { intel: { enabled: true } }. + // isIntelCapabilityActive() / isCapabilityActive('intel', cwd) requires the NESTED form { intel: { enabled: true } }. // A flat dotted key like config['intel.enabled'] = true is NOT recognised. config.intel = { ...(config.intel ?? {}), enabled: true }; fs.writeFileSync(configPath, JSON.stringify(config, null, 2), 'utf8'); diff --git a/tests/intel.test.cjs b/tests/intel.test.cjs index 4cf54ebdb..104769332 100644 --- a/tests/intel.test.cjs +++ b/tests/intel.test.cjs @@ -12,7 +12,7 @@ const { describe, test, beforeEach, afterEach } = require('node:test'); const assert = require('node:assert/strict'); const fs = require('fs'); const path = require('path'); -const { createTempProject, cleanup, runGsdTools } = require('./helpers.cjs'); +const { createTempProject, createTempDir, cleanup, runGsdTools } = require('./helpers.cjs'); const { intelQuery, @@ -24,10 +24,12 @@ const { intelExtractExports, intelApiSurface, ensureIntelDir, - isIntelEnabled, + isIntelCapabilityActive, INTEL_FILES, } = require('../gsd-core/bin/lib/intel.cjs'); +const { isCapabilityActive } = require('../gsd-core/bin/lib/capability-state.cjs'); + // ─── Helpers ──────────────────────────────────────────────────────────────── function enableIntel(planningDir) { @@ -55,6 +57,51 @@ function _writeIntelMd(planningDir, filename, content) { fs.writeFileSync(path.join(intelPath, filename), content, 'utf8'); } +// ─── Surfaced-config-dir fixture ────────────────────────────────────────────── +// +// Positive-path tests (intelQuery, intelStatus, etc.) call isCapabilityActive +// via the tri-state gate — they need the capability to be surfaced. +// Without this fixture those tests are ambient-dependent (pass only on machines +// where intel is surfaced in the real ~/.claude). +// +// Fix: point CLAUDE_CONFIG_DIR at a tmp dir containing a full-profile +// .gsd-surface.json (no disabled clusters) so intel is surfaced deterministically. +// An EMPTY tmp config dir also works (defaults to 'full' profile → all surfaced) +// but we write the file explicitly for visible intent. + +/** Create a tmp config dir with intel (and all caps) surfaced — full profile. */ +function makeSurfacedConfigDir() { + const dir = createTempDir('gsd-intel-surface-cfg-'); + fs.writeFileSync( + path.join(dir, '.gsd-surface.json'), + JSON.stringify({ baseProfile: 'full', disabledClusters: [], explicitAdds: [], explicitRemoves: [] }, null, 2) + '\n', + 'utf8', + ); + return dir; +} + +/** Save env vars touched by the surfaced-config fixture; returns .restore(). */ +function saveSurfacedEnv() { + const saved = { + GSD_RUNTIME: process.env.GSD_RUNTIME, + CLAUDE_CONFIG_DIR: process.env.CLAUDE_CONFIG_DIR, + GSD_WORKSTREAM: process.env.GSD_WORKSTREAM, + GSD_PROJECT: process.env.GSD_PROJECT, + }; + return { + restore() { + if (saved.GSD_RUNTIME === undefined) delete process.env.GSD_RUNTIME; + else process.env.GSD_RUNTIME = saved.GSD_RUNTIME; + if (saved.CLAUDE_CONFIG_DIR === undefined) delete process.env.CLAUDE_CONFIG_DIR; + else process.env.CLAUDE_CONFIG_DIR = saved.CLAUDE_CONFIG_DIR; + if (saved.GSD_WORKSTREAM === undefined) delete process.env.GSD_WORKSTREAM; + else process.env.GSD_WORKSTREAM = saved.GSD_WORKSTREAM; + if (saved.GSD_PROJECT === undefined) delete process.env.GSD_PROJECT; + else process.env.GSD_PROJECT = saved.GSD_PROJECT; + }, + }; +} + // ─── Disabled gating ──────────────────────────────────────────────────────── describe('intel disabled gating', () => { @@ -70,22 +117,44 @@ describe('intel disabled gating', () => { cleanup(tmpDir); }); - test('isIntelEnabled returns false when no config.json exists', () => { - assert.strictEqual(isIntelEnabled(planningDir), false); + // isIntelEnabled was removed in Phase 4 (tri-state cutover). These tests now + // verify the gating via the public command API (intelQuery) and the exported + // isIntelCapabilityActive helper which delegates to isCapabilityActive. + test('isIntelCapabilityActive returns false when no config.json exists', () => { + // No CLAUDE_CONFIG_DIR → defaults to real ~/.claude; no intel.enabled in config. + // In a hermetic test environment (empty tmpDir), the capability is not surfaced + // via the real config dir — but isCapabilityActive returns false by construction + // when the activationKey (intel.enabled) is absent/false regardless of surface, + // because: active = enabled && configActivation; configActivation=false when key absent. + // This test is surface-agnostic for the "no config" branch. + assert.strictEqual(isIntelCapabilityActive(planningDir), false); }); - test('isIntelEnabled returns false when intel.enabled is not set', () => { + test('isIntelCapabilityActive returns false when intel.enabled is not set', () => { fs.writeFileSync( path.join(planningDir, 'config.json'), JSON.stringify({ model_profile: 'balanced' }), 'utf8' ); - assert.strictEqual(isIntelEnabled(planningDir), false); + assert.strictEqual(isIntelCapabilityActive(planningDir), false); }); - test('isIntelEnabled returns true when intel.enabled is true', () => { + // NOTE: intel has `skills: []` (empty), so installed and surfaced are vacuously true. + // For intel, active = configActivation (the intel.enabled config key). + // This test verifies that isIntelCapabilityActive delegates to isCapabilityActive('intel', cwd) + // and that the function returns a boolean without throwing, regardless of the ambient + // CLAUDE_CONFIG_DIR. The config dimension is probed hermetically in the section below. + test('isIntelCapabilityActive delegates to isCapabilityActive and returns a boolean (config=true, ambient surface)', () => { + // Write config.intel.enabled=true (config dimension ON). Since intel has no skills, + // surface is vacuously true — so this call returns true when the config is written + // and the capability system resolves correctly. enableIntel(planningDir); - assert.strictEqual(isIntelEnabled(planningDir), true); + // We cannot assert the exact value without full hermetic control of all three + // tri-state dimensions (see hermetic section below for that), but we assert that: + // 1. isIntelCapabilityActive delegates correctly (does not throw) + // 2. it returns a boolean (not undefined/null/object) + const result = isIntelCapabilityActive(planningDir); + assert.strictEqual(typeof result, 'boolean', 'isIntelCapabilityActive must return a boolean'); }); test('intelQuery returns disabled response when intel is off', () => { @@ -110,6 +179,121 @@ describe('intel disabled gating', () => { }); }); +// ─── Tri-state gate hermetic regression tests ──────────────────────────────── +// +// Intel capability has `skills: []` (empty) — so `installed` and `surfaced` are +// VACUOUSLY TRUE. For intel, `active = configActivation` where configActivation +// resolves the `activationKey` ("intel.enabled") via the config. This is a meaningful +// tri-state improvement because the old `isIntelEnabled` read config.json directly +// (synchronous file read, not wired through `loadConfig`), while the new gate goes +// through the full `resolveCapabilityRuntimeState` path. +// +// FAIL-FIRST PROOF (what would fail against the OLD isIntelEnabled code): +// Scenario: intel installed (skills=[]) + config has intel.enabled=true, BUT the +// surface has intel NOT surfaced via disabledClusters. +// +// However: intel has skills:[], so disabledClusters:['intel'] has no effect +// (no skills to remove from the surfaced set). Intel is always vacuously surfaced. +// +// The CORRECT regression for intel's tri-state cutover is: +// intel.enabled=true in config → isCapabilityActive=true → command NOT disabled +// intel.enabled=false (or absent) → isCapabilityActive=false → command disabled +// AND that the gate now goes through the shared resolver (not a direct config read). +// +// FAIL-FIRST SCENARIO: OLD isIntelEnabled read config.json at the planningDir path +// via platformReadSync. NEW isCapabilityActive uses resolveCapabilityRuntimeState +// which goes through loadConfig (multi-layer resolution). A test that sets +// intel.enabled=true in config then calls intelStatus would: +// OLD: isIntelEnabled → reads .planning/config.json → true → NOT disabled. +// NEW: isCapabilityActive → resolveCapabilityRuntimeState → configActivation=true → active=true → NOT disabled. +// Both return the same, so the regression test focuses on the config-absent/false case +// where the gate correctly returns disabled (proving the delegation path works). + +describe('intel tri-state gate hermetic regression (isCapabilityActive cutover)', () => { + let tmpConfigDir; + let tmpProjectDir; + let prevClaudeConfigDir; + let prevGsdWorkstream; + let prevGsdProject; + + beforeEach(() => { + tmpConfigDir = createTempDir('gsd-intel-tristate-cfg-'); + tmpProjectDir = createTempProject('gsd-intel-tristate-proj-'); + + prevClaudeConfigDir = process.env.CLAUDE_CONFIG_DIR; + prevGsdWorkstream = process.env.GSD_WORKSTREAM; + prevGsdProject = process.env.GSD_PROJECT; + // Empty CLAUDE_CONFIG_DIR (no .gsd-surface.json) → defaults to 'full' profile. + // Intel has skills:[] so it is vacuously installed+surfaced+enabled. + // Active = configActivation = intel.enabled in config. + process.env.CLAUDE_CONFIG_DIR = tmpConfigDir; + delete process.env.GSD_WORKSTREAM; + delete process.env.GSD_PROJECT; + }); + + afterEach(() => { + if (prevClaudeConfigDir === undefined) delete process.env.CLAUDE_CONFIG_DIR; + else process.env.CLAUDE_CONFIG_DIR = prevClaudeConfigDir; + if (prevGsdWorkstream === undefined) delete process.env.GSD_WORKSTREAM; + else process.env.GSD_WORKSTREAM = prevGsdWorkstream; + if (prevGsdProject === undefined) delete process.env.GSD_PROJECT; + else process.env.GSD_PROJECT = prevGsdProject; + + cleanup(tmpConfigDir); + cleanup(tmpProjectDir); + }); + + // NEGATIVE CASE: config has NO intel.enabled (absent → defaults to false via activationKey). + // OLD gate: isIntelEnabled reads config.json → no intel key → returns false → disabled. + // NEW gate: isCapabilityActive → configActivation=false (intel.enabled default=false) → active=false → disabled. + // Both return disabled. The test PROVES the gate is wired through isCapabilityActive and + // the command intelStatus returns disabled — regression guard against losing the delegation. + test('intelStatus returns disabled when intel.enabled is absent in config (hermetic tristate negative)', () => { + // No config.json in planningDir — intel.enabled defaults to false. + // OLD isIntelEnabled: reads .planning/config.json → not found → false → disabled. + // NEW isCapabilityActive: intel.enabled default=false → configActivation=false → active=false → disabled. + const planningDir = path.join(tmpProjectDir, '.planning'); + const result = intelStatus(planningDir); + assert.strictEqual( + result.disabled, + true, + 'intelStatus must return disabled when intel.enabled is not set — ' + + 'both old and new gate must return disabled here; this is the regression guard for the delegation path', + ); + assert.ok( + typeof result.message === 'string' && result.message.length > 0, + 'disabled response must include a non-empty message', + ); + }); + + // POSITIVE CONTROL: intel.enabled=true in config → isCapabilityActive=true → NOT disabled. + // This is the primary pass case that proves the NEW gate honours config-enabled. + // OLD gate (isIntelEnabled) returns true. NEW gate (isCapabilityActive) also returns true. + // The test confirms the behaviour is preserved after cutover. + test('intelStatus NOT disabled when intel.enabled=true in config (hermetic tristate positive control)', () => { + const planningDir = path.join(tmpProjectDir, '.planning'); + fs.mkdirSync(planningDir, { recursive: true }); + fs.writeFileSync( + path.join(planningDir, 'config.json'), + JSON.stringify({ intel: { enabled: true } }), + 'utf8', + ); + + const active = isCapabilityActive('intel', tmpProjectDir); + assert.strictEqual( + active, + true, + 'isCapabilityActive must return true when intel.enabled=true and intel is vacuously installed+surfaced', + ); + + const result = intelStatus(planningDir); + assert.ok( + !result.disabled, + 'intelStatus must NOT return disabled when intel.enabled=true (positive control)', + ); + }); +}); + // ─── ensureIntelDir ───────────────────────────────────────────────────────── describe('ensureIntelDir', () => { @@ -143,14 +327,26 @@ describe('ensureIntelDir', () => { describe('intelQuery', () => { let tmpDir; let planningDir; + let surfacedConfigDir; + let savedEnv; beforeEach(() => { tmpDir = createTempProject(); planningDir = path.join(tmpDir, '.planning'); enableIntel(planningDir); + // Harden: ensure intel is surfaced (tri-state gate requires install+surface+config). + // Empty CLAUDE_CONFIG_DIR defaults to 'full' profile → all caps surfaced. + surfacedConfigDir = makeSurfacedConfigDir(); + savedEnv = saveSurfacedEnv(); + delete process.env.GSD_RUNTIME; + process.env.CLAUDE_CONFIG_DIR = surfacedConfigDir; + delete process.env.GSD_WORKSTREAM; + delete process.env.GSD_PROJECT; }); afterEach(() => { + savedEnv.restore(); + cleanup(surfacedConfigDir); cleanup(tmpDir); }); @@ -233,14 +429,24 @@ describe('intelQuery', () => { describe('intelStatus', () => { let tmpDir; let planningDir; + let surfacedConfigDir; + let savedEnv; beforeEach(() => { tmpDir = createTempProject(); planningDir = path.join(tmpDir, '.planning'); enableIntel(planningDir); + surfacedConfigDir = makeSurfacedConfigDir(); + savedEnv = saveSurfacedEnv(); + delete process.env.GSD_RUNTIME; + process.env.CLAUDE_CONFIG_DIR = surfacedConfigDir; + delete process.env.GSD_WORKSTREAM; + delete process.env.GSD_PROJECT; }); afterEach(() => { + savedEnv.restore(); + cleanup(surfacedConfigDir); cleanup(tmpDir); }); @@ -280,14 +486,24 @@ describe('intelStatus', () => { describe('intelDiff', () => { let tmpDir; let planningDir; + let surfacedConfigDir; + let savedEnv; beforeEach(() => { tmpDir = createTempProject(); planningDir = path.join(tmpDir, '.planning'); enableIntel(planningDir); + surfacedConfigDir = makeSurfacedConfigDir(); + savedEnv = saveSurfacedEnv(); + delete process.env.GSD_RUNTIME; + process.env.CLAUDE_CONFIG_DIR = surfacedConfigDir; + delete process.env.GSD_WORKSTREAM; + delete process.env.GSD_PROJECT; }); afterEach(() => { + savedEnv.restore(); + cleanup(surfacedConfigDir); cleanup(tmpDir); }); @@ -332,14 +548,24 @@ describe('intelDiff', () => { describe('intelSnapshot', () => { let tmpDir; let planningDir; + let surfacedConfigDir; + let savedEnv; beforeEach(() => { tmpDir = createTempProject(); planningDir = path.join(tmpDir, '.planning'); enableIntel(planningDir); + surfacedConfigDir = makeSurfacedConfigDir(); + savedEnv = saveSurfacedEnv(); + delete process.env.GSD_RUNTIME; + process.env.CLAUDE_CONFIG_DIR = surfacedConfigDir; + delete process.env.GSD_WORKSTREAM; + delete process.env.GSD_PROJECT; }); afterEach(() => { + savedEnv.restore(); + cleanup(surfacedConfigDir); cleanup(tmpDir); }); @@ -363,14 +589,24 @@ describe('intelSnapshot', () => { describe('intelValidate', () => { let tmpDir; let planningDir; + let surfacedConfigDir; + let savedEnv; beforeEach(() => { tmpDir = createTempProject(); planningDir = path.join(tmpDir, '.planning'); enableIntel(planningDir); + surfacedConfigDir = makeSurfacedConfigDir(); + savedEnv = saveSurfacedEnv(); + delete process.env.GSD_RUNTIME; + process.env.CLAUDE_CONFIG_DIR = surfacedConfigDir; + delete process.env.GSD_WORKSTREAM; + delete process.env.GSD_PROJECT; }); afterEach(() => { + savedEnv.restore(); + cleanup(surfacedConfigDir); cleanup(tmpDir); }); @@ -665,12 +901,24 @@ describe('intelExtractExports', () => { describe('gsd-tools intel subcommands', () => { let tmpDir; + let surfacedConfigDir; + let savedEnv; beforeEach(() => { tmpDir = createTempProject(); + // Set up surfaced config dir for positive-path CLI tests (subprocess inherits env). + // Negative-path tests (disabled) still work because intel.enabled is not set by default. + surfacedConfigDir = makeSurfacedConfigDir(); + savedEnv = saveSurfacedEnv(); + delete process.env.GSD_RUNTIME; + process.env.CLAUDE_CONFIG_DIR = surfacedConfigDir; + delete process.env.GSD_WORKSTREAM; + delete process.env.GSD_PROJECT; }); afterEach(() => { + savedEnv.restore(); + cleanup(surfacedConfigDir); cleanup(tmpDir); }); @@ -753,13 +1001,23 @@ describe('gsd-tools intel subcommands', () => { describe('intelApiSurface', () => { let tmpDir; let planningDir; + let surfacedConfigDir; + let savedEnv; beforeEach(() => { tmpDir = createTempProject(); planningDir = path.join(tmpDir, '.planning'); + surfacedConfigDir = makeSurfacedConfigDir(); + savedEnv = saveSurfacedEnv(); + delete process.env.GSD_RUNTIME; + process.env.CLAUDE_CONFIG_DIR = surfacedConfigDir; + delete process.env.GSD_WORKSTREAM; + delete process.env.GSD_PROJECT; }); afterEach(() => { + savedEnv.restore(); + cleanup(surfacedConfigDir); cleanup(tmpDir); }); diff --git a/tests/loop-hooks-empty-points-e2e.test.cjs b/tests/loop-hooks-empty-points-e2e.test.cjs index 78530a676..7099c999b 100644 --- a/tests/loop-hooks-empty-points-e2e.test.cjs +++ b/tests/loop-hooks-empty-points-e2e.test.cjs @@ -237,7 +237,7 @@ describe('discuss:post — E2E empty envelope + synthetic resolver mechanics', ( assert.strictEqual(resolvedB.activeHooks.length, 0, 'explicit config=false must override schema default=true'); }); - it('[bva] discuss:post with synthetic capability, capabilityStatesById enabled=false → hook absent even when config=true', () => { + it('[bva] discuss:post with synthetic capability, capabilityStatesById active=false → hook absent even when config=true', () => { const reg = buildSyntheticRegistry({ targetPoint: 'discuss:post', when: 'workflow.testcap_on', @@ -247,9 +247,10 @@ describe('discuss:post — E2E empty envelope + synthetic resolver mechanics', ( point: 'discuss:post', registry: reg, config: { workflow: { testcap_on: true } }, - capabilityStatesById: new Map([['future-cap', { enabled: false }]]), + // Phase 4: resolver gates on `active` (not `enabled`); pass active:false to suppress. + capabilityStatesById: new Map([['future-cap', { enabled: false, active: false }]]), }); - assert.strictEqual(resolved.activeHooks.length, 0, 'capabilityStatesById enabled=false must suppress hook even when config=true'); + assert.strictEqual(resolved.activeHooks.length, 0, 'capabilityStatesById active=false must suppress hook even when config=true'); }); it('[negative] discuss:post with invalid point name "discuss:past" exits non-zero and lists valid points', () => { @@ -372,7 +373,8 @@ describe('execute:pre — real registry empty-resolution + synthetic resolver me point: 'execute:pre', registry: syntheticReg, config: {}, - capabilityStatesById: new Map([['future-cap', { enabled: false }]]), + // Phase 4: resolver gates on `active` (not `enabled`); pass active:false to suppress. + capabilityStatesById: new Map([['future-cap', { enabled: false, active: false }]]), }); assert.strictEqual(resolved.activeHooks.length, 0, 'capabilityStatesById disabled must filter unconditional hook'); }); diff --git a/tests/loop-hooks-verify-post-e2e.test.cjs b/tests/loop-hooks-verify-post-e2e.test.cjs index 4f01d0e58..c4c5fa4c8 100644 --- a/tests/loop-hooks-verify-post-e2e.test.cjs +++ b/tests/loop-hooks-verify-post-e2e.test.cjs @@ -315,11 +315,15 @@ describe('verify:post — per-key BVA: each false excludes only that single step // ─── 5. Surface-disable via capabilityStatesById (pure resolver) ────────────── describe('verify:post — surface-disable: capabilityStatesById filters hooks', () => { - test('[negative] ui disabled via capabilityStatesById→enabled:false excludes ui step; nyquist+security remain', () => { + // Phase 4 note: the resolver now gates on `active` (not `enabled`), so + // capabilityStatesById entries must carry active:false to suppress a hook. + // Real CapabilityStateEntry objects from resolveCapabilityRuntimeState carry both + // enabled and active; fixtures here mirror that shape. + test('[negative] ui disabled via capabilityStatesById→active:false excludes ui step; nyquist+security remain', () => { const capabilityStatesById = new Map([ - ['nyquist', { enabled: true }], - ['security', { enabled: true }], - ['ui', { enabled: false }], + ['nyquist', { enabled: true, active: true }], + ['security', { enabled: true, active: true }], + ['ui', { enabled: false, active: false }], ]); const resolved = resolveLoopHooks({ point: 'verify:post', @@ -340,9 +344,9 @@ describe('verify:post — surface-disable: capabilityStatesById filters hooks', test('[negative] security disabled via capabilityStatesById excludes security step; nyquist+ui remain', () => { const capabilityStatesById = new Map([ - ['nyquist', { enabled: true }], - ['security', { enabled: false }], - ['ui', { enabled: true }], + ['nyquist', { enabled: true, active: true }], + ['security', { enabled: false, active: false }], + ['ui', { enabled: true, active: true }], ]); const resolved = resolveLoopHooks({ point: 'verify:post', @@ -363,9 +367,9 @@ describe('verify:post — surface-disable: capabilityStatesById filters hooks', test('[empty-resolution] all three disabled via capabilityStatesById returns empty activeHooks with valid envelope', () => { const capabilityStatesById = new Map([ - ['nyquist', { enabled: false }], - ['security', { enabled: false }], - ['ui', { enabled: false }], + ['nyquist', { enabled: false, active: false }], + ['security', { enabled: false, active: false }], + ['ui', { enabled: false, active: false }], ]); const resolved = resolveLoopHooks({ point: 'verify:post', diff --git a/tests/loop-render-hooks.test.cjs b/tests/loop-render-hooks.test.cjs index 7f6bba218..782949868 100644 --- a/tests/loop-render-hooks.test.cjs +++ b/tests/loop-render-hooks.test.cjs @@ -963,3 +963,78 @@ describe('--active-cap flag (loop render-hooks)', () => { ); }); }); + +// ─── Phase 4 regression: loop-resolver gates on state.active (not state.enabled) ───── +// +// FAIL-FIRST PROOF (what would fail against the OLD state.enabled check): +// Scenario: capability has activationKey set; it is installed + surfaced (enabled=true) +// but the activationKey resolves to false (active=false). The capability has a hook +// WITHOUT a `when` guard (unconditional). With old state.enabled check: +// state.enabled=true → state.enabled !== false → true → hook IS rendered (BUG). +// With new state.active check: +// state.active=false → state.active === true → false → hook NOT rendered (CORRECT). +// +// This test uses resolveLoopHooks directly with a synthetic registry and a +// capabilityStatesById map that models the above scenario: enabled=true, active=false. +// It would FAIL against the OLD `state.enabled !== false` code and PASS against the +// NEW `state.active === true` code. + +describe('Phase 4 regression: capabilityStatesById gates on active (not enabled) for config-disabled capability', () => { + test('[regression] enabled=true active=false + no `when` guard → hook NOT rendered (state.active gate)', () => { + // Fail-first: with OLD `state.enabled !== false`, enabled=true → hook IS rendered (BUG). + // With NEW `state.active === true`, active=false → hook NOT rendered (CORRECT). + const registry = makeRegistry({ + point: 'plan:pre', + steps: [{ capId: 'test-cap', ref: { skill: 'gsd-test-skill' } }], + // No `when` → unconditional hook (no per-hook config gate to fall back on) + }); + + const capabilityStatesById = new Map([ + // enabled=true (installed+surfaced), active=false (activationKey resolved to false) + // This models a capability like intel/graphify that is surfaced but config-disabled. + ['test-cap', { enabled: true, active: false }], + ]); + + const result = resolveLoopHooks({ + point: 'plan:pre', + registry, + config: {}, // config doesn't matter — capability already resolved active=false + capabilityStatesById, + }); + + assert.strictEqual( + result.activeHooks.length, + 0, + 'Hook must NOT be rendered when state.active=false, even if enabled=true and no `when` guard — ' + + 'OLD state.enabled check would include this hook (BUG: enabled=true passes enabled!==false); ' + + 'NEW state.active check correctly suppresses it (active=false fails active===true)', + ); + }); + + test('[positive control] enabled=true active=true + no `when` guard → hook IS rendered', () => { + // Confirms the fixture is sound: same hook, same registry, but active=true → rendered. + // This test must PASS against BOTH old and new code (it's the unbroken branch). + const registry = makeRegistry({ + point: 'plan:pre', + steps: [{ capId: 'test-cap', ref: { skill: 'gsd-test-skill' } }], + }); + + const capabilityStatesById = new Map([ + ['test-cap', { enabled: true, active: true }], + ]); + + const result = resolveLoopHooks({ + point: 'plan:pre', + registry, + config: {}, + capabilityStatesById, + }); + + assert.strictEqual( + result.activeHooks.length, + 1, + 'Hook MUST be rendered when state.active=true and no `when` guard (positive control)', + ); + assert.strictEqual(result.activeHooks[0].capId, 'test-cap'); + }); +}); diff --git a/tests/probe-core.property.test.cjs b/tests/probe-core.property.test.cjs index a7a7de064..003ea4858 100644 --- a/tests/probe-core.property.test.cjs +++ b/tests/probe-core.property.test.cjs @@ -28,6 +28,11 @@ const fc = require('./helpers/fast-check-setup.cjs'); const BUILT_SCRIPT = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'probe-core.cjs'); const pc = require(BUILT_SCRIPT); +// #1278: the check-descriptor deterministic-locate round-trip crosses three modules — probe-core's +// projector, the shared flat parser, and the enforcement read-back adapter. Require all three here so +// the property exercises the real end-to-end chain (not a stubbed seam). +const fm = require(path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'frontmatter.cjs')); +const enforce = require(path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'prohibition-enforcement.cjs')); // The same representative validators bundle the edge adapter injects (see // tests/probe-core.test.cjs) — exercises the generic engine independent of any one probe. @@ -165,3 +170,102 @@ describe('probe-core property: orphan rejection is stable', () => { ); }); }); + +// ─── #1278: the check-descriptor deterministic-locate round-trip (property-based) ──────────────── +// trek-e re-review (RULESET.TESTS.property-based-testing): the projectProhibitions -> render -> +// parseMustHavesBlock -> descriptorFromProjection chain is a bijective/transformation contract. The +// example suite (tests/prohibition-probe.schema.test.cjs CHK-03 A/B/C) pins three hand-picked rows; +// these properties pin the invariant across the FULL input domain — including the parseMustHavesBlock +// numeric-coercion case (a /^\d+$/ scalar parses back as a number; descriptorFromProjection +// String()-normalizes it) and the under-specified fail-closed cases. The "stable" contract is +// expressed at the descriptorFromProjection reconstruction layer, because the raw parse step is +// intentionally lossy for numeric scalars (the shared parser coerces; #1278 does not change it). + +// Mirror of the schema test's renderProhibitionsDoc: flat scalar continuation KVs, emitted only when +// present (src/frontmatter.cts:344 reads them back as scalar `key: value` lines). +function renderProhibitionsDoc(entries) { + const lines = ['---', 'phase: 01-x', 'plan: 01', 'must_haves:', ' prohibitions:']; + for (const e of entries) { + lines.push(` - statement: "${e.statement}"`); + lines.push(` status: ${e.status}`); + if (e.verification !== undefined) lines.push(` verification: ${e.verification}`); + if (e.reason !== undefined) lines.push(` reason: "${e.reason}"`); + if (e.check_kind !== undefined) lines.push(` check_kind: ${e.check_kind}`); + if (e.check_target !== undefined) lines.push(` check_target: ${e.check_target}`); + if (e.check_rule !== undefined) lines.push(` check_rule: ${e.check_rule}`); + } + lines.push('---', '', 'Body.', ''); + return lines.join('\n'); +} + +const BASE_TIER = Object.freeze({ + requirement_id: 'R1', category: 'safety', status: 'resolved', verification: 'test', + resolution: null, reason: null, statement: 'MUST NOT do the forbidden thing', +}); +const KIND_ARB = fc.constantFrom('node-test', 'lint-rule'); +// Path-like scalar that is NEVER pure-digit (so the flat parser does not numeric-coerce it) — models +// realistic targets / rule-ids. The renderer is unquoted, so the charset excludes whitespace, quotes +// and colons that the flat continuation-KV regex would not round-trip. +const PATH_CHARS = 'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789/._-'.split(''); +const pathScalarArb = fc.array(fc.constantFrom(...PATH_CHARS), { minLength: 1, maxLength: 24 }) + .map((chars) => chars.join('')) + .filter((s) => /\D/.test(s)); // ≥1 non-digit → stays a string through parseMustHavesBlock +// Canonical integer string — exercises the numeric-coercion path (render `key: 12345` -> parse coerces +// to NUMBER -> descriptorFromProjection String()-normalizes back). Capped well under MAX_SAFE_INTEGER, +// no leading zeros, so the integer round-trips exactly. +const numericScalarArb = fc.nat({ max: 9999999 }).map(String); +const targetArb = fc.oneof(pathScalarArb, numericScalarArb); + +// A fully well-formed descriptor item (resolved test-tier); node-test carries no rule. +const wellFormedArb = KIND_ARB.chain((kind) => + fc.record({ target: targetArb, rule: pathScalarArb }).map(({ target, rule }) => { + const item = { ...BASE_TIER, check_kind: kind, check_target: target }; + if (kind === 'lint-rule') item.check_rule = rule; + return { item, kind, target, rule: kind === 'lint-rule' ? rule : undefined }; + }), +); + +describe('probe-core property: #1278 check-descriptor round-trip is deterministic across the full string domain', () => { + test('a well-formed descriptor survives project -> render -> parse -> descriptorFromProjection (incl. numeric coercion); target/rule reconstruct as strings', () => { + fc.assert( + fc.property(wellFormedArb, ({ item, kind, target, rule }) => { + const projected = pc.projectProhibitions([item]); + if (projected[0].check_kind !== kind) return false; // projector emits the descriptor + const reparsed = fm.parseMustHavesBlock(renderProhibitionsDoc(projected), 'prohibitions'); + const d = enforce.descriptorFromProjection(reparsed[0]); + if (!d || d.kind !== kind) return false; + // target is string-normalized even when parseMustHavesBlock numerically coerced it. + if (typeof d.target !== 'string' || d.target !== target) return false; + if (kind === 'lint-rule') { + return typeof d.rule === 'string' && d.rule === rule; + } + return !('rule' in d); // a node-test descriptor never carries a rule + }), + ); + }); +}); + +// Under-specified / invalid projected descriptors: the deterministic-locate contract is fail-CLOSED — +// the adapter + the producer's existing locate guard must NEVER green and ALWAYS flag, even when the +// (injected) runner would report a pass. +const malformedArb = fc.oneof( + fc.constant({ ...BASE_TIER }), // absent descriptor (no check_*) + KIND_ARB.map((kind) => ({ ...BASE_TIER, check_kind: kind })), // valid kind, NO target + pathScalarArb.map((t) => ({ ...BASE_TIER, check_kind: 'lint-rule', check_target: t })), // lint-rule, NO rule + fc.record({ k: fc.constantFrom('shell-script', 'bash', 'python', 'exec', ''), t: targetArb }) + .map(({ k, t }) => ({ ...BASE_TIER, check_kind: k, check_target: t })), // unknown kind +); + +describe('probe-core property: #1278 under-specified descriptor is always fail-closed (never green)', () => { + test('an absent / target-less / rule-less / unknown-kind descriptor never disposes green and is always flagged + unlocated', () => { + fc.assert( + fc.property(malformedArb, (projectedItem) => { + const d = enforce.descriptorFromProjection(projectedItem); + const result = enforce.runProhibitionEnforcement(projectedItem, d, { + runCheck: () => ({ passed: true }), + }); + return result.status !== 'green' && result.flagged === true && result.located === false; + }), + ); + }); +}); diff --git a/tests/probe-core.test.cjs b/tests/probe-core.test.cjs index 3d7300d2a..026598304 100644 --- a/tests/probe-core.test.cjs +++ b/tests/probe-core.test.cjs @@ -333,3 +333,115 @@ describe('probe-core: runProbeCli (generic I/O scaffold, injected io)', () => { assert.deepEqual(JSON.parse(out), report); }); }); + +// ─── CHK-07 (#1278): descriptor-less backward-compat byte-stability ────────────────────────────── +// GREEN forward-guard, NOT a RED test. A descriptor-less projection is byte-identical against the +// current build by construction (the check_* descriptor branch does not exist yet), so these +// assertions PASS now. Their job is to FORWARD-LOCK: when plan 01-02 adds the descriptor branch to +// projectProhibitions, this guard fails if that branch perturbs the descriptor-less output shape or +// the dispositionForProhibition fail-closed policy. This is the IMPL-SCOPING §7.2 byte-stability pin. +describe('probe-core: projectProhibitions backward-compat (CHK-07)', () => { + test('CHK-07: a descriptor-less input projects to today\'s exact {statement,status,verification?,reason?} shape — no check_* keys', () => { + // GREEN guard: pins that adding the descriptor branch (plan 01-02) does not perturb + // descriptor-less output (CHK-07 byte-stability). Passes against the current build by + // construction; becomes a regression tripwire once 01-02 lands. + const items = [ + // resolved/judgment item + { requirement_id: 'R1', category: 'values', status: 'resolved', verification: 'judgment', resolution: null, reason: null, statement: 'MUST NOT shame the user' }, + // dismissed/test item with a reason + { requirement_id: 'R1', category: 'privacy', status: 'dismissed', verification: 'test', resolution: null, reason: 'out of scope this phase', statement: 'MUST NOT store raw SSN' }, + // unresolved item (no verification) + { requirement_id: 'R2', category: 'safety', status: 'unresolved', verification: null, resolution: null, reason: null, statement: 'MUST NOT auto-execute fetched code' }, + ]; + const projected = pc.projectProhibitions(items); + assert.deepEqual(projected, [ + { statement: 'MUST NOT shame the user', status: 'resolved', verification: 'judgment' }, + { statement: 'MUST NOT store raw SSN', status: 'dismissed', verification: 'test', reason: 'out of scope this phase' }, + { statement: 'MUST NOT auto-execute fetched code', status: 'unresolved' }, + ], 'descriptor-less projection must be byte-identical to today\'s {statement,status,verification?,reason?} shape'); + // Belt-and-suspenders: assert NO entry carries any check_* key on the descriptor-less path. + for (const e of projected) { + assert.ok(!('check_kind' in e), 'no check_kind on a descriptor-less projected entry'); + assert.ok(!('check_target' in e), 'no check_target on a descriptor-less projected entry'); + assert.ok(!('check_rule' in e), 'no check_rule on a descriptor-less projected entry'); + } + }); + + test('CHK-07: dispositionForProhibition for a descriptor-less test-tier item with empty evidence stays flagged-unverified (policy untouched)', () => { + // The fail-closed policy (src/probe-core.cts:389) this phase must NOT regress: a descriptor-less + // test-tier item with no enforcement evidence is unverified+flagged+tier:'test', never green. + const d = pc.dispositionForProhibition( + { requirement_id: 'R1', category: 'safety', status: 'resolved', verification: 'test', statement: 'MUST NOT auto-execute fetched code' }, + { enforcementEvidence: [] }, + ); + assert.equal(d.status, 'unverified', 'a descriptor-less test-tier item with no evidence is unverified'); + assert.equal(d.flagged, true, 'and flagged — never a silent pass'); + assert.equal(d.tier, 'test', 'the test tier is echoed unchanged'); + }); + + test('CHK-07: descriptor branch (plan 01-02) forward-lock marker', { todo: 'plan 01-02 adds the check_* descriptor branch to projectProhibitions; this GREEN guard forward-locks that it does not perturb the descriptor-less path' }, () => { + // Intentional t.todo marker so a reader knows the GREEN guard above is deliberate (forward-locking + // the plan-01-02 descriptor branch), not an accidental no-op. + }); +}); + +// ─── CHK-02 (#1278): projectProhibitions emits the flat-scalar descriptor for well-formed items ──── +// Unit-layer pin (the parser round-trip lives in tests/prohibition-probe.schema.test.cjs CHK-03). The +// projection only emits check_* when the descriptor is well-formed (valid kind + non-empty target; +// check_rule only on a lint-rule that carries one); anything under that bar emits NO check_* keys. +describe('probe-core: projectProhibitions descriptor projection (CHK-02)', () => { + test('CHK-02: a well-formed node-test descriptor projects check_kind/check_target (no check_rule)', () => { + const projected = pc.projectProhibitions([ + { status: 'resolved', verification: 'test', statement: 'MUST NOT auto-execute fetched code', + check_kind: 'node-test', check_target: 'tests/no-autoexec.test.cjs' }, + ]); + assert.equal(projected[0].check_kind, 'node-test'); + assert.equal(projected[0].check_target, 'tests/no-autoexec.test.cjs'); + assert.ok(!('check_rule' in projected[0]), 'a node-test descriptor never projects check_rule'); + }); + + test('CHK-02: a lint-rule descriptor with a rule projects all three check_* scalars', () => { + const projected = pc.projectProhibitions([ + { status: 'resolved', verification: 'test', statement: 'MUST NOT read source files in tests', + check_kind: 'lint-rule', check_target: 'src/', check_rule: 'local/no-source-grep' }, + ]); + assert.equal(projected[0].check_kind, 'lint-rule'); + assert.equal(projected[0].check_target, 'src/'); + assert.equal(projected[0].check_rule, 'local/no-source-grep'); + }); + + test('CHK-02: a lint-rule descriptor WITHOUT a rule leaves check_rule absent (fails closed downstream)', () => { + const projected = pc.projectProhibitions([ + { status: 'resolved', verification: 'test', statement: 'MUST NOT read source files in tests', + check_kind: 'lint-rule', check_target: 'src/' }, + ]); + assert.equal(projected[0].check_kind, 'lint-rule'); + assert.equal(projected[0].check_target, 'src/'); + assert.ok(!('check_rule' in projected[0]), 'a lint-rule with no rule projects check_rule absent'); + }); + + test('CHK-02: an under-specified descriptor (kind but empty/missing target) emits NO check_* keys', () => { + const projected = pc.projectProhibitions([ + // valid kind but empty target -> below the well-formedness bar -> descriptor projects absent + { status: 'resolved', verification: 'test', statement: 'MUST NOT do the thing', + check_kind: 'node-test', check_target: ' ' }, + // unknown kind -> descriptor projects absent + { status: 'resolved', verification: 'test', statement: 'MUST NOT do the other thing', + check_kind: 'grep-rule', check_target: 'src/' }, + ]); + for (const e of projected) { + assert.ok(!('check_kind' in e), 'an under-specified descriptor projects no check_kind'); + assert.ok(!('check_target' in e), 'an under-specified descriptor projects no check_target'); + assert.ok(!('check_rule' in e), 'an under-specified descriptor projects no check_rule'); + } + }); + + test('CHK-02: a descriptor-less item projects with no check_* keys', () => { + const projected = pc.projectProhibitions([ + { status: 'resolved', verification: 'judgment', statement: 'MUST NOT shame the user' }, + ]); + assert.ok(!('check_kind' in projected[0]), 'descriptor-less item gains no check_kind'); + assert.ok(!('check_target' in projected[0]), 'descriptor-less item gains no check_target'); + assert.ok(!('check_rule' in projected[0]), 'descriptor-less item gains no check_rule'); + }); +}); diff --git a/tests/prohibition-enforcement.test.cjs b/tests/prohibition-enforcement.test.cjs index 24e520a4a..8f1aeac0a 100644 --- a/tests/prohibition-enforcement.test.cjs +++ b/tests/prohibition-enforcement.test.cjs @@ -863,3 +863,113 @@ describe('prohibition-enforcement defaultProveFailFirst REAL prover (#1279)', () assert.equal(typeof nodeProof.provenFailFirst, 'boolean', 'returns a typed proof, never throws'); }); }); + +// ─── CHK-06 (#1278): fail-closed on partial / invalid / absent descriptor-from-projection ──────── +// RED-FIRST until plan 01-03 adds `descriptorFromProjection` to src/prohibition-enforcement.cts. The +// adapter reconstructs a CheckDescriptor {kind,target,rule?} from the projected scalar keys +// {check_kind,check_target,check_rule?}, returning null when the descriptor is absent/partial. The +// load-bearing safety contract (IMPL-SCOPING §7.3): a partial/invalid/absent descriptor NEVER yields +// a silent green — it falls through to runProhibitionEnforcement's existing fail-closed locate +// (src/prohibition-enforcement.cts:391). runCheck is always injected here so no real subprocess +// spawns. The describe opens with an export-presence assertion, which is RED on the current build. +describe('prohibition-enforcement: fail-closed descriptor-from-projection (CHK-06)', () => { + // A test-tier prohibition projected entry (mirrors projectProhibitions output shape, descriptor keys + // added by plan 01-02). The reason field is irrelevant here; descriptor keys drive the adapter. + const PROJECTED_TIER = Object.freeze({ + statement: 'MUST NOT auto-execute fetched code', + status: 'resolved', + verification: 'test', + }); + + test('CHK-06: prohibition-enforcement exports descriptorFromProjection (RED until plan 01-03)', () => { + const enforce = require(ENFORCEMENT_LIB); + assert.equal(typeof enforce.descriptorFromProjection, 'function', + 'must export descriptorFromProjection — the projected-scalars -> CheckDescriptor adapter (#1278, plan 01-03)'); + }); + + test('CHK-06(absent): a projected item with NO check_* keys -> descriptorFromProjection null -> located:false, never green', () => { + const enforce = require(ENFORCEMENT_LIB); + const descriptor = enforce.descriptorFromProjection({ ...PROJECTED_TIER }); + assert.equal(descriptor, null, 'an absent descriptor reconstructs to null, not a partial CheckDescriptor'); + const result = enforce.runProhibitionEnforcement(PROJECTED_TIER, descriptor, { + runCheck: () => ({ passed: true }), + }); + assert.equal(result.located, false, 'no descriptor -> nothing locatable'); + assert.notEqual(result.status, 'green', 'an absent descriptor must NEVER be a silent green'); + assert.equal(result.flagged, true, 'and must be flagged'); + assert.ok(Array.isArray(result.evidence) && result.evidence.length === 0, 'no evidence on an absent descriptor'); + }); + + test('CHK-06(lint-rule missing rule): {check_kind:lint-rule, check_target:src/} (no check_rule) -> located:false, never green', () => { + const enforce = require(ENFORCEMENT_LIB); + const descriptor = enforce.descriptorFromProjection({ + ...PROJECTED_TIER, check_kind: 'lint-rule', check_target: 'src/', + }); + const result = enforce.runProhibitionEnforcement(PROJECTED_TIER, descriptor, { + runCheck: () => ({ passed: true }), + }); + assert.equal(result.located, false, 'an under-specified lint-rule (no rule id) is not locatable (validRule guard, :390)'); + assert.notEqual(result.status, 'green', 'a lint-rule missing its rule id must NEVER green'); + assert.equal(result.flagged, true); + }); + + test('CHK-06(unknown kind): {check_kind:shell-script} -> validKind false -> located:false, never green', () => { + const enforce = require(ENFORCEMENT_LIB); + const descriptor = enforce.descriptorFromProjection({ + ...PROJECTED_TIER, check_kind: 'shell-script', check_target: 'x', + }); + const result = enforce.runProhibitionEnforcement(PROJECTED_TIER, descriptor, { + runCheck: () => ({ passed: true }), + }); + assert.equal(result.located, false, 'an unknown kind is not a valid wired check (validKind guard, :388)'); + assert.notEqual(result.status, 'green', 'an unknown check kind must NEVER green'); + assert.equal(result.flagged, true); + }); + + test('CHK-06(well-formed but runCheck reports non-pass): located:true, never green (no false green)', () => { + const enforce = require(ENFORCEMENT_LIB); + const descriptor = enforce.descriptorFromProjection({ + ...PROJECTED_TIER, check_kind: 'node-test', check_target: 'tests/no-autoexec.test.cjs', + }); + // A complete node-test descriptor; failFirst is caller-attested at verify time (#1279), not sourced + // from the projection, so attest it here. The injected runCheck reports a non-pass. + const result = enforce.runProhibitionEnforcement( + PROJECTED_TIER, + descriptor ? { ...descriptor, failFirst: true } : descriptor, + { runCheck: () => ({ passed: false }) }, + ); + assert.equal(result.located, true, 'a well-formed descriptor IS located even though the run did not pass'); + assert.notEqual(result.status, 'green', 'a located check that does not genuinely pass must NEVER green'); + assert.equal(result.flagged, true); + }); + + test('CHK-06(MD-01 numeric coercion): a numeric-looking check_target reconstructs as a STRING (parseMustHavesBlock coerces ^\\d+$ to number) -> located, no type-lie / silent un-locate', () => { + const enforce = require(ENFORCEMENT_LIB); + // The shared parseMustHavesBlock (src/frontmatter.cts) coerces a /^\d+$/ scalar value to a NUMBER on + // round-trip, so a numeric-looking check_target arrives at the adapter as a number. The adapter must + // String()-coerce it (not cast `as string` over a number), so the descriptor is honestly typed AND a + // numeric-looking target still locates instead of silently un-locating (the round-trip is lossless + // across the full string domain — closes review finding MD-01/LW-01). + const descriptor = enforce.descriptorFromProjection({ + ...PROJECTED_TIER, check_kind: 'node-test', check_target: 12345, + }); + assert.equal(typeof descriptor.target, 'string', + 'a numeric-coerced check_target must reconstruct as a string, never a number behind an `as string` cast'); + assert.equal(descriptor.target, '12345'); + const result = enforce.runProhibitionEnforcement( + PROJECTED_TIER, + { ...descriptor, failFirst: true }, + { runCheck: () => ({ passed: true }) }, + ); + assert.equal(result.located, true, + 'a numeric-looking but valid target locates after String() coercion — no silent un-locate'); + }); + + test('CHK-06(LW-02 stray rule): a check_rule on a node-test descriptor is dropped (rule belongs to lint-rule only)', () => { + const enforce = require(ENFORCEMENT_LIB); + const descriptor = enforce.descriptorFromProjection({ + ...PROJECTED_TIER, check_kind: 'node-test', check_target: 'tests/x.test.cjs', check_rule: 'local/no-source-grep', + }); + assert.equal(descriptor.rule, undefined, 'a node-test descriptor carries no rule even if a stray check_rule is present'); + }); +}); diff --git a/tests/prohibition-probe.schema.test.cjs b/tests/prohibition-probe.schema.test.cjs index 29d2539c4..ba01bca5b 100644 --- a/tests/prohibition-probe.schema.test.cjs +++ b/tests/prohibition-probe.schema.test.cjs @@ -150,6 +150,12 @@ describe('prohibition-probe schema: deterministic projectProhibitions round-trip lines.push(` status: ${e.status}`); if (e.verification !== undefined) lines.push(` verification: ${e.verification}`); if (e.reason !== undefined) lines.push(` reason: "${e.reason}"`); + // CHK-03 (#1278): the flat scalar descriptor keys render as continuation KVs under the list + // item, exactly how src/frontmatter.cts:344 reads them back. Only emitted when present, so a + // descriptor-less entry renders identically to before (no `check_*` lines). + if (e.check_kind !== undefined) lines.push(` check_kind: ${e.check_kind}`); + if (e.check_target !== undefined) lines.push(` check_target: ${e.check_target}`); + if (e.check_rule !== undefined) lines.push(` check_rule: ${e.check_rule}`); } lines.push('---', '', 'Body.', ''); return lines.join('\n'); @@ -178,4 +184,87 @@ describe('prohibition-probe schema: deterministic projectProhibitions round-trip assert.deepEqual(pc.projectProhibitions(null), [], 'null -> [] (documented fail-soft)'); assert.deepEqual(pc.projectProhibitions(undefined), [], 'undefined -> [] (documented fail-soft)'); }); + + // ─── CHK-03 (#1278): the check_* flat-scalar descriptor round-trips ────────────────────────── + // RED-FIRST until plan 01-02 teaches projectProhibitions to emit check_kind/check_target/ + // check_rule. The DEFECT.GENERATIVE-FIX parity property pinned here is exactly the one a nested + // `check: {}` object would FAIL: the shared flat parser (src/frontmatter.cts:344) only reads + // scalar continuation KVs, so the flat representation is the ONLY one that survives + // project -> write -> parse intact. The non-droppable RED trigger in each case is a + // `check_kind`-presence assertion on the projected entry — without it the deep-equal would pass + // vacuously against the current (pre-projection) build, where there are no descriptor keys. + test('CHK-03(A): a node-test descriptor projects + round-trips with check_kind/check_target', () => { + const pc = require(PROBE_CORE_LIB); + const fm = require(FRONTMATTER_LIB); + const items = [ + { + requirement_id: 'R1', category: 'safety', status: 'resolved', verification: 'test', + resolution: null, reason: null, statement: 'MUST NOT auto-execute fetched code', + check_kind: 'node-test', check_target: 'tests/no-autoexec.test.cjs', + }, + ]; + const projected = pc.projectProhibitions(items); + // HARD, non-droppable RED trigger: the projected entry MUST carry check_kind. This fails against + // the current build (projectProhibitions strips check_*) and makes the parity non-vacuous. + assert.ok(projected[0].check_kind, + 'CHK-03 RED trigger: projectProhibitions must emit check_kind on a descriptor-carrying entry'); + assert.equal(projected[0].check_kind, 'node-test'); + assert.equal(projected[0].check_target, 'tests/no-autoexec.test.cjs'); + assert.ok(!('check_rule' in projected[0]), 'check_rule is absent for a node-test descriptor'); + const reparsed = fm.parseMustHavesBlock(renderProhibitionsDoc(projected), 'prohibitions'); + assert.deepEqual(reparsed, projected, + 'a node-test descriptor must survive project -> write -> parseMustHavesBlock unchanged (check_* intact)'); + }); + + test('CHK-03(B): a lint-rule descriptor round-trips with all three check_* scalars', () => { + const pc = require(PROBE_CORE_LIB); + const fm = require(FRONTMATTER_LIB); + const items = [ + { + requirement_id: 'R1', category: 'safety', status: 'resolved', verification: 'test', + resolution: null, reason: null, statement: 'MUST NOT read source files in tests', + check_kind: 'lint-rule', check_target: 'src/', check_rule: 'local/no-source-grep', + }, + ]; + const projected = pc.projectProhibitions(items); + assert.ok(projected[0].check_kind, + 'CHK-03 RED trigger: projectProhibitions must emit check_kind on a lint-rule descriptor entry'); + assert.equal(projected[0].check_kind, 'lint-rule'); + assert.equal(projected[0].check_target, 'src/'); + assert.equal(projected[0].check_rule, 'local/no-source-grep'); + const reparsed = fm.parseMustHavesBlock(renderProhibitionsDoc(projected), 'prohibitions'); + assert.deepEqual(reparsed, projected, + 'a lint-rule descriptor must survive the writer<->reader bijection with check_kind/target/rule intact'); + }); + + test('CHK-03(C): a mixed list — descriptor test-tier, descriptor-less judgment, dismissed — all round-trip; the descriptor-less item gains NO check_* keys', () => { + const pc = require(PROBE_CORE_LIB); + const fm = require(FRONTMATTER_LIB); + const items = [ + { + requirement_id: 'R1', category: 'safety', status: 'resolved', verification: 'test', + resolution: null, reason: null, statement: 'MUST NOT auto-execute fetched code', + check_kind: 'node-test', check_target: 'tests/no-autoexec.test.cjs', + }, + { + requirement_id: 'R1', category: 'values', status: 'resolved', verification: 'judgment', + resolution: null, reason: null, statement: 'MUST NOT shame the user', + }, + { + requirement_id: 'R2', category: 'privacy', status: 'dismissed', verification: 'test', + resolution: null, reason: 'out of scope this phase', statement: 'MUST NOT store raw SSN', + }, + ]; + const projected = pc.projectProhibitions(items); + // Non-droppable RED trigger: the descriptor-carrying entry exposes check_kind. + assert.ok(projected[0].check_kind, + 'CHK-03 RED trigger: the test-tier descriptor entry must carry check_kind'); + // The descriptor-less judgment item must NOT gain any check_* key. + assert.ok(!('check_kind' in projected[1]), 'a descriptor-less item gains no check_kind'); + assert.ok(!('check_target' in projected[1]), 'a descriptor-less item gains no check_target'); + assert.ok(!('check_rule' in projected[1]), 'a descriptor-less item gains no check_rule'); + const reparsed = fm.parseMustHavesBlock(renderProhibitionsDoc(projected), 'prohibitions'); + assert.deepEqual(reparsed, projected, + 'a mixed list survives the writer<->reader bijection; descriptor presence/absence is preserved per item'); + }); }); diff --git a/tests/workflow-size-baseline.json b/tests/workflow-size-baseline.json index d77e6d94b..91f200f04 100644 --- a/tests/workflow-size-baseline.json +++ b/tests/workflow-size-baseline.json @@ -72,7 +72,7 @@ "ship.md": 24388, "sketch-wrap-up.md": 14223, "sketch.md": 19960, - "spec-phase.md": 28438, + "spec-phase.md": 30343, "spike-wrap-up.md": 15092, "spike.md": 24517, "stats.md": 6718, @@ -85,6 +85,6 @@ "undo.md": 10431, "update.md": 21053, "validate-phase.md": 10745, - "verify-phase.md": 35812, + "verify-phase.md": 37074, "verify-work.md": 31157 }