* feat(#1950): broken-windows ledger — cross-phase defect register gating ship Adds a new capability (#1950) that operationalizes GSD's no-defer discipline as a tracked, enforced artifact: accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths across phases, and /gsd-ship blocks while any entry is open. Implementation: - src/broken-windows.cts → gsd-core/bin/lib/broken-windows.cjs: typed IR + I/O entry points (parseLedger/renderLedger/appendWindow/markWaived/markFixed + cmdWindowsStatus/Append/Waive/MarkFixed). Frozen REASON enum for typed error assertions. Windows-safe atomic rename with retry on transient EPERM/EBUSY/EACCES. - gsd-tools.cjs: new subcommand (status | append | waive | fixed), wired via routeWindows + HOST_COMMAND_ROUTERS.windows. - capabilities/broken-windows/capability.json: one ship:pre gate with artifact-frontmatter-equals predicate on WINDOWS.md open_count == 0. activationKey windows.enabled (default true) + sibling windows.enforce (default true, separate so tracking can precede enforcement). - gsd-core/workflows/ship.md: capId==broken-windows branch in preflight, sibling to security — reads gsd_run windows status --raw, fails closed on open_count > 0 or unreadable ledger. - agents/gsd-executor.md: extends the existing ## Known Stubs instruction to also append to WINDOWS.md via gsd_run windows append (best-effort, never blocks execution). - agents/gsd-verifier.md: new Step 8b — record unmet truths + human-verify items in WINDOWS.md. - gsd-core/workflows/progress.md: surfaces open + waived counts. - docs/COMMANDS.md + CONTEXT.md glossary entry + docs/INVENTORY.md: document the gate, waiver mechanism, and new module. - tests/broken-windows.test.cjs: pure + CLI behavioral coverage + fast-check roundtrip property; fail-closed on malformed ledger; security boundary on path traversal in --file. Backward-compatible: a project with no .planning/WINDOWS.md reports open_count: 0 and ships cleanly. Disable enforcement per-project with gsd config-set windows.enforce false (tracking continues, gate stays open). * chore(#1950): ratchet size baselines, defer verifier integration - Workflow size baseline: ship.md 25575→27928, progress.md 31789→32632 (broken-windows preflight branch + open-windows surface). - Agent size baseline: gsd-executor.md 46644→47951 (Known Stubs → also appends to WINDOWS.md). gsd-verifier.md unchanged. - LARGE_CAP (49152) preempted the planned verifier integration (gsd-verifier.md was at 49140 pre-PR — 12 bytes of headroom, not the documented 'real headroom'). Verifier integration deferred to a follow-up PR that extracts the VERIFICATION.md template (lines 739-859) to gsd-core/references/ — a pre-existing cap-tightness defect this PR exposed but does not expand scope to fix. Verifier integration is not in the issue's acceptance criteria (executor writes is; unmet-truths recording was an enhancement, not a gate). * fix(#1950): gate default-off, rename to workflow.windows_enforce, regen goldens Test-failure-driven fixes after first gsd-test run on db8733c8f failed 44 cases (pre-existing structural tests encoded 'ship:pre has 1 gate' / 'all caps off → empty hooks'): - capability manifest: rename windows.enabled+windows.enforce (default true) → single federated key workflow.windows_enforce (default FALSE, opt-in). Matches security's workflow.security_enforce convention and makes the adr857 all-caps-off test pass without modification (the test's buildAllFalseConfig handles workflow.* out of the box). Default-OFF keeps the gate out of the registry's default ship:pre resolution so existing loop-hooks-ship-pre-e2e structural assertions (exactly 1 gate, capId 'security') stay valid; users opt in via gsd config-set workflow.windows_enforce true. - drop activationKey (security doesn't have one either; workflow.* key doubles as the activation toggle). - regenerate docs/reference/capability-matrix.md to include broken-windows (capability-matrix-sync test). - regenerate tests/fixtures/golden-install-parity/*.json (18 runtimes) — installer now emits the new capability + lib file. - update CONTEXT.md, docs/COMMANDS.md, docs/FEATURES.md, ship.md, agents/gsd-executor.md to use the new key name and /gsd:colon slash syntax (slash-command-namespace test). - restore accidentally-regressed /gsd:capture in progress.md. Tracking-only by default; enforcement is opt-in. Acceptance criterion '/gsd-ship fails while any ledger entry is open' is met when workflow.windows_enforce=true (test fixture enables it). * test(#1950): update ship:pre structural invariants for 2-gate registry - loop-hooks-ship-pre-e2e: the registry now declares 2 gates at ship:pre (security + broken-windows), regardless of activation. Activation tests above still pin security-only or empty behavior via fixtures; these structural tests pin the REGISTRY shape, which has 2 gates as of #1950. - workflow-size-baseline: ship.md 27928→27945 (workflow.windows_enforce rename added 17 bytes). * fix(#1950): review H1+H2+M1+M2+M3 — fence-injection, EACCES fail-closed, cleanup, strict line, stryker Adversarial isolated review (Step 6.3) found 2 HIGH findings that block the PR and 3 mediums. All addressed: H1 (HIGH): description containing the markdown 3-backtick fence would terminate the ledger's JSON code block early inside JSON.stringify output (JSON doesn't escape backticks), corrupting the file and bricking the next parse. Fix: use a 4-backtick fence (json ... ) which JSON.stringify cannot produce on its own, AND validate that no entry text field contains a 4-backtick run (reject at append time with new WINDOWS_INVALID_TEXT reason code). Locked by a regression test. H2 (HIGH): readLedgerOrNull swallowed ALL fs errors as 'no ledger', silently returning open_count:0 on EACCES/EPERM/EIO. The ship gate would then pass on an unreadable ledger — the precise vector the workflow doc claims is impossible. Fix: only ENOENT returns null; every other fs error propagates as WINDOWS_LEDGER_MALFORMED so the gate blocks and the operator sees a real diagnostic. Locked by a regression test that chmod 000s a ledger with open_count=1 and asserts the result is never a false-green 0. M1: writeLedgerAtomic left an orphaned .tmp file on rename failure. Wrapped renameWithRetry in try/catch with best-effort unlink. M2: validateLine silently coerced 'abc' → NaN → null, hiding type drift. Removed the line === 0 special case (was undocumented) and made the error message match the strict check. Now any non-positive- integer line value throws, including strings. M3: tests/broken-windows.test.cjs (with its fast-check property test) was not in stryker.config.mjs DEFAULT_TEST_CMD — Stryker would mutate src/broken-windows.cts but no test would catch the mutations, producing false surviving-mutant scores. Added to the list. L1 (dead throw e after error()), L7 (line boundary tests, H1/H2 regression tests, 4-backtick CLI test) also addressed. * docs(#1950): inline concurrency + busy-wait notes (review L2+L3) * fix(#1950): regen goldens against latest gsd-tools; correct --line 0 boundary test gsd-test v4 caught two issues: - goldens I regenerated earlier (commit 526682084) predated the L1 routeWindows catch-block cleanup (commit dd844d565). Regenerated via 'npm run gen:golden' against current HEAD so the install parity hash for gsd-tools.cjs matches. - 'append --line boundary' test expected --line 0 to succeed with null entry.line, but the M2 fix correctly rejects 0 (lines are 1-indexed; 0 is not a valid source line). Updated the boundary test to assert --line 0 fails alongside -1 and 'abc'. * chore(#1950): regen goldens after rebase onto next * chore(#1950): quick.md baseline 50699→50993 (correct resolution from next rebase) * chore(changeset): backfill pr:2441 in .changeset/broken-windows-ledger.md * fix(#1950): renderTable escapes backslash before pipe (CodeQL incomplete-sanitization) CodeQL flagged the markdown-table cell escaper: String(s ?? '').replace(/\|/g, '\\|') — it escapes pipe but not backslash first. A description containing '\|' would render as '\\|' which markdown parses as 'literal backslash' + 'cell separator', splitting the column. Fix: escape backslash FIRST (each \ → \\), then pipe (each | → \|). Now a description with '\|' renders as '\\\\|' (literal '\\' + escaped pipe), which markdown renders as a single '\|' inside the cell. The JSON code block (the parse source-of-truth) was already correctly escaped via JSON.stringify; only the display-only table was affected. Locked by a regression test that: 1. Verifies the JSON block reparses with the description intact. 2. Walks the rendered table row counting unescaped pipes — must be exactly 11 (the row separators for 10 cells), proving no in-cell pipe added a split.
163 lines
9.5 KiB
Markdown
163 lines
9.5 KiB
Markdown
# Capability matrix reference
|
|
|
|
> **Generated file — do not edit by hand.**
|
|
> This matrix is generated from the capability registry by
|
|
> `scripts/gen-capability-matrix.cjs` and kept honest by a drift guard
|
|
> (`tests/capability-matrix-sync.test.cjs` runs `--check`). Any manual edit is
|
|
> overwritten on the next generation run. To change a capability's declared
|
|
> metadata, edit the corresponding `capabilities/<id>/capability.json` and run
|
|
> `node scripts/gen-capability-matrix.cjs --write`.
|
|
|
|
See also: [ADR-1244](../adr/1244-capability-ecosystem.md) —
|
|
[Capability manifest fields](#manifest-field-reference) —
|
|
[The capability trust model](../explanation/capability-trust-model.md)
|
|
|
|
---
|
|
|
|
## Column definitions
|
|
|
|
| Column | Description |
|
|
|---|---|
|
|
| **id** | Canonical capability identifier; unique across first- and third-party capabilities. Reserved prefixes: `gsd-`, `gsd-core-`, `anthropic-`. |
|
|
| **role** | `feature` — extends what the loop does; `runtime` — adapts GSD to a specific AI runtime/IDE. |
|
|
| **tier** | `core` — always active; `standard` — active when the runtime supports it; `full` — opt-in or runtime-specific. |
|
|
| **engines.gsd** | Semver RANGE expressing host-version compatibility. A hard gate at install and at load. `—` means the capability declares no range. |
|
|
| **extension points** | The loop points this capability registers hooks into (from the registry's `byLoopPoint` index). `—` means it registers none (typical for runtime capabilities, whose job is surface emission). |
|
|
| **hook kinds** | Which of `step`, `contribution`, `gate` the capability's hooks use. `—` means none. |
|
|
| **source** | `first-party` — ships with GSD Core; `third-party` — installed from an external source via `gsd capability install`. |
|
|
|
|
> **On versions.** This matrix intentionally omits a per-capability `version`
|
|
> column. First-party capabilities are versioned **in lockstep** with the GSD
|
|
> Core package (their `capability.json` `version` always equals the GSD release
|
|
> version), so a per-row version would simply repeat the package version and
|
|
> churn the committed file on every release. The stable host-compatibility
|
|
> signal — `engines.gsd` — is shown instead. A third-party capability's exact
|
|
> version is recorded in the per-runtime ledger (`.gsd-capabilities.json`) at
|
|
> install time.
|
|
|
|
---
|
|
|
|
## Native (first-party) capabilities
|
|
|
|
First-party capabilities are implicitly trusted: they ship as part of the GSD
|
|
Core package and are stamped with the package version at release (per
|
|
ADR-1244 D6). They are not subject to the consent or integrity-pin flow applied
|
|
to third-party capabilities.
|
|
|
|
### Feature capabilities (role: feature) — 20
|
|
|
|
Feature capabilities extend what the loop does — contributing research,
|
|
planning, execution, verification, or ship artefacts at the loop extension
|
|
points.
|
|
|
|
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|
|
|---|---|---|---|---|---|---|
|
|
| `ai-integration` | feature | full | `>=1.6.0` | `plan:pre`, `verify:pre` | step, contribution, gate | first-party |
|
|
| `assumption-delta` | feature | full | `>=1.6.0` | `plan:pre` | contribution | first-party |
|
|
| `audit` | feature | full | `>=1.6.0` | — | — | first-party |
|
|
| `broken-windows` | feature | full | `>=1.7.0` | `ship:pre` | gate | first-party |
|
|
| `claude-orchestration` | feature | full | `>=1.7.0` | `plan:post`, `execute:wave:pre` | contribution | first-party |
|
|
| `code-review` | feature | full | `>=1.6.0` | `execute:post` | step | first-party |
|
|
| `drift` | feature | full | `>=1.6.0` | `plan:pre`, `execute:wave:post` | gate | first-party |
|
|
| `external-job` | feature | full | `>=1.7.0` | `plan:post`, `execute:wave:post` | contribution | first-party |
|
|
| `gap-analysis` | feature | standard | `>=1.6.0` | `plan:post` | gate | first-party |
|
|
| `graphify` | feature | full | `>=1.6.0` | — | — | first-party |
|
|
| `intel` | feature | full | `>=1.6.0` | `plan:pre` | step | first-party |
|
|
| `mempalace` | feature | full | `>=1.6.0` | `discuss:pre`, `discuss:post`, `plan:pre`, `plan:post`, `execute:wave:post`, `verify:post`, `ship:post` | step, contribution | first-party |
|
|
| `nyquist` | feature | full | `>=1.6.0` | `verify:post` | step | first-party |
|
|
| `pattern-mapper` | feature | full | `>=1.6.0` | `plan:pre` | step | first-party |
|
|
| `profile-pipeline` | feature | full | `>=1.6.0` | — | — | first-party |
|
|
| `research` | feature | standard | `>=1.6.0` | `plan:pre` | step | first-party |
|
|
| `schema-gate` | feature | full | `>=1.6.0` | `plan:pre` | contribution | first-party |
|
|
| `security` | feature | full | `>=1.6.0` | `plan:pre`, `verify:post`, `ship:pre` | step, contribution, gate | first-party |
|
|
| `tdd` | feature | full | `>=1.6.0` | `plan:pre`, `execute:post` | contribution, gate | first-party |
|
|
| `ui` | feature | full | `>=1.6.0` | `plan:pre`, `execute:wave:post`, `verify:post` | step, gate | first-party |
|
|
|
|
### Runtime capabilities (role: runtime) — 18
|
|
|
|
Runtime capabilities adapt GSD to a specific AI runtime or IDE — emitting
|
|
skills, agents, hooks configuration, and surface files for that host. They
|
|
typically register no loop hooks (their primary responsibility is surface
|
|
emission), so their extension-point and hook-kind cells are `—`.
|
|
|
|
| id | role | tier | engines.gsd | extension points | hook kinds | source |
|
|
|---|---|---|---|---|---|---|
|
|
| `antigravity` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
| `augment` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
| `claude` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
| `cline` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
| `codebuddy` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
| `codex` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
| `copilot` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
| `cursor` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
| `hermes` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
| `kilo` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
| `kimi` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
| `opencode` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
| `pi` | runtime | core | `>=1.7.0` | — | — | first-party |
|
|
| `qwen` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
| `trae` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
| `vscode` | runtime | core | `>=1.7.0` | — | — | first-party |
|
|
| `windsurf` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
| `zcode` | runtime | core | `>=1.6.0` | — | — | first-party |
|
|
|
|
---
|
|
|
|
## Third-party capabilities
|
|
|
|
This matrix is the **first-party catalogue**: it is generated from the committed
|
|
registry and therefore lists only the capabilities that ship with GSD Core.
|
|
Installed third-party capabilities are NOT written into this committed file. Once a
|
|
user installs one via `gsd capability install <spec>` it enters the **runtime
|
|
registry overlay** (ADR-1244 D2); the overlay-aware view of what is installed on a
|
|
given machine is `gsd capability list` (see the
|
|
[`gsd capability` command reference](gsd-capability-command.md)), which reports
|
|
first-party and installed third-party capabilities together using the same column
|
|
fields described below, with `source` = `third-party`.
|
|
|
|
### Column values for third-party rows
|
|
|
|
| Column | Value |
|
|
|---|---|
|
|
| **id** | As declared in `capability.json`. Must not use reserved prefixes (`gsd-`, `gsd-core-`, `anthropic-`). |
|
|
| **role** | `feature` or `runtime`, as declared. |
|
|
| **tier** | `core`, `standard`, or `full`, as declared. |
|
|
| **engines.gsd** | Range from `capability.json`; verified at install and at each load. |
|
|
| **extension points** | The loop points the capability registers into, validated against the known 12 identifiers. |
|
|
| **hook kinds** | `step`, `contribution`, and/or `gate` as declared. Disclosed in the consent summary at install. |
|
|
| **source** | `third-party` |
|
|
|
|
### Community registry
|
|
|
|
Whether GSD operates or advertises a central community registry of third-party
|
|
capabilities is **TBD/TBA** (PRD). The matrix mechanic and all manifest fields
|
|
ship regardless of that decision; URL/git/npm/tarball import does not depend on
|
|
a central registry.
|
|
|
|
---
|
|
|
|
## Manifest field reference
|
|
|
|
The fields below are defined in `capability.json` and govern how a capability
|
|
appears in this matrix. For the full schema, see
|
|
[ADR-1244 D1](../adr/1244-capability-ecosystem.md#d1--versioned-capability-manifest)
|
|
and the [capability manifest reference](capability-manifest.md).
|
|
|
|
| Field | Required | Type | Purpose |
|
|
|---|---|---|---|
|
|
| `version` | **Yes** | semver string | Capability version. The registry rejects manifests without it. |
|
|
| `engines.gsd` | Recommended | semver range | Host-version compatibility gate. Enforced at install and load. |
|
|
| `compatVersions` | No | object: cap-version → gsd-range | Graceful-downgrade table for sources that enumerate versions (git tags, registry, npm). |
|
|
| `integrity` | No | `sha512-<base64>` | SHA-512 digest of the fetched bundle. Verified before extraction when present; mismatch aborts. |
|
|
| `provenance` | No | `{ sourceRepo, commit }` | Source provenance; populated in CI for first-party/curated capabilities. |
|
|
|
|
---
|
|
|
|
## Related documents
|
|
|
|
- [ADR-1244 — Capability Ecosystem](../adr/1244-capability-ecosystem.md)
|
|
- [The capability trust model](../explanation/capability-trust-model.md) — why the trust rules are structured as they are
|
|
- [The phase loop](../explanation/the-phase-loop.md) — the 12 loop extension points in context
|
|
- [Capability manifest reference](capability-manifest.md) — the full `capability.json` schema
|
|
- [ADR-857](../adr/857-capability-system.md) — the original capability architecture (D7/D8 extended by ADR-1244)
|