* chore(#2800): derive reviewer flag lists and gate reviewer lane docs across locales The reviewer lane roster was hand-enumerated across five documentation surfaces and three workflow files that had drifted apart: --kimi-code was missing from all four translated COMMANDS.md mirrors, --coderabbit from every workflow forwarding list, and --antigravity from FEATURES.md. Adds checkReviewerDocsParity, a second pure gate deliberately separate from checkReviewerLaneParity so a stale doc cannot make the runtime checker look red. Workflows now derive their flag lists from a new review-lane flags query instead of hand-enumerating them, which also retires the unanchored grep that matched --agy inside --antigravity. Documents the previously absent reviewer body and hostBehaviors field in the capability manifest reference. Closes #2800 Closes #2781 Closes #2272 * fix(#2800): key the docs parity table arm on first-cell position Review found the flag arm was file-scoped, so the forwarding row that lists every flag in its third cell satisfied it on its own. Deleting a lane's own reviewer-table row -- the #2781 regression this gate exists to prevent -- therefore passed undetected. Arm 4 keys on the FIRST table cell, which separates a lane row from the forwarding row structurally and in every locale. Regression test included. * fix(#2800): shape-filter the flags subcommand output All three consumers read review-lane flags through an unquoted command substitution so the output word-splits into loop items. Phase 2 admits third-party overlay lanes, so an overlay flag containing whitespace would inject a second loop item and one containing a glob would expand against the cwd. Emit only well-formed flags so neither reaches the shell. * fix(#2800): remove the regex length ceiling and count only prose mentions Review found two real defects in the docs parity gate. The never-throws contract was false: building a RegExp from a declared flag or section title throws SyntaxError past ~100k chars, and Phase 2 admits overlay lanes whose declared strings are untrusted in length. Every one of these matches is literal, so String.includes replaces the regex outright, which also deletes escapeLiteral and the llama.cpp escaping it existed for. Arm 1 was context-blind: a flag mentioned only inside a fenced example or a commented-out row counted as documented. Both are stripped before matching. Also advertises all 13 lane flags in the argument-hint and corrects a stale eleven-lane count in the slug grammar note. * test(#2800): repoint the convergence suite off deleted workflow text The derived flag loop deleted the literal per-flag grep lines four tests matched on. Two of those failed loudly. The behavioral and property tests failed SILENTLY instead: their end marker no longer resolved, so the parse block extracted empty and both passed vacuously, and the property test's gsd_run stub had a no-op default that hid it. All now share one extractor and execute the real deployed block through a gsd_run shim backed by the actual binary. The whitelist assertions become an anti-parity check: re-adding a hand-written flag list must fail. Also repairs two vacuous cases in the docs parity suite. The unreadable-doc test called its own mock rather than the reader, and the integration test bounded nothing, so a doc losing its marker would have been silently skipped and still passed green. * fix(#2800): run the derived flag loop after the launcher preamble The remote matrix caught a real runtime bug, not a test artifact. In autonomous.md and plan-review-convergence.md the launcher preamble that defines gsd_run lives in a separate, LATER bash fence than the derived loop. Each fence is its own shell, so gsd_run was undefined where the loop ran: the command substitution yielded nothing and zero reviewer flags would have been forwarded. Worse than the drift this epic fixes, and silent. The whole CONVERGENCE_ARGS construction moves as one unit, because the --max-cycles append sits between the loop and the preamble and would otherwise have run against an uninitialized variable and then been dropped by the relocated initializer. Also documents all 13 lane flags in help/modes/full.md, which the repo gates bidirectionally against each command's argument-hint. * test(#2800): repoint the two converge suites off deleted flag literals Both asserted workflow.includes('--codex') against the hand-enumerated list the derived loop removed. They now assert the derivation itself, keep --all and --text (convergence controls, still literal), and add an anti-parity guard so re-adding a hardcoded list fails. The lost pass-through proof is replaced with a real one: every flag the tests used to hardcode is asserted present in the actual roster emitted by the binary, which is the property the old assertion was protecting. * test(#2800): acknowledge the workflow byte growth from the derived flag loop * chore(#2800): backfill changeset pr number to 2882 * fix(#2800): strip HTML comments to a fixed point in the parity gate CodeQL js/incomplete-multi-character-sanitization (high) on PR #2882: the single-pass <!--...--> strip can leave a live <!-- behind, so a join-trick construction smuggles a commented-out row past the gate and it counts as documented. Not an injection risk here since nothing is rendered, but it is the exact false pass this helper exists to prevent. Strips to a fixed point, then treats any surviving opener as unterminated so the multi-line branch closes it on a later line. Terminates because every pass strictly shortens the string. * test(#2800): pin the comment-smuggling regression with a real reproducer The obvious fixture for this class does not reproduce it: <!--<!---->--> leaves a dangling --> rather than a live <!--, and is caught either way, so it would have passed with and without the fix. The join-trick construction (<!- + <!--DUMMY--> + -...-->), the <scr<script>ipt> shape, genuinely regresses on the single-pass strip and is what the test now uses. --------- Co-authored-by: Test <test@example.com>
309 lines
14 KiB
JavaScript
309 lines
14 KiB
JavaScript
#!/usr/bin/env node
|
|
'use strict';
|
|
|
|
/**
|
|
* gen-capability-matrix.cjs — ADR-1244 Phase 6 (Decision D9).
|
|
*
|
|
* Generates docs/reference/capability-matrix.md FROM the committed capability
|
|
* registry (gsd-core/bin/lib/capability-registry.cjs), so the matrix can never
|
|
* drift from the actual capability set. Kept honest by a drift guard
|
|
* (tests/capability-matrix-sync.test.cjs runs `--check`).
|
|
*
|
|
* The matrix is RELEASE-STABLE by design: it does NOT embed each capability's
|
|
* exact `version` (which tracks the GSD package version in lockstep and would
|
|
* churn the committed file — and trip the drift guard — on every release). It
|
|
* shows `engines.gsd` (the stable host-compatibility RANGE) instead, and notes
|
|
* the version-lockstep rule in prose. The committed matrix therefore changes
|
|
* only on intentional capability edits (add/remove a capability, change its
|
|
* tier/role/engines/extension-points/hook-kinds) — never on a version bump.
|
|
*
|
|
* Usage:
|
|
* node scripts/gen-capability-matrix.cjs # print to stdout
|
|
* node scripts/gen-capability-matrix.cjs --write # write the committed file
|
|
* node scripts/gen-capability-matrix.cjs --check # exit 1 if the committed file is stale
|
|
*/
|
|
|
|
const fs = require('fs');
|
|
const path = require('path');
|
|
const { ExitError, runMain } = require('./lib/cli-exit.cjs');
|
|
|
|
const ROOT = path.resolve(__dirname, '..');
|
|
const REGISTRY_PATH = path.join(ROOT, 'gsd-core', 'bin', 'lib', 'capability-registry.cjs');
|
|
const MATRIX_PATH = path.join(ROOT, 'docs', 'reference', 'capability-matrix.md');
|
|
|
|
/** Canonical loop extension points, in order (mirrors the phase loop). */
|
|
const LOOP_POINTS = [
|
|
'discuss:pre', 'discuss:post',
|
|
'plan:pre', 'plan:post',
|
|
'execute:pre', 'execute:wave:pre', 'execute:wave:post', 'execute:post',
|
|
'verify:pre', 'verify:post',
|
|
'ship:pre', 'ship:post',
|
|
];
|
|
const POINT_ORDER = new Map(LOOP_POINTS.map((p, i) => [p, i]));
|
|
|
|
/**
|
|
* Build a capId → { points:Set, kinds:Set } map from the registry's byLoopPoint
|
|
* index — the authoritative record of which loop points each capability registers
|
|
* into and with which hook kind (step / contribution / gate).
|
|
*/
|
|
function extensionsByCapability(registry) {
|
|
const out = new Map();
|
|
const byPoint = registry.byLoopPoint || {};
|
|
const KIND = { steps: 'step', contributions: 'contribution', gates: 'gate' };
|
|
for (const point of Object.keys(byPoint)) {
|
|
const reg = byPoint[point] || {};
|
|
for (const arrKey of ['steps', 'contributions', 'gates']) {
|
|
for (const hook of reg[arrKey] || []) {
|
|
const capId = hook && hook.capId;
|
|
if (typeof capId !== 'string') continue;
|
|
let e = out.get(capId);
|
|
if (!e) { e = { points: new Set(), kinds: new Set() }; out.set(capId, e); }
|
|
e.points.add(point);
|
|
e.kinds.add(KIND[arrKey]);
|
|
}
|
|
}
|
|
}
|
|
return out;
|
|
}
|
|
|
|
function fmtPoints(set) {
|
|
if (!set || set.size === 0) return '—';
|
|
for (const p of set) {
|
|
// Surface a typo'd/unknown loop point at generation time rather than silently sorting it last.
|
|
// The registry validates point names at load, so this should never fire — but if it does, the
|
|
// generator (not a confused reader) is where it must be caught.
|
|
if (!POINT_ORDER.has(p)) {
|
|
process.stderr.write(`gen-capability-matrix: WARNING — unknown loop point "${p}" (not one of the ${LOOP_POINTS.length} canonical points)\n`);
|
|
}
|
|
}
|
|
return [...set]
|
|
.sort((a, b) => (POINT_ORDER.has(a) ? POINT_ORDER.get(a) : 99) - (POINT_ORDER.has(b) ? POINT_ORDER.get(b) : 99) || a.localeCompare(b))
|
|
.map((p) => '`' + p + '`')
|
|
.join(', ');
|
|
}
|
|
|
|
function fmtKinds(set) {
|
|
if (!set || set.size === 0) return '—';
|
|
const order = { step: 0, contribution: 1, gate: 2 };
|
|
return [...set].sort((a, b) => (order[a] ?? 9) - (order[b] ?? 9)).join(', ');
|
|
}
|
|
|
|
function fmtEngines(cap) {
|
|
const g = cap && cap.engines && cap.engines.gsd;
|
|
return typeof g === 'string' && g ? '`' + g + '`' : '—';
|
|
}
|
|
|
|
/** Render one capability table (rows sorted by id) for the given role. */
|
|
function renderTable(caps, role, extByCap) {
|
|
const rows = caps
|
|
.filter((c) => c.role === role)
|
|
.sort((a, b) => a.id.localeCompare(b.id))
|
|
.map((c) => {
|
|
const ext = extByCap.get(c.id) || { points: null, kinds: null };
|
|
return `| \`${c.id}\` | ${c.role} | ${c.tier || '—'} | ${fmtEngines(c)} | ${fmtPoints(ext.points)} | ${fmtKinds(ext.kinds)} | first-party |`;
|
|
});
|
|
return [
|
|
'| id | role | tier | engines.gsd | extension points | hook kinds | source |',
|
|
'|---|---|---|---|---|---|---|',
|
|
...rows,
|
|
].join('\n');
|
|
}
|
|
|
|
function buildMatrix(registry) {
|
|
const caps = Object.values(registry.capabilities || {});
|
|
const extByCap = extensionsByCapability(registry);
|
|
const featureTable = renderTable(caps, 'feature', extByCap);
|
|
const runtimeTable = renderTable(caps, 'runtime', extByCap);
|
|
// ADR-2782 D3 added a third role. Rendering only feature+runtime silently
|
|
// DROPPED every role:"reviewer" capability from the catalogue — and because
|
|
// `--check` compares generated output against the committed file, both omitted
|
|
// them identically, so the drift guard reported "up to date" while five shipped
|
|
// capabilities were invisible. A guard blind to a whole role is not guarding.
|
|
const reviewerTable = renderTable(caps, 'reviewer', extByCap);
|
|
const featureCount = caps.filter((c) => c.role === 'feature').length;
|
|
const runtimeCount = caps.filter((c) => c.role === 'runtime').length;
|
|
const reviewerCount = caps.filter((c) => c.role === 'reviewer').length;
|
|
|
|
return `# Capability matrix reference
|
|
|
|
> **Generated file — do not edit by hand.**
|
|
> This matrix is generated from the capability registry by
|
|
> \`scripts/gen-capability-matrix.cjs\` and kept honest by a drift guard
|
|
> (\`tests/capability-matrix-sync.test.cjs\` runs \`--check\`). Any manual edit is
|
|
> overwritten on the next generation run. To change a capability's declared
|
|
> metadata, edit the corresponding \`capabilities/<id>/capability.json\` and run
|
|
> \`node scripts/gen-capability-matrix.cjs --write\`.
|
|
|
|
See also: [ADR-1244](../adr/1244-capability-ecosystem.md) —
|
|
[Capability manifest fields](#manifest-field-reference) —
|
|
[The capability trust model](../explanation/capability-trust-model.md)
|
|
|
|
---
|
|
|
|
## Column definitions
|
|
|
|
| Column | Description |
|
|
|---|---|
|
|
| **id** | Canonical capability identifier; unique across first- and third-party capabilities. Reserved prefixes: \`gsd-\`, \`gsd-core-\`, \`anthropic-\`. |
|
|
| **role** | \`feature\` — extends what the loop does; \`runtime\` — adapts GSD to a specific AI runtime/IDE; \`reviewer\` — declares a cross-AI reviewer lane (ADR-2782). A capability may be both a runtime and a reviewer. |
|
|
| **tier** | \`core\` — always active; \`standard\` — active when the runtime supports it; \`full\` — opt-in or runtime-specific. |
|
|
| **engines.gsd** | Semver RANGE expressing host-version compatibility. A hard gate at install and at load. \`—\` means the capability declares no range. |
|
|
| **extension points** | The loop points this capability registers hooks into (from the registry's \`byLoopPoint\` index). \`—\` means it registers none (typical for runtime capabilities, whose job is surface emission). |
|
|
| **hook kinds** | Which of \`step\`, \`contribution\`, \`gate\` the capability's hooks use. \`—\` means none. |
|
|
| **source** | \`first-party\` — ships with GSD Core; \`third-party\` — installed from an external source via \`gsd capability install\`. |
|
|
|
|
> **On versions.** This matrix intentionally omits a per-capability \`version\`
|
|
> column. First-party capabilities are versioned **in lockstep** with the GSD
|
|
> Core package (their \`capability.json\` \`version\` always equals the GSD release
|
|
> version), so a per-row version would simply repeat the package version and
|
|
> churn the committed file on every release. The stable host-compatibility
|
|
> signal — \`engines.gsd\` — is shown instead. A third-party capability's exact
|
|
> version is recorded in the per-runtime ledger (\`.gsd-capabilities.json\`) at
|
|
> install time.
|
|
|
|
---
|
|
|
|
## Native (first-party) capabilities
|
|
|
|
First-party capabilities are implicitly trusted: they ship as part of the GSD
|
|
Core package and are stamped with the package version at release (per
|
|
ADR-1244 D6). They are not subject to the consent or integrity-pin flow applied
|
|
to third-party capabilities.
|
|
|
|
### Feature capabilities (role: feature) — ${featureCount}
|
|
|
|
Feature capabilities extend what the loop does — contributing research,
|
|
planning, execution, verification, or ship artefacts at the loop extension
|
|
points.
|
|
|
|
${featureTable}
|
|
|
|
### Runtime capabilities (role: runtime) — ${runtimeCount}
|
|
|
|
Runtime capabilities adapt GSD to a specific AI runtime or IDE — emitting
|
|
skills, agents, hooks configuration, and surface files for that host. They
|
|
typically register no loop hooks (their primary responsibility is surface
|
|
emission), so their extension-point and hook-kind cells are \`—\`.
|
|
|
|
${runtimeTable}
|
|
|
|
### Reviewer capabilities (role: reviewer) — ${reviewerCount}
|
|
|
|
Reviewer capabilities declare a cross-AI **reviewer lane** — one external CLI or
|
|
model endpoint \`/gsd:review\` hands a plan to (ADR-2782 D3). They are not install
|
|
targets: they emit no skills, agents, hooks or surface files, so their
|
|
extension-point and hook-kind cells are \`—\`. A host that is *also* a reviewer
|
|
(Claude, Codex, Cursor, OpenCode, Qwen, Antigravity) keeps one manifest and
|
|
appears under **runtime** above, carrying its lane alongside its runtime body;
|
|
only lanes that GSD never installs into appear here.
|
|
|
|
Because a lane receives the plan text, requirements, research findings and
|
|
\`CONTEXT.md\` decisions, it is a disclosed executable surface and is consent-gated
|
|
at install like any other — see
|
|
[the trust model](../explanation/capability-trust-model.md).
|
|
|
|
${reviewerTable}
|
|
|
|
---
|
|
|
|
## Third-party capabilities
|
|
|
|
This matrix is the **first-party catalogue**: it is generated from the committed
|
|
registry and therefore lists only the capabilities that ship with GSD Core.
|
|
Installed third-party capabilities are NOT written into this committed file. Once a
|
|
user installs one via \`gsd capability install <spec>\` it enters the **runtime
|
|
registry overlay** (ADR-1244 D2); the overlay-aware view of what is installed on a
|
|
given machine is \`gsd capability list\` (see the
|
|
[\`gsd capability\` command reference](gsd-capability-command.md)), which reports
|
|
first-party and installed third-party capabilities together using the same column
|
|
fields described below, with \`source\` = \`third-party\`.
|
|
|
|
### Column values for third-party rows
|
|
|
|
| Column | Value |
|
|
|---|---|
|
|
| **id** | As declared in \`capability.json\`. Must not use reserved prefixes (\`gsd-\`, \`gsd-core-\`, \`anthropic-\`). |
|
|
| **role** | \`feature\`, \`runtime\`, or \`reviewer\`, as declared. |
|
|
| **tier** | \`core\`, \`standard\`, or \`full\`, as declared. |
|
|
| **engines.gsd** | Range from \`capability.json\`; verified at install and at each load. |
|
|
| **extension points** | The loop points the capability registers into, validated against the known 12 identifiers. |
|
|
| **hook kinds** | \`step\`, \`contribution\`, and/or \`gate\` as declared. Disclosed in the consent summary at install. |
|
|
| **source** | \`third-party\` |
|
|
|
|
### Community registry
|
|
|
|
Whether GSD operates or advertises a central community registry of third-party
|
|
capabilities is **TBD/TBA** (PRD). The matrix mechanic and all manifest fields
|
|
ship regardless of that decision; URL/git/npm/tarball import does not depend on
|
|
a central registry.
|
|
|
|
---
|
|
|
|
## Manifest field reference
|
|
|
|
The fields below are defined in \`capability.json\` and govern how a capability
|
|
appears in this matrix. For the full schema, see
|
|
[ADR-1244 D1](../adr/1244-capability-ecosystem.md#d1--versioned-capability-manifest)
|
|
and the [capability manifest reference](capability-manifest.md).
|
|
|
|
| Field | Required | Type | Purpose |
|
|
|---|---|---|---|
|
|
| \`version\` | **Yes** | semver string | Capability version. The registry rejects manifests without it. |
|
|
| \`engines.gsd\` | Recommended | semver range | Host-version compatibility gate. Enforced at install and load. |
|
|
| \`compatVersions\` | No | object: cap-version → gsd-range | Graceful-downgrade table for sources that enumerate versions (git tags, registry, npm). |
|
|
| \`integrity\` | No | \`sha512-<base64>\` | SHA-512 digest of the fetched bundle. Verified before extraction when present; mismatch aborts. |
|
|
| \`provenance\` | No | \`{ sourceRepo, commit }\` | Source provenance; populated in CI for first-party/curated capabilities. |
|
|
|
|
---
|
|
|
|
## Related documents
|
|
|
|
- [ADR-1244 — Capability Ecosystem](../adr/1244-capability-ecosystem.md)
|
|
- [The capability trust model](../explanation/capability-trust-model.md) — why the trust rules are structured as they are
|
|
- [The phase loop](../explanation/the-phase-loop.md) — the 12 loop extension points in context
|
|
- [Capability manifest reference](capability-manifest.md) — the full \`capability.json\` schema
|
|
- [ADR-857](../adr/857-capability-system.md) — the original capability architecture (D7/D8 extended by ADR-1244)
|
|
`;
|
|
}
|
|
|
|
function loadRegistry() {
|
|
delete require.cache[require.resolve(REGISTRY_PATH)];
|
|
return require(REGISTRY_PATH);
|
|
}
|
|
|
|
/** Normalize CRLF→LF + ensure a single trailing newline, for cross-platform compare. */
|
|
function normalize(s) {
|
|
return s.replace(/\r\n/g, '\n').replace(/\n+$/, '\n');
|
|
}
|
|
|
|
function main() {
|
|
const flag = process.argv[2];
|
|
const registry = loadRegistry();
|
|
const content = buildMatrix(registry);
|
|
|
|
if (flag === '--check') {
|
|
let committed;
|
|
try {
|
|
committed = fs.readFileSync(MATRIX_PATH, 'utf8');
|
|
} catch {
|
|
throw new ExitError(1, `${path.relative(ROOT, MATRIX_PATH)} is missing. Run:\n node scripts/gen-capability-matrix.cjs --write`);
|
|
}
|
|
if (normalize(committed) !== normalize(content)) {
|
|
throw new ExitError(1, `${path.relative(ROOT, MATRIX_PATH)} is stale. Run:\n node scripts/gen-capability-matrix.cjs --write`);
|
|
}
|
|
console.log(`${path.relative(ROOT, MATRIX_PATH)} is up to date.`);
|
|
return;
|
|
}
|
|
if (flag === '--write') {
|
|
fs.mkdirSync(path.dirname(MATRIX_PATH), { recursive: true });
|
|
fs.writeFileSync(MATRIX_PATH, content, 'utf8');
|
|
console.log(`Wrote ${path.relative(ROOT, MATRIX_PATH)}`);
|
|
return;
|
|
}
|
|
process.stdout.write(content);
|
|
}
|
|
|
|
if (require.main === module) runMain(main);
|
|
|
|
module.exports = { buildMatrix, extensionsByCapability };
|