chore(#2371): representative gate-fixture corpus + document-shaped property test (#2380)

* chore(#2371): representative gate-fixture corpus + document-shaped property test

Adds tests/fixtures/representative/ — a permanent corpus of verbatim,
incident-sourced fixtures (never author-invented) from #2286, #2347,
#2365, #2366, each labeled with its expected gate verdict in a
MANIFEST.json and driven through the real CLI gate entrypoint via
tests/representative-corpus.test.cjs.

Adds a document-shaped fast-check property test alongside the existing
writer-seeded bijection test in tests/api-coverage.test.cjs: the existing
generator produces rows and renders them through the writer, so the
document shape is a constant and it cannot fail against a decoy table;
the new one generates the document space instead.

Two gates (#2365, #2347) are still open, so their corpus/property
assertions are marked with node:test's official `todo` option — the test
executes and reports its failure without affecting the process exit code
(https://nodejs.org/api/test.html#test-options). The audit-uat corpus
(#2286, fixed by #2317) is a normal passing assertion, proving the
methodology works end to end and not just cataloguing gaps.

Records the fixture-provenance rule in CONTRIBUTING.md: a gate's fixtures
may not be derived from the gate's own writer, grammar, or docstring
examples; a negative fixture must come from a source that doesn't know
the gate exists.

No production src/*.cts changes — validation only, per #2371's scope.

Closes #2371

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2371): address orthogonal-review findings — dedup generators, fix field naming, wire dead fields

Standards-axis review findings, all fixed:

- Deduplicated the row-shape generators (capabilityGen/rowGen/validRowGen)
  that were copy-pasted between the parse/render bijection test and the
  new document-shaped property test in tests/api-coverage.test.cjs — a
  future edit to one could have silently desynced the two properties.
  Hoisted to a single module-scope declaration both tests reference.

- Renamed decision-coverage-guard/MANIFEST.json's expectedOutcome ->
  expectedReason. It asserted against the gate's `reason` field, but this
  codebase already has a real, different `outcome` field at parser
  altitude (extractDecisions' DecisionOutcome) — naming the manifest
  field after the wrong altitude's term was exactly the ambiguity the
  "Fixture provenance" rule this PR adds exists to eliminate.

- Removed the unused `role` field from three MANIFEST.json files (never
  read by any test) and wired the previously-dead per-fixture
  `expectedMinItems` in audit-uat/MANIFEST.json into a real per-file
  assertion in tests/representative-corpus.test.cjs, using cmdAuditUat's
  `results` array — catches a regression that moves items between the
  two fixture files while preserving the aggregate total, which the
  existing total_items check alone would miss.

No changes to test intent or coverage — same assertions, correctly named
and fully wired.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: remove dead capabilityState/capabilityWriter requires from gsd-tools.cjs

Surfaced by the mandatory pre-PR lint gate (no-unused-vars) while
preparing this PR — unrelated to #2371's own changes, but a defect
found while working is fixed in place rather than deferred.

Leftover from #2368/#2370 (merged just before this branch rebased onto
it): the case 'capability' arm that needed these two requires was
relocated to bin/lib/capability-command-router.cjs, which already
requires both modules directly (lines 24-25) and is their only real
consumer (cmdCapabilityState, resolveCapabilityRuntimeState,
cmdCapabilitySet). The two requires left behind in gsd-tools.cjs had
zero other references in the file and were never re-exported —
confirmed via grep across the file and its module.exports.

Behavior-preserving: Node's require cache means the underlying modules
still load exactly once via capability-command-router.cjs's own
requires; gsd-tools.cjs never used its now-removed local bindings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2371): replace todo-marked assertions with characterization tests

gsd-test's own JSONL result parser (gsd-test-runner's
internal/pipeline/parse.go, verified directly against that repo's
source) has no concept of node:test's `todo` option — it only
recognizes kind:"pass"|"fail" and hard-errors on anything else. A
{ todo: true } test whose body throws is counted as a real failure in
gsd-test's own verdict, exactly as if it weren't marked todo — proven
by an actual gsd-test run against this branch, which reported
outcome:"failed" with all six todo-marked assertions (the property
test plus five representative-corpus fixtures) in the failure list,
each carrying the correct raw node:test `todo` field the tool's parser
simply doesn't read.

Replaces todo with characterization: MANIFEST.json now carries both
the correct target verdict (expected*) and the exact current observed
verdict (currentBuggyOutput, directly verified against live CLI
output for all five fixtures). Tests assert currentBuggyOutput — an
honest, non-vacuous pin of today's known-broken reality that passes
today and will fail loudly the moment the referenced fix changes the
observed output, at which point the assertion should be flipped to
expected* and currentBuggyOutput deleted.

The document-shaped property test switches from throwing fc.assert to
non-throwing fc.check (returns RunDetails per fast-check's own docs)
and asserts report.failed === true directly, for the same reason.

Updates all prose (CONTRIBUTING.md, the fixture READMEs) that
previously claimed todo would be respected — that claim was
factually wrong for this repo's actual tooling and must not ship.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: regenerate golden-install-parity fixtures after rebase onto next

Rebasing onto the current next (which now includes #2381's
todo-severity changes to gsd-core/bin/gsd-tools.cjs) produced real
conflicts in all 18 golden-install-parity fixtures — expected, since
both branches changed the same gsd-tools.cjs hash entry. Resolved by
taking one side to unblock the rebase, then regenerating fresh from
source via npm run gen:golden and verifying the result; every file's
diff is exactly the one hash line for gsd-core/bin/gsd-tools.cjs,
correcting a stale intermediate hash from the arbitrary conflict pick.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Tom Boucher
2026-07-17 15:23:09 -04:00
committed by GitHub
parent 58028eaf56
commit 062f3fda90
39 changed files with 721 additions and 34 deletions

View File

@@ -20,6 +20,25 @@ const fc = require('fast-check');
const MODULE_PATH = path.join(__dirname, '..', 'gsd-core', 'bin', 'lib', 'api-coverage.cjs');
// Shared row-shape generators for the coverage-matrix property tests below
// (the parse/render bijection and the #2371 document-shaped property both
// build matrices from the same canonical row shape — a single declaration
// here means the two properties can't silently desync).
const capabilityGen = fc.stringMatching(/^[a-z][a-z0-9-]{0,14}$/);
const rowGen = fc.record({
capability: capabilityGen,
decision: fc.constantFrom('INTEGRATE', 'OPT-OUT'),
// Reasons are short prose (e.g. "not needed yet"). The matrix is a
// markdown table, so cell text is format-safe: no pipes / newlines.
reason: fc.stringMatching(/^[a-z0-9 ,.\-!?]{0,20}$/),
});
// OPT-OUT rows must carry a non-empty reason for the round-trip to validate.
const validRowGen = rowGen.map((r) =>
r.decision === 'OPT-OUT' && r.reason.trim() === ''
? { ...r, reason: 'because' }
: { ...r, reason: r.reason.trim() }
);
describe('detectApiIntegration — pure detector (#1562)', () => {
let mod;
try {
@@ -346,20 +365,6 @@ describe('coverage matrix — parse/render bijection (fast-check)', () => {
const { renderCoverageMatrix, validateCoverageMatrix } = mod;
test('any valid row set renders and re-validates to the same counts', () => {
const capabilityGen = fc.stringMatching(/^[a-z][a-z0-9-]{0,14}$/);
const rowGen = fc.record({
capability: capabilityGen,
decision: fc.constantFrom('INTEGRATE', 'OPT-OUT'),
// Reasons are short prose (e.g. "not needed yet"). The matrix is a
// markdown table, so cell text is format-safe: no pipes / newlines.
reason: fc.stringMatching(/^[a-z0-9 ,.\-!?]{0,20}$/),
});
// OPT-OUT rows must carry a non-empty reason for the round-trip to validate.
const validRowGen = rowGen.map((r) =>
r.decision === 'OPT-OUT' && r.reason.trim() === ''
? { ...r, reason: 'because' }
: { ...r, reason: r.reason.trim() }
);
const matrixGen = fc.uniqueArray(validRowGen, {
minLength: 1,
maxLength: 8,
@@ -379,6 +384,121 @@ describe('coverage matrix — parse/render bijection (fast-check)', () => {
});
});
// ──────────────────────────────────────────────────────────────────────────────
// Document-shaped property (#2371): the bijection test above generates ROWS and
// renders them through the writer, so the document shape is a constant — it
// cannot generate a second table, a decoy table, or surrounding prose, and so
// cannot fail against #2366's bugs. This property generates the DOCUMENT
// space instead: a canonical matrix interleaved with content a real
// COVERAGE.md may legitimately contain that is NOT the matrix. See
// tests/fixtures/representative/README.md and CONTRIBUTING.md's "Fixture
// provenance" section for the full rationale.
// ──────────────────────────────────────────────────────────────────────────────
describe('coverage matrix — document-shaped fast-check (extract-exactly-canonical, #2371)', () => {
let mod;
try {
mod = require(MODULE_PATH);
} catch (err) {
throw new Error(`Could not require ${MODULE_PATH}. Run "npm run build:lib" first. Underlying: ${err.message}`);
}
const { parseCoverageMatrix, renderCoverageMatrix } = mod;
// Uses the shared capabilityGen/rowGen/validRowGen declared at module scope
// above (same generators the bijection test uses), so the two properties
// exercise the same canonical-row space and can't silently desync.
const canonicalMatrixGen = fc.uniqueArray(validRowGen, {
minLength: 1,
maxLength: 5,
selector: (r) => r.capability.toLowerCase(),
});
// Decoy blocks: content a document may legitimately contain that is NOT the
// canonical matrix. Kept to three explicit, independently-readable shapes
// rather than a generic "random markdown" generator — a combinatorial but
// opaque generator is exactly the kind of cleverness that's unrunnable to
// debug when it fails (Kernighan's Law).
const proseDecoyGen = fc.constantFrom(
'## Notes\n\nSee the ADR for background.',
'This phase also touches the auth helper.',
'## Risks\n\n- Rollout risk is low.',
);
const summaryTableDecoyGen = fc
.record({
label: fc.stringMatching(/^[a-z][a-z0-9 ]{0,10}$/),
integrateCount: fc.nat({ max: 50 }),
optoutCount: fc.nat({ max: 50 }),
})
.map(
({ label, integrateCount, optoutCount }) =>
`## Coverage summary\n\n| tier | INTEGRATE | OPT-OUT |\n|---|---|---|\n` +
`| ${label} | ${integrateCount} | ${optoutCount} |`
);
const secondSectionMatrixGen = fc
.uniqueArray(validRowGen, { minLength: 1, maxLength: 3, selector: (r) => r.capability.toLowerCase() })
.map((rows) => `## Transferred to a later phase\n\n${renderCoverageMatrix(rows)}`);
const decoyGen = fc.oneof(proseDecoyGen, summaryTableDecoyGen, secondSectionMatrixGen);
const documentGen = fc.record({
canonicalRows: canonicalMatrixGen,
decoysBefore: fc.array(decoyGen, { maxLength: 2 }),
decoysAfter: fc.array(decoyGen, { maxLength: 2 }),
});
// #2371's own test-runner (gsd-test / gsd-test-runner v1.6.2) has no concept
// of node:test's `todo` option: its JSONL result parser
// (internal/pipeline/parse.go's parseJSONL, gsd-test-runner repo) only
// recognizes `kind: "pass" | "fail"` and hard-errors on anything else, so a
// `{ todo: true }` test whose body throws is still counted as a failure in
// the tool's own verdict — verified directly against that source, not
// assumed. So this property uses fc's non-throwing `fc.check` (returns
// `RunDetails` instead of throwing — see fast-check's runners docs) and
// asserts on `.failed` directly: today the invariant genuinely does NOT
// hold (that is #2366), so `report.failed === true` is an honest,
// non-vacuous, currently-PASSING characterization of today's known-broken
// reality — not a fake pass. The moment #2366 makes the invariant hold for
// real, `report.failed` becomes `false` and THIS assertion fails loudly,
// forcing whoever's fix landed to notice and flip it. The fix itself stays
// owned by #2366.
test(
'given a document containing exactly one canonical matrix plus arbitrary other content, ' +
'the parser extracts exactly that matrix\'s rows and ignores everything else ' +
'(currently violated — #2366)',
() => {
const report = fc.check(
fc.property(documentGen, ({ canonicalRows, decoysBefore, decoysAfter }) => {
const canonicalBlock = renderCoverageMatrix(canonicalRows);
const doc = [...decoysBefore, canonicalBlock, ...decoysAfter].join('\n\n');
const result = parseCoverageMatrix(doc);
const expectedByCap = new Map(canonicalRows.map((r) => [r.capability.toLowerCase(), r]));
const actualByCap = new Map(result.rows.map((r) => [r.capability.toLowerCase(), r]));
if (actualByCap.size !== expectedByCap.size) return false;
for (const [cap, expected] of expectedByCap) {
const actual = actualByCap.get(cap);
if (!actual || actual.decision !== expected.decision) return false;
}
return result.errors.length === 0;
}),
{ numRuns: 100 }
);
assert.strictEqual(
report.failed,
true,
'This property is expected to be VIOLATED today (#2366 — a decoy summary table or a ' +
'second canonical-schema section corrupts the result or spuriously errors). If this ' +
'assertion fails, the property now HOLDS — #2366 appears fixed; replace this ' +
'characterization with a real fc.assert of the invariant.'
);
}
);
});
// ──────────────────────────────────────────────────────────────────────────────
// CLI entry point (STDIN → exit codes mirror grep, like assumption-delta)
// ──────────────────────────────────────────────────────────────────────────────