* chore(#2371): representative gate-fixture corpus + document-shaped property test Adds tests/fixtures/representative/ — a permanent corpus of verbatim, incident-sourced fixtures (never author-invented) from #2286, #2347, #2365, #2366, each labeled with its expected gate verdict in a MANIFEST.json and driven through the real CLI gate entrypoint via tests/representative-corpus.test.cjs. Adds a document-shaped fast-check property test alongside the existing writer-seeded bijection test in tests/api-coverage.test.cjs: the existing generator produces rows and renders them through the writer, so the document shape is a constant and it cannot fail against a decoy table; the new one generates the document space instead. Two gates (#2365, #2347) are still open, so their corpus/property assertions are marked with node:test's official `todo` option — the test executes and reports its failure without affecting the process exit code (https://nodejs.org/api/test.html#test-options). The audit-uat corpus (#2286, fixed by #2317) is a normal passing assertion, proving the methodology works end to end and not just cataloguing gaps. Records the fixture-provenance rule in CONTRIBUTING.md: a gate's fixtures may not be derived from the gate's own writer, grammar, or docstring examples; a negative fixture must come from a source that doesn't know the gate exists. No production src/*.cts changes — validation only, per #2371's scope. Closes #2371 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2371): address orthogonal-review findings — dedup generators, fix field naming, wire dead fields Standards-axis review findings, all fixed: - Deduplicated the row-shape generators (capabilityGen/rowGen/validRowGen) that were copy-pasted between the parse/render bijection test and the new document-shaped property test in tests/api-coverage.test.cjs — a future edit to one could have silently desynced the two properties. Hoisted to a single module-scope declaration both tests reference. - Renamed decision-coverage-guard/MANIFEST.json's expectedOutcome -> expectedReason. It asserted against the gate's `reason` field, but this codebase already has a real, different `outcome` field at parser altitude (extractDecisions' DecisionOutcome) — naming the manifest field after the wrong altitude's term was exactly the ambiguity the "Fixture provenance" rule this PR adds exists to eliminate. - Removed the unused `role` field from three MANIFEST.json files (never read by any test) and wired the previously-dead per-fixture `expectedMinItems` in audit-uat/MANIFEST.json into a real per-file assertion in tests/representative-corpus.test.cjs, using cmdAuditUat's `results` array — catches a regression that moves items between the two fixture files while preserving the aggregate total, which the existing total_items check alone would miss. No changes to test intent or coverage — same assertions, correctly named and fully wired. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix: remove dead capabilityState/capabilityWriter requires from gsd-tools.cjs Surfaced by the mandatory pre-PR lint gate (no-unused-vars) while preparing this PR — unrelated to #2371's own changes, but a defect found while working is fixed in place rather than deferred. Leftover from #2368/#2370 (merged just before this branch rebased onto it): the case 'capability' arm that needed these two requires was relocated to bin/lib/capability-command-router.cjs, which already requires both modules directly (lines 24-25) and is their only real consumer (cmdCapabilityState, resolveCapabilityRuntimeState, cmdCapabilitySet). The two requires left behind in gsd-tools.cjs had zero other references in the file and were never re-exported — confirmed via grep across the file and its module.exports. Behavior-preserving: Node's require cache means the underlying modules still load exactly once via capability-command-router.cjs's own requires; gsd-tools.cjs never used its now-removed local bindings. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2371): replace todo-marked assertions with characterization tests gsd-test's own JSONL result parser (gsd-test-runner's internal/pipeline/parse.go, verified directly against that repo's source) has no concept of node:test's `todo` option — it only recognizes kind:"pass"|"fail" and hard-errors on anything else. A { todo: true } test whose body throws is counted as a real failure in gsd-test's own verdict, exactly as if it weren't marked todo — proven by an actual gsd-test run against this branch, which reported outcome:"failed" with all six todo-marked assertions (the property test plus five representative-corpus fixtures) in the failure list, each carrying the correct raw node:test `todo` field the tool's parser simply doesn't read. Replaces todo with characterization: MANIFEST.json now carries both the correct target verdict (expected*) and the exact current observed verdict (currentBuggyOutput, directly verified against live CLI output for all five fixtures). Tests assert currentBuggyOutput — an honest, non-vacuous pin of today's known-broken reality that passes today and will fail loudly the moment the referenced fix changes the observed output, at which point the assertion should be flipped to expected* and currentBuggyOutput deleted. The document-shaped property test switches from throwing fc.assert to non-throwing fc.check (returns RunDetails per fast-check's own docs) and asserts report.failed === true directly, for the same reason. Updates all prose (CONTRIBUTING.md, the fixture READMEs) that previously claimed todo would be respected — that claim was factually wrong for this repo's actual tooling and must not ship. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore: regenerate golden-install-parity fixtures after rebase onto next Rebasing onto the current next (which now includes #2381's todo-severity changes to gsd-core/bin/gsd-tools.cjs) produced real conflicts in all 18 golden-install-parity fixtures — expected, since both branches changed the same gsd-tools.cjs hash entry. Resolved by taking one side to unblock the rebase, then regenerating fresh from source via npm run gen:golden and verifying the result; every file's diff is exactly the one hash line for gsd-core/bin/gsd-tools.cjs, correcting a stale intermediate hash from the arbitrary conflict pick. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
83 lines
4.3 KiB
Markdown
83 lines
4.3 KiB
Markdown
# Representative Gate Fixtures (#2371)
|
|
|
|
Every fixture in this tree is **verbatim** (or, where noted, a minimal
|
|
faithful subset) of an artifact that a real user actually produced and
|
|
reported against a real GSD gate. None of it was written by a gate's own
|
|
author to exercise that gate.
|
|
|
|
## Why this directory exists, and why `tests/fixtures/adversarial/` isn't enough
|
|
|
|
`tests/fixtures/adversarial/` covers hostile input — unicode, CRLF, nested
|
|
fences, heredoc breakout. Nobody attacked the gates these fixtures target.
|
|
A developer wrote an ordinary, well-formed artifact — a plan, a coverage
|
|
matrix, a CONTEXT.md — that a gate misjudged anyway. That input is neither
|
|
synthetic-happy nor adversarial; it's simply *real*, and until #2371 no gate
|
|
had coverage for it.
|
|
|
|
The pattern this corpus exists to break: a gate's test fixtures were
|
|
authored by the same person (or model) who wrote the gate, from the same
|
|
mental model, so they can only confirm what the author already believed —
|
|
never surface what the author didn't anticipate. See #2371 for the full
|
|
diagnosis (four incidents across three gates in ten days, including a
|
|
fast-check property test whose generator was seeded from the parser's own
|
|
writer function and therefore could not fail).
|
|
|
|
## Rule
|
|
|
|
**A gate's fixtures may not be derived from the gate's own writer, grammar,
|
|
or docstring examples. A negative fixture must come from a source that
|
|
does not know the gate exists.** Recorded in `CONTRIBUTING.md` under
|
|
"Fixture provenance."
|
|
|
|
## Layout
|
|
|
|
Each subdirectory is one gate:
|
|
|
|
- `api-coverage-detector/` — `detectApiIntegration` (#2365)
|
|
- `api-coverage-matrix/` — `parseCoverageMatrix` (#2366)
|
|
- `audit-uat/` — `parseUatItems` / `parseVerificationItems` (#2286, fixed by #2317)
|
|
- `decision-coverage-guard/` — `extractDecisions`'s could-not-parse guard (#1365 gap, #2347)
|
|
|
|
Each carries its own `README.md` (what each fixture is and where it came
|
|
from) and `MANIFEST.json` (machine-readable: file → source issue → gate →
|
|
expected verdict). `tests/representative-corpus.test.cjs` loads every
|
|
manifest and drives each fixture through the gate's real CLI entrypoint —
|
|
never the parser function directly — so the assertion is at gate-verdict
|
|
altitude (the boolean/JSON a user actually sees), not parse-tree altitude.
|
|
|
|
## Why the still-broken fixtures assert `currentBuggyOutput`, not a red `todo`
|
|
|
|
Three of the four gates here are still open bugs (#2365, #2366, #2347) at
|
|
the time this corpus was added. Their `MANIFEST.json` entries carry BOTH
|
|
the correct target verdict (`expected*` — what the eventual fix must
|
|
produce) and the exact CURRENT observed verdict (`currentBuggyOutput` —
|
|
what today's code actually returns). The test asserts against
|
|
`currentBuggyOutput`: an honest, non-vacuous characterization of today's
|
|
known-broken reality, not a fake pass.
|
|
|
|
This is deliberately NOT node:test's `todo` option. `todo` looked like the
|
|
right tool — a todo test executes and reports its failure without
|
|
affecting Node's own process exit code
|
|
(https://nodejs.org/api/test.html#test-options) — but this repo's actual
|
|
test-runner (`gsd-test` / `gsd-test-runner` v1.6.2) has no concept of it:
|
|
its JSONL result parser (`internal/pipeline/parse.go`'s `parseJSONL`, in
|
|
the separate `gsd-test-runner` repo) only recognizes `kind: "pass" | "fail"`
|
|
and hard-errors on anything else — verified directly against that source,
|
|
not assumed. A `{ todo: true }` test whose body throws is still counted as
|
|
a real failure in `gsd-test`'s own verdict, which would block the push
|
|
gate exactly as if it weren't marked todo at all.
|
|
|
|
Asserting `currentBuggyOutput` sidesteps this because the test genuinely
|
|
passes today — no runner-level "expected failure" feature required. The
|
|
fixes belong to #2365 / #2366 / #2347, not to this corpus. When one of
|
|
those lands, the corresponding assertion will fail (the gate now returns
|
|
something other than the pinned buggy value) — at that point, flip the
|
|
test to assert `expected*` instead and delete the stale
|
|
`currentBuggyOutput`.
|
|
|
|
The `audit-uat/` corpus has no `todo`: #2286 was fixed by #2317 before this
|
|
corpus was written, so its assertions are ordinary, currently-passing
|
|
tests — the proof that a representative fixture, driven through the real
|
|
gate, is not automatically doomed to fail. It demonstrates the methodology
|
|
working, not just the gaps it finds.
|