Files
msd-core/tests/fixtures/representative
Tom Boucher 06eba5fdb0 fix(#4130): parse phase-prefixed decision IDs (D4-01) (#4357)
* test(#4130): failing-first regression for phase-prefixed decision IDs

Add the #4130 matrix: D4-01/D12-01 across all three bullet forms, tags,
discretion, wrapped lead-ins, gate-level plan/verify end-to-end rows, and
parity properties (well-formed digit-prefixed ids parse to their exact id;
a non-digit injected into the prefix fails loud). Update the #2347
non-D-prefix fixture from D5-NN (now a legal grammar) to DEC-NN, and
graduate the representative d5-prefix corpus fixture from could-not-parse
to parsed-but-uncovered.

All new rows are RED against origin/next; they go green with the parser
fix in the next commit.

* fix(#4130): parse phase-prefixed decision IDs (D4-01)

The three declaration grammars, the parse-miss guard, the #3939 join
regexes, and the token evidence all anchored on the literal 'D-' (or
'**D-'), so an ID carrying a digit-run phase prefix between the leading
letter and the hyphen matched nothing — while the #2347 shape detector
correctly called those bullets decision-shaped, collapsing the whole
CONTEXT.md to could-not-parse with 0 extracted instead of a coverage
verdict.

Derive the extractor ID grammar from one shared DECISION_ID_SOURCE
('D[0-9]*-' + the existing alnum tail, full id captured), widen the
guard/join anchors to ID_ATTEMPT_SOURCE (bare 'D-' or a digit-initial
prefix run, so a typo'd 'D4x-01' fails loud while letter-initial prose
like 'Deferred-until' stays none-present), and align the bare-token
evidence. Both gates and the gap-checker share the parser, so all three
surfaces read phase-prefixed decisions now; the gate messages name the
accepted forms including the phase-prefixed one.

* docs(#4130): document the phase-prefixed decision identifier form

The canonical CONTEXT.md reference said decisions carry 'a sequential
D-NN identifier' with no mention of the optional phase-number prefix the
parser now accepts (D4-01) or the alphanumeric tail it always accepted
(D-INFRA-01). Name both in the Decision identifier format section, EN
and ja-JP.

* chore(#4130): changeset

* chore(#4130): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-05 21:05:19 -04:00
..

Representative Gate Fixtures (#2371)

Every fixture in this tree is verbatim (or, where noted, a minimal faithful subset) of an artifact that a real user actually produced and reported against a real GSD gate. None of it was written by a gate's own author to exercise that gate.

Why this directory exists, and why tests/fixtures/adversarial/ isn't enough

tests/fixtures/adversarial/ covers hostile input — unicode, CRLF, nested fences, heredoc breakout. Nobody attacked the gates these fixtures target. A developer wrote an ordinary, well-formed artifact — a plan, a coverage matrix, a CONTEXT.md — that a gate misjudged anyway. That input is neither synthetic-happy nor adversarial; it's simply real, and until #2371 no gate had coverage for it.

The pattern this corpus exists to break: a gate's test fixtures were authored by the same person (or model) who wrote the gate, from the same mental model, so they can only confirm what the author already believed — never surface what the author didn't anticipate. See #2371 for the full diagnosis (four incidents across three gates in ten days, including a fast-check property test whose generator was seeded from the parser's own writer function and therefore could not fail).

Rule

A gate's fixtures may not be derived from the gate's own writer, grammar, or docstring examples. A negative fixture must come from a source that does not know the gate exists. Recorded in CONTRIBUTING.md under "Fixture provenance."

Layout

Each subdirectory is one gate:

  • api-coverage-detector/ — detectApiIntegration (#2365)
  • api-coverage-matrix/ — parseCoverageMatrix (#2366)
  • audit-uat/ — parseUatItems / parseVerificationItems (#2286, fixed by #2317)
  • decision-coverage-guard/ — extractDecisions's could-not-parse guard (#1365 gap, #2347)

Each carries its own README.md (what each fixture is and where it came from) and MANIFEST.json (machine-readable: file → source issue → gate → expected verdict). tests/representative-corpus.test.cjs loads every manifest and drives each fixture through the gate's real CLI entrypoint — never the parser function directly — so the assertion is at gate-verdict altitude (the boolean/JSON a user actually sees), not parse-tree altitude.

Why the still-broken fixtures assert currentBuggyOutput, not a red todo

Three of the four gates here are still open bugs (#2365, #2366, #2347) at the time this corpus was added. Their MANIFEST.json entries carry BOTH the correct target verdict (expected* — what the eventual fix must produce) and the exact CURRENT observed verdict (currentBuggyOutput — what today's code actually returns). The test asserts against currentBuggyOutput: an honest, non-vacuous characterization of today's known-broken reality, not a fake pass.

This is deliberately NOT node:test's todo option. todo looked like the right tool — a todo test executes and reports its failure without affecting Node's own process exit code (https://nodejs.org/api/test.html#test-options) — but this repo's actual test-runner (gsd-test / gsd-test-runner v1.6.2) has no concept of it: its JSONL result parser (internal/pipeline/parse.go's parseJSONL, in the separate gsd-test-runner repo) only recognizes kind: "pass" | "fail" and hard-errors on anything else — verified directly against that source, not assumed. A { todo: true } test whose body throws is still counted as a real failure in gsd-test's own verdict, which would block the push gate exactly as if it weren't marked todo at all.

Asserting currentBuggyOutput sidesteps this because the test genuinely passes today — no runner-level "expected failure" feature required. The fixes belong to #2365 / #2366 / #2347, not to this corpus. When one of those lands, the corresponding assertion will fail (the gate now returns something other than the pinned buggy value) — at that point, flip the test to assert expected* instead and delete the stale currentBuggyOutput.

The audit-uat/ corpus has no todo: #2286 was fixed by #2317 before this corpus was written, so its assertions are ordinary, currently-passing tests — the proof that a representative fixture, driven through the real gate, is not automatically doomed to fail. It demonstrates the methodology working, not just the gaps it finds.