enhance(#2721): regenerating merge driver, regen:derived, and a name for the emitted-artifact family (#2730)

* test(#2721): failing-first suite for the gsd-regen driver and CONTEXT.md parity

Tests precede the implementation per the TDD gate. The driver module does not
exist yet, so tests/git-merge-regen-driver.test.cjs fails at require time; the
contributor-standards parity assertions fail against next as it stands today,
where the standards doc names two CONTEXT.md headings that have never existed.

Refs #2721

* feat(#2721): add the gsd-regen merge driver and regen:derived

The golden parity manifests and the two size baselines are pure functions of
the source tree, so their only correct merge is "recompute" -- something git's
ours/theirs interface cannot express. 140 of 143 conflicted-file instances
across the open PR queue are these files.

The driver deliberately does NOT regenerate. Four probes established that at
merge-driver time neither the working tree nor the index reflects the merge:
both hold the ours side, a file added by theirs does not exist yet, and
MERGE_HEAD is unwritten. Git also invokes the driver once per conflicted path
(20 here). A regenerating driver would therefore read the ours-side tree and
emit a plausible-but-wrong hash manifest -- worse than a conflict, because a
conflict is visible. So it accepts %A, runs zero subprocesses, records the
resolved paths, and prints one notice pointing at npm run regen:derived.
Staleness stays caught where it already was, by golden-install-parity in CI.

Every failure path degrades toward today's behaviour (a normal conflict).
install-tree is deliberately excluded per ADR-2719 section 7.

Also folded in, per the no-defer rule: workflow-size.cjs claimed .md files have
no eol=lf in .gitattributes; git check-attr shows eol: lf, set by .gitattributes
line 2 since #1088.

Refs #2721

* docs(#2721): document regen:derived and the gsd-regen merge driver

Adds the how-to a contributor actually reaches for when the generated parity
manifests or size baselines conflict, in both places they would look: the
merge-conflict path in CONTRIBUTING.md and the full guide in TESTING-SUITES.md,
including what the driver deliberately does not do (it does not clear GitHub's
CONFLICTING label, and it does not regenerate mid-merge).

Also scopes the new contributor-standards parity assertion to the doc's own
CONTEXT.md section. Its first run flagged `## Decision`, `## Consequences` and
`## Standards followed`, which the doc attributes to an ADR body and a PR body
rather than to CONTEXT.md -- a doc-wide extractor would have demanded CONTEXT.md
grow headings that do not belong to it.

Refs #2721

* fix(#2721): stop passing %P to the merge driver — shell injection

The isolated adversarial review found, and I independently reproduced, local
arbitrary command execution.

Git does not invoke a merge driver with an argv array. It substitutes %O %A %B
%L %P textually into the configured string and runs the whole thing through a
shell, and $(...) executes inside POSIX double quotes -- so quoting the
placeholder does not neutralise it. %O/%A/%B are git-generated temp names and
%L is an integer, but %P is the file's own path, chosen freely by any
contributor. A branch renaming a covered fixture to
evil$(touch PWNED_SENTINEL).json executed that command on the machine of every
maintainer who merged it, and the merge still reported success.

Fix removes the input rather than filtering it: %P is no longer registered, so
the driver receives no attacker-controlled argument at all. The marker records
a count instead of path names. A metacharacter filter would have been a guess
about shell grammar; passing nothing is a property. Re-ran the identical
exploit against the fixed command: nothing executed, conflict still resolved.

Two regressions guard it -- a platform-independent assertion that the
registered command carries no %P, and a real merge driven by the actual
planInstall output with a $(...) filename.

Also from review: CLI dispatch had no coverage at all (CONTRIBUTING's
"CLI and command routing" matrix), which is why runInstall/runStatus now take
{repoRoot} -- hardcoding REPO_ROOT was what made them untestable. Renamed
planResolution to resolveAndRecord since the plan* prefix promised purity it
did not have. Reconciled the eleven-vs-twelve generator count across
CONTEXT.md, CONTRIBUTING.md and the changeset.

Refs #2721

* test(#2721): scope safe.directory for the check-attr helper

The 66f4d85a run failed 11 assertions, all in the .gitattributes scoping block,
with "fatal: detected dubious ownership in repository at '/work'". The test
container checks the repo out at a path its user does not own, so git refuses
check-attr outright. Everything else passed (27,185).

`check-attr` is a pure read of .gitattributes -- no hooks, no filters -- so the
exemption is scoped to that one invocation. It is deliberately NOT applied to
the driver's own production `git config` calls, which run in the user's own
clone and should keep the protection.

Refs #2721

* test(#2721): delete the stale assertion that the driver command carries %P

The plex2 run on bdfd0856 left exactly two failures, both this test: it still
asserted the pre-fix command string, i.e. the vulnerable behaviour. Deleted
rather than relaxed, per RULESET.TESTS.delete-bad-tests -- its useful half is
already covered, in both directions, by
registeredDriverCommandNeverPassesThePlaceholderForTheFilePath.

Refs #2721

* test(#2721): drive the end-to-end merges from the real planInstall output

The e2e helper hand-rolled its own driver registration, and still carried %P.
That meant the five real-git tests were not exercising the production command
string at all -- planInstall could drift and they would keep passing. They now
register exactly what a contributor gets from npm run setup:merge-driver.

Refs #2721

* chore(#2721): backfill changeset pr number to 2730
This commit is contained in:
Tom Boucher
2026-07-27 19:55:37 -04:00
committed by GitHub
parent e48eb44003
commit a613caaeef
11 changed files with 1491 additions and 8 deletions

View File

@@ -117,12 +117,56 @@ the workflow and agent guards). To resolve:
If a hard cap (not the baseline) is what failed, regeneration will **not** help —
that is the signal to extract, per step 3.
### How-to: the baselines or golden fixtures conflict on merge
`tests/workflow-size-baseline.json`, `tests/agent-size-baseline.json` and
`tests/fixtures/golden-install-parity/*.json` are **Emitted Artifact Provenance**
files (`CONTEXT.md` → `RULESET.EMITTED_ATTRIBUTION`): pure functions of the source
tree. Their correct merge is always *recompute*, which git's ours/theirs interface
cannot express — so a conflict here is never something to hand-resolve.
Register the merge driver once per clone:
```bash
npm run setup:merge-driver
```
Afterwards a conflicting merge, rebase or cherry-pick keeps your branch's copy and
prints a one-line notice. Recompute the artifacts before committing:
```bash
npm run regen:derived
```
That one command runs every generator in dependency order (`gen:golden` last,
because it hashes installed output). On an unmodified tree it produces no diff.
Two things it deliberately does **not** do:
- **It does not clear GitHub's `CONFLICTING` label.** Merge drivers live in
`.git/config`, so forks do not have one and github.com's own merge never runs a
custom driver. The driver removes the labour, not the label.
- **It does not regenerate during the merge.** At the moment git invokes a merge
driver, neither the working tree nor the index reflects the merge yet — so
regenerating there would compute the artifact from the *pre-merge* tree and write
a confidently wrong answer. Running `regen:derived` afterwards is what makes it
correct.
This driver is a bridge introduced by [#2721](https://github.com/open-gsd/gsd-core/issues/2721)
and retired by [#2724](https://github.com/open-gsd/gsd-core/issues/2724), which
replaces these committed artifacts with a computed attribution check (ADR-2719).
`tests/fixtures/install-tree/*.json` is deliberately excluded and keeps normal merge
semantics — its diffs are readable and it must stay an absolute "the installer
stopped shipping X" failure.
### Reference
| Artifact | Role |
|---|---|
| `scripts/workflow-size.cjs` | Single source of truth — LF-normalized byte counter (`lfByteCount`) + generic `measureMdFiles(dir, predicate)` (backs both workflows and agents) + workflow enumeration (`listWorkflowStems`, `measureWorkflows`). Imported by **both** the guards and the generator so they can never measure differently. |
| `scripts/update-size-baseline.cjs` (`npm run size:baseline`) | Regenerates **both** `tests/workflow-size-baseline.json` and `tests/agent-size-baseline.json` — sorted keys, trailing newline, idempotent. |
| `npm run regen:derived` | Runs every generator in dependency order (build → registry → ADR index → capability matrix → inventory manifest → manifest versions → size baselines → golden fixtures). Use it instead of remembering which generator owns which artifact. |
| `scripts/git-merge-regen-driver.cjs` (`npm run setup:merge-driver`) | Registers the `gsd-regen` merge driver in this clone. Keeps your branch's copy of a conflicting generated artifact and points you at `regen:derived`. Bridge for #2721; retired by #2724. |
| `tests/workflow-size-baseline.json` | The committed per-workflow snapshot (one entry per workflow). |
| `tests/agent-size-baseline.json` | The committed per-agent snapshot (one entry per `gsd-*` agent). |
| `tests/workflow-size-budget.test.cjs` | The three workflow guards above, plus the `discuss-phase` progressive-disclosure checks. |

View File

@@ -18,23 +18,23 @@ These apply to every PR — fix, enhancement, or feature. They are part of the m
`CONTEXT.md` is the single source of truth for domain vocabulary. It defines:
- **Domain terms** — canonical Module names, seam vocabulary, and Interface names (e.g. Dispatch Policy Module, Command Contract Validation Module, Planning Workspace Module)
- **Domain modules and seams** — canonical Module names, seam vocabulary, and Interface names (e.g. Dispatch Policy Module, Command Contract Validation Module, Planning Workspace Module)
- **Recurring PR mistakes** — CodeRabbit findings that recur; covers tests, shell guards, changesets, docs
- **Workflow learnings** — patterns distilled from triage + PR cycles
### Format
`CONTEXT.md` is written as flat named sections under `## Domain terms` (for Modules/seams) and `##` sections for recurring rules. Machine-oriented predicates use `KEY.SUBKEY=value` flat format in code blocks under `## AI Ops Memory`.
`CONTEXT.md` is written as flat named sections under `## Glossary — Domain modules and seams` (for Modules/seams) and `##` sections for recurring rules. Machine-oriented predicates use `KEY.SUBKEY=value` flat format, grouped into the `##` section that owns the topic — `## Test rules and lint`, `## CodeRabbit + repo-process guards (machine-oriented predicates)`, or `## Workspace seams (machine-oriented predicates)`.
Adding a new Module or seam:
- Add a `### <Module Name>` entry under `## Domain terms`.
- Add a `### <Module Name>` entry under `## Glossary — Domain modules and seams`.
- Write one paragraph. State what the Module owns. Be concrete — list the Interface names and policy boundaries it covers.
- Do not add synonyms; pick one name and use it everywhere.
Extending an existing predicate:
- Add a `KEY.SUBKEY=value` line inside the relevant `## AI Ops Memory` block.
- Add a `KEY.SUBKEY=value` line inside the relevant predicate section — for a test or lint rule that is `## Test rules and lint`.
- Do not create a new top-level section for a variation on an existing concept.
When to add a new predicate vs extend an existing one: