enh(#3565): sentinel/contract registry + check:contract-drift lint (#3571)

* enh(#3565): sentinel/contract registry + check:contract-drift lint

* fix(#3565): report artifact-row markers once and dedupe per marker

* docs(#3565): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
This commit is contained in:
Tom Boucher
2026-08-16 12:59:08 -04:00
committed by GitHub
parent a4a02a7a01
commit 3c61b4a838
19 changed files with 2587 additions and 40 deletions

View File

@@ -0,0 +1,5 @@
---
type: Added
pr: 3571
---
**`check:contract-drift` — a machine-enforced agent-contract registry** — sentinel markers, read-tag gates, and deleted-file test references can no longer drift silently: the Agent Registry table in `gsd-core/references/agent-contracts.md` is now linted against what agents emit and what workflows consume, and `lint-removed-but-needed` catches tests that pin files your PR deleted. (#3565)

View File

@@ -933,4 +933,4 @@ Full detail in `~/.claude/skills/gsd-pr-fix-discipline/SKILL.md`. AI agents MUST
- **Fix:** Merge manually by hand in dependency order once CI greens; `gh pr merge <N> --squash --repo open-gsd/gsd-core`
### Defect enforcement (ADR-2143 follow-on)
Prose defect entries retired in favour of gates; the gate IS the record. `DEFECT.UNBOUNDED-SUBPROCESS` → `eslint-rules/require-subprocess-timeout.cjs` (error, `src/**`; options literal must carry `timeout` — git 5-30s, npm 60s; the rule only requires the call be bounded, not what the caller does after — 7 of the 8 sites this surfaced degrade to an empty/false/null result on failure, and the 1 that guards a destructive real-run migration (`roadmap-upgrade.cts`'s pre-mutation clean-tree check) correctly still throws rather than proceed against an unverified working tree). `DEFECT.CANARY-VERSION-LEAK` → `scripts/lint-canary-version-leak.cjs` + the `canary-version-leak` job in `.github/workflows/version-gate.yml` (PRs whose base is `main`). `DEFECT.CHANGESET-PR-FIELD-DRIFT` → `findPrFieldDrift` in `scripts/changeset/lint.cjs`. Entries whose condition no automated check can evaluate were deleted rather than kept as unenforceable prose. `DEFECT.FRONTMATTER-SCALAR-BROAD-GREP` → `scripts/lint-frontmatter-scalar-broad-grep.cjs` (in `lint:ci`; flags an unscoped `grep "^key:"` over a whole planning doc with no frontmatter slice and no `-m1`/`head -1` guard). `DEFECT.REMOVED-BUT-NEEDED` → `scripts/lint-removed-but-needed.cjs` (in `lint:ci`; a deleted file whose basename still appears in `.github/workflows/`, `gsd-core/`, `docs/` or `package.json`). `DEFECT.DEFAULT-FLIP-DOCUMENTATION` → `scripts/lint-default-flip-documentation.cjs` + `.github/workflows/default-flip-documentation.yml`, covering `gsd-core/bin/shared/config-defaults.manifest.json` only: a changed value for an EXISTING key requires a `## Breaking Changes` PR section. The six Windows-portability entries (`WINDOWS-TEST-PORTABILITY`, `WINDOWS-POSIX-MODE-BIT-ASSERT`, `WINDOWS-PATH-LEAK-IN-MARKDOWN-CONTENT`, `WINDOWS-PATH-LITERAL-IN-ASSERT`, `WINDOWS-FS-OPS`, `TEST-SHELL-PIPELINE-NONPORTABLE`) were already enforced by ADR-1703's own `local/*` AST ESLint rules (`no-posix-mode-bit-assert`, `normalize-path-in-content`, `no-path-literal-in-assert`, `require-fs-op-fallback`, `no-crlf-fragile-split`, `no-unguarded-nonportable-exec` — see `eslint.config.mjs`) before this PR; their prose duplicated the rules' own doc comments, so it was deleted rather than kept as a second copy. **Residual, deliberately unenforced:** `buildNewProjectConfig`'s hardcoded literal in `src/config.cts` has env-derived branches and `CONFIG_DEFAULTS` spreads, so no reliable resolved-value diff exists without executing the compiled module at both refs — a line-diff there false-fires on any refactor that merely moves the object, so it was left unchecked rather than shipped noisy.
Prose defect entries retired in favour of gates; the gate IS the record. `DEFECT.UNBOUNDED-SUBPROCESS` → `eslint-rules/require-subprocess-timeout.cjs` (error, `src/**`; options literal must carry `timeout` — git 5-30s, npm 60s; the rule only requires the call be bounded, not what the caller does after — 7 of the 8 sites this surfaced degrade to an empty/false/null result on failure, and the 1 that guards a destructive real-run migration (`roadmap-upgrade.cts`'s pre-mutation clean-tree check) correctly still throws rather than proceed against an unverified working tree). `DEFECT.CANARY-VERSION-LEAK` → `scripts/lint-canary-version-leak.cjs` + the `canary-version-leak` job in `.github/workflows/version-gate.yml` (PRs whose base is `main`). `DEFECT.CHANGESET-PR-FIELD-DRIFT` → `findPrFieldDrift` in `scripts/changeset/lint.cjs`. Entries whose condition no automated check can evaluate were deleted rather than kept as unenforceable prose. `DEFECT.FRONTMATTER-SCALAR-BROAD-GREP` → `scripts/lint-frontmatter-scalar-broad-grep.cjs` (in `lint:ci`; flags an unscoped `grep "^key:"` over a whole planning doc with no frontmatter slice and no `-m1`/`head -1` guard). `DEFECT.REMOVED-BUT-NEEDED` → `scripts/lint-removed-but-needed.cjs` (in `lint:ci`; a deleted file whose basename still appears in `.github/workflows/`, `gsd-core/`, `docs/` or `package.json`; and since #3565, in `tests/` behind a `pins-existence` vs `asserts-absence` discriminator — `fs.existsSync`/`readFileSync`/`require`/quoted-object-key on the deleted basename fails, a negated `!…includes`/`!fs.existsSync` absence assertion is the correct post-deletion state and passes; a bare mention that is neither is not flagged, because the undiscriminated widening was tried in #3560 and reverted for firing on the very tests that prove a deletion worked). `DEFECT.DEFAULT-FLIP-DOCUMENTATION` → `scripts/lint-default-flip-documentation.cjs` + `.github/workflows/default-flip-documentation.yml`, covering `gsd-core/bin/shared/config-defaults.manifest.json` only: a changed value for an EXISTING key requires a `## Breaking Changes` PR section. The six Windows-portability entries (`WINDOWS-TEST-PORTABILITY`, `WINDOWS-POSIX-MODE-BIT-ASSERT`, `WINDOWS-PATH-LEAK-IN-MARKDOWN-CONTENT`, `WINDOWS-PATH-LITERAL-IN-ASSERT`, `WINDOWS-FS-OPS`, `TEST-SHELL-PIPELINE-NONPORTABLE`) were already enforced by ADR-1703's own `local/*` AST ESLint rules (`no-posix-mode-bit-assert`, `normalize-path-in-content`, `no-path-literal-in-assert`, `require-fs-op-fallback`, `no-crlf-fragile-split`, `no-unguarded-nonportable-exec` — see `eslint.config.mjs`) before this PR; their prose duplicated the rules' own doc comments, so it was deleted rather than kept as a second copy. **Residual, deliberately unenforced:** `buildNewProjectConfig`'s hardcoded literal in `src/config.cts` has env-derived branches and `CONFIG_DEFAULTS` spreads, so no reliable resolved-value diff exists without executing the compiled module at both refs — a line-diff there false-fires on any refactor that merely moves the object, so it was left unchecked rather than shipped noisy.

View File

@@ -224,8 +224,6 @@ This is the single entry point `gsd-roadmapper` reads.
Return ≤ 10 lines to the orchestrator:
```
## Synthesis Complete
Docs synthesized: {N} ({breakdown})
Decisions locked: {N}
Requirements: {N}

View File

@@ -8,6 +8,9 @@ color: cyan
<role>
You are the MemPalace curator. You run once per phase at `ship:post`, after verification has passed, to consolidate the phase's memory into the palace. Everything you do is best-effort and wing-scoped: a MemPalace failure must never fail the ship step (`onError: skip`), and you must never touch drawers outside this project's wing.
**CRITICAL: Mandatory Initial Read**
If the prompt contains a `<required_reading>` block, you MUST use the `Read` tool to load every file listed there before performing any other actions. This is your primary context.
</role>
<inputs>

View File

@@ -915,6 +915,40 @@ See @~/.claude/gsd-core/references/planner-chunked.md for `## OUTLINE COMPLETE`
<success_criteria>
## Return Markers
Your orchestrator dispatches on exact marker strings in your final output. Emit exactly one of:
```markdown
## PLANNING COMPLETE
```
(final plans committed, ready for verification)
```markdown
## OUTLINE COMPLETE
```
(outline produced, awaiting confirmation — chunked planning mode)
```markdown
## PHASE SPLIT RECOMMENDED
```
(phase too large to plan as one unit, include the proposed split)
```markdown
## ⚠ Source Audit
```
(unplanned items found in the requirements, include the options)
```markdown
## CHECKPOINT REACHED
```
(paused at a user checkpoint, include resume instructions)
```markdown
## PLANNING INCONCLUSIVE
```
(cannot produce a plan, include exactly what is missing)
## Standard Mode
Phase planning complete when:

View File

@@ -13,6 +13,9 @@ You are spawned by the profile orchestration workflow (Phase 3) or by write-prof
Your job: Apply the heuristics defined in the user-profiling reference document to score each dimension with evidence and confidence. Return structured JSON analysis.
CRITICAL: You must apply the rubric defined in the reference document. Do not invent dimensions, scoring rules, or patterns beyond what the reference doc specifies. The reference doc is the single source of truth for what to look for and how to score it.
**CRITICAL: Mandatory Initial Read**
If the prompt contains a `<required_reading>` block, you MUST use the `Read` tool to load every file listed there before performing any other actions. This is your primary context.
</role>
<input>

View File

@@ -819,3 +819,21 @@ Twelve additional agents ship under `agents/gsd-*.md` and are used by specialty
- Researchers have web access — they need current ecosystem information
- Executors have Edit — they modify code but not web access
- Mappers have Write — they write analysis documents but not Edit (no code changes)
## Completion Contracts (machine-enforced)
Every agent's return contract is declared in [`gsd-core/references/agent-contracts.md`](../gsd-core/references/agent-contracts.md)'s **Agent Registry** table — `(Agent, Completion Markers, Consumed by, Kind)` — and enforced by `npm run check:contract-drift` (part of `lint:ci`).
The `Kind` column records how a caller actually detects the agent's completion:
| Kind | Detection mechanism |
|---|---|
| `sentinel-match` | Exact-case string match against a declared marker (by a workflow, command, or another agent) |
| `artifact+query` | The agent writes a file; the caller reads or queries that artifact |
| `structured-return` | The agent returns parseable sections/JSON inline; the caller reads the return text |
When you add an agent or change what it returns, update its registry row in the same change — a stale row is a build failure, not a documentation cleanup for later. Markers are extracted **fence-aware** (a heading inside a fenced block is the emitted template; the same words outside a fence are prose documentation), producer scope includes `@`-included `gsd-core/references/**` files, and consumers are matched **exact-case** (a case-insensitive hit is reported as a collision, never accepted). A marker that is deliberately emitted but matched by nothing carries an `(unconsumed: <reason>)` annotation — an auditable exemption that waives only the consumer requirement.
The same check also enforces the read-tag pairing: whenever a declared consumer emits `<required_reading>`, the producing agent's instructions must reference the gate (directly or via an `@`-included reference) — and the retired `<files_to_read>` vocabulary may not reappear under `workflows/`, `commands/`, or `agents/`.
For acting on a specific finding, see [How to resolve a contract-drift finding](how-to/resolve-contract-drift-findings.md).

View File

@@ -24,6 +24,7 @@ Language versions: [English](README.md) · [Português (pt-BR)](pt-BR/README.md)
- [Resolve edge-coverage findings](how-to/resolve-edge-coverage-findings.md) — turn the spec phase's surfaced domain-boundary edges into covered, dismissed, or backstopped spec decisions
- [Resolve prohibition findings](how-to/resolve-prohibition-findings.md) — turn the spec phase's surfaced must-NOT constraints into resolved, dismissed, or deferred spec decisions
- [Resolve an unreachable-workflow finding](how-to/resolve-unreachable-workflow-findings.md) — wire or fully sweep a shipped workflow that no command, agent, or skill references
- [Resolve a contract-drift finding](how-to/resolve-contract-drift-findings.md) — bring an agent's completion contract, read-tag gate, or deleted-file test reference back into agreement with the registry
- [Resolve unreachable-guard findings](how-to/resolve-unreachable-guard-findings.md) — fix shell guards whose fallback arm cannot run, and tell "nothing to report" apart from "could not look"
- [Resolve an ESLint glob-coverage finding](how-to/resolve-eslint-coverage-findings.md) — bring a source file that matches no lint rule under coverage, or record a reasoned exemption
- [Plan a phase](how-to/plan-a-phase.md) — run research, decompose work, and verify plan quality

View File

@@ -0,0 +1,77 @@
# How to resolve a contract-drift finding
**Goal:** Bring `gsd-core/references/agent-contracts.md`'s Agent Registry back into agreement with what agents actually emit and what workflows, commands, and agents actually consume — so a completion marker, a read-tag gate, or a deleted-file reference can never silently drift apart again.
**Prerequisites:** A failing `npm run lint:ci` (or a direct `npm run check:contract-drift`) reporting one or more violations. The check runs automatically as part of `lint:ci` — you do not invoke it separately.
For why the registry exists (five shipped defects, all one root cause: two surfaces sharing a contract with nothing enforcing agreement), see [issue #3565](https://github.com/open-gsd/gsd-core/issues/3565) and epic [#1891](https://github.com/open-gsd/gsd-core/issues/1891). This guide covers only how to *act* on a finding.
---
## Read a finding
```
ERROR check-contract-drift: 2 violation(s) across 2 kind(s)
no_consumer (1):
- gsd-example marker "EXAMPLE COMPLETE": no consumer text contains "EXAMPLE COMPLETE" (exact case)
remedy: add an exact-case consumer for this marker, or reclassify the row Kind to artifact+query/structured-return
```
Every finding names its **kind**, the agent (or registry), the marker where one is involved, and a **remedy** line. The kinds and what they mean:
| Kind | Meaning | Correct resolution |
|---|---|---|
| `no_consumer` | A declared `sentinel-match` marker no workflow/command/agent matches exact-case | Wire the consumer, fix the casing on either side, or reclassify the row's `Kind` |
| `declared_consumer_no_match` | The marker is consumed somewhere, but by none of the files the row's `Consumed by` cell names | Fix the cell to name a real consumer — the cell is what the read-tag arm and humans navigate by |
| `case_only_match` | A marker matches a consumer only case-insensitively | Fix the casing to match exactly — never relax the matcher; lowercase prose is a coincidence, not a contract |
| `case_collision` | Two declared markers differ only by case | Rename one deliberately; a case-insensitive runtime match cannot tell them apart |
| `declared_marker_not_emitted` | The registry declares a marker the agent never emits in-fence | Emit it as a fenced example in the agent file (a `@`-included `gsd-core/references/**` file counts), or remove it from the row |
| `emitted_marker_not_declared` | The agent emits an in-fence marker-shaped heading absent from its row | Add it to the row — after verifying a real consumer exists (it may be your next `no_consumer`) |
| `vestigial_marker` | An `artifact+query`/`structured-return` row's agent still emits a marker | Delete the marker from the agent — unless it carries an `(unconsumed: …)` annotation (see below) |
| `agent_without_contract` / `duplicate_registry_row` | An agent file with no row, or two rows | Add exactly one row; the registry is the declaration of record for every agent |
| `unknown_producer` / `unknown_consumer` | A row names an agent or a `Consumed by` file that does not exist | Fix or drop the name — a lint over fictional paths guards nothing |
| `read_tag_gate_missing` | A declared consumer emits `<required_reading>` but the agent never references the gate | Add the MUST-Read gate clause to the agent (an `@`-included reference carrying it counts) — this is the F8 defect one layer up |
| `legacy_read_tag` | A `<files_to_read>` survived under workflows/commands/agents | Rename to `<required_reading>` ([#3423](https://github.com/open-gsd/gsd-core/issues/3423) standardized on the gate tag) |
| `unmatched_consumer_token` | A workflow/command matches a quoted `## TOKEN` no agent declares or emits | Declare the producer, or delete the dispatch — this is F9's shape from the consumer side |
| `parse_error` / `unclosed_fence` | The registry table or an agent file is malformed | Fix the row / close the fence; extraction over an unclosed fence is unreliable |
| `pins-existence` (from `lint-removed-but-needed`) | A test depends on a file this PR deleted | Update the test to assert absence instead — `assert.ok(!content.includes('x.md'))` is the correct post-deletion state |
---
## Pick the right `Kind`
Every row declares exactly one of three kinds, and the kind is a **measured fact about the consumer**, not a style preference:
- **`sentinel-match`** — a workflow, command, or another agent detects completion by an exact-case string match. Only kind whose markers need a consumer.
- **`artifact+query`** — the agent writes a file and the caller reads or queries it (`*-VERIFICATION.md` + `gsd_run query verification.status`). No marker required.
- **`structured-return`** — the agent cannot write (no `Write` tool) or returns parseable sections/JSON inline, and the caller reads the return text.
To decide, open the consumer and find the line that waits for the agent. If it greps a marker string → `sentinel-match`. If it reads a file or queries a CLI → `artifact+query`. If it parses the return message → `structured-return`.
## The `(unconsumed: …)` annotation — an audit trail, not a mute button
Some markers are emitted deliberately while nothing matches them: Marker Rule 2's title-case markers are a recorded decision, and a draft-presentation heading may serve a human, not a spawner. Declare those as:
```markdown
`## ROADMAP DRAFT` (unconsumed: draft-presentation format the shipped execution flow never invokes — Step 8 returns `## ROADMAP CREATED`)
```
The annotation waives **only** the consumer requirement. The marker must still be declared *and* emitted, and it still counts for case-collision detection. A non-empty reason is required — the reason is what makes the exemption auditable rather than a silent pass. Never use it to silence a real orphan; if no reason survives scrutiny, the marker is dead and should be deleted.
## After editing the registry
Nothing regenerates from `agent-contracts.md` — it is the source of truth, read directly by the check. Re-run:
```bash
npm run check:contract-drift
```
If the failing finding came from the `tests/` arm of `lint-removed-but-needed`, re-run `npm run lint:removed-but-needed` with `GSD_REMOVED_BUT_NEEDED_BASE` pointing at your base branch (CI sets this automatically).
## Silent cases are clean by design
- A marker heading **outside** a code fence is prose documentation, not a contract — the check ignores it (fence-awareness is what keeps 7 legitimate prose headings from being reported as case-drift).
- An `(unconsumed: …)` marker with no consumer is not an orphan — that is the annotation working.
- A test that **asserts absence** of a deleted file is not a surviving reference — that is the discriminator working.
- Nothing to report and nothing to fix look identical in the output: `ok check:contract-drift: N agents, M markers, 0 violations`. The agent/marker counts moving between runs is your signal the registry actually changed.

View File

@@ -8,38 +8,55 @@ This doc describes what IS, not what should be. Casing inconsistencies are docum
## Agent Registry
| Agent | Role | Completion Markers |
|-------|------|--------------------|
| gsd-planner | Plan creation | `## PLANNING COMPLETE` |
| gsd-executor | Plan execution | `## PLAN COMPLETE`, `## CHECKPOINT REACHED` |
| gsd-phase-researcher | Phase-scoped research | `## RESEARCH COMPLETE`, `## RESEARCH BLOCKED` |
| gsd-project-researcher | Project-wide research | `## RESEARCH COMPLETE`, `## RESEARCH BLOCKED` |
| gsd-plan-checker | Plan validation | `## VERIFICATION PASSED`, `## ISSUES FOUND` |
| gsd-research-synthesizer | Multi-research synthesis | `## SYNTHESIS COMPLETE`, `## SYNTHESIS BLOCKED` |
| gsd-debugger | Debug investigation | `## DEBUG COMPLETE`, `## ROOT CAUSE FOUND`, `## CHECKPOINT REACHED` |
| gsd-roadmapper | Roadmap creation/revision | `## ROADMAP CREATED`, `## ROADMAP REVISED`, `## ROADMAP BLOCKED` |
| gsd-ui-auditor | UI review | `## UI REVIEW COMPLETE` |
| gsd-ui-checker | UI validation | `## ISSUES FOUND` |
| gsd-ui-researcher | UI spec creation | `## UI-SPEC COMPLETE`, `## UI-SPEC BLOCKED` |
| gsd-verifier | Post-execution verification | `## Verification Complete` (title case) |
| gsd-integration-checker | Cross-phase integration check | `## Integration Check Complete` (title case) |
| gsd-nyquist-auditor | Sampling audit | `## PARTIAL`, `## ESCALATE` (non-standard) |
| gsd-security-auditor | Security audit | `## OPEN_THREATS`, `## ESCALATE` (non-standard) |
| gsd-codebase-mapper | Codebase analysis | No marker (writes docs directly) |
| gsd-assumptions-analyzer | Assumption extraction | No marker (returns `## Assumptions` sections) |
| gsd-doc-verifier | Doc validation | No marker (writes JSON to `.planning/tmp/`) |
| gsd-doc-writer | Doc generation | No marker (writes docs directly) |
| gsd-advisor-researcher | Advisory research | No marker (utility agent) |
| gsd-user-profiler | User profiling | No marker (returns JSON in analysis tags) |
| gsd-intel-updater | Codebase intelligence analysis | `## INTEL UPDATE COMPLETE`, `## INTEL UPDATE FAILED` |
| Agent | Role | Completion Markers | Consumed by | Kind |
|-------|------|--------------------|--------------|------|
| gsd-ai-researcher | AI framework research | No marker (writes the AI-SPEC.md framework section via Edit) | `gsd-core/workflows/ai-integration-phase.md` reads the AI-SPEC.md section after the agent returns | artifact+query |
| gsd-planner | Plan creation | `## PLANNING COMPLETE`, `## OUTLINE COMPLETE`, `## PHASE SPLIT RECOMMENDED`, `## ⚠ Source Audit`, `## CHECKPOINT REACHED`, `## PLANNING INCONCLUSIVE` | `gsd-core/workflows/plan-phase.md`, `gsd-core/workflows/plan-phase/steps/chunked-planning-mode.md`, `gsd-core/workflows/plan-review-convergence.md`, `gsd-core/workflows/quick.md` | sentinel-match |
| gsd-executor | Plan execution | `## PLAN COMPLETE`, `## CHECKPOINT REACHED` | `gsd-core/workflows/plan-phase.md`, `gsd-core/workflows/plan-phase/steps/chunked-planning-mode.md`, `agents/gsd-debug-session-manager.md`, `agents/gsd-debugger.md` | sentinel-match |
| gsd-phase-researcher | Phase-scoped research | `## RESEARCH COMPLETE`, `## RESEARCH BLOCKED` | `gsd-core/workflows/plan-phase.md`, `gsd-core/workflows/quick/steps/research-phase.md`, `agents/gsd-project-researcher.md` | sentinel-match |
| gsd-project-researcher | Project-wide research | `## RESEARCH COMPLETE`, `## RESEARCH BLOCKED` | `gsd-core/workflows/plan-phase.md`, `gsd-core/workflows/quick/steps/research-phase.md`, `agents/gsd-phase-researcher.md` | sentinel-match |
| gsd-plan-checker | Plan validation | `## VERIFICATION PASSED`, `## ISSUES FOUND` | `gsd-core/workflows/plan-phase.md`, `gsd-core/workflows/plan-phase/steps/stall-detection-helpers.md`, `gsd-core/workflows/quick/steps/plan-checker-loop.md`, `gsd-core/workflows/ui-phase.md`, `gsd-core/workflows/verify-work.md`, `agents/gsd-ui-checker.md` | sentinel-match |
| gsd-research-synthesizer | Multi-research synthesis | `## SYNTHESIS COMPLETE`, `## SYNTHESIS BLOCKED` (unconsumed: blocked-research return — spawners detect failure via the #222 SUMMARY.md-on-disk check, no dispatch branch keys on the marker) | `gsd-core/workflows/new-milestone.md`, `gsd-core/workflows/new-project.md` | sentinel-match |
| gsd-debugger | Debug investigation | `## DEBUG COMPLETE`, `## ROOT CAUSE FOUND`, `## CHECKPOINT REACHED`, `## INVESTIGATION INCONCLUSIVE`, `## TDD CHECKPOINT`, `## FIX REJECTED BY GUARDRAIL` | `agents/gsd-debug-session-manager.md`, `gsd-core/workflows/diagnose-issues.md`, `gsd-core/workflows/plan-phase.md`, `agents/gsd-executor.md` | sentinel-match |
| gsd-debug-session-manager | Debug checkpoint loop | `## DEBUG SESSION COMPLETE`, `## CONTINUE_REQUIRED` | `gsd-core/workflows/debug.md` | sentinel-match |
| gsd-roadmapper | Roadmap creation/revision | `## ROADMAP CREATED`, `## ROADMAP REVISED`, `## ROADMAP BLOCKED`, `## ROADMAP DRAFT` (unconsumed: draft-presentation format the shipped execution flow never invokes — Step 8 returns `## ROADMAP CREATED`; retained for interactive draft review) | `gsd-core/workflows/new-milestone.md`, `gsd-core/workflows/new-project.md` | sentinel-match |
| gsd-ui-auditor | UI review | `## UI REVIEW COMPLETE` | `gsd-core/workflows/ui-review.md` | sentinel-match |
| gsd-ui-checker | UI validation | `## ISSUES FOUND`, `## UI-SPEC VERIFIED` | `gsd-core/workflows/plan-phase.md`, `gsd-core/workflows/quick/steps/plan-checker-loop.md`, `gsd-core/workflows/ui-phase.md`, `gsd-core/workflows/verify-work.md`, `agents/gsd-plan-checker.md` | sentinel-match |
| gsd-ui-researcher | UI spec creation | `## UI-SPEC COMPLETE`, `## UI-SPEC BLOCKED` | `gsd-core/workflows/ui-phase.md` | sentinel-match |
| gsd-verifier | Post-execution verification | `## Verification Complete` (unconsumed: Marker Rule 2 recorded decision — intentional title-case marker; completion is detected via the artifact route, nothing matches the marker) | `*-VERIFICATION.md` artifact + `gsd_run query verification.status` in `gsd-core/workflows/verify-work.md` | artifact+query |
| gsd-integration-checker | Cross-phase integration check | `## Integration Check Complete` (unconsumed: Marker Rule 2 recorded decision — intentional title-case marker; the auditor reads the inline report, nothing matches the marker) | `gsd-core/workflows/audit-milestone.md` reads the agent's inline return text directly (agent has no Write tool -- it cannot write an artifact) | structured-return |
| gsd-nyquist-auditor | Sampling audit | `## PARTIAL`, `## ESCALATE`, `## GAPS FILLED` (non-standard) | `gsd-core/workflows/validate-phase.md`, `gsd-core/workflows/secure-phase.md`, `agents/gsd-security-auditor.md` | sentinel-match |
| gsd-security-auditor | Security audit | `## OPEN_THREATS`, `## ESCALATE`, `## SECURED` (non-standard) | `gsd-core/workflows/secure-phase.md`, `gsd-core/workflows/validate-phase.md`, `agents/gsd-nyquist-auditor.md` | sentinel-match |
| gsd-codebase-mapper | Codebase analysis | No marker (writes docs directly) | `.planning/codebase/*.md` artifacts, checked via `ls`/`wc -l` in `gsd-core/workflows/map-codebase.md` | artifact+query |
| gsd-code-fixer | Applies code-review fixes | No marker (fix commits + REVIEW.md updates) | `gsd-core/workflows/code-review-fix.md` reads REVIEW.md resolution state + git log | artifact+query |
| gsd-code-reviewer | Source-code review | No marker (writes REVIEW.md) | `gsd-core/workflows/code-review.md` reads REVIEW.md | artifact+query |
| gsd-assumptions-analyzer | Assumption extraction | No marker (returns `## Assumptions` sections) | `gsd-core/workflows/discuss-phase-assumptions.md` reads the inline `## Assumptions` sections from the agent's return | structured-return |
| gsd-doc-classifier | Planning-doc classification | No marker (writes `.planning/intel/classifications/*.json`) | `gsd-core/workflows/ingest-docs.md` reads the classification JSON | artifact+query |
| gsd-doc-verifier | Doc validation | No marker (writes JSON to `.planning/tmp/`) | `.planning/tmp/verify-{doc_filename}.json` artifact, read by `gsd-core/workflows/docs-update.md` | artifact+query |
| gsd-doc-writer | Doc generation | No marker (writes docs directly) | generated doc files, consumed by `gsd-core/workflows/docs-update.md` and `gsd-core/workflows/docs-update/steps/dispatch-monorepo-packages.md` | artifact+query |
| gsd-domain-researcher | Domain research | No marker (writes the AI-SPEC.md domain section via Edit) | `gsd-core/workflows/ai-integration-phase.md` reads the AI-SPEC.md section after the agent returns | artifact+query |
| gsd-eval-auditor | Evaluation coverage audit | No marker (writes the REVIEW.md audit section) | `gsd-core/workflows/eval-review.md` reads REVIEW.md | artifact+query |
| gsd-eval-planner | Evaluation strategy design | No marker (writes the AI-SPEC.md evaluation section via Edit) | `gsd-core/workflows/ai-integration-phase.md` reads the AI-SPEC.md section after the agent returns | artifact+query |
| gsd-framework-selector | Framework decision matrix | No marker (returns the interactive decision matrix inline) | `gsd-core/workflows/ai-integration-phase.md` reads the returned matrix | structured-return |
| gsd-advisor-researcher | Advisory research | No marker (utility agent) | `gsd-core/workflows/discuss-phase/modes/advisor.md` reads the inline comparison table from the agent's return | structured-return |
| gsd-user-profiler | User profiling | No marker (returns JSON in analysis tags) | `gsd-core/workflows/profile-user.md` extracts the inline `<analysis>` JSON block from the agent's return | structured-return |
| gsd-intel-updater | Codebase intelligence analysis | No marker (`.planning/intel/*.json` artifacts) | `.planning/intel/*.json` artifacts, read via `gsd_run intel query` / `intel validate` (no `*.md` workflow currently spawns this agent -- see `docs/adr/22-plan-drift-guard.md`, "never auto-spawned") | artifact+query |
| gsd-mempalace-curator | Ship-time MemPalace curation | No marker (writes the session diary + cross-links) | `gsd-core/workflows/ship.md` reads the diary artifacts | artifact+query |
| gsd-pattern-mapper | Codebase pattern mapping | `## PATTERN MAPPING COMPLETE` | `gsd-core/workflows/plan-phase.md` (also spawned by `gsd-core/workflows/settings.md`) | sentinel-match |
| gsd-doc-synthesizer | Doc synthesis for `/gsd:ingest-docs` | No marker (SYNTHESIS.md and INGEST-CONFLICTS.md artifacts) | `.planning/intel/SYNTHESIS.md` and `.planning/INGEST-CONFLICTS.md` artifacts, read by `gsd-core/workflows/ingest-docs.md` | artifact+query |
## Marker Rules
1. **ALL-CAPS markers** (e.g., `## PLANNING COMPLETE`) are the standard convention
2. **Title-case markers** (e.g., `## Verification Complete`) exist in gsd-verifier and gsd-integration-checker -- these are intentional as-is, not bugs
2. **Title-case markers in gsd-verifier and gsd-integration-checker are intentional as-is, not bugs — a recorded decision.** Their rows are `artifact+query`/`structured-return` (completion is detected through the row's `Kind` route), and the markers are carried as `(unconsumed: Marker Rule 2 recorded decision …)` annotations: an auditable exemption, never deleted and never silently passed. `## Synthesis Complete` in gsd-doc-synthesizer was NOT covered by this rule; #3565 deleted it deliberately because it case-collides with gsd-research-synthesizer's `## SYNTHESIS COMPLETE` and nothing matched it
3. **Non-standard markers** (e.g., `## PARTIAL`, `## ESCALATE`) in audit agents indicate partial results requiring orchestrator judgment
4. **Agents without markers** either write artifacts directly to disk or return structured data (JSON/sections) that the caller parses
4. **`Kind` describes how a caller actually detects an agent's completion, and is exactly one of:**
- `sentinel-match` -- a workflow, command, or another agent detects completion by an exact-case string match against a declared marker
- `artifact+query` -- the agent writes a file (report, JSON, generated doc) and the caller reads or queries that artifact instead of matching any marker text
- `structured-return` -- the agent has no way to write files (no `Write` tool) or simply doesn't; it returns parseable sections, a table, or JSON inline, and the caller reads that return text directly
5. Markers must appear as H2 headings (`## `) at the start of a line in the agent's final output
6. The `Consumed by` / `Kind` columns are machine-enforced by `check:contract-drift` (`scripts/check-contract-drift.cjs`), which cross-checks this table against what each `agents/*.md` file actually emits in-fence and what every `gsd-core/workflows/**`, `commands/**`, and `agents/**` file actually consumes. Update this table whenever an agent's return contract changes -- a stale row is a violation the check will report, not something to leave for later.
7. A marker entry annotated `(unconsumed: <reason>)` is emitted deliberately but matched by no workflow, command, or agent — e.g. `## ROADMAP DRAFT`, a presentation format a human approves interactively. The check still verifies the marker is declared **and** emitted, and still counts it for case-collision purposes; only the consumer requirement is waived. Use it for display formats, never to silence a real orphan.
## Key Handoff Contracts

View File

@@ -89,6 +89,7 @@
},
"scripts": {
"sync:launcher": "node scripts/sync-runtime-launcher.cjs",
"check:contract-drift": "node scripts/check-contract-drift.cjs",
"check:env": "node scripts/check-env.cjs",
"check:alias-drift": "node scripts/check-alias-drift.cjs",
"check:identity-drift": "node scripts/lint-package-identity-drift.cjs",
@@ -118,7 +119,7 @@
"lint:table-schema-drift": "node scripts/lint-table-schema-drift.cjs",
"lint:frontmatter-scalar-broad-grep": "node scripts/lint-frontmatter-scalar-broad-grep.cjs",
"lint:removed-but-needed": "node scripts/lint-removed-but-needed.cjs",
"lint:ci": "npm run lint && npm run lint:skill-deps && npm run lint:generated-sync && node scripts/lint-test-file-count.cjs && node scripts/lint-command-contract.cjs && node scripts/lint-pr-check-project-dir.cjs && npm run lint:legacy-name && node scripts/lint-regression-test-names.cjs && node scripts/lint-allow-test-rule-refs.cjs && node scripts/lint-resolution-provenance.cjs && node scripts/lint-emitted-drift-ack.cjs && node scripts/lint-portable-timeout.cjs && node scripts/validate-registry.cjs && node scripts/lint-table-schema-drift.cjs && node scripts/lint-fix-has-regression-test.cjs && node scripts/lint-example-parser-parity.cjs && node scripts/lint-docs-command-form.cjs && node scripts/lint-plan-count-drift.cjs && node scripts/lint-milestone-window-drift.cjs && node scripts/lint-phase-enumeration-drift.cjs && node scripts/lint-planning-prompt-drift.cjs && node scripts/lint-unreachable-guard-drift.cjs && node scripts/lint-completion-ratio-drift.cjs && node scripts/lint-state-field-drift.cjs && node scripts/lint-state-write-path-drift.cjs && node scripts/lint-completion-predicate-drift.cjs && node scripts/lint-planning-snapshot-bypass-drift.cjs && node scripts/lint-health-diagnostic-rule-table.cjs && node scripts/lint-planning-artifact-writer-drift.cjs && node scripts/lint-frontmatter-scalar-broad-grep.cjs && node scripts/lint-removed-but-needed.cjs && node scripts/lint-no-adhoc-regex-escape.cjs && node scripts/lint-vendored-deps.cjs",
"lint:ci": "npm run lint && npm run lint:skill-deps && npm run lint:generated-sync && node scripts/lint-test-file-count.cjs && node scripts/lint-command-contract.cjs && node scripts/lint-pr-check-project-dir.cjs && npm run lint:legacy-name && node scripts/lint-regression-test-names.cjs && node scripts/lint-allow-test-rule-refs.cjs && node scripts/lint-resolution-provenance.cjs && node scripts/lint-emitted-drift-ack.cjs && node scripts/lint-portable-timeout.cjs && node scripts/validate-registry.cjs && node scripts/lint-table-schema-drift.cjs && node scripts/lint-fix-has-regression-test.cjs && node scripts/lint-example-parser-parity.cjs && node scripts/lint-docs-command-form.cjs && node scripts/lint-plan-count-drift.cjs && node scripts/lint-milestone-window-drift.cjs && node scripts/lint-phase-enumeration-drift.cjs && node scripts/lint-planning-prompt-drift.cjs && node scripts/lint-unreachable-guard-drift.cjs && node scripts/lint-completion-ratio-drift.cjs && node scripts/lint-state-field-drift.cjs && node scripts/lint-state-write-path-drift.cjs && node scripts/lint-completion-predicate-drift.cjs && node scripts/lint-planning-snapshot-bypass-drift.cjs && node scripts/lint-health-diagnostic-rule-table.cjs && node scripts/lint-planning-artifact-writer-drift.cjs && node scripts/lint-frontmatter-scalar-broad-grep.cjs && node scripts/lint-removed-but-needed.cjs && node scripts/lint-no-adhoc-regex-escape.cjs && node scripts/lint-vendored-deps.cjs && node scripts/check-contract-drift.cjs",
"lint:allow-test-rule-refs": "node scripts/lint-allow-test-rule-refs.cjs",
"lint:regression-names": "node scripts/lint-regression-test-names.cjs",
"lint:descriptions": "node scripts/lint-descriptions.cjs",

View File

@@ -0,0 +1,297 @@
#!/usr/bin/env node
/**
* check-contract-drift.cjs
*
* Enforces that gsd-core/references/agent-contracts.md's `## Agent Registry`
* table stays in sync with reality:
*
* 1. The table itself must parse cleanly (no malformed rows).
* 2. Every agents/*.md file must have its fenced code blocks properly
* closed (an unclosed fence makes in-fence marker detection unreliable).
* 3. Every marker an agent actually emits in-fence, and every marker the
* registry declares for it, must agree (contractViolations' declared/
* emitted checks) -- unless the row opts out via `kind:
* artifact+query`/`structured-return`, in which case any emitted
* marker is itself a violation (vestigial_marker).
* 4. Every `sentinel-match` row's declared markers must have at least one
* exact-case consumer somewhere under gsd-core/workflows/, commands/,
* or agents/ (excluding the producing agent's own file).
* 5. Registry roster coverage: every agents/*.md file has exactly one
* row, every row names an agent file that exists
* (agent_without_contract / duplicate_registry_row / unknown_producer),
* and every file-shaped `Consumed by` entry resolves (unknown_consumer).
* 6. Read-tag arm: no `<files_to_read>` survives anywhere in the consumer
* corpus (legacy_read_tag), and whenever a declared consumer emits
* `<required_reading>` the producing agent's file must reference the
* gate (read_tag_gate_missing).
* 7. Reverse direction: a workflow/command matching a quoted `## TOKEN`
* no agent declares or emits is dispatch-on-phantom
* (unmatched_consumer_token) — F9's shape from the consumer side.
*
* Exit 0 = clean. Exit 1 = violations (with diagnostics on stderr).
*/
'use strict';
const fs = require('fs');
const path = require('path');
function resolveRoot(argv) {
const idx = argv.indexOf('--root');
if (idx === -1) return path.join(__dirname, '..');
const value = argv[idx + 1];
if (!value) {
throw new Error('check-contract-drift: --root requires a directory argument');
}
return path.resolve(value);
}
const ROOT = resolveRoot(process.argv.slice(2));
const CONTRACTS_FILE = path.join(ROOT, 'gsd-core', 'references', 'agent-contracts.md');
const AGENTS_DIR = path.join(ROOT, 'agents');
const WORKFLOWS_DIR = path.join(ROOT, 'gsd-core', 'workflows');
const COMMANDS_DIR = path.join(ROOT, 'commands');
const {
extractMarkers,
parseAgentContracts,
contractViolations,
readTagViolations,
parseConsumedByCell,
unmatchedConsumerTokens,
sanitizeEcho,
REMEDIES,
} = require('./command-contract-helpers.cjs');
const { runMain } = require('./lib/cli-exit.cjs');
// ─── helpers ────────────────────────────────────────────────────────────────
function walkMarkdownFiles(dir, acc) {
let entries;
try {
entries = fs.readdirSync(dir, { withFileTypes: true });
} catch {
return acc;
}
for (const entry of entries) {
const full = path.join(dir, entry.name);
if (entry.isDirectory()) {
walkMarkdownFiles(full, acc);
} else if (entry.isFile() && entry.name.endsWith('.md')) {
acc.push(full);
}
}
return acc;
}
function toRepoRelative(absPath) {
return path.relative(ROOT, absPath).split(path.sep).join('/');
}
/**
* referenceIncludes(content)
*
* Plain scan for `@~/.claude/gsd-core/references/*.md` tokens anywhere in an
* agent file's content -- inside an `<execution_context>` block (already
* covered structurally by `executionContextRefs` in command-contract-helpers,
* but a raw regex over the whole string picks those up too) and, just as
* importantly, OUTSIDE one: agents frequently point at a reference doc from
* plain prose (e.g. "See @~/.claude/gsd-core/references/planner-guidance.md
* for ...") rather than from the eager `<execution_context>` include list.
* An agent's completion-marker contract can be authored in such a reference
* file rather than the agent file itself -- gsd-planner declares
* `PLANNING COMPLETE` in its registry row, but the example heading itself
* lives in `gsd-core/references/planner-guidance.md`, which the agent only
* `@`-includes -- so the producer scan below must follow these includes to
* see markers an agent's contract legitimately delegates to a reference doc.
* Returns ROOT-relative paths (`gsd-core/references/foo.md`), de-duplicated.
*/
function referenceIncludes(content) {
const seen = new Set();
const re = /@~\/\.claude\/gsd-core\/references\/[A-Za-z0-9._-]+\.md/g;
let m;
while ((m = re.exec(content)) !== null) {
const relPath = 'gsd-core/references/' + m[0].slice('@~/.claude/gsd-core/references/'.length);
seen.add(relPath);
}
return [...seen];
}
function remedyFor(kind) {
return REMEDIES[kind] || 'review the registry row and agent file for drift';
}
// ─── run ─────────────────────────────────────────────────────────────────────
function main() {
if (!fs.existsSync(CONTRACTS_FILE)) {
process.stderr.write(`\nERROR check-contract-drift: contracts file not found at ${toRepoRelative(CONTRACTS_FILE)}\n\n`);
return 1;
}
const contractsMd = fs.readFileSync(CONTRACTS_FILE, 'utf-8');
const { rows: registry, errors: parseErrors } = parseAgentContracts(contractsMd);
// knownMarkers: every marker string declared anywhere in the registry --
// extractMarkers only ever resolves a heading against this vocabulary, it
// never falls back to guessing a shape. Includes `(unconsumed: …)` entries
// so their declared↔emitted agreement is checked too.
const knownMarkers = new Set();
for (const row of registry) {
for (const m of row.completion_markers || []) knownMarkers.add(m);
for (const m of row.unconsumed_markers || []) knownMarkers.add(m);
}
// producerMarkers: agent -> [in-fence marker strings that matched knownMarkers]
// candidateMarkers: agent -> [in-fence marker-shaped headings NOT in knownMarkers,
// deduped per marker — "emitted but undeclared" is a fact about the
// marker, not about each line or file it appears in]
// agentTexts: agent -> file content CONCATENATED with every
// references/*.md file the agent @-includes, for the read-tag arm's gate
// check. The fold is load-bearing: planner/executor/phase-researcher
// deliver the MUST-Read gate via the shared mandatory-initial-read.md
// include, so the agent file alone would report a gate that actually
// arrives (the same producer-scope gap referenceIncludes() fixes for
// markers, one layer up).
const producerMarkers = new Map();
const candidateMarkers = new Map();
const agentTexts = new Map();
const unclosedFenceViolations = [];
const agentFiles = fs.existsSync(AGENTS_DIR)
? fs.readdirSync(AGENTS_DIR).filter(f => f.endsWith('.md'))
: [];
for (const file of agentFiles) {
const agent = file.replace(/\.md$/, '');
const abs = path.join(AGENTS_DIR, file);
const content = fs.readFileSync(abs, 'utf-8');
// Single pass over the agent's @-included references: each file is read
// once and feeds BOTH the read-tag fold (agentTexts) and marker
// extraction (producer/candidate attribution).
const includeTexts = [];
for (const refRelPath of referenceIncludes(content)) {
try {
includeTexts.push(fs.readFileSync(path.join(ROOT, refRelPath), 'utf-8'));
} catch {
// include miss — lint-command-contract rule 4 owns @-ref existence
}
}
agentTexts.set(agent, [content, ...includeTexts].join('\n'));
const { markers, candidates, unclosedFence } = extractMarkers(agentTexts.get(agent), knownMarkers);
const inFenceMarkers = markers.filter(m => m.inFence).map(m => m.marker);
const inFenceCandidates = candidates.filter(m => m.inFence).map(m => m.marker);
producerMarkers.set(agent, inFenceMarkers);
candidateMarkers.set(agent, [...new Set(inFenceCandidates)]);
if (unclosedFence) {
unclosedFenceViolations.push({
kind: 'unclosed_fence',
agent,
marker: null,
detail: `${toRepoRelative(abs)} has an unterminated code fence`,
});
}
}
// consumerTexts: every *.md under gsd-core/workflows/, commands/, agents/
// — plus every file-shaped `Consumed by` entry that resolves on disk, so
// a row citing a reference doc or an ADR (e.g. intel-updater's
// docs/adr/22-plan-drift-guard.md) is validated against the real file and
// its text participates in consumer matching, not just workflows.
const consumerTexts = new Map();
const consumerDirs = [WORKFLOWS_DIR, COMMANDS_DIR, AGENTS_DIR];
for (const dir of consumerDirs) {
for (const abs of walkMarkdownFiles(dir, [])) {
consumerTexts.set(toRepoRelative(abs), fs.readFileSync(abs, 'utf-8'));
}
}
for (const row of registry) {
for (const rel of parseConsumedByCell(row.consumed_by)) {
if (consumerTexts.has(rel)) continue;
const abs = path.join(ROOT, rel);
try {
consumerTexts.set(rel, fs.readFileSync(abs, 'utf-8'));
} catch {
// absent — contractViolations reports it as unknown_consumer
}
}
}
const contractViolationsList = contractViolations({ registry, producerMarkers, candidateMarkers, consumerTexts });
const readTagViolationsList = readTagViolations({ registry, agentTexts, consumerTexts });
const reverseViolationsList = unmatchedConsumerTokens({ consumerTexts, vocabulary: knownMarkers });
const parseViolations = parseErrors.map(e => ({
kind: 'parse_error',
agent: null,
marker: null,
detail: `agent-contracts.md:${e.line} — ${e.reason}`,
}));
const allViolations = [
...parseViolations,
...unclosedFenceViolations,
...contractViolationsList,
...readTagViolationsList,
...reverseViolationsList,
];
const agentCount = registry.length;
const markerCount = registry.reduce((n, r) => n + (r.completion_markers || []).length, 0);
// --json: the typed surface tests consume (CONTRIBUTING "Raw Text
// Matching" rule — the human formatter below is for operators only).
if (process.argv.includes('--json')) {
console.log(
JSON.stringify({
check: 'check-contract-drift',
status: allViolations.length === 0 ? 'ok' : 'violations',
agents: agentCount,
markers: markerCount,
violations: allViolations.map((v) => ({
kind: v.kind,
agent: v.agent ?? null,
marker: v.marker ?? null,
detail: sanitizeEcho(v.detail),
})),
}),
);
return allViolations.length === 0 ? 0 : 1;
}
if (allViolations.length === 0) {
console.log(`ok check-contract-drift: ${agentCount} agents, ${markerCount} markers, 0 violations`);
return 0;
}
// group by violation kind
const byKind = new Map();
for (const v of allViolations) {
if (!byKind.has(v.kind)) byKind.set(v.kind, []);
byKind.get(v.kind).push(v);
}
process.stderr.write(
`\nERROR check-contract-drift: ${allViolations.length} violation(s) across ${byKind.size} kind(s)\n\n`,
);
for (const [kind, violations] of byKind) {
process.stderr.write(` ${kind} (${violations.length}):\n`);
for (const v of violations) {
const agentLabel = v.agent ? sanitizeEcho(v.agent) : '(registry)';
const markerLabel = v.marker ? ` marker "${sanitizeEcho(v.marker)}"` : '';
process.stderr.write(` - ${agentLabel}${markerLabel}: ${sanitizeEcho(v.detail)}\n`);
process.stderr.write(` remedy: ${remedyFor(kind)}\n`);
}
process.stderr.write('\n');
}
process.stderr.write('See gsd-core/references/agent-contracts.md for the registry contract spec.\n\n');
return 1;
}
runMain(main);

View File

@@ -172,10 +172,795 @@ function unreachableWorkflows(loaderContents, gsdFiles, workflowPaths) {
return workflowPaths.filter(p => !visited.has(p));
}
/**
* splitOutsideDelimiter(str, delimiter)
*
* Splits `str` on `delimiter` characters that fall outside any backtick
* span AND outside any parenthetical span. A backtick toggles an "inside a
* code span" flag; `(` / `)` track nesting depth. Delimiters seen while
* inside either construct are treated as literal content rather than split
* points. Used both for comma-splitting a completion-markers cell and for
* pipe-splitting a table row, so a marker annotation
* (`(unconsumed: draft presentation, approved interactively)`) or a path
* that happens to contain the delimiter inside backticks or parens is never
* torn in half.
*/
function splitOutsideDelimiter(str, delimiter) {
const parts = [];
let cur = '';
let inBacktick = false;
let parenDepth = 0;
for (const ch of str) {
if (ch === '`') inBacktick = !inBacktick;
if (!inBacktick) {
if (ch === '(') parenDepth += 1;
else if (ch === ')' && parenDepth > 0) parenDepth -= 1;
}
if (ch === delimiter && !inBacktick && parenDepth === 0) {
parts.push(cur);
cur = '';
} else {
cur += ch;
}
}
parts.push(cur);
return parts;
}
function splitTableRow(line) {
let s = line.trim();
if (s.startsWith('|')) s = s.slice(1);
if (s.endsWith('|')) s = s.slice(0, -1);
return splitOutsideDelimiter(s, '|').map(c => c.trim());
}
function isSeparatorRow(line) {
const t = line.trim();
return t.length > 0 && /^[|:\- ]+$/.test(t) && t.includes('-');
}
/**
* parseCompletionMarkersCell(value)
*
* Parses one `Completion Markers` table cell into `{markers, notes}`.
* Entries are comma-separated OUTSIDE backticks (so a marker's own text
* never leaks a split point). Each entry is stripped of a trailing
* parenthetical annotation (`(title case)`, `(non-standard)`, ...) — that
* annotation is captured into `notes` rather than the marker text — then
* stripped of surrounding backticks and a leading `## ` heading prefix. A
* bare `No marker` entry (post-annotation-strip) contributes nothing to
* `markers`; its parenthetical still lands in `notes`.
*/
function parseCompletionMarkersCell(value) {
const markers = [];
const notes = [];
const unconsumedMarkers = [];
if (!value) return { markers, notes, unconsumedMarkers };
const entries = splitOutsideDelimiter(value, ',');
for (const rawEntry of entries) {
const entry = rawEntry.trim();
if (entry === '') continue;
let markerPortion = entry;
let annotation = null;
const noteMatch = entry.match(/^(.*?)\s*\(([^)]*)\)\s*$/);
if (noteMatch) {
markerPortion = noteMatch[1].trim();
annotation = noteMatch[2].trim();
notes.push(annotation);
}
markerPortion = markerPortion.replace(/`/g, '').trim();
markerPortion = markerPortion.replace(/^#{1,6}\s+/, '').trim();
if (markerPortion === '' || markerPortion === 'No marker') continue;
// An `(unconsumed: <reason>)` annotation declares a marker that is
// emitted deliberately but matched by no consumer (e.g. a user-facing
// presentation format). It stays in the declared+emitted contract and in
// case-collision scope; only the consumer requirement is waived.
if (annotation !== null && /^unconsumed\b/i.test(annotation)) {
unconsumedMarkers.push(markerPortion);
continue;
}
markers.push(markerPortion);
}
return { markers, notes, unconsumedMarkers };
}
/**
* extractMarkers(content, knownMarkers)
*
* Scans a markdown string line by line for H1-H6 headings, and reports
* whether each heading sits inside a fenced code block. A heading counts as
* a marker only when its captured text EXACTLY matches (case-sensitively)
* one of `knownMarkers` — the declared vocabulary from the agent-contracts
* registry, supplied by the caller. There is deliberately no shape
* heuristic (e.g. "ends in COMPLETE") here: the registry already declares
* every marker string, several of which do NOT end in COMPLETE (`##
* ISSUES FOUND`, `## ROADMAP CREATED`, `## CHECKPOINT REACHED`, `##
* ESCALATE`, ...), so guessing the shape both over-matches ordinary prose
* headings that happen to look marker-shaped and under-matches every real
* marker that doesn't end in the word COMPLETE. When `knownMarkers` is
* omitted or empty, `markers` is always `[]` — this function never falls
* back to guessing.
*
* `candidates` separately reports in-fence headings that are NOT in
* `knownMarkers` but still look marker-shaped (`/^[A-Z][A-Z0-9 _:'\/-]*$/`,
* no backticks, <= 60 chars) — these feed the `emitted_marker_not_declared`
* check. They are reported as candidates, never silently treated as
* markers.
*
* Fence tracking is a single boolean, not a stack: a line whose trimmed
* form starts with a run of 3+ backticks or tildes toggles it, but only a
* closing run of the SAME character and AT LEAST the same length actually
* closes an open fence — a shorter or differently-charred run nested inside
* (e.g. a 3-backtick span shown inside a 4-backtick outer fence) is inert
* and leaves the outer fence open. `unclosedFence` reports when the file
* ends without ever closing the last-opened fence, so callers can tell a
* truncated/malformed file from one that's actually well-formed.
*/
function extractMarkers(content, knownMarkers) {
const markers = [];
const candidates = [];
if (content == null || typeof content !== 'string' || content.trim() === '') {
return { markers, candidates, unclosedFence: false };
}
const known = new Set(knownMarkers || []);
const lines = content.split(/\r?\n/);
const headingRe = /^(#{1,6})\s+(\S.*?)\s*$/;
const candidateShapeRe = /^[A-Z][A-Z0-9 _:'/-]*$/;
let inFence = false;
let fenceChar = null;
let fenceLen = 0;
for (let i = 0; i < lines.length; i++) {
const line = lines[i];
const trimmed = line.trim();
const stateForThisLine = inFence;
const fenceMatch = trimmed.match(/^(`{3,}|~{3,})/);
if (fenceMatch) {
const runChar = fenceMatch[1][0];
const runLen = fenceMatch[1].length;
if (!inFence) {
inFence = true;
fenceChar = runChar;
fenceLen = runLen;
} else if (runChar === fenceChar && runLen >= fenceLen) {
inFence = false;
fenceChar = null;
fenceLen = 0;
}
// else: nested/mismatched run — does not close the outer fence.
}
const hm = line.match(headingRe);
if (hm) {
const text = hm[2].trim();
if (known.has(text)) {
markers.push({ marker: text, inFence: stateForThisLine, line: i + 1 });
} else if (stateForThisLine && text.length <= 60 && candidateShapeRe.test(text)) {
candidates.push({ marker: text, inFence: stateForThisLine, line: i + 1 });
}
}
}
return { markers, candidates, unclosedFence: inFence };
}
/**
* parseAgentContracts(markdown)
*
* Parses the `## Agent Registry` pipe table in
* gsd-core/references/agent-contracts.md into structured rows, keyed by
* lowercased, underscore-joined column headers (`Completion Markers` ->
* `completion_markers`). Column set is read from the header row itself, so
* the table can grow a `Consumed By` / `Kind` column later without this
* parser needing an update — rows simply gain the corresponding key.
*
* The `completion_markers` column gets special handling via
* parseCompletionMarkersCell: it is never a plain string on the returned
* row, always the parsed `{markers}` array (with any parenthetical
* annotation split out into a row-level `notes` array instead).
*
* Malformed rows (wrong cell count for the header, or an empty agent name)
* are reported in `errors` as `{line, reason}` and excluded from `rows`
* rather than thrown.
*/
function parseAgentContracts(markdown) {
const rows = [];
const errors = [];
if (markdown == null || typeof markdown !== 'string') return { rows, errors };
const lines = markdown.split(/\r?\n/);
const registryIdx = lines.findIndex(l => /^##\s+Agent Registry\s*$/.test(l.trim()));
if (registryIdx === -1) return { rows, errors };
let i = registryIdx + 1;
while (i < lines.length && lines[i].trim() === '') i++;
if (i >= lines.length || !lines[i].trim().startsWith('|')) return { rows, errors };
const columnKeys = splitTableRow(lines[i]).map(h =>
h.trim().toLowerCase().replace(/\s+/g, '_')
);
i++;
if (i < lines.length && isSeparatorRow(lines[i])) i++;
for (; i < lines.length; i++) {
const raw = lines[i];
if (!raw.trim().startsWith('|')) break;
const lineNum = i + 1;
const cells = splitTableRow(raw);
if (cells.length !== columnKeys.length) {
errors.push({
line: lineNum,
reason: `column count mismatch: expected ${columnKeys.length}, got ${cells.length}`,
});
continue;
}
const row = {};
for (let j = 0; j < columnKeys.length; j++) {
const key = columnKeys[j];
const value = cells[j];
if (key === 'completion_markers') {
const parsed = parseCompletionMarkersCell(value);
row.completion_markers = parsed.markers;
row.notes = parsed.notes;
row.unconsumed_markers = parsed.unconsumedMarkers;
} else {
row[key] = value;
}
}
if (!row.agent) {
errors.push({ line: lineNum, reason: 'empty agent name' });
continue;
}
rows.push(row);
}
return { rows, errors };
}
const CONTRACT_KINDS = new Set(['sentinel-match', 'artifact+query', 'structured-return']);
/**
* contractViolations({registry, producerMarkers, candidateMarkers, consumerTexts})
*
* Cross-checks the agent-contracts registry (from parseAgentContracts)
* against what agent files actually emit in-fence (from extractMarkers) and
* what every workflow/command/agent file actually consumes, exact-case.
* Rows are evaluated independently; `registry`, `producerMarkers`,
* `candidateMarkers`, and `consumerTexts` default to empty when omitted so a
* partial call never throws.
*
* `emitted_marker_not_declared` is driven by `candidateMarkers` (agent ->
* [in-fence heading text]) rather than `producerMarkers` -- a heading only
* lands in `producerMarkers` once it already matches a KNOWN registry
* marker (see extractMarkers), so it can never be "emitted but undeclared";
* `candidateMarkers` is exactly the set of marker-shaped in-fence headings
* extractMarkers could NOT resolve against the known vocabulary.
*
* A row declaring `kind: 'artifact+query'` or `'structured-return'` opts
* out of the sentinel-marker contract entirely — its only possible
* violation is `vestigial_marker` (it still emits an in-fence marker
* despite the row saying completion isn't detected that way). Every other
* row (kind `sentinel-match`, or the column simply absent/empty) is held to
* the full contract: declared markers must be emitted, emitted markers must
* be declared, and — for sentinel-match rows specifically — every declared
* marker must have at least one EXACT-CASE consumer somewhere in
* `consumerTexts`, excluding the producing agent's own file (an agent can
* never satisfy its own marker, even when the same map entry also serves as
* a consumer of some OTHER agent's marker). A case-insensitive-only match
* is reported (`case_only_match`), never treated as satisfying the
* contract. `case_collision` is computed once across the whole registry,
* independent of kind, since two markers differing only by case is a
* registry-wide authoring hazard.
*/
function contractViolations({ registry, producerMarkers, candidateMarkers, consumerTexts } = {}) {
const violations = [];
try {
const rows = Array.isArray(registry) ? registry : [];
const producers = producerMarkers instanceof Map ? producerMarkers : new Map();
const candidatesByAgent = candidateMarkers instanceof Map ? candidateMarkers : new Map();
const consumers = consumerTexts instanceof Map ? consumerTexts : new Map();
// case_collision — global, across every declared marker in the registry,
// including `(unconsumed: …)` entries: a case-variant of an unconsumed
// marker is still an authoring hazard even though no consumer matches it.
const byLower = new Map();
for (const row of rows) {
for (const m of [...(row.completion_markers || []), ...(row.unconsumed_markers || [])]) {
const lower = m.toLowerCase();
if (!byLower.has(lower)) byLower.set(lower, new Set());
byLower.get(lower).add(m);
}
}
for (const row of rows) {
for (const m of [...(row.completion_markers || []), ...(row.unconsumed_markers || [])]) {
const variants = byLower.get(m.toLowerCase());
if (variants && variants.size > 1) {
const others = [...variants].filter(v => v !== m);
violations.push({
kind: 'case_collision',
agent: row.agent,
marker: m,
detail: `case-insensitive collision with: ${others.join(', ')}`,
});
}
}
}
// roster coverage — the registry is the declaration of record for every
// agents/*.md file: an agent with no row, two rows for one agent, or a
// row naming an agent that does not exist are all contract drift. The
// roster is the producer map's key set; guards on map size keep empty
// inputs inert rather than spuriously "missing".
const seenAgents = new Set();
for (const row of rows) {
const agent = row.agent;
if (seenAgents.has(agent)) {
violations.push({
kind: 'duplicate_registry_row',
agent,
marker: null,
detail: `agent "${agent}" has more than one registry row`,
});
}
seenAgents.add(agent);
if (producers.size > 0 && !producers.has(agent)) {
violations.push({
kind: 'unknown_producer',
agent,
marker: null,
detail: `registry row for "${agent}" but no agents/${agent}.md exists`,
});
}
}
for (const agent of producers.keys()) {
if (!seenAgents.has(agent)) {
violations.push({
kind: 'agent_without_contract',
agent,
marker: null,
detail: `agents/${agent}.md exists but has no registry row in agent-contracts.md`,
});
}
}
for (const row of rows) {
const agent = row.agent;
const declared = row.completion_markers || [];
const unconsumed = row.unconsumed_markers || [];
const declaredAll = [...declared, ...unconsumed];
// Deduped: the same marker emitted on several lines/files is one
// emission fact, never several violations.
const produced = [...new Set(producers.get(agent) || [])];
const kind = row.kind;
const kindProvided = kind !== undefined && kind !== null && kind !== '';
// A row naming a nonexistent agent file: unknown_producer (above) is
// the violation; per-marker noise on top of it hides nothing.
if (producers.size > 0 && !producers.has(agent)) continue;
// unknown_consumer — the `Consumed by` column feeds readTagViolations,
// so a file-shaped entry that resolves to nothing would make that arm a
// lint over fiction. Only checked when a consumer corpus was supplied.
for (const p of parseConsumedByCell(row.consumed_by)) {
if (consumers.size > 0 && !consumers.has(p)) {
violations.push({
kind: 'unknown_consumer',
agent,
marker: null,
detail: `Consumed by names "${p}" but no such file exists`,
});
}
}
if (kindProvided && !CONTRACT_KINDS.has(kind)) {
violations.push({
kind: 'unknown_kind',
agent,
marker: null,
detail: `row declares unknown kind "${kind}"`,
});
}
const isArtifactOrStructured =
kindProvided && (kind === 'artifact+query' || kind === 'structured-return');
// declaredAll includes unconsumed entries: an `(unconsumed: …)` marker
// must still be emitted — the exemption waives the consumer
// requirement only, never the declared↔emitted agreement (Marker
// Rule 7). This runs for EVERY row kind.
for (const m of declaredAll) {
if (!produced.includes(m)) {
violations.push({
kind: 'declared_marker_not_emitted',
agent,
marker: m,
detail: `declared in registry but not emitted in-fence by ${agent}`,
});
}
}
if (isArtifactOrStructured) {
// An `(unconsumed: …)` annotation on an artifact/structured row is a
// recorded decision to keep emitting a marker the completion route
// doesn't need (Marker Rule 2's intentional title-case markers) —
// exempt from vestigial_marker. Every OTHER emitted marker on such a
// row is reported ONCE, as vestigial_marker: the candidate check
// below is skipped for these rows because emitted_marker_not_declared
// on top of vestigial_marker double-reports the same fact, and the
// consumer contract cannot apply to a row whose completion is not
// detected by marker matching at all.
const unconsumedSet = new Set(unconsumed);
for (const m of produced) {
if (unconsumedSet.has(m)) continue;
violations.push({
kind: 'vestigial_marker',
agent,
marker: m,
detail: `row kind "${kind}" expects no marker, but agent still emits "${m}" in-fence`,
});
}
continue;
}
// Deduped per marker: "emitted but undeclared" is a fact about the
// marker, and a caller-supplied list carrying the same marker twice
// (two template lines, or agent file + @-included reference) is one
// violation, not two.
const candidates = [...new Set(candidatesByAgent.get(agent) || [])];
for (const m of candidates) {
if (!declaredAll.includes(m)) {
violations.push({
kind: 'emitted_marker_not_declared',
agent,
marker: m,
detail: `emitted in-fence by ${agent} but absent from its registry row`,
});
}
}
const isSentinelMatchKind = !kindProvided || kind === 'sentinel-match';
if (isSentinelMatchKind) {
// Only markers the agent actually emits participate in the consumer
// contract — a declared_marker_not_emitted finding already covers a
// phantom declaration, and stacking no_consumer on top of it is
// double-reporting the same dead row.
const declaredConsumerPaths = parseConsumedByCell(row.consumed_by);
for (const m of declared) {
if (!produced.includes(m)) continue;
let exactMatch = false;
let caseInsensitiveMatch = false;
let declaredConsumerMatch = false;
for (const [consumerId, text] of consumers.entries()) {
// consumerTexts is keyed by repo-relative path — an agent's own
// file is 'agents/<name>.md'. The producer's own template can
// never satisfy its marker.
if (consumerId === `agents/${agent}.md`) continue;
if (typeof text !== 'string') continue;
if (text.includes(m)) {
exactMatch = true;
if (declaredConsumerPaths.includes(consumerId)) declaredConsumerMatch = true;
} else if (text.toLowerCase().includes(m.toLowerCase())) {
caseInsensitiveMatch = true;
}
}
if (!exactMatch) {
if (caseInsensitiveMatch) {
violations.push({
kind: 'case_only_match',
agent,
marker: m,
detail: `no exact-case consumer for "${m}"; case-insensitive match found only`,
});
} else {
violations.push({
kind: 'no_consumer',
agent,
marker: m,
detail: `no consumer text contains "${m}" (exact case)`,
});
}
} else if (
declaredConsumerPaths.length > 0 &&
!declaredConsumerMatch
) {
// The marker IS consumed somewhere, but by none of the consumers
// the row cites — the `Consumed by` cell is documentation the
// read-tag arm and humans both navigate by, so a cell naming
// files that never match is registry drift even when the corpus
// at large happens to contain the token.
violations.push({
kind: 'declared_consumer_no_match',
agent,
marker: m,
detail: `consumed somewhere, but none of the row's Consumed by entries ( ${declaredConsumerPaths.join(', ')} ) contain "${m}"`,
});
}
}
}
}
} catch (e) {
// A drift gate must fail RED, never green: an internal error here must
// surface, not be swallowed into '0 violations'.
e.message = `contractViolations (internal failure — failing red): ${e.message}`;
throw e;
}
return violations;
}
/**
* parseConsumedByCell(value)
*
* Extracts every file-shaped path from a `Consumed by` table cell. Cell
* entries are comma-separated outside backticks, and an entry may carry
* several backtick-wrapped tokens (a path, a command, a marker string) in
* prose — so every backtick span is inspected and only tokens that look like
* repo-relative markdown paths (`^[\w][\w./-]*\.md$` — no spaces, no glob
* stars) are returned. A glob (`*-VERIFICATION.md`) or a command
* (`gsd_run query verification.status`) describes the consumption mechanism
* in prose; it is not a file the check can open, and is ignored.
*/
function parseConsumedByCell(value) {
if (!value || typeof value !== 'string') return [];
const paths = [];
const re = /`([^`]+)`/g;
let m;
while ((m = re.exec(value)) !== null) {
const token = m[1].trim();
if (!/^[\w][\w./-]*\.md$/.test(token) || token.includes('*')) continue;
// Traversal segments are dropped, never surfaced — the same rule
// workflowPathRefs applies to its refs: this resolver only ever reports
// paths inside the repo, never something a `..` could walk outside of
// (the driver reads these paths under --root, so an interior `..` would
// be an out-of-root file read).
if (token.split('/').includes('..')) continue;
paths.push(token);
}
return paths;
}
/**
* readTagViolations({registry, agentTexts, consumerTexts})
*
* The read-tag arm — F8 one layer up. Two invariants:
*
* legacy_read_tag — no `<files_to_read>` may survive anywhere in the
* consumer corpus (workflows/, commands/, agents/). #3423 standardized on
* `<required_reading>` and retired the old vocabulary via a test; this
* folds that assertion into the gate itself.
*
* read_tag_gate_missing — for every registry row, every file-shaped
* `Consumed by` consumer that itself emits `<required_reading>` requires
* the producing agent's file to reference `<required_reading>`. A spawner
* that sends a required-reading block to an agent whose instructions never
* mention the gate is exactly the F8 shape: the emit side is correct, the
* gate side never fires — and #3423's per-agent test cannot see it,
* because an agent carrying no reading instruction at all is skipped by
* it. The registry supplies the producer↔consumer pair, so this check
* needs no spawn-site parsing heuristics.
*/
function readTagViolations({ registry, agentTexts, consumerTexts } = {}) {
const violations = [];
try {
const rows = Array.isArray(registry) ? registry : [];
const agents = agentTexts instanceof Map ? agentTexts : new Map();
const consumers = consumerTexts instanceof Map ? consumerTexts : new Map();
for (const [file, text] of consumers.entries()) {
if (typeof text === 'string' && text.includes('<files_to_read>')) {
violations.push({
kind: 'legacy_read_tag',
agent: null,
marker: null,
detail: `${file} still contains the retired <files_to_read> tag`,
});
}
}
for (const row of rows) {
const agent = row.agent;
const agentText = agents.get(agent);
if (typeof agentText !== 'string') continue; // unknown_producer reports this
if (agentText.includes('<required_reading>')) continue;
for (const p of parseConsumedByCell(row.consumed_by)) {
const text = consumers.get(p);
if (typeof text === 'string' && text.includes('<required_reading>')) {
violations.push({
kind: 'read_tag_gate_missing',
agent,
marker: null,
detail: `consumer ${p} emits <required_reading> but agents/${agent}.md never references the gate`,
});
break; // one violation per agent is the fact; more is noise
}
}
}
} catch (e) {
// A drift gate must fail RED, never green: an internal error here must
// surface, not be swallowed into '0 violations'.
e.message = `readTagViolations (internal failure — failing red): ${e.message}`;
throw e;
}
return violations;
}
/**
* Tokens that look like completion markers but are NOT agent contracts, and
* therefore exempt from the reverse-direction check. Each entry carries its
* justification inline; a test locks the set's membership so it cannot grow
* silently.
*/
const NON_AGENT_TOKENS = new Set([
// Display headings the /gsd:graphify command itself renders from
// `gsd_run graphify build` CLI output — tool output shown to the user,
// not an agent return any spawner dispatches on.
'GRAPHIFY BUILD COMPLETE',
'GRAPHIFY BUILD FAILED',
]);
/**
* unmatchedConsumerTokens({consumerTexts, vocabulary})
*
* The reverse direction of the registry contract: a consumer that matches a
* `## TOKEN` string no producer emits is dispatch-on-phantom — F9's shape
* (a sanctioned return that no dispatch branch ever keyed on), seen from
* the consumer side. Scans ONLY workflow/command files (the dispatch
* surfaces); `agents/` is excluded because an agent mentioning a token in
* its own file is a self-description, not a spawner's match instruction —
* the forward direction already validates agent-to-agent consumption.
*
* A token is extracted when it appears quoted (`## X`, "## X", '## X') —
* the shape a match/dispatch instruction actually writes — is ALL-CAPS
* marker-shaped (the same candidate shape extractMarkers uses, ≤60 chars),
* is absent from the declared vocabulary, and is not exempt via
* NON_AGENT_TOKENS.
*/
function unmatchedConsumerTokens({ consumerTexts, vocabulary } = {}) {
const violations = [];
try {
const consumers = consumerTexts instanceof Map ? consumerTexts : new Map();
const known = vocabulary instanceof Set ? vocabulary : new Set(knownMarkersArray(vocabulary));
const shapeRe = /^[A-Z][A-Z0-9 _:'/-]*$/;
const tokenRe = /[`'"]## ([^`'"\n]+)[`'"]/g;
const reported = new Set();
for (const [file, text] of consumers.entries()) {
if (!file.startsWith('gsd-core/workflows/') && !file.startsWith('commands/')) continue;
if (typeof text !== 'string') continue;
tokenRe.lastIndex = 0;
let m;
while ((m = tokenRe.exec(text)) !== null) {
const token = m[1].trim();
if (known.has(token)) continue;
if (NON_AGENT_TOKENS.has(token)) continue;
if (token.length > 60 || !shapeRe.test(token)) continue;
const key = `${file}\0${token}`;
if (reported.has(key)) continue;
reported.add(key);
violations.push({
kind: 'unmatched_consumer_token',
agent: null,
marker: token,
detail: `${file} matches \`## ${token}\` but no agent declares or emits it`,
});
}
}
} catch (e) {
// A drift gate must fail RED, never green: an internal error here must
// surface, not be swallowed into '0 violations'.
e.message = `unmatchedConsumerTokens (internal failure — failing red): ${e.message}`;
throw e;
}
return violations;
}
function knownMarkersArray(vocabulary) {
// vocabulary may arrive as an array (tests) — normalize; the Set branch
// above handles the driver's Set.
return Array.isArray(vocabulary) ? vocabulary : [];
}
/**
* The frozen violation-kind vocabulary — one entry per finding the check
* can emit. Tests lock this set and the driver's REMEDIES parity, so a new
* kind is three coordinated changes: enum + remedy + test, the same
* discipline verify-reapply-patches' REASON enum follows.
*/
const VIOLATION_KINDS = Object.freeze({
PARSE_ERROR: 'parse_error',
UNCLOSED_FENCE: 'unclosed_fence',
DECLARED_MARKER_NOT_EMITTED: 'declared_marker_not_emitted',
EMITTED_MARKER_NOT_DECLARED: 'emitted_marker_not_declared',
CASE_COLLISION: 'case_collision',
CASE_ONLY_MATCH: 'case_only_match',
NO_CONSUMER: 'no_consumer',
DECLARED_CONSUMER_NO_MATCH: 'declared_consumer_no_match',
UNKNOWN_KIND: 'unknown_kind',
VESTIGIAL_MARKER: 'vestigial_marker',
AGENT_WITHOUT_CONTRACT: 'agent_without_contract',
DUPLICATE_REGISTRY_ROW: 'duplicate_registry_row',
UNKNOWN_PRODUCER: 'unknown_producer',
UNKNOWN_CONSUMER: 'unknown_consumer',
LEGACY_READ_TAG: 'legacy_read_tag',
READ_TAG_GATE_MISSING: 'read_tag_gate_missing',
UNMATCHED_CONSUMER_TOKEN: 'unmatched_consumer_token',
});
/**
* The remedy shown per violation kind — one entry per VIOLATION_KINDS value.
* The parity test locks REMEDIES keys against VIOLATION_KINDS values, so a
* new kind cannot ship without its remedy.
*/
const REMEDIES = Object.freeze({
parse_error: 'fix the malformed row in gsd-core/references/agent-contracts.md',
unclosed_fence: 'close the unterminated code fence in this agent file',
declared_marker_not_emitted:
'remove the marker from the registry row, or emit it as an in-fence example in the agent file',
emitted_marker_not_declared:
"add this marker to the agent's Completion Markers cell in the registry table",
case_collision: 'rename one of the colliding markers so they no longer differ only by case',
case_only_match:
'fix the consumer file to match this marker exact-case, or update the marker/registry to match the consumer',
no_consumer:
'add an exact-case consumer for this marker, or reclassify the row Kind to artifact+query/structured-return',
declared_consumer_no_match:
"the marker is consumed, but not by any file the row's Consumed by cell names — fix the cell to name a real consumer",
unknown_kind: 'set the Kind column to one of sentinel-match, artifact+query, structured-return',
vestigial_marker:
"remove this marker heading from the agent file, since the row's Kind says no marker is expected",
agent_without_contract: 'add a registry row for this agent in agent-contracts.md',
duplicate_registry_row: 'merge the two rows for this agent into one',
unknown_producer: 'remove the row, or create the agents/<name>.md file it names',
unknown_consumer:
'fix the path in the Consumed by cell, or drop it — only real files the check can open belong there',
legacy_read_tag:
'replace <files_to_read> with <required_reading> (#3423 standardized on the gate tag)',
read_tag_gate_missing:
"add the <required_reading> MUST-Read gate clause to this agent's instructions (see docs/AGENTS.md)",
unmatched_consumer_token:
'declare the producing agent and marker in the registry, or remove the dispatch — nothing emits this token',
});
/**
* sanitizeEcho(text)
*
* File-derived strings (registry cells, marker tokens, workflow text) are
* echoed into lint output that AI agents read as trusted instructions —
* every interpolated field is control-char-stripped and length-capped
* before it is written, so a malicious repo file cannot make the lint emit
* instruction-shaped text. Implemented as a code-point filter rather than a
* regex so no control-character regex literal exists to lint.
*/
function sanitizeEcho(text) {
let out = '';
for (const ch of String(text)) {
const code = ch.codePointAt(0);
if (code <= 0x1f || code === 0x7f) continue;
out += ch;
}
return out.length > 200 ? out.slice(0, 200) : out;
}
module.exports = {
CANONICAL_TOOLS,
parseFrontmatter,
executionContextRefs,
workflowPathRefs,
unreachableWorkflows,
extractMarkers,
parseAgentContracts,
contractViolations,
parseConsumedByCell,
readTagViolations,
unmatchedConsumerTokens,
NON_AGENT_TOKENS,
VIOLATION_KINDS,
REMEDIES,
sanitizeEcho,
};

View File

@@ -7,19 +7,38 @@
* ## Why
*
* A file/key gets removed because "no longer used" without verifying every
* consumer (workflows, docs, manifests, npm scripts). #3316: root
* consumer (workflows, docs, manifests, npm scripts, tests). #3316: root
* `package-lock.json` was deleted while `package.json` still declares deps
* and workflows still use `cache: 'npm'` + `npm ci` (which require a
* lockfile). e3b52c70: docs referenced a removed `/gsd-new-workspace`
* workflow after it was deleted.
* workflow after it was deleted. #3560: a deleted workflow was pinned by an
* existence assertion in `tests/phase.test.cjs` — the lint passed clean and
* the breakage surfaced only as four red tests on the remote runner,
* because `tests/` was not scanned at all.
*
* ## What this checks
*
* For every file deleted (`git diff --name-status <base>...HEAD`, status
* `D`), grep the post-diff tree (`.github/workflows/`, `gsd-core/`, `docs/`,
* `package.json`) for the deleted file's basename. Fails if any reference
* survives. `package-lock.json` deletions additionally fail if any workflow
* still uses `npm ci` or `cache: 'npm'`/`cache: "npm"` — those depend on a
* `D`), grep the post-diff tree for the deleted file's basename:
*
* - `.github/workflows/`, `gsd-core/`, `docs/`, `package.json` — ANY
* surviving reference fails (the original rule).
* - `tests/` — scanned with a discriminator (#3565): a reference that PINS
* existence (`fs.existsSync(path)`, `readFileSync`, `require`, or the
* basename as a quoted object key) fails; a reference that ASSERTS
* ABSENCE (`assert.ok(!content.includes('x.md'))`, `!fs.existsSync(...)`)
* is the correct post-deletion state and passes. The naive widening —
* flagging every mention — was tried in #3560 and reverted: an absence
* assertion must contain the basename to assert the file is gone, so
* without the two kinds separated the guard fires on correct code and
* gets disabled. A mention that is neither a pin nor an absence
* assertion (prose, a path assembled on another line) is not flagged —
* a documented known limit, chosen because a false violation reds a
* correct tree while a missed prose mention only loses one detection
* channel.
*
* `package-lock.json` deletions additionally fail if any workflow still
* uses `npm ci` or `cache: 'npm'`/`cache: "npm"` — those depend on a
* lockfile even though they never spell out its filename.
*
* ## False-positive risk (moderate, per audit)
@@ -34,10 +53,12 @@ const path = require('node:path');
const cp = require('node:child_process');
const { ExitError, runMain } = require('./lib/cli-exit.cjs');
const { escapeRegex } = require('../gsd-core/bin/lib/pattern.cjs');
const { sanitizeEcho } = require('./command-contract-helpers.cjs');
const ROOT = path.join(__dirname, '..');
const SCAN_ROOTS = ['.github/workflows', 'gsd-core', 'docs'];
const EXTRA_FILES = ['package.json'];
const TESTS_ROOT = 'tests';
// Skip these when walking SCAN_ROOTS — binary/generated content that can
// never carry a meaningful basename reference, and is often large.
@@ -115,6 +136,81 @@ function findSurvivingReferences(deletedFiles, corpus) {
return violations;
}
/**
* Pure: classify ONE line of a test file that references `basename`.
* Returns 'asserts-absence' | 'pins-existence' | null.
*
* - asserts-absence: a NEGATED check carrying the basename —
* `assert.ok(!content.includes('x.md'))`, `!fs.existsSync(p)`,
* `!/x\.md/.test(s)`. The basename must appear for the assertion to
* work, so this is the CORRECT post-deletion state (#3560's reverted
* naive widening fired exactly here).
* - pins-existence: a direct dependency on the file existing —
* `fs.existsSync/readFileSync/statSync/readdirSync/require(...)` on the
* same line, or the basename as a quoted object key (`'x.md': ...`,
* the allowlist trap — a key means the test resolves against the file).
* - null: neither (prose mention, path assembled on another line). Not a
* violation — see the header's known-limit note.
*
* @param {string} line
* @param {string} basename
* @returns {'asserts-absence'|'pins-existence'|null}
*/
function classifyTestReference(line, basename) {
const negatedCheck =
/!\s*[A-Za-z_$][\w.$]*\.(includes|indexOf|search|startsWith|endsWith|match)\s*\(/.test(line) ||
/!\s*(fs\.)?(existsSync|statSync)\s*\(/.test(line) ||
// [^/\n]* (not .*) so an adversarial line cannot backtrack the regex
/!\s*\/[^/\n]*\/\s*\.(test|match)\s*\(/.test(line);
if (negatedCheck) return 'asserts-absence';
const pinsByCall =
/\b(existsSync|readFileSync|statSync|readdirSync|require|accessSync)\s*\(/.test(line) ||
/\brequire\s*\(?['"]/.test(line);
if (pinsByCall) return 'pins-existence';
// Template literal needs \\s so the RegExp constructor receives \s — a
// bare \s in a template literal collapses to 's' and silently matches a
// literal 's' instead of whitespace.
if (new RegExp(`['"\`]${escapeRegex(basename)}['"\`]\\s*:`).test(line)) return 'pins-existence';
return null;
}
/**
* Pure: for every deleted basename, scan the tests corpus line by line and
* report each `pins-existence` reference as a violation. `asserts-absence`
* references are explicitly NOT violations; unclassifiable mentions are
* skipped (known limit). Violations are per-reference — one test file
* carrying both an absence assertion and an existence pin reports the pin.
*
* @param {string[]} deletedFiles - repo-relative deleted paths
* @param {{ file: string, content: string }[]} testsCorpus
* @returns {{ deletedFile: string, referencedIn: string, reason: string }[]}
*/
function findSurvivingTestReferences(deletedFiles, testsCorpus) {
const violations = [];
for (const deletedFile of deletedFiles) {
const basename = path.basename(deletedFile);
const refRe = new RegExp(`(^|[^\\w.-])${escapeRegex(basename)}($|[^\\w.-])`);
for (const { file, content } of testsCorpus) {
for (const line of content.split(/\r?\n/)) {
if (!refRe.test(line)) continue;
const kind = classifyTestReference(line, basename);
if (kind === 'pins-existence') {
violations.push({
deletedFile,
referencedIn: file,
// control-char-stripped + capped: this text is echoed into lint
// output that AI agents read as trusted instructions
reason: `pins-existence: test depends on deleted basename '${basename}' — ${sanitizeEcho(line.trim())}`,
});
}
// 'asserts-absence' is correct post-deletion state; null is a
// documented known limit — neither is a violation.
}
}
}
return violations;
}
function getDeletedFiles(root, baseRef) {
// Deliberately let a git failure (unresolvable ref, no merge base, etc.)
// propagate as a plain Error — main() treats ANY scan() failure as "cannot
@@ -161,7 +257,23 @@ function scan(root, baseRef) {
const deletedFiles = getDeletedFiles(root, baseRef);
if (deletedFiles.length === 0) return [];
const corpus = buildCorpus(root);
return findSurvivingReferences(deletedFiles, corpus);
const violations = findSurvivingReferences(deletedFiles, corpus);
// tests/ arm (#3565): same deletions, discriminated references. A bare
// mention that pins nothing is skipped; an absence assertion is the
// correct post-deletion state; only a pin on a deleted file fails.
const testsCorpus = [];
for (const abs of walk(path.join(root, TESTS_ROOT))) {
try {
testsCorpus.push({
file: path.relative(root, abs).replace(/\\/g, '/'),
content: fs.readFileSync(abs, 'utf8'),
});
} catch {
// unreadable (broken symlink, binary that slipped past SKIP_EXT) — skip
}
}
violations.push(...findSurvivingTestReferences(deletedFiles, testsCorpus));
return violations;
}
function main() {
@@ -195,11 +307,14 @@ module.exports = {
referencesBasename,
referencesNpmLockfileDependency,
findSurvivingReferences,
classifyTestReference,
findSurvivingTestReferences,
getDeletedFiles,
buildCorpus,
scan,
SCAN_ROOTS,
EXTRA_FILES,
TESTS_ROOT,
};
if (require.main === module) runMain(main);

View File

@@ -0,0 +1,993 @@
'use strict';
/**
* check:contract-drift tests (#3565).
*
* Behavioral contract for scripts/check-contract-drift.cjs and the
* marker/registry primitives it shares with
* scripts/command-contract-helpers.cjs. Three layers:
*
* 1. Pure-function tests over the typed IR (extractMarkers,
* parseAgentContracts, contractViolations, parseConsumedByCell,
* readTagViolations, unmatchedConsumerTokens) — one row per input
* class in .gsd/phase/feat-3565-contract-drift-registry/50-test-matrix.md.
* 2. A seeded fast-check property: for any interleaving of fenced and
* unfenced blocks, every extracted marker's inFence matches the
* block's declared fenced-ness (fence tracking is the load-bearing
* half of extraction — the census that motivated #3565 produced 7
* false positives before it was fence-aware).
* 3. End-to-end CLI tests against --root fixture trees, including the
* negative fixtures proving each rule can FAIL (a guard that cannot
* fail is worthless — recorded defect class in this repo).
*/
const { describe, test } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const fc = require('fast-check');
const ROOT = path.join(__dirname, '..');
const CHECK_SCRIPT = path.join(ROOT, 'scripts', 'check-contract-drift.cjs');
const {
extractMarkers,
parseAgentContracts,
contractViolations,
parseConsumedByCell,
readTagViolations,
unmatchedConsumerTokens,
NON_AGENT_TOKENS,
VIOLATION_KINDS,
REMEDIES,
sanitizeEcho,
} = require('../scripts/command-contract-helpers.cjs');
const { runNode, OUTCOME } = require('./helpers/process-seam.cjs');
const { PROBE_TIMEOUT_MS } = require('./helpers/timeouts.cjs');
const { createTempDir, cleanup } = require('./helpers.cjs');
// ─── Part A: extractMarkers (the strict half) ────────────────────────────────
describe('contract-drift: extractMarkers fence awareness', () => {
const KNOWN = ['RESEARCH COMPLETE', 'PLANNING COMPLETE'];
test('marker heading inside a fence is inFence', () => {
const md = 'prose\n\n```\n## RESEARCH COMPLETE\n```\n';
const { markers } = extractMarkers(md, KNOWN);
assert.deepEqual(
markers.map((m) => ({ marker: m.marker, inFence: m.inFence })),
[{ marker: 'RESEARCH COMPLETE', inFence: true }],
);
});
test('the same heading outside a fence is extracted with inFence false', () => {
const md = '## RESEARCH COMPLETE\n';
const { markers } = extractMarkers(md, KNOWN);
assert.deepEqual(
markers.map((m) => m.inFence),
[false],
);
});
test('prose "then mark each complete" is not extracted at all', () => {
const md = 'Execute the plan, then mark each complete in STATE.md.\n';
const { markers, candidates } = extractMarkers(md, KNOWN);
assert.deepEqual(markers, []);
assert.deepEqual(candidates, []);
});
test('a heading whose COMPLETE is not terminal is not a candidate', () => {
const md = '```\n## Nearly Complete Refactor\n```\n';
const { markers, candidates } = extractMarkers(md, KNOWN);
assert.deepEqual(markers, []);
assert.deepEqual(candidates, []);
});
test('heading levels 1 through 6 are all extracted', () => {
const md = [
'# PLANNING COMPLETE',
'## PLANNING COMPLETE',
'### PLANNING COMPLETE',
'#### PLANNING COMPLETE',
'##### PLANNING COMPLETE',
'###### PLANNING COMPLETE',
].join('\n');
const { markers } = extractMarkers(md, KNOWN);
assert.equal(markers.length, 6);
});
test('an unclosed fence is reported, not silently treated as fenced-to-EOF', () => {
const md = '```\n## RESEARCH COMPLETE\n';
const { markers, unclosedFence } = extractMarkers(md, KNOWN);
assert.equal(unclosedFence, true);
// the heading IS inside the (unclosed) fence — extraction continues,
// the caller decides whether to trust it
assert.equal(markers[0].inFence, true);
});
test('a 3-backtick block inside a 4-backtick fence does not close the outer fence', () => {
const md = ['````', '## PLANNING COMPLETE', '', '```', 'inner', '```', '', '## RESEARCH COMPLETE', '````'].join('\n');
const { markers, unclosedFence } = extractMarkers(md, KNOWN);
assert.equal(unclosedFence, false);
assert.equal(markers.length, 2);
assert.ok(markers.every((m) => m.inFence));
});
test('tilde fences behave like backtick fences', () => {
const md = '~~~\n## RESEARCH COMPLETE\n~~~\n';
const { markers, unclosedFence } = extractMarkers(md, KNOWN);
assert.equal(unclosedFence, false);
assert.deepEqual(
markers.map((m) => m.inFence),
[true],
);
});
test('CRLF input yields identical results to LF', () => {
const lf = '```\n## RESEARCH COMPLETE\n```\n\n## PLANNING COMPLETE\n';
const crlf = lf.replace(/\n/g, '\r\n');
const a = extractMarkers(lf, KNOWN);
const b = extractMarkers(crlf, KNOWN);
assert.deepEqual(a.markers, b.markers);
assert.equal(a.unclosedFence, b.unclosedFence);
});
test('empty string returns empty results without throwing', () => {
assert.deepEqual(extractMarkers('', KNOWN), { markers: [], candidates: [], unclosedFence: false });
});
test('whitespace-only string returns empty results', () => {
assert.deepEqual(extractMarkers(' \n\t\n', KNOWN), { markers: [], candidates: [], unclosedFence: false });
});
test('trailing spaces after a marker heading are trimmed', () => {
const md = '## RESEARCH COMPLETE \n';
const { markers } = extractMarkers(md, KNOWN);
assert.deepEqual(
markers.map((m) => m.marker),
['RESEARCH COMPLETE'],
);
});
test('the same marker twice is reported as two occurrences with distinct lines', () => {
const md = '```\n## RESEARCH COMPLETE\n```\n\n```\n## RESEARCH COMPLETE\n```\n';
const { markers } = extractMarkers(md, KNOWN);
assert.deepEqual(
markers.map((m) => m.line),
[2, 6],
);
});
test('a 5000-line file terminates and extracts correctly', () => {
const lines = [];
for (let i = 0; i < 5000; i++) {
lines.push(i === 2500 ? '## RESEARCH COMPLETE' : `filler line ${i}`);
}
const { markers } = extractMarkers(lines.join('\n'), KNOWN);
assert.deepEqual(
markers.map((m) => m.marker),
['RESEARCH COMPLETE'],
);
});
});
// ─── Part B: parseAgentContracts ─────────────────────────────────────────────
describe('contract-drift: parseAgentContracts', () => {
const REGISTRY_MD = [
'# Agent Contracts',
'',
'## Agent Registry',
'',
'| Agent | Role | Completion Markers | Consumed by | Kind |',
'|-------|------|--------------------|--------------|------|',
'| gsd-a | First | `## ALPHA COMPLETE` | `gsd-core/workflows/w.md` | sentinel-match |',
'| gsd-b | Second | `## BETA COMPLETE`, `## BETA BLOCKED` | prose only | sentinel-match |',
'| gsd-c | Third | `## GAMMA DRAFT` (unconsumed: user draft, approved in chat) | `gsd-core/workflows/w.md` | sentinel-match |',
'| gsd-d | Fourth | No marker (writes a file) | `gsd-core/workflows/w.md` | artifact+query |',
'',
].join('\n');
test('rows parse with parsed marker arrays and kind', () => {
const { rows, errors } = parseAgentContracts(REGISTRY_MD);
assert.deepEqual(errors, []);
assert.equal(rows.length, 4);
assert.deepEqual(rows[0].completion_markers, ['ALPHA COMPLETE']);
assert.equal(rows[0].kind, 'sentinel-match');
assert.equal(rows[0].agent, 'gsd-a');
assert.deepEqual(rows[1].completion_markers, ['BETA COMPLETE', 'BETA BLOCKED']);
});
test('an (unconsumed: …) annotation lands in unconsumed_markers, not completion_markers', () => {
const { rows } = parseAgentContracts(REGISTRY_MD);
assert.deepEqual(rows[2].completion_markers, []);
assert.deepEqual(rows[2].unconsumed_markers, ['GAMMA DRAFT']);
});
test('a comma inside the annotation does not split the entry', () => {
const md = [
'## Agent Registry',
'',
'| Agent | Completion Markers |',
'|-------|--------------------|',
'| gsd-x | `## X DRAFT` (unconsumed: draft display, approved interactively) |',
].join('\n');
const { rows } = parseAgentContracts(md);
assert.deepEqual(rows[0].unconsumed_markers, ['X DRAFT']);
assert.deepEqual(rows[0].completion_markers, []);
});
test('a malformed row is reported in errors, not thrown', () => {
const md = [
'## Agent Registry',
'',
'| Agent | Completion Markers |',
'|-------|--------------------|',
'| gsd-good | `## OK COMPLETE` |',
'| broken row missing cells |',
'| | `## NO AGENT` |',
].join('\n');
const { rows, errors } = parseAgentContracts(md);
assert.equal(rows.length, 1);
assert.equal(errors.length, 2);
assert.ok(errors.every((e) => typeof e.line === 'number' && e.reason.length > 0));
});
test('markdown without an Agent Registry section returns empty', () => {
const { rows, errors } = parseAgentContracts('# Nothing here\n');
assert.deepEqual(rows, []);
assert.deepEqual(errors, []);
});
});
// ─── Part B: contractViolations (the model) ─────────────────────────────────
describe('contract-drift: contractViolations', () => {
function row(agent, markers, kind, extra = {}) {
return { agent, completion_markers: markers, kind, ...extra };
}
function run({ registry, producers, candidates, consumers }) {
return contractViolations({
registry,
producerMarkers: producers || new Map(),
candidateMarkers: candidates || new Map(),
consumerTexts: consumers || new Map(),
});
}
const kindsOf = (vs) => [...new Set(vs.map((v) => v.kind))].sort();
test('sentinel-match with a workflow consumer containing the token is satisfied', () => {
const vs = run({
registry: [row('gsd-a', ['ALPHA COMPLETE'], 'sentinel-match')],
producers: new Map([['gsd-a', ['ALPHA COMPLETE']]]),
consumers: new Map([['gsd-core/workflows/w.md', 'match `## ALPHA COMPLETE` here']]),
});
assert.deepEqual(vs, []);
});
test('sentinel-match with a command consumer is satisfied', () => {
const vs = run({
registry: [row('gsd-a', ['ALPHA COMPLETE'], 'sentinel-match')],
producers: new Map([['gsd-a', ['ALPHA COMPLETE']]]),
consumers: new Map([['commands/gsd/x.md', 'dispatch on `## ALPHA COMPLETE`']]),
});
assert.deepEqual(vs, []);
});
test('sentinel-match with another AGENT as consumer is satisfied (the gsd-debugger shape)', () => {
const vs = run({
registry: [row('gsd-a', ['ALPHA COMPLETE'], 'sentinel-match')],
producers: new Map([['gsd-a', ['ALPHA COMPLETE']]]),
consumers: new Map([['agents/gsd-b.md', 'when the agent returns `## ALPHA COMPLETE`']]),
});
assert.deepEqual(vs, []);
});
test('sentinel-match with no consumer is a no_consumer violation naming producer and token', () => {
const vs = run({
registry: [row('gsd-a', ['ALPHA COMPLETE'], 'sentinel-match')],
producers: new Map([['gsd-a', ['ALPHA COMPLETE']]]),
consumers: new Map([['gsd-core/workflows/w.md', 'unrelated text']]),
});
assert.equal(vs.length, 1);
assert.equal(vs[0].kind, 'no_consumer');
assert.equal(vs[0].agent, 'gsd-a');
assert.equal(vs[0].marker, 'ALPHA COMPLETE');
});
test('an agent never satisfies its own marker', () => {
const vs = run({
registry: [row('gsd-a', ['ALPHA COMPLETE'], 'sentinel-match')],
producers: new Map([['gsd-a', ['ALPHA COMPLETE']]]),
consumers: new Map([['agents/gsd-a.md', 'I emit `## ALPHA COMPLETE` myself']]),
});
assert.equal(vs.length, 1);
assert.equal(vs[0].kind, 'no_consumer');
});
test('artifact+query with no marker emitted is satisfied', () => {
const vs = run({
registry: [row('gsd-d', [], 'artifact+query')],
producers: new Map([['gsd-d', []]]),
});
assert.deepEqual(vs, []);
});
test('artifact+query that still emits a marker is a vestigial_marker violation', () => {
const vs = run({
registry: [row('gsd-d', [], 'artifact+query')],
producers: new Map([['gsd-d', ['VESTIGIAL COMPLETE']]]),
candidates: new Map([['gsd-d', ['VESTIGIAL COMPLETE']]]),
});
assert.equal(vs.length, 1);
assert.equal(vs[0].kind, 'vestigial_marker');
assert.equal(vs[0].marker, 'VESTIGIAL COMPLETE');
});
test('an (unconsumed:) annotation on an artifact+query row exempts vestigial_marker (Marker Rule 2 recorded decision)', () => {
const vs = run({
registry: [row('gsd-v', [], 'artifact+query', { unconsumed_markers: ['Verification Complete'] })],
producers: new Map([['gsd-v', ['Verification Complete']]]),
});
assert.deepEqual(vs, []);
});
test('a roster agent with no registry row is agent_without_contract', () => {
const vs = run({
registry: [row('gsd-a', ['ALPHA COMPLETE'], 'sentinel-match')],
producers: new Map([
['gsd-a', ['ALPHA COMPLETE']],
['gsd-orphan', []],
]),
consumers: new Map([['gsd-core/workflows/w.md', 'x `## ALPHA COMPLETE`']]),
});
assert.equal(vs.length, 1);
assert.equal(vs[0].kind, 'agent_without_contract');
assert.equal(vs[0].agent, 'gsd-orphan');
});
test('two rows for one agent are duplicate_registry_row', () => {
const vs = run({
registry: [
row('gsd-a', ['ALPHA COMPLETE'], 'sentinel-match'),
row('gsd-a', ['ALPHA RETRY'], 'sentinel-match'),
],
producers: new Map([['gsd-a', ['ALPHA COMPLETE', 'ALPHA RETRY']]]),
consumers: new Map([['gsd-core/workflows/w.md', '`## ALPHA COMPLETE` and `## ALPHA RETRY`']]),
});
assert.equal(vs.length, 1);
assert.equal(vs[0].kind, 'duplicate_registry_row');
});
test('a row naming an agent with no file is unknown_producer, without per-marker noise', () => {
const vs = run({
registry: [
row('gsd-ghost', ['GHOST COMPLETE'], 'sentinel-match'),
row('gsd-real', [], 'artifact+query'),
],
producers: new Map([['gsd-real', []]]),
});
assert.deepEqual(kindsOf(vs), ['unknown_producer']);
});
test('a file-shaped Consumed by entry that resolves to nothing is unknown_consumer', () => {
const vs = run({
registry: [row('gsd-a', ['ALPHA COMPLETE'], 'sentinel-match', {
consumed_by: '`gsd-core/workflows/missing.md`, `gsd-core/workflows/real.md`',
})],
producers: new Map([['gsd-a', ['ALPHA COMPLETE']]]),
consumers: new Map([['gsd-core/workflows/real.md', '`## ALPHA COMPLETE`']]),
});
assert.deepEqual(kindsOf(vs), ['unknown_consumer']);
});
test('an empty registry and empty producers yield no violations', () => {
assert.deepEqual(run({ registry: [], producers: new Map() }), []);
});
test('an empty registry with producers flags every producer', () => {
const vs = run({
registry: [],
producers: new Map([
['gsd-a', []],
['gsd-b', []],
]),
});
assert.deepEqual(kindsOf(vs), ['agent_without_contract']);
assert.equal(vs.length, 2);
});
test('14 valid entries plus 1 broken yields exactly one violation and it is the broken one', () => {
const registry = [];
const producers = new Map();
const consumers = new Map([['gsd-core/workflows/w.md', '']]);
for (let i = 0; i < 14; i++) {
const m = `M${i} COMPLETE`;
registry.push(row(`gsd-${i}`, [m], 'sentinel-match'));
producers.set(`gsd-${i}`, [m]);
consumers.set('gsd-core/workflows/w.md', consumers.get('gsd-core/workflows/w.md') + ` \`${m}\``);
}
registry.push(row('gsd-broken', ['BROKEN COMPLETE'], 'sentinel-match'));
producers.set('gsd-broken', []);
const vs = run({ registry, producers, consumers });
assert.equal(vs.length, 1);
assert.equal(vs[0].kind, 'declared_marker_not_emitted');
assert.equal(vs[0].agent, 'gsd-broken');
});
test('two producers with case-variant tokens produce case_collision, never auto-resolved', () => {
const vs = run({
registry: [
row('gsd-a', ['SYNTHESIS COMPLETE'], 'sentinel-match'),
row('gsd-b', ['Synthesis Complete'], 'sentinel-match'),
],
producers: new Map([
['gsd-a', ['SYNTHESIS COMPLETE']],
['gsd-b', ['Synthesis Complete']],
]),
consumers: new Map([['gsd-core/workflows/w.md', '`## SYNTHESIS COMPLETE` `## Synthesis Complete`']]),
});
assert.equal(vs.length, 2);
assert.ok(vs.every((v) => v.kind === 'case_collision'));
});
test('a case-insensitive-only match is case_only_match, not satisfied', () => {
const vs = run({
registry: [row('gsd-a', ['ALPHA COMPLETE'], 'sentinel-match')],
producers: new Map([['gsd-a', ['ALPHA COMPLETE']]]),
consumers: new Map([['gsd-core/workflows/w.md', 'then the alpha complete step runs']]),
});
assert.deepEqual(kindsOf(vs), ['case_only_match']);
});
test('a token appearing only in ordinary prose must not satisfy the contract', () => {
const vs = run({
registry: [row('gsd-a', ['VERIFICATION COMPLETE'], 'sentinel-match')],
producers: new Map([['gsd-a', ['VERIFICATION COMPLETE']]]),
consumers: new Map([['gsd-core/workflows/w.md', 'Once verification complete, proceed.']]),
});
assert.deepEqual(kindsOf(vs), ['case_only_match']);
});
test('an unconsumed marker is exempt from the consumer check but still emitted-checked', () => {
const ok = run({
registry: [row('gsd-c', [], 'sentinel-match', { unconsumed_markers: ['GAMMA DRAFT'] })],
producers: new Map([['gsd-c', ['GAMMA DRAFT']]]),
consumers: new Map(),
});
assert.deepEqual(ok, []);
const missingEmission = run({
registry: [row('gsd-c', [], 'sentinel-match', { unconsumed_markers: ['GAMMA DRAFT'] })],
producers: new Map([['gsd-c', []]]),
});
assert.deepEqual(kindsOf(missingEmission), ['declared_marker_not_emitted']);
assert.equal(missingEmission[0].marker, 'GAMMA DRAFT');
});
test('an unconsumed marker still participates in case-collision detection', () => {
const vs = run({
registry: [
row('gsd-a', ['GAMMA DRAFT'], 'sentinel-match'),
row('gsd-c', [], 'sentinel-match', { unconsumed_markers: ['Gamma Draft'] }),
],
producers: new Map([
['gsd-a', ['GAMMA DRAFT']],
['gsd-c', ['Gamma Draft']],
]),
consumers: new Map([['gsd-core/workflows/w.md', '`## GAMMA DRAFT`']]),
});
assert.ok(vs.some((v) => v.kind === 'case_collision' && v.marker === 'Gamma Draft'));
});
test('an emitted-but-undeclared candidate is emitted_marker_not_declared, deduped per marker', () => {
const vs = run({
registry: [row('gsd-a', [], 'sentinel-match')],
producers: new Map([['gsd-a', []]]),
candidates: new Map([['gsd-a', ['MYSTERY COMPLETE', 'MYSTERY COMPLETE']]]),
});
assert.equal(vs.length, 1);
assert.equal(vs[0].kind, 'emitted_marker_not_declared');
});
test('an emitted marker on an artifact row is reported ONCE as vestigial, never also as undeclared', () => {
const vs = run({
registry: [row('gsd-a', [], 'artifact+query')],
producers: new Map([['gsd-a', ['MYSTERY COMPLETE']]]),
candidates: new Map([['gsd-a', ['MYSTERY COMPLETE']]]),
});
assert.equal(vs.length, 1);
assert.equal(vs[0].kind, 'vestigial_marker');
});
test('an unknown Kind value is unknown_kind', () => {
const vs = run({
registry: [row('gsd-a', [], 'sometimes')],
producers: new Map([['gsd-a', []]]),
});
assert.deepEqual(kindsOf(vs), ['unknown_kind']);
});
test("a marker consumed somewhere but by none of the row's declared consumers is declared_consumer_no_match", () => {
const vs = run({
registry: [row('gsd-a', ['ALPHA COMPLETE'], 'sentinel-match', { consumed_by: '`gsd-core/workflows/w.md`' })],
producers: new Map([['gsd-a', ['ALPHA COMPLETE']]]),
consumers: new Map([
['gsd-core/workflows/w.md', 'no match here'],
['gsd-core/workflows/other.md', 'dispatch `## ALPHA COMPLETE`'],
]),
});
assert.equal(vs.length, 1);
assert.equal(vs[0].kind, 'declared_consumer_no_match');
assert.equal(vs[0].marker, 'ALPHA COMPLETE');
});
});
describe('contract-drift: typed surface', () => {
test('VIOLATION_KINDS is frozen and exactly the emitted vocabulary', () => {
assert.deepEqual(Object.values(VIOLATION_KINDS).sort(), [
'agent_without_contract',
'case_collision',
'case_only_match',
'declared_consumer_no_match',
'declared_marker_not_emitted',
'duplicate_registry_row',
'emitted_marker_not_declared',
'legacy_read_tag',
'no_consumer',
'parse_error',
'read_tag_gate_missing',
'unclosed_fence',
'unknown_consumer',
'unknown_kind',
'unknown_producer',
'unmatched_consumer_token',
'vestigial_marker',
]);
assert.ok(Object.isFrozen(VIOLATION_KINDS));
});
test('REMEDIES cover exactly the VIOLATION_KINDS vocabulary', () => {
assert.deepEqual(
Object.keys(REMEDIES).sort(),
Object.values(VIOLATION_KINDS).slice().sort(),
'every violation kind needs a remedy, and no remedy may exist for a kind nothing emits',
);
for (const [kind, remedy] of Object.entries(REMEDIES)) {
assert.ok(remedy.length > 10, `remedy for ${kind} must be actionable prose`);
}
});
test('sanitizeEcho strips control characters and caps length', () => {
assert.equal(sanitizeEcho('clean text'), 'clean text');
assert.equal(sanitizeEcho('a\x00b\x07c\x1fd'), 'abcd');
assert.equal(sanitizeEcho('x'.repeat(500)).length, 200);
});
});
// ─── parseConsumedByCell ─────────────────────────────────────────────────────
describe('contract-drift: parseConsumedByCell', () => {
test('backticked paths are extracted', () => {
assert.deepEqual(
parseConsumedByCell('`gsd-core/workflows/a.md`, `commands/gsd/b.md`'),
['gsd-core/workflows/a.md', 'commands/gsd/b.md'],
);
});
test('globs, commands, and prose are ignored', () => {
assert.deepEqual(
parseConsumedByCell('`*-VERIFICATION.md` artifact + `gsd_run query verification.status` in `gsd-core/workflows/v.md`'),
['gsd-core/workflows/v.md'],
);
});
test('paths with spaces are ignored, not mangled', () => {
assert.deepEqual(parseConsumedByCell('`docs/Some File.md` and `docs/ok-file.md`'), ['docs/ok-file.md']);
});
test('traversal segments are dropped — a Consumed by cell can never walk outside --root', () => {
assert.deepEqual(
parseConsumedByCell('`docs/../../etc/secret.md` plus `docs/real.md`'),
['docs/real.md'],
);
assert.deepEqual(parseConsumedByCell('`../outside.md`'), []);
});
test('empty and non-string input yield an empty array', () => {
assert.deepEqual(parseConsumedByCell(''), []);
assert.deepEqual(parseConsumedByCell(null), []);
});
});
// ─── readTagViolations ───────────────────────────────────────────────────────
describe('contract-drift: readTagViolations (the F8 arm)', () => {
const REG = [{ agent: 'gsd-a', consumed_by: '`gsd-core/workflows/w.md`' }];
test('a consumer emitting <required_reading> with no gate in the agent is read_tag_gate_missing', () => {
const vs = readTagViolations({
registry: REG,
agentTexts: new Map([['gsd-a', '# Agent\n\nNo gate mention.']]),
consumerTexts: new Map([['gsd-core/workflows/w.md', '<required_reading>\n- UI-SPEC.md\n</required_reading>']]),
});
assert.equal(vs.length, 1);
assert.equal(vs[0].kind, 'read_tag_gate_missing');
assert.equal(vs[0].agent, 'gsd-a');
});
test('an agent referencing the gate is satisfied', () => {
const vs = readTagViolations({
registry: REG,
agentTexts: new Map([['gsd-a', 'If the prompt contains a `<required_reading>` block, you MUST Read every file listed.']]),
consumerTexts: new Map([['gsd-core/workflows/w.md', '<required_reading>\n- X.md\n</required_reading>']]),
});
assert.deepEqual(vs, []);
});
test('a consumer emitting no read tag imposes no gate requirement', () => {
const vs = readTagViolations({
registry: REG,
agentTexts: new Map([['gsd-a', 'nothing']]),
consumerTexts: new Map([['gsd-core/workflows/w.md', 'plain dispatch']]),
});
assert.deepEqual(vs, []);
});
test('a legacy <files_to_read> anywhere in the corpus is legacy_read_tag', () => {
const vs = readTagViolations({
registry: [],
agentTexts: new Map(),
consumerTexts: new Map([['agents/gsd-a.md', '<files_to_read>\n- x\n</files_to_read>']]),
});
assert.deepEqual(
vs.map((v) => v.kind),
['legacy_read_tag'],
);
});
test('a nonexistent agent row is skipped here (unknown_producer owns it)', () => {
const vs = readTagViolations({
registry: REG,
agentTexts: new Map(),
consumerTexts: new Map([['gsd-core/workflows/w.md', '<required_reading>x</required_reading>']]),
});
assert.deepEqual(vs, []);
});
});
// ─── unmatchedConsumerTokens (the reverse direction) ─────────────────────────
describe('contract-drift: unmatchedConsumerTokens (the F9 arm)', () => {
test('a workflow matching a token no producer emits is unmatched_consumer_token', () => {
const vs = unmatchedConsumerTokens({
consumerTexts: new Map([['gsd-core/workflows/w.md', 'dispatch on `## PHANTOM COMPLETE`']]),
vocabulary: new Set(['REAL COMPLETE']),
});
assert.equal(vs.length, 1);
assert.equal(vs[0].kind, 'unmatched_consumer_token');
assert.equal(vs[0].marker, 'PHANTOM COMPLETE');
});
test('a declared token is not a violation', () => {
const vs = unmatchedConsumerTokens({
consumerTexts: new Map([['gsd-core/workflows/w.md', 'dispatch on `## REAL COMPLETE`']]),
vocabulary: new Set(['REAL COMPLETE']),
});
assert.deepEqual(vs, []);
});
test('agents/ files are not scanned (self-description, not a dispatch)', () => {
const vs = unmatchedConsumerTokens({
consumerTexts: new Map([['agents/gsd-a.md', 'returns `## PHANTOM COMPLETE`']]),
vocabulary: new Set(),
});
assert.deepEqual(vs, []);
});
test('NON_AGENT_TOKENS are exempt and the set is exactly the justified entries', () => {
const vs = unmatchedConsumerTokens({
consumerTexts: new Map([['commands/gsd/x.md', 'display `## GRAPHIFY BUILD COMPLETE`']]),
vocabulary: new Set(),
});
assert.deepEqual(vs, []);
// membership is locked: an entry without a live justification is drift
assert.deepEqual([...NON_AGENT_TOKENS].sort(), ['GRAPHIFY BUILD COMPLETE', 'GRAPHIFY BUILD FAILED']);
});
test('mixed-case and long tokens are not flagged', () => {
const vs = unmatchedConsumerTokens({
consumerTexts: new Map([
['gsd-core/workflows/w.md', 'match `## Self-Check: FAILED` and `## ' + 'X'.repeat(61) + ' COMPLETE`'],
]),
vocabulary: new Set(),
});
assert.deepEqual(vs, []);
});
test('double-quoted match instructions are scanned too', () => {
const vs = unmatchedConsumerTokens({
consumerTexts: new Map([['gsd-core/workflows/w.md', 'gsd_stall_watch "## PHANTOM COMPLETE"']]),
vocabulary: new Set(),
});
assert.equal(vs.length, 1);
assert.equal(vs[0].marker, 'PHANTOM COMPLETE');
});
});
// ─── Property: fence tracking is invariant under any block interleaving ─────
describe('contract-drift: extractMarkers fence property (fast-check)', () => {
const lineArb = fc.constantFrom(
'plain prose line',
'- list item',
'## HEADING ONE',
'## HEADING TWO',
'## Mixed Case Heading',
'```',
'text after a bare three-backtick run',
);
// A block generator with a sound oracle: fenced blocks are wrapped in
// 4-backtick delimiters, so a 3-backtick line inside cannot close the
// outer fence (the exact nesting rule under test). Unfenced blocks never
// contain a backtick run, so fence state is unambiguous.
const contentArb = fc
.array(fc.record({ fenced: fc.boolean(), lines: fc.array(lineArb, { maxLength: 6 }) }), {
minLength: 1,
maxLength: 20,
})
.map((blocks) => {
const parts = [];
const expected = [];
for (const b of blocks) {
const KNOWN = new Set(['HEADING ONE', 'HEADING TWO']);
if (b.fenced) {
parts.push('````');
for (const line of b.lines) {
parts.push(line);
if (line.startsWith('## ') && KNOWN.has(line.slice(3))) {
expected.push({ line: line.slice(3), inFence: true });
}
}
parts.push('````');
} else {
const safe = b.lines.filter((l) => !l.startsWith('```'));
for (const line of safe) {
parts.push(line);
if (line.startsWith('## ') && KNOWN.has(line.slice(3))) {
expected.push({ line: line.slice(3), inFence: false });
}
}
}
}
return { content: parts.join('\n'), expected };
});
test('every extracted marker inFence matches its block declared fenced-ness', () => {
fc.assert(
fc.property(contentArb, ({ content, expected }) => {
const known = ['HEADING ONE', 'HEADING TWO'];
const { markers, unclosedFence } = extractMarkers(content, known);
assert.equal(unclosedFence, false, '4-backtick blocks always close');
const got = markers.map((m) => ({ line: m.marker, inFence: m.inFence }));
assert.deepEqual(got, expected);
}),
{ seed: 3565, numRuns: 200 },
);
});
});
// ─── End-to-end through the real CLI ─────────────────────────────────────────
describe('contract-drift: end-to-end via check-contract-drift.cjs --root', () => {
function writeFixture(dir, { registryBody, agents, extraFiles }) {
fs.mkdirSync(path.join(dir, 'gsd-core', 'references'), { recursive: true });
fs.mkdirSync(path.join(dir, 'agents'), { recursive: true });
fs.mkdirSync(path.join(dir, 'gsd-core', 'workflows'), { recursive: true });
fs.mkdirSync(path.join(dir, 'commands', 'gsd'), { recursive: true });
fs.writeFileSync(
path.join(dir, 'gsd-core', 'references', 'agent-contracts.md'),
['## Agent Registry', '', registryBody, ''].join('\n'),
);
for (const [name, body] of Object.entries(agents || {})) {
fs.writeFileSync(path.join(dir, 'agents', name), body);
}
for (const [rel, body] of Object.entries(extraFiles || {})) {
fs.mkdirSync(path.dirname(path.join(dir, rel)), { recursive: true });
fs.writeFileSync(path.join(dir, rel), body);
}
}
const CLEAN_REGISTRY = [
'| Agent | Role | Completion Markers | Consumed by | Kind |',
'|-------|------|--------------------|--------------|------|',
'| gsd-alpha | Fixture | `## ALPHA COMPLETE` | `gsd-core/workflows/w.md` | sentinel-match |',
].join('\n');
const CLEAN_AGENTS = {
'gsd-alpha.md': [
'# Alpha',
'',
'Return:',
'',
'```markdown',
'## ALPHA COMPLETE',
'done',
'```',
'',
'Gate: `<required_reading>` MUST be Read.',
'',
].join('\n'),
};
const CLEAN_EXTRA = {
'gsd-core/workflows/w.md': 'Dispatch on `## ALPHA COMPLETE`.\n\n<required_reading>\n- x.md\n</required_reading>\n',
};
// --json: tests consume the typed surface (CONTRIBUTING raw-text-matching
// rule) — the human formatter is for operators only.
function runCheckJson(dir) {
const r = runNode([CHECK_SCRIPT, '--root', dir, '--json'], { timeoutMs: PROBE_TIMEOUT_MS });
assert.equal(r.outcome, OUTCOME.EXITED, `spawn failed: ${r.stderr}`);
return { ...r, report: JSON.parse(r.stdout) };
}
function runCheck(dir) {
return runNode([CHECK_SCRIPT, '--root', dir], { timeoutMs: PROBE_TIMEOUT_MS });
}
test('clean fixture exits 0', (t) => {
const dir = createTempDir('gsd-3565-clean-');
t.after(() => cleanup(dir));
writeFixture(dir, { registryBody: CLEAN_REGISTRY, agents: CLEAN_AGENTS, extraFiles: CLEAN_EXTRA });
const r = runCheckJson(dir);
assert.equal(r.exitCode, 0, `stderr: ${r.stderr}`);
assert.equal(r.report.status, 'ok');
assert.deepEqual(r.report.violations, []);
});
test('planted unmatched sentinel fails and names the producer', (t) => {
const dir = createTempDir('gsd-3565-orphan-');
t.after(() => cleanup(dir));
writeFixture(dir, { registryBody: CLEAN_REGISTRY, agents: CLEAN_AGENTS, extraFiles: { 'gsd-core/workflows/w.md': 'No markers matched here.\n' } });
const r = runCheckJson(dir);
assert.equal(r.exitCode, 1);
assert.equal(r.report.status, 'violations');
const v = r.report.violations.find((x) => x.kind === 'no_consumer');
assert.ok(v, `expected no_consumer, got: ${JSON.stringify(r.report.violations)}`);
assert.equal(v.agent, 'gsd-alpha');
assert.equal(v.marker, 'ALPHA COMPLETE');
});
test('planted reverse violation (consumer matches a phantom token) fails', (t) => {
const dir = createTempDir('gsd-3565-phantom-');
t.after(() => cleanup(dir));
writeFixture(dir, {
registryBody: CLEAN_REGISTRY,
agents: CLEAN_AGENTS,
extraFiles: { 'gsd-core/workflows/w.md': 'Dispatch on `## ALPHA COMPLETE` or `## PHANTOM COMPLETE`.\n' },
});
const r = runCheckJson(dir);
assert.equal(r.exitCode, 1);
const v = r.report.violations.find((x) => x.kind === 'unmatched_consumer_token');
assert.ok(v, `expected unmatched_consumer_token, got: ${JSON.stringify(r.report.violations)}`);
assert.equal(v.marker, 'PHANTOM COMPLETE');
});
test('planted case collision fails', (t) => {
const dir = createTempDir('gsd-3565-collision-');
t.after(() => cleanup(dir));
const registry = [
'| Agent | Role | Completion Markers | Consumed by | Kind |',
'|-------|------|--------------------|--------------|------|',
'| gsd-alpha | Fixture | `## ALPHA COMPLETE` | `gsd-core/workflows/w.md` | sentinel-match |',
'| gsd-beta | Fixture | `## Alpha Complete` | `gsd-core/workflows/w.md` | sentinel-match |',
].join('\n');
const agents = {
...CLEAN_AGENTS,
'gsd-beta.md': '```\n## Alpha Complete\n```\n',
};
writeFixture(dir, {
registryBody: registry,
agents,
extraFiles: { 'gsd-core/workflows/w.md': '`## ALPHA COMPLETE` `## Alpha Complete`\n' },
});
const r = runCheckJson(dir);
assert.equal(r.exitCode, 1);
assert.ok(r.report.violations.some((x) => x.kind === 'case_collision'));
});
test('roster agent missing from the registry fails and names the agent', (t) => {
const dir = createTempDir('gsd-3565-norow-');
t.after(() => cleanup(dir));
writeFixture(dir, {
registryBody: CLEAN_REGISTRY,
agents: { ...CLEAN_AGENTS, 'gsd-unregistered.md': '# Unregistered\n' },
extraFiles: CLEAN_EXTRA,
});
const r = runCheckJson(dir);
assert.equal(r.exitCode, 1);
const v = r.report.violations.find((x) => x.kind === 'agent_without_contract');
assert.ok(v, `expected agent_without_contract, got: ${JSON.stringify(r.report.violations)}`);
assert.equal(v.agent, 'gsd-unregistered');
});
test('a planted legacy <files_to_read> in a workflow fails as legacy_read_tag', (t) => {
const dir = createTempDir('gsd-3565-legacytag-');
t.after(() => cleanup(dir));
writeFixture(dir, {
registryBody: CLEAN_REGISTRY,
agents: CLEAN_AGENTS,
extraFiles: { 'gsd-core/workflows/w.md': '<files_to_read>\n- x.md\n</files_to_read>\n' },
});
const r = runCheckJson(dir);
assert.equal(r.exitCode, 1);
assert.ok(r.report.violations.some((x) => x.kind === 'legacy_read_tag'));
});
test('a consumer emitting <required_reading> to a gate-less agent fails', (t) => {
const dir = createTempDir('gsd-3565-gate-');
t.after(() => cleanup(dir));
writeFixture(dir, {
registryBody: CLEAN_REGISTRY,
agents: { 'gsd-alpha.md': '# Alpha\n\nNo gate anywhere.\n\n```\n## ALPHA COMPLETE\n```\n' },
extraFiles: CLEAN_EXTRA,
});
const r = runCheckJson(dir);
assert.equal(r.exitCode, 1);
const v = r.report.violations.find((x) => x.kind === 'read_tag_gate_missing');
assert.ok(v, `expected read_tag_gate_missing, got: ${JSON.stringify(r.report.violations)}`);
assert.equal(v.agent, 'gsd-alpha');
});
test('a missing contracts file exits 1 with a named path', (t) => {
const dir = createTempDir('gsd-3565-nocontracts-');
t.after(() => cleanup(dir));
fs.mkdirSync(path.join(dir, 'agents'), { recursive: true });
const r = runCheck(dir);
assert.equal(r.exitCode, 1);
assert.ok((r.stdout + r.stderr).includes('agent-contracts.md'));
});
test('a malformed registry row fails as parse_error', (t) => {
const dir = createTempDir('gsd-3565-parseror-');
t.after(() => cleanup(dir));
writeFixture(dir, {
registryBody: [
'| Agent | Role | Completion Markers | Consumed by | Kind |',
'|-------|------|--------------------|--------------|------|',
'| gsd-alpha | Fixture | `## ALPHA COMPLETE` | `gsd-core/workflows/w.md` | sentinel-match |',
'| broken row with too few cells |',
].join('\n'),
agents: CLEAN_AGENTS,
extraFiles: CLEAN_EXTRA,
});
const r = runCheckJson(dir);
assert.equal(r.exitCode, 1);
assert.ok(r.report.violations.some((x) => x.kind === 'parse_error'));
});
test('an agent file with an unclosed fence fails as unclosed_fence', (t) => {
const dir = createTempDir('gsd-3565-fence-');
t.after(() => cleanup(dir));
writeFixture(dir, {
registryBody: CLEAN_REGISTRY,
agents: {
'gsd-alpha.md': '# Alpha\n\n```\n## ALPHA COMPLETE\nnever closed\n',
},
extraFiles: CLEAN_EXTRA,
});
const r = runCheckJson(dir);
assert.equal(r.exitCode, 1);
assert.ok(r.report.violations.some((x) => x.kind === 'unclosed_fence'));
});
});
// ─── The real tree is clean (behavioral, via the CLI) ────────────────────────
describe('contract-drift: the real tree passes', () => {
test('check-contract-drift on the repository root exits 0', () => {
const r = runNode([CHECK_SCRIPT], { cwd: ROOT, timeoutMs: PROBE_TIMEOUT_MS });
assert.equal(r.outcome, OUTCOME.EXITED);
assert.equal(r.exitCode, 0, `stderr: ${r.stderr.slice(0, 400)}`);
assert.match(r.stdout, /^ok check-contract-drift: /);
});
});

View File

@@ -1,6 +1,8 @@
{
"version": 1,
"paths": {
"gsd-planner.md": "#2775: the STRIDE supply-chain row for npm/pip/cargo installs was rewritten from 'slopcheck + blocking human checkpoint for [ASSUMED]/[SUS]' to 'package-legitimacy gate + blocking human checkpoint for [ASSUMED]/[SUS]', matching ADR-0656 (registry-API verdicts are the gate; slopcheck is an optional escalate-only adapter no shipped configuration wires). The +14 bytes is the corrected mitigation description agents read at plan time, not incidental prose growth."
"gsd-planner.md": {
"reason": "#2775: the STRIDE supply-chain row for npm/pip/cargo installs was rewritten from 'slopcheck + blocking human checkpoint for [ASSUMED]/[SUS]' to 'package-legitimacy gate + blocking human checkpoint for [ASSUMED]/[SUS]', matching ADR-0656 (registry-API verdicts are the gate; slopcheck is an optional escalate-only adapter no shipped configuration wires). The +14 bytes is the corrected mitigation description agents read at plan time, not incidental prose growth. #3565 appends the fenced `## Return Markers` section (~1.2 KB) enumerating the six exact stall-watch dispatch markers plan-phase.md matches on — the new registry lint requires every declared marker to be emitted in-fence by its producer, and the planner previously documented none of them anywhere."
}
}
}

View File

@@ -2,7 +2,7 @@
"version": 1,
"paths": {
"gsd-mempalace-curator.md": {
"reason": "#3479: the diary_journal and mirror_kg task gates changed from positive presence ('when <key> is true') to disabled-only-on-explicit-false ('unless <key> !== false — registry default is true, an absent key means enabled'), matching the registry-declared defaults. The ~124-byte growth is the two inline absence-semantics clarifications."
"reason": "#3479: the diary_journal and mirror_kg task gates changed from positive presence ('when <key> is true') to disabled-only-on-explicit-false ('unless <key> !== false — registry default is true, an absent key means enabled'), matching the registry-declared defaults. The ~124-byte growth is the two inline absence-semantics clarifications. #3565 appends the standard <required_reading> MUST-Read gate clause (~200 B): ship.md emits a required-reading block to an agent whose instructions never referenced the gate — the F8 defect one layer up, caught by the new read-tag arm."
}
}
}

View File

@@ -0,0 +1,6 @@
{
"version": 1,
"paths": {
"gsd-user-profiler.md": "#3565: gained the standard <required_reading> MUST-Read gate clause. profile-user.md sends a required-reading block; the agent's instructions never referenced the gate, so the clause could not fire — the F8 defect one layer up, caught by the new read-tag arm. +~200 bytes."
}
}

View File

@@ -19,7 +19,14 @@ const path = require('node:path');
const ROOT = path.join(__dirname, '..');
const LINT_SCRIPT = path.join(ROOT, 'scripts', 'lint-removed-but-needed.cjs');
const { referencesBasename, referencesNpmLockfileDependency, findSurvivingReferences, scan } = require(LINT_SCRIPT);
const {
referencesBasename,
referencesNpmLockfileDependency,
findSurvivingReferences,
classifyTestReference,
findSurvivingTestReferences,
scan,
} = require(LINT_SCRIPT);
const { cleanup } = require('./helpers.cjs');
const { runNode } = require('./helpers/process-seam.cjs');
const { gitOrThrow } = require('./helpers/git-fixture.cjs');
@@ -222,3 +229,188 @@ describe('removed-but-needed lint: main() end-to-end wiring', () => {
assert.match(result.stdout, /skipping/);
});
});
// ─── #3565: tests/ arm — pins-existence vs asserts-absence ──────────────────
describe('removed-but-needed lint: classifyTestReference (pure, #3565)', () => {
test('a negated includes carrying the basename is asserts-absence', () => {
assert.equal(
classifyTestReference("assert.ok(!content.includes('discovery-phase.md'))", 'discovery-phase.md'),
'asserts-absence',
);
});
test('a negated existsSync is asserts-absence', () => {
assert.equal(
classifyTestReference("assert.ok(!fs.existsSync(path.join(dir, 'x.md')))", 'x.md'),
'asserts-absence',
);
});
test('a negated regex test is asserts-absence', () => {
assert.equal(
classifyTestReference("assert.ok(!/x\\.md/.test(out));", 'x.md'),
'asserts-absence',
);
});
test('an unnegated existsSync is pins-existence', () => {
assert.equal(
classifyTestReference("assert.ok(fs.existsSync(path.join(dir, 'x.md')))", 'x.md'),
'pins-existence',
);
});
test('a readFileSync on the basename is pins-existence', () => {
assert.equal(
classifyTestReference("const c = fs.readFileSync(fixture('x.md'));", 'x.md'),
'pins-existence',
);
});
test('a require of the basename is pins-existence', () => {
assert.equal(
classifyTestReference("const data = require('./fixtures/x.md');", 'x.md'),
'pins-existence',
);
});
test('the basename as a quoted object key is pins-existence (the allowlist trap)', () => {
assert.equal(classifyTestReference(" 'x.md': true,", 'x.md'), 'pins-existence');
});
test('a bare prose mention is unclassifiable (known limit) — never a violation', () => {
assert.equal(classifyTestReference('// see the old x.md workflow', 'x.md'), null);
});
});
describe('removed-but-needed lint: findSurvivingTestReferences (pure, #3565)', () => {
const CORPUS = [
{
file: 'tests/absence.test.cjs',
content: [
"test('workflow is gone', () => {",
" const content = fs.readFileSync(INVENTORY, 'utf8');",
" assert.ok(!content.includes('discovery-phase.md'));",
'});',
'',
].join('\n'),
},
{
file: 'tests/pin.test.cjs',
content: [
"test('workflow ships', () => {",
" assert.ok(fs.existsSync(path.join(WF, 'discovery-phase.md')));",
'});',
'',
].join('\n'),
},
{
file: 'tests/both.test.cjs',
content: [
"test('both', () => {",
" assert.ok(!content.includes('discovery-phase.md'));",
" assert.ok(fs.existsSync(path.join(WF, 'discovery-phase.md')));",
'});',
'',
].join('\n'),
},
];
test('an absence assertion alone is NOT a violation — the case that reverted the naive #3560 widening', () => {
const vs = findSurvivingTestReferences(
['gsd-core/workflows/discovery-phase.md'],
[CORPUS[0]],
);
assert.deepEqual(vs, []);
});
test('an existence pin on the deleted file IS a violation — the #3560 red-runner case', () => {
const vs = findSurvivingTestReferences(
['gsd-core/workflows/discovery-phase.md'],
[CORPUS[1]],
);
assert.equal(vs.length, 1);
assert.equal(vs[0].deletedFile, 'gsd-core/workflows/discovery-phase.md');
assert.equal(vs[0].referencedIn, 'tests/pin.test.cjs');
assert.match(vs[0].reason, /pins-existence/);
});
test('a file with BOTH shapes reports the pin only, per-reference', () => {
const vs = findSurvivingTestReferences(
['gsd-core/workflows/discovery-phase.md'],
[CORPUS[2]],
);
assert.equal(vs.length, 1);
assert.equal(vs[0].referencedIn, 'tests/both.test.cjs');
});
test('a basename referenced for a DIFFERENT (undeleted) file is independent', () => {
const vs = findSurvivingTestReferences(
['gsd-core/workflows/other.md'],
[CORPUS[1]],
);
assert.deepEqual(vs, []);
});
});
describe('removed-but-needed lint: tests/ arm end-to-end (#3565)', () => {
test('a deleted workflow pinned by a test existsSync fails with the pins-existence reason', (t) => {
const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-rbn-tests-pin-'));
t.after(() => cleanup(tmpDir));
buildTempRepo(
tmpDir,
[
{ file: 'gsd-core/workflows/gone.md', content: '# gone\n' },
{
file: 'tests/pin.test.cjs',
content: [
"const fs = require('node:fs');",
"const path = require('node:path');",
"test('gone.md still ships', () => {",
" assert.ok(fs.existsSync(path.join('gsd-core', 'workflows', 'gone.md')));",
'});',
'',
].join('\n'),
},
],
[{ file: 'gsd-core/workflows/gone.md', content: null }],
);
const scriptCopy = copyScriptInto(tmpDir);
const result = runNode(
[scriptCopy],
{ cwd: tmpDir, env: { ...process.env, GSD_REMOVED_BUT_NEEDED_BASE: 'main' } },
);
assert.equal(result.exitCode, 1, `expected exit 1, got ${result.exitCode}: ${result.stderr}`);
assert.match(result.stderr, /pins-existence/);
assert.match(result.stderr, /tests[\\/]pin[\\.]test[\\.]cjs/);
});
test('a deleted workflow asserted ABSENT by a test passes — the discriminator is the whole point', (t) => {
const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-rbn-tests-absence-'));
t.after(() => cleanup(tmpDir));
buildTempRepo(
tmpDir,
[
{ file: 'gsd-core/workflows/gone.md', content: '# gone\n' },
{
file: 'tests/gone.test.cjs',
content: [
"const content = 'no trace of the deleted workflow';",
"test('gone.md is not referenced', () => {",
" assert.ok(!content.includes('gone.md'));",
'});',
'',
].join('\n'),
},
],
[{ file: 'gsd-core/workflows/gone.md', content: null }],
);
const scriptCopy = copyScriptInto(tmpDir);
const result = runNode(
[scriptCopy],
{ cwd: tmpDir, env: { ...process.env, GSD_REMOVED_BUT_NEEDED_BASE: 'main' } },
);
assert.equal(result.exitCode, 0, `expected exit 0, got ${result.exitCode}: ${result.stderr}`);
});
});