Files
msd-core/tests/agent-classification-parity.test.cjs
0xdhx 3274db2757 fix(#2526): remove gsd-ui-auditor's uncallable Playwright-MCP block (#2594)
* fix(#2526): drop gsd-ui-auditor's uncallable Playwright-MCP block

The agent declares `tools: Read, Write, Bash, Grep, Glob, Skill` — no
`mcp__*` grant of any kind — while its body presented a
`<playwright_mcp_approach>` block as the *preferred* capture path. That
branch was unreachable by construction: the availability check had a
fixed answer, the three `mcp__playwright__*` calls could never dispatch,
and the "when Playwright-MCP is NOT available" fallback was the only
branch that ever ran — 39 lines of instruction loaded on every
/gsd-ui-review spawn that also invited the model to claim a capture path
it could not take.

Remove the dead block, leaving the CLI screenshot path as the sole
documented approach.

Guard the class in tests/mcp-tool-inheritance.test.cjs, which already
owns agent MCP-grant parity: the new block generalizes the #1284
researcher check from two agents and one dispatch table to every
agents/*.md and its whole body — no agent may document an
`mcp__<server>__*` namespace absent from its own `tools:` declaration.

Frontmatter is read through the canonical parser
(gsd-core/bin/lib/frontmatter.cjs) rather than a hand-rolled scan, so
inline CSV, block sequences, flow arrays, quoted scalars and full-line
comments are handled by construction; inline comments inside a scalar
survive that parser, so they are stripped explicitly. The check is
server-level by design, ignores prose metavariables like `mcp__X__*`,
matches hyphenated server ids, and carries a discovery guard plus
negative controls for every documented boundary so it cannot decay into
a vacuous pass.

The session-level Playwright-MCP pass in gsd-core/workflows/ui-review.md
is deliberately untouched — workflow files carry no fixed allowlist, so
their availability check is genuinely runtime-detected and honest.

Fixes #2526

* chore(#2526): set changeset fragment pr to 2594

The fragment shipped with the documented `pr: 0` placeholder because the
PR number does not exist until the PR is opened, and
scripts/changeset/parse.cjs rejects `pr <= 0`. Now that the PR is open,
set the real number so changeset-lint passes.

* test(#2526): cover the two-char server-id boundary of the metavariable exclusion

The length-1 "prose metavariable" exclusion was tested at length=1 and at
real ids (>=3 chars), but never at length=2 — the limit+1 boundary where a
server id starts being recognized. Review finding on #2594: `mcp__ab__foo`
in a body with no grant must flag `['ab']`.

* fix(#2526): treat a bare mcp__* grant as covering every server

`grantedServers()` stripped `mcp__*` to the empty string and dropped it via
`if (server)`, so an allowlist that grants every MCP server read as granting
none — and the guard then fired against a body the grant plainly covered.
That is the one input shape that inverts the check, turning it against a
correct agent rather than merely missing a bad one.

A `/^mcp__\*+$/` token now sets a GRANT_ALL sentinel that short-circuits
`ungrantedServers()`. The sentinel `*` is outside REFERENCE_RE's character
class, so no body reference can collide with it. A bare `mcp__` with no
wildcard stays a typo rather than a grant and keeps failing closed.

No agent uses the `mcp__*` spelling today, so this was latent rather than
live. Two negative controls pin both halves.

* fix(#2526): scan the frontmatter description for MCP references too

`ungrantedServers()` scanned `stripFrontmatter(content)` only, so an
`mcp__foo__bar` reference in the `description` field escaped the check. That
field ships with the agent and the dispatcher reads it, which makes a dead
reference there exactly as dead as one in the body.

Only `description` is added to the scanned surface, never the whole
frontmatter: `tools:` is the grant list itself, so scanning it would let every
allowlist satisfy itself and turn the guard vacuous. A negative control pins
that boundary alongside the new positive case.

All 34 per-agent tests still pass with the wider surface, so no live agent
verdict changes — this was latent.

* test(#2526): give multi-character placeholders a convention the checker knows

The metavariable exclusion is `length === 1`, so the natural placeholders
`mcp__SRV__*` and `mcp__SERVER__*` were flagged as real references — and the
failure message then offered an author two remedies ("grant the namespace or
drop the block") that both misread what they wrote.

Adopts the angle-bracket half of the suggested fix: `mcp__<SERVER>__*` is the
sanctioned multi-character placeholder, exempt by construction because `<` is
outside the reference pattern's character class. This pins an existing
property rather than adding a special case.

Declines the all-caps half. An all-caps exemption would be a false NEGATIVE
for any real server spelled in caps, and a guard that misses a dead reference
fails in exactly the direction this check exists to prevent. The bare-caps
form keeps firing; the message now names the convention as a third remedy.

Also corrects "grants neither" in that message, which was wrong for any count
other than two.

* test(#2526): pin the zero-length server id, completing the boundary triple

`mcp____foo` yields `[]`, but for a different reason than the length-1 case:
it is unrepresentable by `/mcp__([A-Za-z0-9_-]+?)__/g` since `+?` requires at
least one character, so the pattern skips it before the metavariable
exclusion is ever consulted. Pinning limit-1 completes the 0/1/2 boundary
rule on its own terms and records which mechanism owns the case.

* docs(#2526): correct every drifted AGENTS.md Tools row, not just the one

The review asked for the one-line `gsd-ui-auditor` correction (missing
`Skill`). Sweeping the defect class first — every `**Tools**` row in
docs/AGENTS.md against its agent's `tools:` frontmatter — found it was 26 of
34 rows, so the one-line framing was the reviewer's premise rather than the
population.

Breakdown of the 26:
  * 21 omitted `Skill`, 6 omitted `Edit` (overlapping) — under-promises, the
    same drift class as #2526 but in the harmless direction.
  * 8 wrote `mcp (context7)` as shorthand while frontmatter granted up to 8
    servers (firecrawl, exa, tavily, ref, jina, perplexity, both context7s).
  * 1 was actively wrong: gsd-debug-session-manager documented `Task`, a tool
    name that no longer exists — the #2526 shape at the doc layer, naming a
    capability that cannot dispatch.

Every row is now the frontmatter `tools:` value verbatim, which is also what
makes the parity guard in the following commit non-brittle. The diff is
26 insertions / 26 deletions, all Tools rows.

* test(#2526): guard AGENTS.md Tools rows against agent frontmatter

The 26 corrected rows in the previous commit were free to drift because
nothing asserted the role card and the frontmatter agreed — the same reason
the #2526 block itself survived. Correcting them without an invariant just
resets the clock.

Lands in agent-classification-parity.test.cjs rather than a new file: that
suite already owns docs/AGENTS.md as a contract surface, already carries the
`allow-test-rule` exemption for treating the doc as the product, and file
count is the unit of CI overhead (docs/TESTING-SUITES.md).

Compares the row to the frontmatter value VERBATIM, not as a set — a set
comparison would keep accepting the "mcp (context7)" shorthand that hid eight
grants behind one, which is the under-documentation half of the drift.

Carries the same discovery guard #2526's own check uses: a section with a
granted `tools:` but no **Tools** row fails loudly, so deleting a row cannot
silently retire its assertion. Both halves are negative-controlled — against
the pre-fix doc it fails naming 26 rows (gsd-ui-auditor:339 among them), and
with a row deleted it fails on the missing-row assertion.

* chore(#2526): note the AGENTS.md drift correction in the changeset

The role cards are user-visible, and 26 of them documented a tool set the
agent did not have. Type, `pr: 2594`, and the trailing `(#2526)` are
unchanged.

* fix(#2526): use a CRLF-safe split in the AGENTS.md Tools-row guard

`lint-tests` (npm run lint:ci) rejected `rawAgentsMd.split('\n')` under the
repo's local/no-crlf-fragile-split rule: Windows autocrlf yields CRLF, so a
trailing \r rides into the parsed line. Switched to `/\r?\n/`.

Caught by CI on the round-3 push before the response comment went out.

* docs(#2526): correct the drift tallies stated in c5a9607e

Re-derived the census programmatically from the pre-fix doc instead of by eye.
The 26-of-34 headline was right; the breakdown was not.

  Skill omitted   21 -> 22
  Edit omitted     6 ->  7
  "mcp (context7)" shorthand   8 rows -> 7 rows

The 8 was conflating two things: 8 rows omitted MCP grants entirely, but only
7 of them used the "mcp (context7)" shorthand — gsd-executor listed no MCP at
all. Also names the one `Agent` omission (gsd-debug-session-manager, the row
that still read `Task`).

c5a9607e's message keeps the wrong numbers rather than rewriting a pushed
branch mid-review; the test comment and changeset are the durable statements
and both are corrected here.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-28 18:02:01 -04:00

455 lines
18 KiB
JavaScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
// allow-test-rule: runtime-contract-is-the-product — docs/AGENTS.md section layout + docs/INVENTORY.md table ARE the classification surface being validated
'use strict';
/**
* Agent classification parity test (#1171)
*
* Makes docs/AGENTS.md section structure the single source of truth for
* primary-vs-advanced agent classification, and fails when docs/INVENTORY.md
* or the AGENTS.md prose counts drift from it.
*
* Classification rules (derived from AGENTS.md section placement):
* - ### gsd-<name> headings BEFORE "## Advanced and Specialized Agents" → "primary"
* - ### gsd-<name> headings INSIDE/AFTER that section → "advanced stub"
* - agents/gsd-*.md files with NO heading in AGENTS.md → "inventory only"
*/
const { describe, test } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const { listAgentFiles } = require('./helpers/agent-roster.cjs');
const ROOT = path.resolve(__dirname, '..');
const AGENTS_MD = path.join(ROOT, 'docs', 'AGENTS.md');
const INVENTORY_MD = path.join(ROOT, 'docs', 'INVENTORY.md');
const AGENTS_DIR = path.join(ROOT, 'agents');
// ---------------------------------------------------------------------------
// Parsing helpers
// ---------------------------------------------------------------------------
/**
* Parse AGENTS.md and return two arrays:
* primaryHeadings — agent names (without "gsd-" prefix) before the advanced section
* advancedHeadings — agent names inside "## Advanced and Specialized Agents"
*
* We work with the full "gsd-<name>" slug so the names are unambiguous.
*
* State machine (three-state):
* 'before' — before "## Advanced and Specialized Agents"
* 'advanced' — inside that section
* 'past' — after a subsequent ## heading that follows the advanced section
*
* A "## " heading (h2, NOT h3) that appears BEFORE the advanced section does NOT
* trigger 'past'. Only a "## " heading encountered WHILE in 'advanced' does.
* ### gsd-* headings are counted: primary if 'before', advanced if 'advanced',
* ignored if 'past'.
*/
function parseAgentsMd(raw) {
const lines = raw.split('\n');
const ADVANCED_SECTION = /^##\s+Advanced and Specialized Agents\s*$/;
// Matches an h2 heading (exactly two hashes, not three or more)
const H2 = /^##(?!#)\s+\S/;
const GSD_H3 = /^###\s+(gsd-[\w-]+)\s*$/;
const primaryHeadings = [];
const advancedHeadings = [];
// 'before' | 'advanced' | 'past'
let state = 'before';
for (const line of lines) {
if (ADVANCED_SECTION.test(line)) {
state = 'advanced';
continue;
}
// Any other h2 heading while in 'advanced' terminates the advanced region
if (state === 'advanced' && H2.test(line)) {
state = 'past';
continue;
}
const m = GSD_H3.exec(line);
if (m) {
if (state === 'before') {
primaryHeadings.push(m[1]);
} else if (state === 'advanced') {
advancedHeadings.push(m[1]);
}
// state === 'past': silently ignored
}
}
return { primaryHeadings, advancedHeadings };
}
/**
* Parse INVENTORY.md agent table.
* Returns a Map<agentSlug, primaryDocValue> e.g. "primary" | "advanced stub" | "inventory only"
*
* Scoping: only rows that fall between the "## Agents" heading and the NEXT
* "## " heading are considered. This prevents a future non-agent table that
* happens to contain a "| gsd-..." row from polluting the result.
*
* Column resolution: the header row "| Agent | ... | Primary doc |" is parsed
* to find the 0-based index of the "Primary doc" column. Rows are split on "|"
* and only the leading/trailing empty edge cells are dropped (slice(1,-1)) so
* that empty middle cells do NOT shift column positions.
*/
function parseInventoryMd(raw) {
const lines = raw.split('\n');
// Matches any h2 heading (exactly two hashes, not three or more)
const H2 = /^##(?!#)\s+\S/;
// Matches the Agents section heading (e.g. "## Agents (33 shipped)")
const AGENTS_SECTION = /^##\s+Agents\b/;
const result = new Map();
let inAgentsSection = false;
let primaryDocColIndex = -1; // column index within the trimmed, edge-stripped cell array
for (const line of lines) {
if (AGENTS_SECTION.test(line)) {
inAgentsSection = true;
primaryDocColIndex = -1; // reset in case file is re-parsed
continue;
}
// Any subsequent h2 heading ends the agents section
if (inAgentsSection && H2.test(line)) {
break;
}
if (!inAgentsSection) continue;
// Every table row starts and ends with "|"
if (!line.startsWith('|')) continue;
// Split on "|", drop the leading and trailing empty strings that result
// from the leading/trailing "|", but preserve empty middle cells so column
// indices stay stable.
const rawCells = line.split('|');
// rawCells[0] is '' (before the first |), rawCells[last] is '' (after last |)
const cells = rawCells.slice(1, -1).map((c) => c.trim());
// Detect the header row by looking for an "Agent" cell followed by a "Primary doc" cell
if (primaryDocColIndex === -1) {
const pdIdx = cells.findIndex((c) => c === 'Primary doc');
if (pdIdx !== -1 && cells[0] === 'Agent') {
primaryDocColIndex = pdIdx;
}
continue; // header row (or rows before the header is found) — not a data row
}
// Skip separator rows (---|---|...)
if (cells.every((c) => /^[-: ]+$/.test(c))) continue;
// Data rows: first cell must be a gsd-* slug
if (!cells[0].startsWith('gsd-')) continue;
if (cells.length <= primaryDocColIndex) continue;
const agentSlug = cells[0];
const primaryDoc = cells[primaryDocColIndex];
result.set(agentSlug, primaryDoc);
}
assert.ok(
primaryDocColIndex !== -1,
'INVENTORY.md: "Primary doc" column header not found in the Agents table — check the ## Agents section heading and table header row',
);
return result;
}
// ---------------------------------------------------------------------------
// Load and parse
// ---------------------------------------------------------------------------
const rawAgentsMd = fs.readFileSync(AGENTS_MD, 'utf8');
const rawInventoryMd = fs.readFileSync(INVENTORY_MD, 'utf8');
const { primaryHeadings, advancedHeadings } = parseAgentsMd(rawAgentsMd);
const inventoryMap = parseInventoryMd(rawInventoryMd);
// Canonical source roster (sorted gsd-* basenames without .md) — shared helper.
const agentFiles = listAgentFiles(AGENTS_DIR);
// ---------------------------------------------------------------------------
// Robustness guards — must pass before any assertion block runs
// ---------------------------------------------------------------------------
assert.ok(
primaryHeadings.length > 0,
'AGENTS.md: no ### gsd-* headings found before Advanced section — raw head:\n' + rawAgentsMd.slice(0, 300),
);
assert.ok(
advancedHeadings.length > 0,
'AGENTS.md: no ### gsd-* headings found inside Advanced section — raw head:\n' + rawAgentsMd.slice(0, 300),
);
assert.ok(
inventoryMap.size > 0,
'INVENTORY.md: no | gsd-* | rows parsed — raw head:\n' + rawInventoryMd.slice(0, 300),
);
assert.ok(
agentFiles.length > 0,
'agents/: no gsd-*.md files found — check AGENTS_DIR path: ' + AGENTS_DIR,
);
// ---------------------------------------------------------------------------
// Derived sets
// ---------------------------------------------------------------------------
const primarySet = new Set(primaryHeadings);
const advancedSet = new Set(advancedHeadings);
const agentFileSet = new Set(agentFiles);
// Agents that exist on disk but have no heading in AGENTS.md
const inventoryOnly = agentFiles.filter(
(a) => !primarySet.has(a) && !advancedSet.has(a),
);
const inventoryOnlySet = new Set(inventoryOnly);
// ---------------------------------------------------------------------------
// Tests
// ---------------------------------------------------------------------------
describe('agent-classification-parity: AGENTS.md section structure is the single source of truth', () => {
/**
* Test 1 — INVENTORY.md "Primary doc" column matches AGENTS.md-derived class.
* For every agent, the column value must equal:
* "primary" if the agent's ### heading is before ## Advanced and Specialized Agents
* "advanced stub" if the agent's ### heading is inside/after that section
* "inventory only" if the agent has no ### heading in AGENTS.md
*/
test('INVENTORY.md "Primary doc" values match AGENTS.md section placement', () => {
const mismatches = [];
for (const [slug, inventoryValue] of inventoryMap) {
let expected;
if (primarySet.has(slug)) {
expected = 'primary';
} else if (advancedSet.has(slug)) {
expected = 'advanced stub';
} else if (inventoryOnlySet.has(slug)) {
expected = 'inventory only';
} else {
// Row in INVENTORY.md for an agent with no file — handled in test 3
continue;
}
if (inventoryValue !== expected) {
mismatches.push(` ${slug}: INVENTORY.md="${inventoryValue}" but AGENTS.md section says "${expected}"`);
}
}
assert.strictEqual(
mismatches.length,
0,
'INVENTORY.md "Primary doc" column disagrees with AGENTS.md section placement:\n' + mismatches.join('\n'),
);
});
/**
* Test 2 — AGENTS.md prose counts and parenthetical list are accurate.
*
* Checks three sub-facts from line ~13:
* a) "21 primary agents" → count of primary headings
* b) "Twelve additional" → count of advanced headings
* c) The parenthetical slug list → equals the set of advanced heading names
* (slugs without the "gsd-" prefix, as the prose uses)
*/
test('AGENTS.md prose counts and advanced-agent parenthetical list are accurate', () => {
// --- (a) prose primary count ---
const primaryCountMatch = rawAgentsMd.match(/\*\*(\d+)\s+primary\s+agents?\*\*/);
assert.ok(
primaryCountMatch,
'AGENTS.md: could not find "**N primary agents**" in prose — raw head:\n' + rawAgentsMd.slice(0, 500),
);
const prosePrimaryCount = parseInt(primaryCountMatch[1], 10);
assert.strictEqual(
prosePrimaryCount,
primaryHeadings.length,
`AGENTS.md prose says "${prosePrimaryCount} primary agents" but there are ${primaryHeadings.length} ### gsd-* headings before the Advanced section`,
);
// --- (b) prose advanced count (cardinal word or digit) ---
// Look for "Twelve additional" or "12 additional" (case-insensitive cardinal)
const CARDINALS = {
one: 1, two: 2, three: 3, four: 4, five: 5, six: 6, seven: 7,
eight: 8, nine: 9, ten: 10, eleven: 11, twelve: 12, thirteen: 13,
};
const advancedCountMatch = rawAgentsMd.match(
/\b((?:\d+|one|two|three|four|five|six|seven|eight|nine|ten|eleven|twelve|thirteen))\s+additional\s+shipped\s+agents?\b/i,
);
assert.ok(
advancedCountMatch,
'AGENTS.md: could not find "N additional shipped agents" in prose — raw head:\n' + rawAgentsMd.slice(0, 500),
);
const advancedCountRaw = advancedCountMatch[1].toLowerCase();
const proseAdvancedCount = CARDINALS[advancedCountRaw] !== undefined
? CARDINALS[advancedCountRaw]
: parseInt(advancedCountRaw, 10);
assert.strictEqual(
proseAdvancedCount,
advancedHeadings.length,
`AGENTS.md prose says "${advancedCountRaw} additional shipped agents" but there are ${advancedHeadings.length} ### gsd-* headings in the Advanced section`,
);
// --- (c) parenthetical slug list ---
// The prose lists short slugs without "gsd-" prefix, e.g.:
// (pattern-mapper, debug-session-manager, ...)
const parenMatch = rawAgentsMd.match(/\(([^)]+)\)\s+have concise stubs/);
assert.ok(
parenMatch,
'AGENTS.md: could not find parenthetical advanced-agent list "(slug, slug, ...) have concise stubs" — raw head:\n' + rawAgentsMd.slice(0, 500),
);
const proseSlugs = parenMatch[1].split(',').map((s) => s.trim().toLowerCase());
const proseSlugsSet = new Set(proseSlugs);
// Derive expected slugs from AGENTS.md headings (strip "gsd-" prefix)
const expectedSlugs = new Set(advancedHeadings.map((h) => h.replace(/^gsd-/, '')));
const missingFromProse = [...expectedSlugs].filter((s) => !proseSlugsSet.has(s));
const extraInProse = [...proseSlugsSet].filter((s) => !expectedSlugs.has(s));
assert.deepStrictEqual(
{ missingFromProse, extraInProse },
{ missingFromProse: [], extraInProse: [] },
'AGENTS.md parenthetical slug list disagrees with ### headings in the Advanced section.\n' +
` Missing from prose: ${JSON.stringify(missingFromProse)}\n` +
` Extra in prose: ${JSON.stringify(extraInProse)}`,
);
});
/**
* Test 3 — Roster completeness.
*
* Sub-checks:
* a) Every agents/gsd-*.md appears exactly once in (primary ∪ advanced ∪ inventoryOnly)
* — i.e. no agent is double-classified.
* b) INVENTORY.md has a row for every agent file.
* c) INVENTORY.md has no row for a non-existent agent file.
* d) primary.length + advanced.length + inventoryOnly.length === total agent file count
*/
test('Roster completeness: every agent file is classified exactly once', () => {
// (a) no agent appears in more than one classification bucket
const overlap = primaryHeadings.filter((a) => advancedSet.has(a));
assert.deepStrictEqual(
overlap,
[],
'Agents appear in BOTH primary and advanced heading sections: ' + JSON.stringify(overlap),
);
// (b) INVENTORY.md has a row for every agent file
const missingFromInventory = agentFiles.filter((a) => !inventoryMap.has(a));
assert.deepStrictEqual(
missingFromInventory,
[],
'agents/gsd-*.md files missing from INVENTORY.md: ' + JSON.stringify(missingFromInventory),
);
// (c) INVENTORY.md has no row for a non-existent agent file
const phantomRows = [...inventoryMap.keys()].filter((slug) => !agentFileSet.has(slug));
assert.deepStrictEqual(
phantomRows,
[],
'INVENTORY.md rows reference agents that have no agents/gsd-*.md file: ' + JSON.stringify(phantomRows),
);
// (d) counts add up
const total = primaryHeadings.length + advancedHeadings.length + inventoryOnly.length;
assert.strictEqual(
total,
agentFiles.length,
`primary(${primaryHeadings.length}) + advanced(${advancedHeadings.length}) + inventoryOnly(${inventoryOnly.length}) = ${total} ≠ ${agentFiles.length} agent files`,
);
});
/**
* Test 4 — AGENTS.md **Tools** rows match agent frontmatter verbatim (#2526).
*
* Same defect class as #2526 itself, one layer out: there, an agent's body
* documented a capability its `tools:` allowlist withheld; here, the role
* card documents a tool set its frontmatter disagrees with. When the review
* that found it was written, 26 of 34 rows had drifted: 22 omitted `Skill`
* and 7 omitted `Edit`; 8 omitted MCP grants entirely (7 of them writing
* "mcp (context7)" for what was up to eight distinct servers, and
* gsd-executor naming none at all); and one still read `Task`, a tool that
* no longer exists. Nothing asserted the two agreed, so the drift was free.
*
* The row must equal the frontmatter value verbatim rather than as a set:
* a set comparison would accept the "mcp (context7)" shorthand class of
* under-documentation this test exists to stop, and an exact string is what
* makes the check cheap to satisfy — copy the line.
*
* The row-count assertion is the discovery guard (mirroring #2526's own):
* without it, deleting a **Tools** row would silently retire its assertion
* while the remaining rows kept the test green.
*/
test('AGENTS.md **Tools** rows match each agent\'s tools: frontmatter (#2526)', () => {
const { parseFrontmatter } = require('../gsd-core/bin/lib/frontmatter.cjs');
const TOOLS_ROW = /^\|\s*\*\*Tools\*\*\s*\|\s*(.*?)\s*\|\s*$/;
const H3 = /^###\s+(gsd-[\w-]+)\s*$/;
// /\r?\n/, not '\n': Windows autocrlf yields CRLF, and a trailing \r would
// survive into the row's last cell and fail every comparison (local/no-crlf-fragile-split).
const lines = rawAgentsMd.split(/\r?\n/);
const sections = [];
lines.forEach((line, i) => {
const m = H3.exec(line);
if (m) sections.push({ slug: m[1], start: i });
});
sections.forEach((s, i) => {
s.end = i + 1 < sections.length ? sections[i + 1].start : lines.length;
});
const mismatches = [];
const missingRow = [];
for (const section of sections) {
const agentPath = path.join(AGENTS_DIR, `${section.slug}.md`);
if (!fs.existsSync(agentPath)) continue; // phantom headings are test 3's job
const fm = parseFrontmatter(fs.readFileSync(agentPath, 'utf8')) || {};
// No `tools:` key at all means "inherits everything" — there is no
// declared set for the row to agree with, so nothing to assert.
if (fm.tools === undefined || fm.tools === null) continue;
const declared = Array.isArray(fm.tools) ? fm.tools.join(', ') : String(fm.tools);
let row = null;
for (let i = section.start; i < section.end; i += 1) {
const m = TOOLS_ROW.exec(lines[i]);
if (m) { row = { value: m[1], line: i + 1 }; break; }
}
if (row === null) { missingRow.push(section.slug); continue; }
if (row.value !== declared) {
mismatches.push(
` ${section.slug} (docs/AGENTS.md:${row.line})\n` +
` doc: ${row.value}\n` +
` frontmatter: ${declared}`,
);
}
}
assert.deepStrictEqual(
missingRow,
[],
'AGENTS.md sections with a granted tools: frontmatter but no "| **Tools** |" row — ' +
'the row cannot be allowed to vanish, or its parity assertion vanishes with it: ' +
JSON.stringify(missingRow),
);
assert.strictEqual(
mismatches.length,
0,
'docs/AGENTS.md **Tools** rows disagree with the agents\' tools: frontmatter.\n' +
'The role card documents a tool set the agent does not have (or omits one it does) — ' +
'the #2526 drift class at the doc layer. Copy the frontmatter value verbatim:\n' +
mismatches.join('\n'),
);
});
});