Commit Graph

8 Commits

Author SHA1 Message Date
Jakub Zych
a9a7a328e6 refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD
across contents and paths, upstream package/repo coordinates -> @golem15/msd-core
and golem15com/msd-core. Deep links into upstream history, sibling upstream
packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is.

Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line,
package/plugin identity, regenerated lockfile, install-tree fixtures, derived
registries and benchmark baseline; migration checksum baseline re-locked
(MSD keeps its own install state, so no install had applied the old sums);
sort-order and regex-escaped expectations in tests adjusted.
2026-10-06 01:47:40 +02:00
Tom Boucher
37b965c0d1 enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code (#4553)
* enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code

ADR-4139 stream 2. The non-Claude `#2454` persona fallback in cmdAgentSkills
(src/init.cts) now selects between a canonical agents/<name>.md and a
token-minimized agents/<name>.compact.md sibling based on
workflow.compact_content, resolved in code (a real function call with a real
exit code) rather than a prose config-get gate — the same precedent stream 1's
spine/detail split established for a load-bearing seam, applied here because
this seam already runs through TypeScript instead of an eager @-include.

A missing compact sibling falls back to the canonical persona and discloses
the fallback in the served payload itself (a leading HTML-comment provenance
line), so the Done-when contract — compact when on, canonical when off, never
silent or empty — holds even for an agent nobody has compacted yet.

Authored a .compact.md sibling for all 35 shipped agents (agents/gsd-*.md),
each an independent, complete rewrite (not an extraction — nothing is "moved"
the way spine/detail moves text) that preserves frontmatter, every @-include,
every output-format contract, and every guardrail verbatim while cutting
restatement and verbose framing. Verified mechanically: every pair registers
(a canonical sibling exists), every compact file is strictly smaller, and the
full @-include set matches canonical's — including which references are
standalone eager-load lines versus inline prose mentions, since demoting one
to inline changes what the host actually substitutes.

Traced the install path before writing any code (.gsd/phase/.../40-design.md):
stageAgentsForRuntimeWithConverter glob-copies every agents/*.md file with no
stem filtering under the default full profile, so the new .compact.md files
install for free with zero installer changes — matching issue #4407's stated
scope. A tiered agent profile that doesn't stage a compact sibling degrades
through the same fallback-with-provenance path already required for an
unauthored one, so no installer change is needed there either.

Extends tests/helpers/compact-content-variant.cjs with an AGENTS_ROOT export
(deliberately not folded into DEFAULT_VARIANT_ROOTS, since agent variants are
reached by a generic code construction rather than a literal path in prose,
and checkReachability's markdown-search shape has nothing to find there).
Reachability is instead proven behaviorally: tests/agent-skills.test.cjs's new
"#4407 compact payload selection" describe block spawns gsd_run agent-skills
against real compact/canonical fixture pairs and asserts on the served
payload, which can only pass if the seam genuinely wires through.

Fixed a pre-existing test whose agents/*.md glob incidentally matched the new
.compact.md siblings (tests/agent-skills.test.cjs's Skill-frontmatter drift
guard) and added the 35 new agents/*.compact.md entries to docs/INVENTORY.md's
roster, both real, unrelated-to-content defects the new files' mere existence
surfaced.

Regenerated: install-tree fixtures (19 runtimes now ship 35 more agent files
under the full profile), INVENTORY-MANIFEST.json, and the variant-swap token
benchmark baseline (npm run benchmark:compact-content-variants --write).

Closes #4407.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): apply orthogonal review findings from the compact-payload seam

Standards axis of /code-review: extracted readNonEmptyFileOrNull(filePath)
to collapse the duplicated read-and-empty-check shape between the compact
and canonical branches in cmdAgentSkills, and updated the adjacent comment
enumerating flat JSON extras to name agent_payload_variant alongside
source/degraded (added by the prior commit, comment left stale).

Security review and the Spec axis found no defects requiring a code change;
their non-blocking observations (a pre-existing, unmodified path-construction
pattern; the reasoned, documented substitution of a behavioral test for the
literal reachability check) are recorded in
.gsd/phase/enhance-4407-agent-skill-seam/60-review.json.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): repo-wide roster/cap fixes surfaced by shipping .compact.md agents

Root-caused via a real gsd-test run (93 failures) rather than guessing which
tests glob agents/ naively. Two classes of defect, both genuine:

1. Identity-roster confusion (11 files/areas): many tests and one production
   script derive "the set of GSD agents" from `readdirSync(agentsDir).filter(f
   => f.endsWith('.md'))`, which incidentally matched the new .compact.md
   variant siblings too — a compact file is a rendering of an EXISTING agent
   identity, not a new one. Fixed at the shared root
   (tests/helpers/agent-roster.cjs's listAgentFiles, which several tests
   already consolidated on) and at each independent glob that didn't use it:
   agent-size-budget.test.cjs (tier-cap lookup now strips the .compact suffix
   before checking XL/LARGE membership, so a compact file inherits its
   canonical sibling's tier instead of silently falling through to DEFAULT),
   agent-skills-bootstrap.test.cjs, check-contract-drift.test.cjs (the actual
   script, not just its test), codex-config.test.cjs (confirmed directly
   against generateCodexAgentToml that a compact role's derived sandbox_mode
   is byte-identical to its canonical sibling's before excluding it — not
   assumed), and copilot-install.test.cjs (two counts that legitimately DO
   need both files — an installed-file count and a full-conversion smoke test
   — fixed to expect 70, not stay pinned to 35).

   no-bare-gsd-tools-command-position.test.cjs needed the opposite kind of fix:
   two compact files reproduce descriptive prose already allowlisted at their
   canonical file's line number; added matching entries at the compact files'
   own line numbers rather than excluding them from the scan (a genuine bare
   gsd-tools command-position bug in a compact file would be as real a defect
   as in canonical).

2. A hard, non-ackable cap (found via emitted-attribution.test.cjs's real-tree
   run): six agents' compact renditions (gsd-debugger, gsd-executor,
   gsd-phase-researcher, gsd-plan-checker, gsd-planner, gsd-verifier) exceed
   the 32,768-byte NEW_FILE_CAP (ADR-1610) even after aggressive compaction —
   confirmed structural, not a compaction-quality gap: each is dominated by
   content this phase's own rules require verbatim (the ~2.6 KB gsd_run
   bootstrap preamble runtime-launcher-parity.test.cjs requires inlined in
   every agent that calls gsd_run, output-format contracts, guardrails).
   ADR-4139's prescribed remedy (spine + lazily-read parts) has no landing
   spot in cmdAgentSkills's single-file synchronous read. Removed these 6
   compact files rather than ship an over-cap file or invent a multi-part
   read mechanism out of scope for this phase; recorded by name with the
   reason in .gsd/phase/enhance-4407-agent-skill-seam/40-design.md and
   50-test-matrix.md, per #4407's own "or explicitly recorded as not worth
   covering" allowance. Their canonical personas are served correctly today
   via the fallback-with-disclosed-provenance path this phase's own Done-when
   #2 already requires — 29 of 35 agents now have a compact variant.

Also fixes an unrelated, genuinely pre-existing defect this gsd-test run
surfaced: gsd-core/workflows/execute-plan.md sat 21 bytes over its own
DEFAULT-tier hard cap (40,960 bytes) at the branch point, before any change in
this PR touched it — confirmed via `git show <merge-base>:...execute-plan.md
| wc -c`. Per CLAUDE.md's no-deferral rule, fixed inline rather than filed:
two meaning-preserving trims in the <success_criteria> block (a repeated
parenthetical replaced with a same-exception reference; one redundant
qualifier dropped) bring it to 40,940 bytes.

Regenerated install-tree fixtures, INVENTORY-MANIFEST.json, and the variant
benchmark baseline to reflect the 6 removed files. Docs/INVENTORY.md's 6
now-orphaned roster rows removed alongside them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): make .compact.md-aware roster checks resilient to partial coverage

Round 2 of the gsd-test-driven roster fixes: two checks assumed every agent
has a compact sibling (true for 29 of 35 after the NEW_FILE_CAP exception),
breaking once 6 stems legitimately have none.

- tests/agent-classification-parity.test.cjs: the INVENTORY.md parser was
  picking up the "### Compact Payload Variants" subsection's rows as
  phantom/uncounted entries in the primary/advanced/inventory-only
  classification this test validates — a compact row documents an existing
  agent's alternate rendition and never gets its own AGENTS.md heading, so it
  was never meant to participate in that classification. Excluded at the
  parser, not per-assertion.
- tests/copilot-install.test.cjs: the derived expected-file-list generator
  assumed every listAgentFiles() stem has a .compact.md source sibling;
  checks disk per stem now instead.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4407): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 12:38:59 -04:00
Tom Boucher
69e7afd0c7 chore(#3212): bounded quantifiers over document content — prohibition with teeth — Phase 4 (#3441)
* feat(#3415): ship local/no-unbounded-quantifier, burn down ReDoS class

Phase 4 of epic #3212 (ADR-3212 §5/§7, the final phase). New rule flags
an unbounded */+/{n,} quantifier over a broad character class
([\s\S], dotAll ., or a 1-2-unit negated class like [^\n]/[^)\n] — the
exact #2128-fixed shape) applied to a regex whose match target is
data-flow-traced to readFileSync content.

eslint-rules/lib/readfilesync-trace.cjs extracts the data-flow tracer
shared with no-crlf-fragile-split (Phase 2) rather than a second copy
— no-crlf-fragile-split refactored onto it with zero behavior change,
parity-tested.

Real triage, not 798 mechanical edits: the ADR's census (2026-08-08)
screened every unbounded quantifier in the tree unscoped. Correctly
scoped to readFileSync-derived content (matching Phase 2's own G2/G3
scoping), the rule found 162 real hits across two detection waves — the
second wave (93) surfaced only after a genuine off-by-one bug in this
rule's own first draft was caught while writing its RuleTester tests
and fixed (the bug silently missed every directly-quantified [\s\S]*
with no gap before the quantifier — exactly the class this rule exists
to catch). 3 hits landed in production src/ (commands.cts, milestone.cts,
roadmap.cts) and were each empirically timed against adversarial input
(matching #2128's own measured-not-assumed precedent) — all confirmed
linear-time/benign, left unbounded with a measured-evidence comment
rather than mechanically bounded. The remaining 159 are test-file
fixture parsing (test-author-controlled, fixed-size content, not
adversarial input) — each suppressed with a specific, non-generic
reason. Zero functional behavior changed anywhere in this diff.

tests/no-pending-3212-markers.test.cjs locks the epic's own closing
invariant (ADR §7: "assert zero pending #3212 markers remain") — ground
truth confirmed trivially true today (no phase left any such marker
behind), now regression-locked going forward.

Design: .gsd/phase/chore-3415-prohibition-with-teeth/40-design.md
Test matrix: .gsd/phase/chore-3415-prohibition-with-teeth/50-test-matrix.md

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3415): correct rule category mislabel, add CI test-scope entry

An orthogonal Standards-axis review found eslint-rules/no-unbounded-quantifier.cjs
mistakenly carried meta.docs.category: 'Portability', copied from a sibling
rule without realizing what that implied: docs/contributing/cross-platform-
portability-rules.md governs an ADR-1703 rule family under a hard "zero
escape hatches" contract (tests/portability-rule-disable-ban.test.cjs's
PROTECTED_RULES bans eslint-disable for those rules entirely). This rule is
not part of that family — it's ADR-3212 (ReDoS/CWE-1333), a different epic —
and its eslint-disable-next-line suppressions (159 of them, added earlier
this same phase after empirical benign-verification) are an intentional,
correct design, not a bypass. Corrected to category: 'Best Practices',
matching the actual precedent (no-adhoc-regex-escape.cjs, Phase 1 of the
same epic, which is also correctly outside PROTECTED_RULES), and the rule's
own docstring now states this explicitly so a future reader doesn't have to
re-derive it.

Also registers a new scripts/ci-test-scope.cjs bucket so editing this rule
or the shared eslint-rules/lib/readfilesync-trace.cjs helper re-runs their
own test suites under targeted CI selection — was previously unregistered
and invisible to that fast-path (this PR's own gsd-test checkpoint runs the
full suite regardless, so this only affects future narrowly-scoped PRs).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3415): bound no-unbounded-quantifier's own scanner (CWE-1333, ironic)

Security review found the rule meant to catch algorithmic-complexity bugs
had one of its own: hasUnboundedBroadQuantifier's negated-class inner
scan walked from each `[^` occurrence to the next `]` (or EOF) with no
bound, while the outer loop only ever advanced by one character — O(n²)
total work on a pattern with many unclosed `[^` runs. Runs unconditionally
inside checkPattern on any `new RegExp('literal string')` argument in any
linted file, before the (cheap) readFileSync data-flow gate — so a single
crafted string literal, no valid regex syntax required, could make
`npm run lint` / CI hang.

Empirically confirmed both the bug and the fix: pre-fix, n=4000/8000/
16000/32000 chars took 30.8/115.6/463.8/1874.3ms (~4x work per 2x n,
quadratic); extrapolated, the 300000-char repro from the finding would
run ~165s. Post-fix (bail the inner scan once units exceeds the rule's
own 1-2-unit scope, rather than continuing to hunt for a closing `]`),
the same 300000-char input runs in 8.7ms via the real rule module,
independently reconfirmed at 18ms via a fresh Linter.verify() call.

New regression row in tests/no-unbounded-quantifier.rule.test.cjs
asserts the RuleTester run on a 50000-char adversarial pattern
completes and returns a defined result — no wall-clock assertion
(CLAUDE.md Clock Seams / local/no-elapsed-assertion).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3415): triage 3 new sites, re-raise ceiling after upstream batch

next merged 12 more PRs during this PR's review. Two consequences:

- tests/edit-phase.test.cjs (fix #3262, unrelated) added 3 new
  content.match(/<tag>([\s\S]*?)<\/tag>/) reads of this repo's own
  workflow .md content — the same Class A pattern as the ~159 sites
  already triaged elsewhere in this PR. Suppressed with the same
  established reason.
- lint-allow-test-rule-refs' ratchet ceiling needed re-raising again
  (301 -> 303) for the same reason as the two prior bumps: organic
  growth from unrelated, already-reviewed PRs landing concurrently,
  not a defect in this branch's own diff.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-14 10:02:28 -04:00
sim
6128f73003 test(#3090): normalize the whole reason line, not just the category token
The allow-test-rule gate keys on identity, and the identity it records is
everything after the colon on the annotation line — not the category token.
Ten annotations carried the canonical category plus a trailing justification
on the same line, so the recorded identity was a prose blob, and where the
prose wrapped it was a sentence fragment: `source-text-is-the-product — the
workflow .md content IS`.

Seven of those ten are ones this branch already rewrote. That pass renamed the
token and left the prose, which is the same error this wave exists to correct,
one level down: the label was fixed without checking what the machine reads.

Justifications move to the following comment line, which the scanner ignores
because it lacks the token. No annotation gains or loses an issue reference, so
no exemption changes compliance status; the allowlist goes 161 to 159 as two
files' duplicate identities collapse.

git-base-branch.test.cjs carried the token twice — once as the real annotation,
once echoed in docblock prose that the line scanner parsed as a second
exemption with a truncated identity. The echo is reworded to drop the literal
token.

intel.test.cjs:1360 was cut off mid-clause with an issue ref appended after the
break; its sentence is restored and the ref kept on the annotation line so it
stays compliant.

Every remaining non-canonical identity is an ESLint RuleTester fixture inside a
`code:` template literal, which the line scanner cannot tell apart from an
annotation. Those two stay grandfathered.

Refs #3057

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 18:28:04 -04:00
sim
7dd9e59f6b test(#3090): stop exempting violations under categories that do not fit
An allow-test-rule annotation citing a category that does not apply is worse
than no annotation, because it reads as reviewed. Eight were confirmed by
reading the assertions each one covered, and auditing the rest found five more
plus one refutation — a converter test whose wording described the wrong
mechanism while the covered assertion genuinely was deployed-text.

The instructive one used the CANONICAL string for the same mistake: STATE.md
command output labelled as a deployed artifact. A canonical string is not
evidence the category fits, which is why normalising strings alone would have
laundered the problem rather than fixed it. Every mapping the audit had inferred
rather than code-verified was spot-checked before rewriting, and the ones that
turned out not to fit were re-annotated rather than relabelled.

Fourteen STATE.md assertions had a typed extractor available all along and now
use it; their annotations came out because nothing needs exempting. Eight
assertions genuinely need a production change first — CLI stdout and stderr with
no structured mode — and are tagged pending-migration-to-typed-ir citing #3090,
which is what that category is for. It had zero real uses before this, while one
file carried a real citation to migration issue #2974 under a non-canonical tag.

Six annotations covered assertions that do no text matching at all. An exemption
for a violation that does not exist is noise that makes the real ones harder to
audit; those are removed.

atomic-write-coverage gains the annotation it always warranted — its own
docstring describes a structural-regression-guard while the file carried none.

Fifty-nine non-canonical strings across roughly thirty files are normalised, and
the allow-test-rule allowlist is regenerated to match. 472 annotations became
463: every one now uses a canonical category, and the two remaining
non-canonical strings are ESLint RuleTester fixtures, not annotations.

Refs #3057

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 17:20:56 -04:00
0xdhx
3274db2757 fix(#2526): remove gsd-ui-auditor's uncallable Playwright-MCP block (#2594)
* fix(#2526): drop gsd-ui-auditor's uncallable Playwright-MCP block

The agent declares `tools: Read, Write, Bash, Grep, Glob, Skill` — no
`mcp__*` grant of any kind — while its body presented a
`<playwright_mcp_approach>` block as the *preferred* capture path. That
branch was unreachable by construction: the availability check had a
fixed answer, the three `mcp__playwright__*` calls could never dispatch,
and the "when Playwright-MCP is NOT available" fallback was the only
branch that ever ran — 39 lines of instruction loaded on every
/gsd-ui-review spawn that also invited the model to claim a capture path
it could not take.

Remove the dead block, leaving the CLI screenshot path as the sole
documented approach.

Guard the class in tests/mcp-tool-inheritance.test.cjs, which already
owns agent MCP-grant parity: the new block generalizes the #1284
researcher check from two agents and one dispatch table to every
agents/*.md and its whole body — no agent may document an
`mcp__<server>__*` namespace absent from its own `tools:` declaration.

Frontmatter is read through the canonical parser
(gsd-core/bin/lib/frontmatter.cjs) rather than a hand-rolled scan, so
inline CSV, block sequences, flow arrays, quoted scalars and full-line
comments are handled by construction; inline comments inside a scalar
survive that parser, so they are stripped explicitly. The check is
server-level by design, ignores prose metavariables like `mcp__X__*`,
matches hyphenated server ids, and carries a discovery guard plus
negative controls for every documented boundary so it cannot decay into
a vacuous pass.

The session-level Playwright-MCP pass in gsd-core/workflows/ui-review.md
is deliberately untouched — workflow files carry no fixed allowlist, so
their availability check is genuinely runtime-detected and honest.

Fixes #2526

* chore(#2526): set changeset fragment pr to 2594

The fragment shipped with the documented `pr: 0` placeholder because the
PR number does not exist until the PR is opened, and
scripts/changeset/parse.cjs rejects `pr <= 0`. Now that the PR is open,
set the real number so changeset-lint passes.

* test(#2526): cover the two-char server-id boundary of the metavariable exclusion

The length-1 "prose metavariable" exclusion was tested at length=1 and at
real ids (>=3 chars), but never at length=2 — the limit+1 boundary where a
server id starts being recognized. Review finding on #2594: `mcp__ab__foo`
in a body with no grant must flag `['ab']`.

* fix(#2526): treat a bare mcp__* grant as covering every server

`grantedServers()` stripped `mcp__*` to the empty string and dropped it via
`if (server)`, so an allowlist that grants every MCP server read as granting
none — and the guard then fired against a body the grant plainly covered.
That is the one input shape that inverts the check, turning it against a
correct agent rather than merely missing a bad one.

A `/^mcp__\*+$/` token now sets a GRANT_ALL sentinel that short-circuits
`ungrantedServers()`. The sentinel `*` is outside REFERENCE_RE's character
class, so no body reference can collide with it. A bare `mcp__` with no
wildcard stays a typo rather than a grant and keeps failing closed.

No agent uses the `mcp__*` spelling today, so this was latent rather than
live. Two negative controls pin both halves.

* fix(#2526): scan the frontmatter description for MCP references too

`ungrantedServers()` scanned `stripFrontmatter(content)` only, so an
`mcp__foo__bar` reference in the `description` field escaped the check. That
field ships with the agent and the dispatcher reads it, which makes a dead
reference there exactly as dead as one in the body.

Only `description` is added to the scanned surface, never the whole
frontmatter: `tools:` is the grant list itself, so scanning it would let every
allowlist satisfy itself and turn the guard vacuous. A negative control pins
that boundary alongside the new positive case.

All 34 per-agent tests still pass with the wider surface, so no live agent
verdict changes — this was latent.

* test(#2526): give multi-character placeholders a convention the checker knows

The metavariable exclusion is `length === 1`, so the natural placeholders
`mcp__SRV__*` and `mcp__SERVER__*` were flagged as real references — and the
failure message then offered an author two remedies ("grant the namespace or
drop the block") that both misread what they wrote.

Adopts the angle-bracket half of the suggested fix: `mcp__<SERVER>__*` is the
sanctioned multi-character placeholder, exempt by construction because `<` is
outside the reference pattern's character class. This pins an existing
property rather than adding a special case.

Declines the all-caps half. An all-caps exemption would be a false NEGATIVE
for any real server spelled in caps, and a guard that misses a dead reference
fails in exactly the direction this check exists to prevent. The bare-caps
form keeps firing; the message now names the convention as a third remedy.

Also corrects "grants neither" in that message, which was wrong for any count
other than two.

* test(#2526): pin the zero-length server id, completing the boundary triple

`mcp____foo` yields `[]`, but for a different reason than the length-1 case:
it is unrepresentable by `/mcp__([A-Za-z0-9_-]+?)__/g` since `+?` requires at
least one character, so the pattern skips it before the metavariable
exclusion is ever consulted. Pinning limit-1 completes the 0/1/2 boundary
rule on its own terms and records which mechanism owns the case.

* docs(#2526): correct every drifted AGENTS.md Tools row, not just the one

The review asked for the one-line `gsd-ui-auditor` correction (missing
`Skill`). Sweeping the defect class first — every `**Tools**` row in
docs/AGENTS.md against its agent's `tools:` frontmatter — found it was 26 of
34 rows, so the one-line framing was the reviewer's premise rather than the
population.

Breakdown of the 26:
  * 21 omitted `Skill`, 6 omitted `Edit` (overlapping) — under-promises, the
    same drift class as #2526 but in the harmless direction.
  * 8 wrote `mcp (context7)` as shorthand while frontmatter granted up to 8
    servers (firecrawl, exa, tavily, ref, jina, perplexity, both context7s).
  * 1 was actively wrong: gsd-debug-session-manager documented `Task`, a tool
    name that no longer exists — the #2526 shape at the doc layer, naming a
    capability that cannot dispatch.

Every row is now the frontmatter `tools:` value verbatim, which is also what
makes the parity guard in the following commit non-brittle. The diff is
26 insertions / 26 deletions, all Tools rows.

* test(#2526): guard AGENTS.md Tools rows against agent frontmatter

The 26 corrected rows in the previous commit were free to drift because
nothing asserted the role card and the frontmatter agreed — the same reason
the #2526 block itself survived. Correcting them without an invariant just
resets the clock.

Lands in agent-classification-parity.test.cjs rather than a new file: that
suite already owns docs/AGENTS.md as a contract surface, already carries the
`allow-test-rule` exemption for treating the doc as the product, and file
count is the unit of CI overhead (docs/TESTING-SUITES.md).

Compares the row to the frontmatter value VERBATIM, not as a set — a set
comparison would keep accepting the "mcp (context7)" shorthand that hid eight
grants behind one, which is the under-documentation half of the drift.

Carries the same discovery guard #2526's own check uses: a section with a
granted `tools:` but no **Tools** row fails loudly, so deleting a row cannot
silently retire its assertion. Both halves are negative-controlled — against
the pre-fix doc it fails naming 26 rows (gsd-ui-auditor:339 among them), and
with a row deleted it fails on the missing-row assertion.

* chore(#2526): note the AGENTS.md drift correction in the changeset

The role cards are user-visible, and 26 of them documented a tool set the
agent did not have. Type, `pr: 2594`, and the trailing `(#2526)` are
unchanged.

* fix(#2526): use a CRLF-safe split in the AGENTS.md Tools-row guard

`lint-tests` (npm run lint:ci) rejected `rawAgentsMd.split('\n')` under the
repo's local/no-crlf-fragile-split rule: Windows autocrlf yields CRLF, so a
trailing \r rides into the parsed line. Switched to `/\r?\n/`.

Caught by CI on the round-3 push before the response comment went out.

* docs(#2526): correct the drift tallies stated in c5a9607e

Re-derived the census programmatically from the pre-fix doc instead of by eye.
The 26-of-34 headline was right; the breakdown was not.

  Skill omitted   21 -> 22
  Edit omitted     6 ->  7
  "mcp (context7)" shorthand   8 rows -> 7 rows

The 8 was conflating two things: 8 rows omitted MCP grants entirely, but only
7 of them used the "mcp (context7)" shorthand — gsd-executor listed no MCP at
all. Also names the one `Agent` omission (gsd-debug-session-manager, the row
that still read `Task`).

c5a9607e's message keeps the wrong numbers rather than rewriting a pushed
branch mid-review; the test comment and changeset are the durable statements
and both are corrected here.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-28 18:02:01 -04:00
Behruz Nassre Esfahani
fc5ca178a2 test(#1178): consolidate duplicated agent-roster helper into tests/helpers (#1420)
* test(#1178): consolidate duplicated agent-roster helper into tests/helpers

The "list gsd-*.md agent files, strip .md, sort" derivation was hand-duplicated
across the suite (two listAgentFiles(), an identical agentFilesOnDisk(), and
inline readdir blocks). Add tests/helpers/agent-roster.cjs exporting
listAgentFiles(agentsDir?) and route the genuinely-identical source-roster sites
through it. Semantically-different sites (installed-dest dirs, absolute-path
returns, .toml-inclusive Codex rosters, full-.md-filename readers, the uniform
multi-family inventory table) are left intact, each with a one-line comment.

Test-only; no production code touched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1178): note AGENTS_DIR export is for future call sites

Review nit: clarify that the currently-unused AGENTS_DIR export is intentional
— available for future tests needing the canonical source agents path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-23 19:34:08 -04:00
Tom Boucher
a19a709e62 test(#1171): add agent-classification parity guard (#1176)
Makes docs/AGENTS.md section structure the single source of truth for the
primary-vs-advanced agent classification and fails when docs/INVENTORY.md's
"Primary doc" column or the AGENTS.md prose counts drift from it.

The classification ("primary" = full role card, "advanced stub" = concise
stub) was hand-duplicated across three doc surfaces with no enforcement.
This drift guard derives the expected class from AGENTS.md section placement
and cross-checks the INVENTORY.md column, the prose counts, the parenthetical
advanced-agent list, and full roster completeness (21 primary / 12 advanced
/ 0 inventory-only today).

No agent frontmatter field is added (avoids the capability-registry `tier`
collision and the research-profiles.cjs ripple); the classification stays
documentary, enforced from where it is defined.

Closes #1171

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 11:28:05 -04:00