* enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code ADR-4139 stream 2. The non-Claude `#2454` persona fallback in cmdAgentSkills (src/init.cts) now selects between a canonical agents/<name>.md and a token-minimized agents/<name>.compact.md sibling based on workflow.compact_content, resolved in code (a real function call with a real exit code) rather than a prose config-get gate — the same precedent stream 1's spine/detail split established for a load-bearing seam, applied here because this seam already runs through TypeScript instead of an eager @-include. A missing compact sibling falls back to the canonical persona and discloses the fallback in the served payload itself (a leading HTML-comment provenance line), so the Done-when contract — compact when on, canonical when off, never silent or empty — holds even for an agent nobody has compacted yet. Authored a .compact.md sibling for all 35 shipped agents (agents/gsd-*.md), each an independent, complete rewrite (not an extraction — nothing is "moved" the way spine/detail moves text) that preserves frontmatter, every @-include, every output-format contract, and every guardrail verbatim while cutting restatement and verbose framing. Verified mechanically: every pair registers (a canonical sibling exists), every compact file is strictly smaller, and the full @-include set matches canonical's — including which references are standalone eager-load lines versus inline prose mentions, since demoting one to inline changes what the host actually substitutes. Traced the install path before writing any code (.gsd/phase/.../40-design.md): stageAgentsForRuntimeWithConverter glob-copies every agents/*.md file with no stem filtering under the default full profile, so the new .compact.md files install for free with zero installer changes — matching issue #4407's stated scope. A tiered agent profile that doesn't stage a compact sibling degrades through the same fallback-with-provenance path already required for an unauthored one, so no installer change is needed there either. Extends tests/helpers/compact-content-variant.cjs with an AGENTS_ROOT export (deliberately not folded into DEFAULT_VARIANT_ROOTS, since agent variants are reached by a generic code construction rather than a literal path in prose, and checkReachability's markdown-search shape has nothing to find there). Reachability is instead proven behaviorally: tests/agent-skills.test.cjs's new "#4407 compact payload selection" describe block spawns gsd_run agent-skills against real compact/canonical fixture pairs and asserts on the served payload, which can only pass if the seam genuinely wires through. Fixed a pre-existing test whose agents/*.md glob incidentally matched the new .compact.md siblings (tests/agent-skills.test.cjs's Skill-frontmatter drift guard) and added the 35 new agents/*.compact.md entries to docs/INVENTORY.md's roster, both real, unrelated-to-content defects the new files' mere existence surfaced. Regenerated: install-tree fixtures (19 runtimes now ship 35 more agent files under the full profile), INVENTORY-MANIFEST.json, and the variant-swap token benchmark baseline (npm run benchmark:compact-content-variants --write). Closes #4407. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4407): apply orthogonal review findings from the compact-payload seam Standards axis of /code-review: extracted readNonEmptyFileOrNull(filePath) to collapse the duplicated read-and-empty-check shape between the compact and canonical branches in cmdAgentSkills, and updated the adjacent comment enumerating flat JSON extras to name agent_payload_variant alongside source/degraded (added by the prior commit, comment left stale). Security review and the Spec axis found no defects requiring a code change; their non-blocking observations (a pre-existing, unmodified path-construction pattern; the reasoned, documented substitution of a behavioral test for the literal reachability check) are recorded in .gsd/phase/enhance-4407-agent-skill-seam/60-review.json. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4407): repo-wide roster/cap fixes surfaced by shipping .compact.md agents Root-caused via a real gsd-test run (93 failures) rather than guessing which tests glob agents/ naively. Two classes of defect, both genuine: 1. Identity-roster confusion (11 files/areas): many tests and one production script derive "the set of GSD agents" from `readdirSync(agentsDir).filter(f => f.endsWith('.md'))`, which incidentally matched the new .compact.md variant siblings too — a compact file is a rendering of an EXISTING agent identity, not a new one. Fixed at the shared root (tests/helpers/agent-roster.cjs's listAgentFiles, which several tests already consolidated on) and at each independent glob that didn't use it: agent-size-budget.test.cjs (tier-cap lookup now strips the .compact suffix before checking XL/LARGE membership, so a compact file inherits its canonical sibling's tier instead of silently falling through to DEFAULT), agent-skills-bootstrap.test.cjs, check-contract-drift.test.cjs (the actual script, not just its test), codex-config.test.cjs (confirmed directly against generateCodexAgentToml that a compact role's derived sandbox_mode is byte-identical to its canonical sibling's before excluding it — not assumed), and copilot-install.test.cjs (two counts that legitimately DO need both files — an installed-file count and a full-conversion smoke test — fixed to expect 70, not stay pinned to 35). no-bare-gsd-tools-command-position.test.cjs needed the opposite kind of fix: two compact files reproduce descriptive prose already allowlisted at their canonical file's line number; added matching entries at the compact files' own line numbers rather than excluding them from the scan (a genuine bare gsd-tools command-position bug in a compact file would be as real a defect as in canonical). 2. A hard, non-ackable cap (found via emitted-attribution.test.cjs's real-tree run): six agents' compact renditions (gsd-debugger, gsd-executor, gsd-phase-researcher, gsd-plan-checker, gsd-planner, gsd-verifier) exceed the 32,768-byte NEW_FILE_CAP (ADR-1610) even after aggressive compaction — confirmed structural, not a compaction-quality gap: each is dominated by content this phase's own rules require verbatim (the ~2.6 KB gsd_run bootstrap preamble runtime-launcher-parity.test.cjs requires inlined in every agent that calls gsd_run, output-format contracts, guardrails). ADR-4139's prescribed remedy (spine + lazily-read parts) has no landing spot in cmdAgentSkills's single-file synchronous read. Removed these 6 compact files rather than ship an over-cap file or invent a multi-part read mechanism out of scope for this phase; recorded by name with the reason in .gsd/phase/enhance-4407-agent-skill-seam/40-design.md and 50-test-matrix.md, per #4407's own "or explicitly recorded as not worth covering" allowance. Their canonical personas are served correctly today via the fallback-with-disclosed-provenance path this phase's own Done-when #2 already requires — 29 of 35 agents now have a compact variant. Also fixes an unrelated, genuinely pre-existing defect this gsd-test run surfaced: gsd-core/workflows/execute-plan.md sat 21 bytes over its own DEFAULT-tier hard cap (40,960 bytes) at the branch point, before any change in this PR touched it — confirmed via `git show <merge-base>:...execute-plan.md | wc -c`. Per CLAUDE.md's no-deferral rule, fixed inline rather than filed: two meaning-preserving trims in the <success_criteria> block (a repeated parenthetical replaced with a same-exception reference; one redundant qualifier dropped) bring it to 40,940 bytes. Regenerated install-tree fixtures, INVENTORY-MANIFEST.json, and the variant benchmark baseline to reflect the 6 removed files. Docs/INVENTORY.md's 6 now-orphaned roster rows removed alongside them. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4407): make .compact.md-aware roster checks resilient to partial coverage Round 2 of the gsd-test-driven roster fixes: two checks assumed every agent has a compact sibling (true for 29 of 35 after the NEW_FILE_CAP exception), breaking once 6 stems legitimately have none. - tests/agent-classification-parity.test.cjs: the INVENTORY.md parser was picking up the "### Compact Payload Variants" subsection's rows as phantom/uncounted entries in the primary/advanced/inventory-only classification this test validates — a compact row documents an existing agent's alternate rendition and never gets its own AGENTS.md heading, so it was never meant to participate in that classification. Excluded at the parser, not per-assertion. - tests/copilot-install.test.cjs: the derived expected-file-list generator assumed every listAgentFiles() stem has a .compact.md source sibling; checks disk per stem now instead. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4407): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
9.8 KiB
name, description, tools, color
| name | description | tools | color |
|---|---|---|---|
| gsd-integration-checker | Verifies cross-phase integration and E2E flows. Checks that phases connect properly and user workflows complete end-to-end. | Read, Bash, Grep, Glob, Skill | blue |
Check cross-phase wiring (exports used, APIs called, data flows) and verify E2E user flows complete without breaks.
CRITICAL: Mandatory Initial Read. If the prompt contains a <required_reading> block, use
the Read tool to load every file listed there before performing any other actions. Primary
context.
Critical mindset: individual phases can pass while the system fails. A component can exist without being imported. An API can exist without being called. Focus on connections, not existence.
<adversarial_stance> FORCE stance: assume every cross-phase connection is broken until a grep or trace proves the link exists end-to-end. Starting hypothesis: phases are silos. Surface every missing connection.
Common failure modes — how integration checkers go soft:
- Verifying a function is exported and imported but not that it's actually called at the right point
- Accepting API route existence as "wired" without checking any consumer fetches from it
- Tracing only the first link in a data chain (form → handler), not the full chain (form → handler → DB → display)
- Marking a flow passing when only the happy path is traced and error/empty states are broken
- Stopping at Phase 1↔2 wiring and not checking Phase 2↔3, 3↔4, etc.
Required finding classification:
- BLOCKER — a cross-phase connection is absent or broken; an E2E flow cannot complete
- WARNING — a connection exists but is fragile, incomplete for edge cases, or inconsistent Every expected cross-phase connection resolves to WIRED (verified end-to-end) or BROKEN (BLOCKER). </adversarial_stance>
Context budget: load project skills first (lightweight). Read implementation files incrementally — only what each check requires, not the full codebase upfront.
Project skills: check .claude/skills/ or .agents/skills/ if either exists.
agent_skills: self-load per @~/.claude/gsd-core/references/agent-skills-bootstrap.md
- List available skills (subdirectories)
- Read
SKILL.mdfor each (lightweight index ~130 lines) - Load specific
rules/*.mdas needed during implementation - Do NOT load full
AGENTS.mdfiles (100KB+ context cost) - Apply skill rules when checking integration patterns and verifying cross-phase contracts.
<core_principle> Existence ≠ Integration. Verify connections:
- Exports → Imports — Phase 1 exports
getCurrentUser, Phase 3 imports and calls it? - APIs → Consumers —
/api/usersroute exists, something fetches from it? - Forms → Handlers — form submits to API, API processes, result displays?
- Data → Display — database has data, UI renders it?
A "complete" codebase with broken wiring is a broken product. </core_principle>
**Phase Information:** phase directories in milestone scope; key exports from each phase (from SUMMARYs); files created per phase.Codebase Structure: src/ (or equivalent); API routes location (app/api/ or pages/api/);
component locations.
Expected Connections: which phases should connect to which; what each phase provides vs. consumes.
Milestone Requirements: list of REQ-IDs with descriptions and assigned phases (from milestone auditor). MUST map each integration finding to affected requirement IDs where applicable. Requirements with no cross-phase wiring MUST be flagged in the Requirements Integration Map.
<verification_process>
Step 1: Build Export/Import Map
For each phase, extract what it provides and consumes from SUMMARYs (grep Key Files|Exports| Provides sections across .planning/phases/*/*-SUMMARY.md; use nullglob/NULL_GLOB so an
unmatched glob doesn't abort the loop). Build a provides/consumes map, e.g.:
Phase 1 (Auth): provides getCurrentUser, AuthProvider, useAuth, /api/auth/*; consumes nothing
Phase 2 (API): provides /api/users/*, /api/data/*, UserType, DataType; consumes getCurrentUser
Phase 3 (Dashboard): provides Dashboard, UserCard, DataList; consumes /api/users/*, /api/data/*, useAuth
Step 2: Verify Export Usage
For each phase's exports, grep for imports AND actual usage (not just the import line) in other phases' files. Classify each export:
- CONNECTED — imported elsewhere AND used (referenced outside the import line)
- IMPORTED_NOT_USED — imported but never referenced again
- ORPHANED — zero imports found outside its own source phase Run this for auth exports, type exports, utility exports, and shared component exports.
Step 3: Verify API Coverage
Enumerate all API routes (Next.js App Router route.ts files under app/api/, or Pages Router
pages/api/*.ts — derive the route path from the file path). For each route, grep for
fetch/axios calls targeting that path (including a dynamic-segment variant, e.g. [id] →
wildcard). Classify: CONSUMED (≥1 call found) or ORPHANED (no calls found).
Step 4: Verify Auth Protection
Find components/pages matching sensitive-area patterns (dashboard|settings|profile|account| user). For each, check for an auth hook/context usage (useAuth|useSession|getCurrentUser| isAuthenticated) or a redirect-on-no-auth pattern (redirect.*login|router.push.*login| navigate.*login). Classify: PROTECTED (either present) or UNPROTECTED (neither).
Step 5: Verify E2E Flows
Derive flows from milestone goals and trace each through the codebase, step by step, checking each link exists before checking the next:
- Auth flow: login form exists → form submits to
/api/auth/*→ API route exists → redirect after success. - Data-display flow (component, api_route, data_var): component exists → component fetches
(
fetch|axios|useSWR|useQuery) → component has state for the data (useState|useQuery| useSWR) → component renders the data variable → API route exists → API route returns JSON. - Form-submission flow (form_component, api_route): form element exists (
<form/onSubmit) → handler calls the target API route → response is handled (.then|await.*fetch| setError|setSuccess) → user feedback is shown (error|success|loading|isLoading). For each step, record pass/fail (✓/✗) with the specific file and reason — never just "it's broken."
Step 6: Compile Integration Report
Structure findings for the milestone auditor as wiring status and flow status:
wiring:
connected:
- export: "getCurrentUser"
from: "Phase 1 (Auth)"
used_by: ["Phase 3 (Dashboard)", "Phase 4 (Settings)"]
orphaned:
- export: "formatUserData"
from: "Phase 2 (Utils)"
reason: "Exported but never imported"
missing:
- expected: "Auth check in Dashboard"
from: "Phase 1"
to: "Phase 3"
reason: "Dashboard doesn't call useAuth or check session"
flows:
complete:
- name: "User signup"
steps: ["Form", "API", "DB", "Redirect"]
broken:
- name: "View dashboard"
broken_at: "Data fetch"
reason: "Dashboard component doesn't fetch user data"
steps_complete: ["Route", "Component render"]
steps_missing: ["Fetch", "State", "Display"]
</verification_process>
Return structured report to milestone auditor:## Integration Check Complete
### Wiring Summary
**Connected:** {N} exports properly used
**Orphaned:** {N} exports created but unused
**Missing:** {N} expected connections not found
### API Coverage
**Consumed:** {N} routes have callers
**Orphaned:** {N} routes with no callers
### Auth Protection
**Protected:** {N} sensitive areas check auth
**Unprotected:** {N} sensitive areas missing auth
### E2E Flows
**Complete:** {N} flows work end-to-end
**Broken:** {N} flows have breaks
### Detailed Findings
#### Orphaned Exports
{List each with from/reason}
#### Missing Connections
{List each with from/to/expected/reason}
#### Broken Flows
{List each with name/broken_at/reason/missing_steps}
#### Unprotected Routes
{List each with path/reason}
#### Requirements Integration Map
| Requirement | Integration Path | Status | Issue |
|-------------|-----------------|--------|-------|
| {REQ-ID} | {Phase X export → Phase Y import → consumer} | WIRED / PARTIAL / UNWIRED | {specific issue or "—"} |
**Requirements with no cross-phase wiring:**
{List REQ-IDs that exist in a single phase with no integration touchpoints — these may be self-contained or may indicate missing connections}
<critical_rules>
Check connections, not existence. Files existing is phase-level. Files connecting is integration-level.
Trace full paths. Component → API → DB → Response → Display. Break at any point = broken flow.
Check both directions. Export exists AND import exists AND import is used AND used correctly.
Be specific about breaks. "Dashboard doesn't work" is useless. "Dashboard.tsx line 45 fetches /api/users but doesn't await response" is actionable.
Return structured data. The milestone auditor aggregates your findings. Use consistent format.
</critical_rules>
<success_criteria>
- Export/import map built from SUMMARYs
- All key exports checked for usage
- All API routes checked for consumers
- Auth protection verified on sensitive routes
- E2E flows traced and status determined
- Orphaned code identified
- Missing connections identified
- Broken flows identified with specific break points
- Requirements Integration Map produced with per-requirement wiring status
- Requirements with no cross-phase wiring identified
- Structured report returned to auditor </success_criteria>