* enhance(#4406): the lazily-read remainder and the artifact templates ADR-4139 Decision 3, Phase 6 of the #4139 Compact Content epic. Covers stream 1b (gsd-core/workflows/<name>/{modes,steps,templates}/*.md) and stream 4 (gsd-core/templates/**) with a variant-swap mechanism, confirmed with the user: two independent, complete files per covered path (canonical + .compact.md sibling), with the gate picking which one gets Read at the call site. This is a different shape from Phase 5's spine+detail partition, and is safe here specifically because these files are already reached only by a runtime Read — a missed Read already means zero overlay content today, with or without workflow.compact_content, so selecting between two independently-complete files introduces no new failure mode (documented in gsd-core/references/compact-content-gate.md's new "Streams 1b and 4" section). Disposition, after inspecting every candidate rather than trusting a byte-size threshold (same rigor Phase 5 applied to review.md): - Stream 1b: 1 of 78 files compacted (help/modes/full.md, a user-facing reference doc emitted verbatim, not orchestrator instruction). The other 9 size-threshold candidates are dominated by fail-closed guards, exact CLI invocations, or output-format contracts (AskUserQuestion blocks) — recorded not-worth-compacting, same reasoning as Phase 5's review.md. - Stream 4: a ground-truth reachability audit replaced the initial size-only candidate list. Two files (summary.md, user-setup.md) got compact variants; a third (spec.md) was drafted, then dropped after discovering its only two call sites are eager @-includes, not a runtime Read — stream-1 material hiding under gsd-core/templates/, not stream-4's actual mechanism. summary.md itself has 3 eager call sites and only 1 genuine runtime-Read call site (execute-plan.md); only that one was wired, so the compact variant's savings apply to the sequential single-plan execution path only. - Discovered while auditing reachability: 12 gsd-core/templates/** files with zero references anywhere in workflow/agent/command prose, compiled source, or tests — dead scaffolding predating this phase. Deleted in this same PR per this repo's no-defer policy, after re-verifying against a computed path.join(...) pattern (not just a plain-string search) that nearly caused two genuinely load-bearing templates (user-profile.md, dev-preferences.md) to be misclassified as dead. New checker (tests/helpers/compact-content-variant.cjs): registration, reachability, protected-content-preserved, size-smaller — replacing Phase 3/5's disjointness/completeness checks, which assume a partition rather than two deliberately-overlapping documents. The reachability check's own "unprefixed match" guard had a real bug (rejected the repo's own `~/.claude/gsd-core/...` convention), caught by running it against the already-wired help/modes/full.compact.md pair rather than only synthetic fixtures — fixed to anchor on the nearest `gsd-core` path segment instead. Template consumer parity (tests/compact-content-template-variant-parity.test.cjs): proves each compact variant's `## File Template` fenced block — the actual output-format contract a generated SUMMARY.md/USER-SETUP.md is parsed against — is byte-identical to the canonical file, then runs the one real deterministic consumer (gsd-core/bin/lib/coverage.cjs's classifyContent, backing `gsd-tools uat classify-coverage`) against content built from that shared contract. Added a sibling benchmark script (scripts/benchmark-compact-content-variants.cjs) rather than extending the existing spine/detail one — different data shape, and the existing script's own contract deliberately isolates it from a test-only helper's shape changing. Emitted-drift acknowledgement: not needed. Every changed/added path in this diff is hand-authored and present in the diff itself, so diffEmitted's attribution loop resolves `via` to the path's own source before reaching the ack-lookup branch (same reasoning Phase 5 verified for its own diff). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * enhance(#4406): address code-review findings on the variant-swap gate - docs/CONFIGURATION.md and gsd-core/references/planning-config.md's workflow.compact_content entries described only the spine+detail mechanism (Phase 5) and were missing this phase's variant-swap mechanism and its benchmark:compact-content-variants script entirely — required since this PR's changeset is type Added (CLAUDE.md's "Missing Docs for Changesets" rule). Both now describe both mechanisms and which call sites are wired. - Added the missing RED^-1/no-op fixture for checkProtectedContentPreserved: a canonical file with zero <!-- gsd:protected --> blocks must be a no-op, not a violation — the only branch of that function the existing fixtures didn't exercise. - Collapsed findCompactFiles/findMarkdownFiles in tests/helpers/compact-content-variant.cjs into one findFilesWithSuffix helper — the two were identical recursive walks differing only in the extension predicate (minor Duplicated-Code finding). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4406): restore copilot-instructions.md, a false-positive dead-template classification gsd-test caught this, not static analysis: 10 real failures in tests/copilot-install.test.cjs, tests/installer-migration-install.integration.test.cjs, and tests/repo-layout.test.cjs — all downstream of bin/install.js's Copilot install path, which does fs.readFileSync(path.join(targetDir, 'gsd-core', 'templates', 'copilot-instructions.md')) after copying gsd-core/templates/** into the target project, then merges it into both .github/copilot-instructions.md and (local installs) AGENTS.md. The reachability audit that flagged this file as dead checked src/*.cts and gsd-core/bin/*.cjs but never the repo-root bin/install.js — a separately maintained installer bundle outside the src/-to-gsd-core/bin/lib/ compiled-output convention. The fs.existsSync guard around that read degrades to a silent skip rather than a crash when the template is missing, which is why this surfaced only once the real E2E install test ran, not from any static check. Re-verified the remaining 11 deleted filenames against bin/install.js specifically (plain substring and quoted-filename search) before trusting that list — all 11 have zero hits there, confirmed dead by the same standard this one file failed. Regenerated the installer emitted-tree goldens (tests/fixtures/install-tree/*.json) to reflect the restored file, and corrected the "Removed" changeset (jolly-lynx-sprint.md) and the phase design doc from 12 to 11 deleted files. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Emitted-Drift-Ack-Growth: execute-plan.md — call-site wiring for the summary.md and user-setup.md .compact.md variants Emitted-Drift-Ack-Growth: help.md — call-site wiring for full.compact.md, same variant-resolution rule * docs(#4406): backfill changeset PR numbers Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4406): resolve removed-but-needed lint findings on the dead-template deletion CI's own full-test matrix (not gsd-test's matrix, which does not run this check) caught 4 more false-positive dead-template classifications via tests/removed-but-needed-lint.test.cjs / scripts/lint-removed-but-needed.cjs — a literal, word-boundary basename check across .github/workflows/, gsd-core/, and docs/ (excluding docs/adr/** and docs/research/**) for every file a PR deletes. It has no semantic awareness, so a deleted template's basename colliding with something else entirely still fires: - claude-md.md: gsd-core/templates/README.md had a stale table row claiming /gsd-profile reads this template to generate CLAUDE.md. Verified false (no code reads it anywhere, same search that already covered bin/install.js) — fixed the row to *(inline)*, matching every other command-generated artifact in that table. File stays deleted. - codebase/testing.md: collided with docs/guides/testing.md, an illustrative example row in docs-update.md's sample output table (an unrelated real generated-docs path). Swapped the example topic to "contributing" — the row is illustrative, any topic works. File stays deleted. - codebase/architecture.md, codebase/stack.md: collided with docs/reference/ planning-artifacts.md's directory listing of a user's own generated .planning/codebase/architecture.md and stack.md output — the same semantic mismatch already investigated and dismissed as unrelated earlier in this phase's audit, now caught by a gate instead of judgment. That listing repeats across 5 locale copies of the doc. - continue-here.md: collided with the real .continue-here.md pause-work artifact, referenced across 15+ locale and workflow files. For the last two, the lint's own error message offers "restore the file or update every consumer in the same commit." Rewording 15+ files across languages I cannot verify translation quality for, to shave 2 already-tiny templates that were merely presumed dead, is disproportionate to this PR's actual scope — restored codebase/architecture.md, codebase/stack.md, and continue-here.md instead, and corrected docs/ARCHITECTURE.md's Templates section accordingly. Final confirmed-dead set: claude-md.md, codebase/concerns.md, codebase/conventions.md, codebase/integrations.md, codebase/structure.md, codebase/testing.md, debug-subagent-prompt.md, discovery.md — 8 files, down from the original 12. Verified locally: GSD_REMOVED_BUT_NEEDED_BASE=next node scripts/lint-removed-but-needed.cjs now passes clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Emitted-Drift-Ack-Growth: docs-update.md — swapped an illustrative example-table topic (testing -> contributing) to avoid a removed-but-needed basename collision with the deleted codebase/testing.md template; net +10 bytes * fix(#4406): split codex-config.test.cjs to fix a genuine Windows CI timeout Root cause of the `full test (windows-latest, 24, shard 2/3)` failure the user asked to be actually fixed, not just re-run past: PR #4497 (landed 2026-09-07, one day before this PR's CI run) isolated tests/codex-config.test.cjs into its own dedicated chunk because its measured weight (17.87, ~45% of the post-cut Windows budget) made it unsafe to share a chunk with any other file. That isolation was necessary but not sufficient — even alone, with zero companion-file contention, the file's real Windows execution time sits right at the 600s per-chunk ceiling. Two independent CI runs on two unrelated PRs (this one and #4154) were both killed within ~1.4s of the identical 600000ms mark — not random contention, a deterministic near-miss the isolation fix couldn't address because it never reduced the file's own cost, only removed the risk of a companion file's cost stacking on top of it (which the PR #4497 comment explicitly anticipated: "if a future profiling pass genuinely speeds up codex-config.test.cjs itself, this isolation can be revisited"). The file itself explains why it's this heavy: 11,262 lines / 433 tests / 79 describe blocks, accumulated over dozens of bug-fix PRs (#2695, #2760, #3245, #3285, #3346, #3426, #3427, #3562, #3566, #3582, #3808, and more), several of which are explicitly documented as "folded" in from separate files that were never actually split back out ("Verified non-duplicate against both the pre-existing target and the other three folded sources"). Split into 4 files by top-level AST statement boundaries (never a naive column-0 regex — an early attempt at that overcounted 79 apparent "describe(" matches when only 21 are genuinely top-level; the rest are nested inside a handful of large folded-in blocks, which a regex can't tell apart from real top-level statements). Verified lossless twice: the split script asserts byte-for-byte reconstruction of every source character, and independently, total test()/describe() call counts match exactly between the original file and the sum across all 4 new files (433/79 both sides). Each new file carries the complete original shared header (imports/helpers) for safety; per-file unused-import warnings from that duplication are resolved via ESLint-precise alias renames (`{ foo: _foo }`, the standard form for an intentionally-unused destructured binding — never a bare `{ _foo }`, which would destructure a different, nonexistent property). No change needed to scripts/run-tests.cjs's ISOLATED_HEAVY_FILES or its pinned test in tests/run-tests-harness.test.cjs: the file that keeps the original name (tests/codex-config.test.cjs) is now only ~28% of the original's size and safely isolated in its own chunk as before; the other three new files re-enter normal weight-balanced packing, none individually close to disproportionate. Confirmed no other file hardcodes the hardcoded filename anywhere that would silently stop these tests from running (the CI test-selection scripts determine scope algorithmically, not by literal filename). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
9.4 KiB
Summary Template
Template for .planning/phases/XX-name/{phase}-{plan}-SUMMARY.md - phase completion documentation.
File Template
---
phase: XX-name
plan: YY
subsystem: [primary category: auth, payments, ui, api, database, infra, testing, etc.]
tags: [searchable tech: jwt, stripe, react, postgres, prisma]
# Dependency graph
requires:
- phase: [prior phase this depends on]
provides: [what that phase built that this uses]
provides:
- [bullet list of what this phase built/delivered]
affects: [list of phase names or keywords that will need this context]
# Actuals (#2632) — pairs with the plan's `estimate` to calibrate future estimates.
# Same estimateTokens scale (chars/4 over the realized diff), never a harness token count.
actuals:
tokens: [chars/4 over files actually changed]
tasks: [tasks completed]
commits: [commits made]
# Tech tracking
tech-stack:
added: [libraries/tools added in this phase]
patterns: [architectural/code patterns established]
key-files:
created: [important files created]
modified: [important files modified]
key-decisions:
- "Decision 1"
- "Decision 2"
patterns-established:
- "Pattern 1: description"
- "Pattern 2: description"
requirements-completed: [] # REQUIRED — Copy ALL requirement IDs from this plan's `requirements` frontmatter field.
# Coverage metadata (#1602) — one entry per shipped deliverable. Drives DETERMINISTIC UAT routing in verify-work.
# OMIT this whole block for legacy/prose-only SUMMARYs — verify-work then falls back to the ## Accomplishments bullets
# (byte-identical behavior for un-migrated phases). See <coverage_guidance> below for the contract.
coverage:
- id: D1
description: "[deliverable in human-readable form — what would have been a prose ## Accomplishments bullet]"
requirement: "[REQ-ID from this plan's `requirements`, or omit if none]"
verification:
- kind: unit # unit | integration | e2e | automated_ui | manual_procedural | other
ref: "[tests/path.test.ts#test name | playwright:shot.png | command invocation]"
status: pass # pass | fail | unknown — from the latest run
human_judgment: false # REQUIRED boolean. false => may auto-pass IF every verification status is `pass`.
- id: D2
description: "[a deliverable that needs a human to sign off]"
verification: []
human_judgment: true
rationale: "[REQUIRED when human_judgment: true — why automation is insufficient]"
# Metrics
duration: Xmin
completed: YYYY-MM-DD
status: complete
---
# Phase [X]: [Name] Summary
**[Substantive one-liner describing outcome - NOT "phase complete" or "implementation finished"]**
## Performance
- **Duration:** [time] (e.g., 23 min, 1h 15m)
- **Started:** [ISO timestamp]
- **Completed:** [ISO timestamp]
- **Tasks:** [count completed]
- **Files modified:** [count]
## Accomplishments
- [Most important outcome]
- [Second key accomplishment]
- [Third if applicable]
## Task Commits
Each task was committed atomically:
1. **Task 1: [task name]** - `abc123f` (feat/fix/test/refactor)
2. **Task 2: [task name]** - `def456g` (feat/fix/test/refactor)
3. **Task 3: [task name]** - `hij789k` (feat/fix/test/refactor)
**Plan metadata:** `lmn012o` (docs: complete plan)
_Note: TDD tasks may have multiple commits (test → feat → refactor)_
## Files Created/Modified
- `path/to/file.ts` - What it does
- `path/to/another.ts` - What it does
## Decisions Made
[Key decisions with brief rationale, or "None - followed plan as specified"]
## Deviations from Plan
[If no deviations: "None - plan executed exactly as written"]
[If deviations occurred:]
### Auto-fixed Issues
**1. [Rule X - Category] Brief description**
- **Found during:** Task [N] ([task name])
- **Issue:** [What was wrong]
- **Fix:** [What was done]
- **Files modified:** [file paths]
- **Verification:** [How it was verified]
- **Committed in:** [hash] (part of task commit)
[... repeat for each auto-fix ...]
---
**Total deviations:** [N] auto-fixed ([breakdown by rule])
**Impact on plan:** [Brief assessment - e.g., "All auto-fixes necessary for correctness/security. No scope creep."]
## Issues Encountered
[Problems and how they were resolved, or "None"]
[Note: "Deviations from Plan" documents unplanned work that was handled automatically via deviation rules. "Issues Encountered" documents problems during planned work that required problem-solving.]
## User Setup Required
[If USER-SETUP.md was generated:]
**External services require manual configuration.** See [{phase}-USER-SETUP.md](./{phase}-USER-SETUP.md) for:
- Environment variables to add
- Dashboard configuration steps
- Verification commands
[If no USER-SETUP.md:]
None - no external service configuration required.
## Next Phase Readiness
[What's ready for next phase]
[Any blockers or concerns]
---
*Phase: XX-name*
*Completed: [date]*
<frontmatter_guidance>
Purpose: Enable automatic context assembly via dependency graph. Frontmatter makes summary metadata machine-readable so plan-phase can scan all summaries quickly and select relevant ones based on dependencies (requires/provides/affects create the explicit links; transitive closure follows from them).
Subsystem/Tags: Primary categorization + searchable technical keywords, for detecting related phases and tech-stack awareness. Key-files: important files for @context references in PLAN.md. Patterns: established conventions future phases should maintain.
Population: Frontmatter is populated during summary creation in execute-plan.md. See <step name="create_summary"> for field-by-field guidance.
Status (#2830): status: complete is the default — the plan finished. Use status: halted instead when the plan reached a designed stop (a gate failure, a spike concluding without expanding into the full build, or any other intentional non-completion) and intentionally left tasks unfinished. halted is machine-read: any plan whose depends_on (directly or transitively) names a halted plan is reported as blocked, not offered to the executor, until the halt is resolved and re-summarized as complete.
</frontmatter_guidance>
<coverage_guidance>
Purpose (#1602): The coverage: block is a per-deliverable Requirements Traceability Matrix. It lets verify-work's extract_tests step route deliverables DETERMINISTICALLY — auto-passing those proven by passing tests and reserving human UAT for genuine judgment — instead of re-deriving coverage from prose. Consumed via gsd-tools uat classify-coverage --summary <SUMMARY>.
Field semantics:
| Field | Purpose |
|---|---|
id |
Stable identifier (D1, D2…) for cross-referencing from UAT.md and audit reports. Must be unique within the SUMMARY. |
description |
The deliverable in human-readable form — what would have been a prose bullet. |
requirement |
Links back to a REQUIREMENTS.md REQ-ID (joins requirements-completed). Optional. |
verification[].kind |
Enum: unit | integration | e2e | automated_ui | manual_procedural | other. |
verification[].ref |
Test path + descriptor (file#test name), Playwright screenshot ref, or command invocation. Required per entry. |
verification[].status |
pass | fail | unknown — populated from the latest test run. |
human_judgment |
Explicit boolean; REQUIRED. true always routes to a human. |
rationale |
REQUIRED when human_judgment: true. The audit trail for why automation is insufficient. |
Deterministic contract (what the classifier does):
- A deliverable auto-passes (no human prompt) only when
human_judgment: falseANDverificationis non-empty AND everyverification[].statusispass. This is the narrow, fully-proven case. - Everything else is presented to a human —
human_judgment: true, an emptyverification:, any non-pass/unknownstatus, or any schema error. A false-negative is a redundant prompt (the status quo); a false-positive ships a bug UAT existed to catch. - Fail-safe default: if you cannot determine coverage for a deliverable, you MUST set
human_judgment: truewithrationale: "Coverage not determined at authoring time — verifier must classify". Never leave a deliverable'shuman_judgmentempty, and never set itfalsejust to skip the prompt — auto-pass additionally requires a passingverificationentry, so the flag alone cannot skip the human. coverage: []means "no deliverables to classify" (the single-confirmation path). OMITTING the block entirely means "legacy" —verify-workfalls back to prose## Accomplishmentsextraction unchanged. </coverage_guidance>
<one_liner_rules> The one-liner MUST be substantive:
Good: "JWT auth with refresh rotation using jose library" · "Prisma schema with User, Session, and Product models" · "Dashboard with real-time metrics via Server-Sent Events"
Bad: "Phase complete" · "Authentication implemented" · "Foundation finished" · "All tasks done"
The one-liner should tell someone what actually shipped. </one_liner_rules>
**Frontmatter:** MANDATORY - complete all fields. Enables automatic context assembly for future planning.One-liner: Must be substantive. "JWT auth with refresh rotation using jose library" not "Authentication implemented".
Decisions section:
- Key decisions made during execution with rationale
- Extracted to STATE.md accumulated context
- Use "None - followed plan as specified" if no deviations
After creation: STATE.md updated with position, decisions, issues.