Files
msd-core/docs/adr/3180-planning-semantic-model-single-owner.md
Tom Boucher b9f51836e6 refactor(#3180): ADR-3180 behavior contract + cross-surface drift guardrails (#3223)
* refactor(#3180): one owner for completion ratio, a prompt-layer drift guard, and a written behavior contract

The 2026-08-08 coverage audit on #3180 found the epic's copy counts were a
lower bound for the third consecutive time, and that two derivation families
had never been named at all.

ADR-3180 gains Decision 7 — a normative behavior contract that says what the
right answer IS for each derivation, not merely who owns it. A reviewer with
no written rule can only ask "does this look like the others", which is how a
fifth copy passes review. Decision 4 gains (d) scan surface is every authored
surface and an owner FILE is never exempt, only its named functions; and (e)
a surface that cannot be consolidated today ships ratcheted, never unguarded.

Completion ratio: `clampPercent` sat exported and unused beside six hand-inlined
copies of its own body across five modules. All six now route through it;
`clampPercentFromFraction` is added for the one caller that already held a
fraction. Every migration is behaviour-identical — clampPercent's first line IS
the `total > 0 ? … : 0` ternary each copy carried. Guarded by
lint-completion-ratio-drift.cjs, which reports zero re-derivations with no
file-level exemption.

Prompt layer: workflow markdown re-derives live-plan counting in raw shell
(#1762), invisible to every `src/`-scoped guard. lint-planning-prompt-drift.cjs
scans it with a shrink-only baseline of the 7 sites that exist today — new
sites fail, and a baseline entry that stops firing fails too, so an
acknowledgment can never outlive the thing it describes.

lint-milestone-window-drift.cjs stops exempting its owner file wholesale; only
the four named canonical functions are exempt now. The blanket exemption was
pointed at the one file most likely to grow the next copy, and it had.

Refs #3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3180): link Phases 6-8 sub-issues (#3216, #3217, #3218) from ADR-3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3180): address orthogonal review — consumer-output identity tests, count-keyed ratchet, property coverage

Five findings from the two orthogonal review passes, all fixed.

Decision 4(c) breach: the completion-ratio identity test asserted at the
OWNER, which is exactly the bypass that decision exists to close — a consumer
can call clampPercent and then post-process locally, leaving both the lint and
an owner-level test green. It now drives `roadmap analyze`, `query progress`
and `stats` and asserts on their own output, over a fixture containing a
`status: superseded` plan so a consumer that re-counted raw files would report
60 where the owner reports 75.

Decision 4(e) breach: ratchet entries named the epic (#3180) rather than the
issue that removes them. They name Phase 8 (#3218) now.

The ratchet keyed on (file, text) alone, so plan-phase.md's two byte-identical
sites were one indistinguishable key and migrating either would have left the
guard green with the other alive. Entries carry an occurrence count; fewer than
acknowledged fails as a partial migration, more fails as a new copy.

Adds the missing MAX_REGEX_LITERAL_LEN boundary coverage the sibling guard's
test already had, and the fast-check property tests CONTRIBUTING requires for
clamp/budget-limit functions.

Refs #3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: stop wrapping a nested double-spawn in a 15s wall-clock budget (bug #641 probes)

`tests/ci-test-scope.test.cjs`'s `bug #641` block spawned `run-tests.cjs`
under PROBE_TIMEOUT_MS=15000; that child then spawned a nested `node --test`.
A fixed wall-clock budget around a double spawn, running inside a container
that is concurrently executing the full ~31k-test suite, fails by construction
under load.

Confirmed against three full matrix runs. Every failure was shaped
`null !== 0` — the child was KILLED, never an assertion about the thing under
test. One captured probe had already printed the correct resolution
(`suite="all" files=2: a.test.cjs b.test.cjs`) and was killed anyway. It
reproduces on `next` alone: 5 failures on linux-node22, 0 on linux-node24. The
victim subset varies by run and by lane.

What these tests are actually about is suite-token RESOLUTION — `unit` as a
bare token in --files/--files-from. Executing the seeded trivial files is
incidental and is the entire timeout surface, so the assertions move
in-process against the same functions `main()` calls, in the same order.
`parseArgs`, `selectExplicitFiles`, `selectFiles` and `walkTestFiles` are
exported for that; no behavior, signature or logic changed.

No coverage lost: `tests/run-tests-harness.test.cjs` already spawns the
harness for real and asserts exit codes end to end, on a 120s budget.

Pre-existing on `next`, fixed here rather than deferred.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: delete the three elapsed-time assertions

CLAUDE.md forbids asserting on wall-clock time. Three assertions did, and all
three are load-sensitive: on a saturated bench each can fail while the code
under test is correct. In every case the load-bearing assertion sits on the
line above and the timing line adds no discrimination.

run-with-timeout: the stated worry — "was this 124 the cap firing or the 30s
harness backstop?" — is already answered by the assertion above it. A backstop
kills by signal, which surfaces as status null, never 124. Observed directly
this session: three matrix runs produced exactly that null shape from killed
children.

normalize-test-command and context-predicates: both bounded a ReDoS check.
A threshold only ever separates "fast" from "slightly slow", which is bench
load, not correctness — catastrophic backtracking on 800 KB of input does not
take 251ms, it does not finish at all. A real regression therefore shows up as
the suite being killed on that test, which is louder and more reliable than a
number. The structural assertions (returned unchanged; cleanly rejected) are
what actually carry those tests, and they stay.

The sweep now reports zero elapsed-time assertions in tests/. The remaining
Date.now() uses are unique-path suffixes, barrier deadlines, fixture
timestamps and fake mtimes — none of them assertions.

Pre-existing on `next`, fixed here rather than deferred.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3180): backfill changeset PR number (#3223)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3180): key the prompt-drift ratchet on POSIX paths so it works on Windows

The baseline keys on (file, trimmed text). `file` came from scanTree's
`path.relative()`, which uses NATIVE separators, while the committed baseline
stores POSIX. On Windows every violation was therefore unmatched — reported as
FRESH — and every baseline entry matched nothing — reported as STALE. The guard
failed 100% of the time there, on both CI shards:

  ✖ scanRepo(repoRoot) matches the baseline exactly: zero fresh AND zero stale
    + { file: 'gsd-core\\workflows\\execute-plan.md', ... }

The remote runner this repo gates on is Linux-only and cannot see this class at
all; the GitHub Actions Windows lane is what caught it.

Normalization is unconditional — never gated on process.platform. A
platform-conditional normalizer makes the POSIX path the special case and
leaves the Windows branch unexercised on every other OS, which is the same
blind spot in a different place. It is applied at one seam inside
findPromptDrift, which builds `file` on every returned violation, so the
baseline key, the --update writer, the stderr report and the tests all consume
one normalized value.

The regression tests drive a Windows-shaped relPath directly and run on every
OS rather than skipping off-Windows — a test that only runs on the platform
where the bug lives is why this escaped. They include a sanity check that
un-normalized input does NOT match, so the assertion cannot pass vacuously.

Audited the three sibling guards: none keys against a committed cross-platform
baseline, and their exemption keys are path.join-built, so producer and
consumer share the native convention. Left correct code alone rather than
making them look alike. scripts/lib/drift-scan.cjs is untouched — normalizing
there would break those three on Windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 16:05:17 -04:00

67 KiB
Raw Blame History

ADR-3180: Planning Semantic Model — Single Owner per Derivation

  • Status: Accepted. Phase 0 shipped this file alone; Amendment 4 (2026-08-08) adds Decision 7 — the normative behavior contract — and lands the guards and the completion-ratio consolidation described there.
  • Date: 2026-08-07
  • Issue: #3180 is the scope authority (epic + approved-enhancement + type: chore), which is why this ADR carries its number. #3182 is the Phase-0 tracking sub-issue this PR closes — the epic stays open until the final phase merges. This follows ADR-3128, whose filename likewise tracks its scope-authority issue while its PR referenced a separate docs sub-issue.
  • Supersedes: nothing
  • Relationship to prior work: extends ADR-2121, which consolidated phase-identifier syntax and proved the mechanism (scripts/lint-phase-id-drift.cjs reports 0 independent re-derivations). This ADR applies the same mechanism one layer up, to the semantics of the .planning/ model. Distinct from #1879 (absent-vs-corrupt across I/O read paths), #2143 (the document-parsing layer beneath), and #3051 (why the suite did not catch these).

Symbol names are the durable anchors throughout. Line references, where given, are as of next @ cbd180c5c and will drift.

Context

ADR-2121 consolidated what a phase is called. Nothing consolidated which phases exist, which milestone owns them, which are done, and how many plans are live. Those derivations are re-implemented independently at every call site.

The divergent surface

In every row a correct implementation already exists beside the broken one. These are not gaps in knowledge; they are fixes that landed on one copy.

Derivation Copies Canonical (correct) Divergent
Milestone windowing 3 currentMilestoneRawRanges::computeSectionEnd — the only copy carrying a "keep in sync" comment extractCurrentMilestone::computeSectionEnd; an undocumented inline copy in getMilestonePhaseFilter's versionOverride branch
Phase enumeration 4 cmdRoadmapAnalyze — scopes via extractCurrentMilestone and filters sentinels cmdProgressRender, cmdStats, cmdPhasesList
Phase completion 2, disagreeing cmdPhaseComplete — calls readVerificationStatus unconditionally buildPhaseCompletionProjection — gates it behind planCount > 0
Live-plan counting 3 scanPhasePlans — excludes status: superseded (#2349) cmdFindPhase; findPhaseInternal/searchPhaseInDir
State field extraction 2+ stateExtractField consumers carrying the #1760 fallback chain cmdStateValidate; cmdStateCompletePhase's idempotency guard

The milestone-windowing duplication is verifiable structurally, not just textually: roadmap-parser.cts contains two distinct computeSectionEnd function nodes — extractCurrentMilestone::computeSectionEnd and currentMilestoneRawRanges::computeSectionEnd — separate definitions with separate call sites, not one function referenced twice.

The failure mode that hides all of it

Every divergent path returns a well-formed, plausible value. None throws, none logs, none returns a sentinel a caller can branch on — the failure and the success are output-identical:

Path Returns on failure
extractCurrentMilestone, truncated window phases: [], phase_count: 0, no error
getMilestonePhaseFilter, empty result a pass-all filter → archives every phase dir on disk
cmdStateValidate {valid: true, warnings: [], drift: {}}, unconditionally
aggregate percent 100 while plans are outstanding
buildPhaseCompletionProjection not_required, ignoring a real passing *-VERIFICATION.md

This is why 13 of the 14 defects in the 2026-08-07 sweep were found by a contributor dogfooding downstream rather than by the suite: a test asserting "returns a number" or "does not throw" passes against every row above.

The derivations are one coupled cluster

get_impact(extractCurrentMilestone, direction=both, depth=10) against next @ cbd180c5c returns risk CRITICAL — 200 affected symbols with total_affected_is_lower_bound: true and truncated: true, spanning 43 distinct affected_files and 24 affected_processes. The counts are depth- and truncation-sensitive: a shallower query returns fewer files and is not a contradiction. Every symbol this epic names sits inside that single blast radius.

Symbol Rating Direct callers
extractCurrentMilestone CRITICAL 20
stateExtractField high 20
scanPhasePlans medium 11
getMilestonePhaseFilter medium 10
buildPhaseCompletionProjection low 3

This bounds the decomposition: the phases are stacked and sequential, never parallel, because a parallel phase would edit symbols inside a sibling's radius.

Correction to #3180's text. The epic states stateExtractField has "five call sites." find_symbol reports 20 direct callers. Phase 5's call-site sweep must be driven from the graph, not from that count.

The hypothesis is falsifiable, and it held twice

Predicted: fixes land on one copy while siblings stay broken. Confirmed — #1760 fixed 3 of 5 stateExtractField sites; #3165's fix provably does not reach #3166's inline copy.

Predicted: new readers arrive carrying new copies. Confirmed — cmdPhasesList was found during triage as a fourth unscoped phasesDir reader that no issue had reported.

Per CONTRIBUTING.md: One issue = one ADR-or-PRD = one PR. This ADR is that one file. It ships no production code.

Decision

Give each semantic derivation a single canonical owner, enforced mechanically the way ADR-2121 enforced identifier syntax, and give every derivation a distinguishable failure signal instead of a plausible default. Six decisions, locked below.

1. One canonical owner per derivation; the duplicates are DELETED

The surviving owner per derivation:

Every owner is named, with a locked module and signature. Phases consume them verbatim.

Derivation Canonical owner (module · symbol) Deleted
Milestone windowing src/roadmap-parser.cts · currentMilestoneRawRanges::computeSectionEnd, lifted to a module-level export computeMilestoneSectionEnd extractCurrentMilestone::computeSectionEnd; the getMilestonePhaseFilter versionOverride inline copy
Phase enumeration src/phase-locator.cts · listMilestonePhaseDirs (new export; the Phase Locator Module already owns on-disk phase discovery) the direct phases-dir reads in cmdProgressRender, cmdStats, cmdPhasesList; the nested cmdRoadmapAnalyze::isSentinelPhase closure
Phase completion src/verification.cts · isPhaseComplete (new export, sited beside readVerificationStatus, which it wraps) the planCount > 0 gate in buildPhaseCompletionProjection
Live-plan counting src/plan-scan.cts · scanPhasePlans filename re-derivation in cmdFindPhase, findPhaseInternal/searchPhaseInDir
State field extraction src/state-document.cts · stateExtractField carrying the #1760 fallback chain per-site re-derivation at all remaining call sites

Locked signatures for the two new owners (ScopedResult<T> is defined in Decision 2):

listMilestonePhaseDirs(roadmapContent: string, phasesDir: string, deps?): ScopedResult<string[]> Applies the milestone window and the sentinel filter in that order, and returns the surviving phase directory names. The sentinel predicate delegates to the existing canonical isSentinelPhaseId in src/phase-id.cts — it does not re-implement the nested cmdRoadmapAnalyze::isSentinelPhase closure, which is itself a sixth instance of this epic's divergence class and is deleted by Phase 3.

isPhaseComplete(phaseDir: string, deps?): ScopedResult<{ complete: boolean; verification: VerificationStatus }> The single predicate for both the read path (buildPhaseCompletionProjection) and the write path (cmdPhaseComplete). It calls readVerificationStatus unconditionally — there is no plan-count precondition. A phase with zero plans and a passing *-VERIFICATION.md is complete.

Deleted, not kept in sync by comment. The "keep in sync" comment on the canonical windowing copy is already in place and already failed; it is evidence the risk was known, not that it was controlled.

Rejected: keeping N copies with a parity assertion test. A parity test proves the copies agree today; it does not stop copy N+1, and cmdPhasesList demonstrates copy N+1 arriving unreported.

2. The shared result contract — PROVISIONAL, validated by Phase 1

Home module — locked. SCOPE and the ScopedResult<T> shape live in a new pure leaf module, src/planning-scope.cts, exporting nothing else. It follows the src/phase-id.cts precedent: pure, no Node built-ins, no config, no other core dependency, so every consumer above it can import it without a cycle. Phase 1 creates it.

Creating a new .cts module carries this repo's six-gate ripple — .gitignore, eslint config, docs/INVENTORY.md, the inventory manifest (regenerate after build:lib, never before, or modules are silently dropped), the CONTEXT.md Glossary — Domain modules and seams entry (a PR gate), and size:baseline. Phase 1 owns all six.

Every consolidated derivation returns a result carrying a scope discriminator drawn from a frozen enum:

const SCOPE = Object.freeze({
  COMPLETE:   'complete',    // computed over the whole intended input
  TRUNCATED:  'truncated',   // input window was cut short
  UNSCOPED:   'unscoped',    // ran without the scoping it required
  UNREADABLE: 'unreadable',  // input absent or unparseable
});

/** @typedef {{ value: T, scope: typeof SCOPE[keyof typeof SCOPE] }} ScopedResult */

ScopedResult<T> carries the derivation's own payload in value (an array for the list-shaped derivations, an object for isPhaseComplete, a nullable string for stateExtractField) plus the scope discriminator. Nothing else is added to the shape — a caller needing more asks for an amendment rather than widening it locally.

COMPLETE with zero items is a real answer — a freshly-declared milestone genuinely has no phases. TRUNCATED/UNSCOPED/UNREADABLE with zero items is a non-answer. Today those are the same value, and that identity is the epic.

The enum is frozen and asserted on directly (result.scope === SCOPE.TRUNCATED). It is not a message string: CONTRIBUTING.md § Prohibited: Raw Text Matching on Test Outputs requires a typed IR and forbids assert.match against rendered prose.

This contract is provisional until Phase 1 validates it. Phase 1 (live-plan counting) is the first and smallest real implementation. If the contract does not fit, this ADR is amended before Phase 2 begins — the contract is not worked around in code. Amendments are recorded in the Amendments section below. This is a deliberate Gall's Law concession: a five-derivation contract locked before a single consolidation exists is a design that has never met production.

Rejected: a boolean ok/degraded — rows TRUNCATED/UNSCOPED/UNREADABLE need three distinct caller responses, and a boolean recreates the collapse this epic removes. Rejected: throwing instead of returning a scope — these paths are read during normal progress rendering, and throwing converts a display degradation into a command failure.

3. Two-tier change policy (Hyrum's Law)

getMilestonePhaseFilter's pass-all degrade is documented in its own comment as deliberate and safe ("over-inclusive, never under-inclusive"). That is not an accidental behavior someone latched onto — it is a written promise. Changing it needs an explicit policy:

  • Tier 1 — internal function contracts (the five owners and their duplicates). Freely changed; duplicates deleted. The consumers are gsd-core's own call sites, enumerable from the graph. No deprecation cycle.

  • Tier 2 — observable command output. These reach downstream projects that cannot be enumerated. Every Tier-2 change requires an explicit breaking-change call-out in its PR, a .changeset/ fragment, and a docs/ update. The complete list, by phase:

    Phase Command surface Output change
    1 phase find (cmdFindPhase, findPhaseInternal/searchPhaseInDir) a phase whose plans are all status: superseded reports zero live plans, not a positive count
    2 roadmap analyze, roadmap get-phase a truncated window stops reporting phase_count: 0 as if it were a real empty; milestone complete stops pass-all archiving on a truncated window
    3 query progress, stats, phases list 999.* backlog directories no longer listed as current-milestone phases; aggregate percent stops reading 100 while plans are outstanding
    4 init manager a zero-plan phase with a passing *-VERIFICATION.md reports complete instead of not_required
    5 state validate reports invalid for genuinely invalid documents instead of unconditional valid: true

    This list is contingent on Decision 2's contract surviving Phase 1. If the contract is amended, this table is re-derived in the same amendment rather than inherited unchanged.

The pass-all degrade is preserved where it is correct (a genuinely-empty new milestone, scope: COMPLETE) and refused where it is destructive (a truncated window, scope: TRUNCATED). Decision 2's contract is what makes that distinction expressible; without it the code cannot tell the two apart, which is exactly why the degrade is dangerous today.

Phase 5 is the sharpest Tier-2 change in the epic: state validate moves from unconditionally valid: true to able to fail, which will surface pre-existing invalid STATE.md documents in downstream CI that currently passes. That is the intended outcome — a gate that cannot fail is worse than no gate — but it ships with an explicit warning.

4. The anti-divergence contract — structural guard PLUS behavioral identity test

Each derivation ships a scripts/lint-<derivation>-drift.cjs guard modelled on the five existing precedents (lint-phase-id-drift.cjs, lint-package-identity-drift.cjs, lint-shell-command-projection-drift.cjs, lint-table-schema-drift.cjs, check-alias-drift.cjs), reporting 0 independent re-derivations, plus a matching identity guard test.

Two constraints are locked, both consequences of Goodhart's Law — "0 re-derivations" is a measure about to become a target:

(a) Guards discover call sites by whole-repo scan, never by an allowlist of known files. An allowlist-driven guard measures "re-derivations in files we remembered to list." cmdPhasesList is the proof: a guard scanning only the three reported unscoped readers would have reported 0 while a fourth existed.

(b) The structural guard and the behavioral identity test are both required, and neither alone is sufficient. The lint is gameable by indirection — route a re-derivation through a wrapper, a differently-named local, or a test helper, and the count stays 0 while the divergence returns. The identity test is gameable the other way: it only covers the input shapes its author imagined, the fixture-provenance trap CONTRIBUTING.md §2371 already names. The lint is the output metric; the identity test is the outcome metric; Goodhart's prescribed defense is to pair them and never report either alone.

(c) The identity test asserts at the CONSUMER's output, not at the owner's return value. This closes the one bypass that defeats both (a) and (b) together: a consumer calls the canonical owner — satisfying the lint, since there is no re-implementation, and satisfying an owner-level identity test, since the owner is untouched — and then post-processes the result locally, e.g. re-applying its own "exclude superseded" pass after scanPhasePlans returns. Divergence is fully restored and both guards stay green.

Therefore each derivation's identity test compares each consumer's observable output against the canonical owner's result for the same input, and fails on any difference. Post-filtering a canonical result is then indistinguishable from re-deriving it, which is the correct equivalence: both produce a second answer to a question that is supposed to have one owner. Where a consumer legitimately needs a narrower set, it passes an argument to the owner — it does not filter the owner's output.

Rejected: a lint that asserts all sites match a golden regex — it enforces textual sameness, not single ownership, and cannot see semantic divergence (this is ADR-2121 Decision 1's rejected option (C), and it applies unchanged here).

(d) A guard's scan surface is every AUTHORED surface that can express the derivation — not src/. Constraint (a) said "whole-repo scan, never an allowlist of known files", and both guards built under it read that as the whole src/ tree. That is itself an allowlist, one directory wide. #1762's second reproduction traced a wrong 30 plans, 24 summaries figure to a raw ls -1 … *-PLAN.md | wc -l snippet inside gsd-core/workflows/progress.md — a live-plan re-derivation lint-plan-count-drift.cjs reported clean because it was not looking there.

A derivation is re-derived wherever it is expressed, and this product expresses these four in two languages: TypeScript under src/, and shell embedded in the workflow/command markdown that ships to every runtime. A guard covering one of the two measures half the surface and reports zero. Each derivation's guard therefore declares its scan surface explicitly, and any derivation reachable from the prompt layer is covered there too — scripts/lint-planning-prompt-drift.cjs scans gsd-core/workflows, commands, agents and skills.

The same constraint applies inward: a guard's owner FILE is not exempt, only its named canonical FUNCTIONS are. A whole-file exemption on the owner is constraint (a)'s forbidden allowlist aimed at the one file most likely to grow the next copy, and it did — see Amendment 4.

(e) Where a surface cannot be consolidated in the same change, the guard ships RATCHETED — never absent. A derivation expressed in the prompt layer cannot be routed onto its .cts owner by an import; it needs a CLI surface to call, which is a phase of its own. The guard still lands, carrying a baseline of the sites that exist at that moment, and follows scripts/qa-smell-ratchet.cjs's invariants exactly:

  • a recorded site never fails — it is acknowledged, in writing, with the issue that owns its removal;
  • an unrecorded site fails — nobody has looked at it;
  • a recorded site that no longer fires also fails, so the baseline can only shrink and an acknowledgment can never outlive the thing it describes;
  • entries are keyed on (file, trimmed source text) plus an occurrence count, never on a line number, which churns on every unrelated edit to the same file. The count is what makes a partial migration visible: two byte-identical sites in one file would otherwise be one indistinguishable key, so migrating one of them would leave the ratchet green while the other survived. Fewer occurrences than acknowledged fails as a partial migration; more fails as a new copy planted beside an acknowledged one.

Rejected: land the guard later, together with the migration. That is the "found it, wrote it down, moved on" posture this epic exists to remove — between the finding and the migration the surface is known-broken and unwatched, which is strictly worse than unknown. Rejected: a bare eslint-disable-style suppression. A suppressed guard and a green guard are indistinguishable at a glance; a ratchet reports its own remaining debt on every run.

5. Migration order — live-plan counting ships BEFORE milestone windowing

Locked order: Phase 1 (live-plan counting) → Phase 2 (milestone windowing) → Phase 3 (enumeration) → Phase 4 (completion) → Phase 5 (state field extraction).

Phase 1 before Phase 2 is not a preference. roadmap.analyze calls extractCurrentMilestone directly, so repairing the window repopulates Route 0's loop and converts #3164 from a silent no-op into a live misroute that re-executes a closed phase. The epic states the constraint as "#3165 must not ship ahead of #3164"; since Phase 2 is the #3165 repair, Phase 1 must precede it. This inverts the order the derivations are listed in #3180's own table.

Phase 3 follows Phase 2 because the enumeration owner must carry the window Phase 2 consolidates. Phases 4 and 5 are order-independent relative to each other but follow the cluster.

6. Scope boundaries

In scope: the five derivations above; their guards and identity tests; the scope contract; boundary coverage per CONTRIBUTING.md and at least one test per derivation asserting the path can fail.

Out of scope: any change to .planning/ on-disk formats; the document-parsing layer (#2143); the I/O-failure layer (#1879).

The child defects — stated precisely, because the epic's shorthand is ambiguous. #3180 says this epic "removes the class, it does not gate the instances." Does not gate means the epic does not wait on them and does not take responsibility for closing them. It does not mean the consolidation leaves their symptoms intact — several are subsumed as a direct consequence of giving the derivation one owner, because the divergent copy that produced the symptom ceases to exist:

Phase Subsumes Why unavoidable
1 #3164 routing the plan count through scanPhasePlans is the superseded-exclusion fix
2 #3165, #3166 deleting the divergent windowing copies removes both the truncated-window report and the pass-all archive degrade
3 #3167, #3161 one enumeration carrying the sentinel filter removes the backlog-dir listing and the 100-percent aggregate
4 #3168 deleting the planCount > 0 gate is that defect's fix
5 #3162 routing cmdStateValidate through the fallback chain is that defect's fix

Each phase's PR names the child issues it subsumes and records the evidence that the symptom is gone. It does not unilaterally close them: #3180 explicitly declined ownership of the instances, so whether a subsumed issue is closed, re-scoped, or left open for its own regression test is the maintainer's call at merge time, made with the evidence in front of them. The remaining child defects (#3169, #3170, #3171, #3174, #3156) are not touched — they sit in adjacent parse/format paths this epic does not consolidate — and stay independently actionable.

Recording the subsumption explicitly because both silences are failures: a phase that demonstrably removes a defect's symptom while claiming to change nothing is a shipped lie, and a phase that closes an issue the epic disclaimed is scope it never had. Naming the effect without claiming the disposition is the only honest position available here.

Scope note on Phase 5. State field extraction is not one of #3180's seven "Done when" items — the epic describes it in evidence as "a fifth instance of the same shape" and lists #3162 among the out-of-scope child defects, while its Goal ("one canonical owner per semantic derivation") covers it. That inconsistency was surfaced during planning and resolved by maintainer decision to include it. #3180's Done-when list should be amended to match, or Phase 5 reads as unclaimed scope.

7. The behavior contract — this section is the SOURCE OF TRUTH

Decisions 1–6 answer who owns each derivation. They do not say what the right answer is, and that omission is why the 2026-08-08 coverage audit could find six copies the epic had never named: a reviewer with no written rule to check a call site against can only ask "does this look like the others", which is how a fifth copy passes review.

This section is that written rule. It is normative, and it is what the guards and identity tests of Decision 4 test against:

  • Where this section and the code disagree, the code is the defect — not this section, and not a caller's local expectation.
  • A behavior not stated here is not decided. It is recorded below as an open question with a forcing function, never resolved silently inside an implementation PR.
  • Amending a rule here is an amendment to this ADR (Amendments section), not a code change with a comment.
  • Each rule carries a status: Enforced (owner exists, guard green) or Required — Phase N (contract locked, migration outstanding). A Required rule is as binding as an Enforced one; the only difference is whether the tree satisfies it yet.

7.1 Milestone windowing — Enforced (Phase 2)

Question. Which byte range of ROADMAP.md belongs to milestone M?

Owner. src/roadmap-parser.cts — locateMilestoneHeadings, computeMilestoneSectionEnd, isMilestoneBoundedInRoadmap, and the composition sliceMilestoneWindow.

Rule. The window opens at the heading locateMilestoneHeadings selects for M and closes where computeMilestoneSectionEnd says. A ### Phase N: … heading never opens or closes a milestone window. The version token's boundary is \b, not (?![\w.-]) — a milestone STATE of v8.0 legitimately selects ## v8.0-B … over a closed v8.0-A sibling (#730; Amendment 2 tried the stricter boundary and reverted it). A free-form legacy ROADMAP carrying no versioned milestone heading is COMPLETE, not UNSCOPED: whole-document genuinely is the milestone there. A composition of these primitives is itself an owner — assembling locate → pick → computeEnd → slice at a call site is a re-derivation even though every step calls the owner (Amendment 2).

Failure signal. ScopedResult.scope per Decision 2.

Guard. scripts/lint-milestone-window-drift.cjs.

7.2 Milestone identity — Required — Phase 6

Question. Which milestone is current, and what is it called?

Owner (to be). getMilestoneInfo binds to locateMilestoneHeadings and deletes its own heading regexes. It is a sixth derivation family — the coverage audit's gap 2 — that no phase of the original decomposition touches.

Rule.

  1. STATE.md's milestone: field selects the version when present; the ROADMAP heuristics are the fallback, not the primary.
  2. The heading is located by the canonical locator of §7.1, which already excludes phase headings. A ### Phase N: Close v3.3 gaps heading is never the milestone heading (#3197 — reproduced live, writing a wrong milestone: to disk).
  3. The name is the heading text following the version token with a leading delimiter (—, –, :, -) stripped. ( is an ordinary name character: the name is not truncated at a parenthetical (#3171).
  4. A failure returns a scope other than COMPLETE. It does not return {version: 'v1.0', name: 'milestone'} presented as an answer — that default is output-identical to a successful read of a genuine v1.0 project, which is this epic's defining failure mode.

Guard. lint-milestone-window-drift.cjs today keys on the #{N,M} heading-level quantifier; getMilestoneInfo's regexes anchor on a literal ## and therefore slip past it. Phase 6 ships the token widening together with the consolidation, never after — a guard added later measures a surface already cleaned and reports a zero it did not earn.

7.3 Phase enumeration — Enforced for the four named consumers (Phase 3, #3222); the fifth copy is unowned

Question. Which directories under <planning>/phases/ are phases of milestone M?

Owner. src/phase-locator.cts · listMilestonePhaseDirs (Decision 1).

Rule. A directory counts iff all three hold: its identifier parses per src/phase-id.cts; it is not a sentinel per isSentinelPhaseId; and its ROADMAP entry falls inside M's window per §7.1. Both filters, in that order. Any surface answering "how many phases does this milestone have" reports the same set for the same input — the progress renderer, the roadmap analysis, the statistics command and the phase listing are not allowed to disagree.

Consumers that must route through the owner. cmdRoadmapAnalyze, cmdProgressRender, cmdStats, cmdPhasesList, and buildStateFrontmatter / syncStateFrontmatter — the fifth copy, reached by state.record-session, state.sync, phase.complete and every other state-mutating verb, which the epic's original scope did not name (coverage-audit gap 1).

Status, precisely. Phase 3 (#3222) enforced this rule for cmdRoadmapAnalyze, cmdProgressRender, cmdStats and cmdPhasesList, and its guard (scripts/lint-phase-enumeration-drift.cjs) found 54 violations where the epic scoped 4. It did not reach buildStateFrontmatter / syncStateFrontmatter. That copy is still live, still writes its answer to disk, and is now unowned by any phase — see Amendment 4's scope table, row 1.

Note on #3204. Routing buildStateFrontmatter through the owner will not by itself fix #3204: its defect is the discriminator one layer above enumeration — "is the ROADMAP's phase count safe to trust" — which misclassifies ordinary ## Overview / ## Progress headings as milestone sectioning. That discriminator is §7.1's isMilestoneBoundedInRoadmap. The enumeration routing and the discriminator replacement must ship together or the defect survives the consolidation.

7.4 Phase completion — Required — Phase 4, blocked

Question. Is phase P complete?

Owner. src/verification.cts · isPhaseComplete (Decision 1).

Rule. readVerificationStatus is called unconditionally. Plan count is not a precondition: a phase with zero plans and a passing *-VERIFICATION.md is complete. The read path and the write path share this predicate, so "phase.complete succeeds while init.manager reports incomplete" is unrepresentable for identical input.

OPEN QUESTION — does a ROADMAP checkbox override disk state? (#2957). There are three completion implementations, not the two the epic recorded: cmdPhaseComplete, buildPhaseCompletionProjection, and buildStateFrontmatter, which computes completed phases from plan scanning alone and never consults the ROADMAP checkbox that cmdRoadmapAnalyze deliberately honors. Checkbox-override versus disk-strict is a product decision, and it is not made here.

Forcing function. Phase 4's drift guard fails while more than one completion predicate exists. It cannot be satisfied by consolidating two of three and leaving the third, and Phase 4 must not ship before #2957 is decided — a shared predicate that silently adopts whichever semantics its author happened to hold is a product decision made by typing order.

7.5 Live-plan counting — Enforced (Phase 1), with a known representation gap

Question. How many plans in phase P are outstanding?

Owner. src/plan-scan.cts · scanPhasePlans, exposing planFiles (live) and allPlanFiles (every plan on disk, pre-supersession).

Rule. A plan is live unless it carries a machine-readable terminal state. The only terminal state today is frontmatter status: superseded (#2349). Choosing between the two sets is explicit per call site, never mechanical: a diagnostic about file naming takes allPlanFiles; a question about outstanding work takes planFiles (Amendment 1 — passing the filtered set into describeNonCanonicalPlans made a superseded-but-correctly-named plan report as a naming violation).

GAP — the lifecycle has exactly one machine-readable terminal state (#1762, coverage-audit gap 6). Plans retired through ROADMAP prose or HTML-comment fences carry no status key, so the canonical owner counts them live. Consolidation cannot fix this; it needs a representation that does not exist yet. The contract, locked now so no surface invents its own: the plan lifecycle's terminal states are a closed, frontmatter-expressed vocabulary. Prose is not a lifecycle signal. Until the vocabulary is extended, a plan a human considers retired but that carries no status key is live, and every surface reports it that way — a caller may not compensate by reading prose locally.

7.6 Completion ratio — arithmetic Enforced; rule 3 Enforced for query progress / stats; rule 4 Required — Phase 7

Question. What percentage of a scoped set is complete?

Owner. src/phase-lifecycle.cts · clampPercent(completed, total) and clampPercentFromFraction(fraction). A seventh derivation family, absent from the epic's table: the identical expression total > 0 ? Math.min(100, Math.round((completed / total) * 100)) : 0 was hand-inlined at six call sites across five modules while the owner sat exported beside them, unused by any of them.

Rule.

  1. Exactly one expression of fraction → integer percent exists: round-half-up, ceiling 100.
  2. A non-positive or absent denominator yields 0. "Nothing to complete" is 0%, never 100%.
  3. The numerator and the denominator come from the same scoped set. A percentage inherits the scope of the counts that produced it.
  4. A derivation whose scope is not COMPLETE does not render a percentage at all.

Status.

  • Rules 1 and 2 — enforced. Six sites migrated onto the owner, guarded by scripts/lint-completion-ratio-drift.cjs.
  • Rule 3 — enforced for query progress and stats. Phase 3 (#3222) routed cmdStats's and cmdProgressRender's totalPlans/totalSummaries accumulation through listMilestonePhaseDirs's scoped, sentinel-filtered set, so a backlog directory's already-summarized plans can no longer inflate the numerator against a milestone that has not finished. That is what closed #3161, and it is rule 3 by another name.
  • Rule 4 — Phase 7 (#3217). Withholding a percentage entirely when the scope is not COMPLETE is not implemented anywhere. Phase 7 also re-checks any consumer Phase 3 did not reach, cmdRoadmapAnalyze first.

Correction, recorded rather than quietly dropped. Amendment 4 originally asserted that #3161 was not fixed by the arithmetic consolidation and that enumeration consolidation "changes nothing here" — and the second half was wrong. The two changes were authored concurrently; Phase 3 merged first and subsumed #3161 through the scoped set. The first half stands: the arithmetic consolidation alone would not have fixed it. Kept visible because a green ratio guard beside an unfixed #3161 would have been exactly the "measure became the target" outcome Decision 4 exists to prevent, and the reason it is not that outcome is Phase 3, not this change.

7.7 State field extraction — Required — Phase 5

Question. What is the value of field F in a .planning/ state document?

Owner. src/state-document.cts · stateExtractField, carrying the #1760 fallback chain.

Rule. Every consumer calls the owner; none re-derives the field's location or its fallback order locally. state validate reports invalid for a genuinely invalid document — an unconditional {valid: true, warnings: [], drift: {}} is a gate that cannot fail, which is worse than no gate (Decision 3).

Call-site sweep is driven from the graph, not from the epic's text — find_symbol reports 20 direct callers where #3180 says five.

Consequences

Positive. "Fixed on one copy, missed on the siblings" becomes unrepresentable for five derivations. phase.complete succeeding while init.manager reports incomplete becomes structurally impossible rather than merely fixed. Failure paths stop being output-identical to success, so the suite can assert on them and the class of bug that required downstream dogfooding to find becomes detectable in CI.

Negative / accepted costs. Five new scripts/ files, which ship in the npm package and installer (inventory and manifest ripples per phase). One new .cts module (src/planning-scope.cts) carrying the full six-gate ripple in Phase 1. Tier-2 output changes will break downstream consumers parsing current command output — deliberately. Phase 2 carries a CRITICAL blast radius and cannot be de-risked by slicing further without breaking the derivation in half. The stacked ordering means the epic cannot be parallelized, so wall-clock is the sum of five phases.

Risks. The contract is validated by exactly one consolidation (Phase 1) before four more build on it; the amendment path in Decision 2 is the mitigation. The guards cannot see re-derivation through dynamic dispatch; the identity tests are the backstop, and they only cover the shapes they were written for.

Alternatives considered

  1. One mega-PR consolidating all five. Rejected — 200+ affected symbols in a single reviewable unit, gsd-test failures unattributable to a derivation, and a violation of one-concern-per-PR.
  2. Fix the 13 child defects individually, no consolidation. Rejected — that is the status quo whose failure mode this epic documents: each fix lands on one copy, and the sweep found a fourth reader nobody had reported.
  3. Consolidate without guards. Rejected — ADR-2121 demonstrated that mechanical enforcement is what makes consolidation durable. Without a guard, copy N+1 arrives with the next reader.
  4. Design the contract inside Phase 1's PR, skip this ADR. Rejected — Phases 2–5 all depend on the contract, so it would be set by whatever was convenient in the first code PR, with no reviewable design step. (Offered during planning and declined by the maintainer.)
  5. Per-derivation bespoke result shapes instead of one scope contract. Rejected — five shapes inside one coupled blast radius reintroduces the divergence one level up.

Software laws applied

Cross-referenced via /skills-from-the-artificer. Four fired; two materially changed this ADR.

  • Gall's Law — changed the design. A five-derivation contract locked before any consolidation exists is a complex system built from scratch. Mitigated by Decision 2's provisional status plus the amendment path, and by sequencing the smallest, lowest-risk derivation (scanPhasePlans, complexity 3) first so the contract meets production before four phases depend on it.
  • Goodhart's Law — changed the design. "0 independent re-derivations" is a measure becoming a target, gameable by indirection or by an allowlist-scoped scan. Produced Decision 4's two locked constraints: whole-repo discovery, and a paired structural + behavioral metric.
  • Hyrum's Law — confirmed, and sharper than expected. The pass-all degrade is documented as intentional in its own comment, so this is a written promise being revoked, not an accident being corrected. Produced Decision 3's two-tier policy. This mirrors ADR-2121 Decision 2, which invoked the same law for normalizePhaseName's CRITICAL radius.
  • Postel's Law — the epic's own stated lens ("Postel / fail-loud"). The defect is not leniency but leniency with no signal that it engaged. Decision 2 makes the degrade decidable rather than removing it.

Considered and not applicable: choose-boring-technology (no new dependency; five in-repo guard precedents), conways-law (no ownership boundary at stake), zawinskis-law (scope grew by one phase by explicit maintainer decision, not creep).

Cross-references

  • ADR-2121 — the proven precedent this extends
  • ADR-2143 — the document-parsing layer beneath
  • scripts/lint-phase-id-drift.cjs — the guard pattern Decision 4 models
  • scripts/qa-smell-ratchet.cjs — the ratchet invariants Decision 4(e) adopts verbatim
  • scripts/lib/drift-scan.cjs — the one tree-walk/confinement/sanitizer implementation every guard shares
  • CONTRIBUTING.md § Prohibited: Raw Text Matching on Test Outputs — why scope is a frozen enum
  • CONTRIBUTING.md § Fixture provenance (#2371) — why the identity test alone is insufficient
  • Phase sub-issues: #3183, #3184, #3185, #3186, #3187, and the three added by Amendment 4 — #3216 (Phase 6, §7.2), #3217 (Phase 7, §7.6), #3218 (Phase 8, §7.5 + Decision 4(d)/(e)).

Guard roster

One row per derivation. A blank owner is a derivation whose contract is locked (§7) but whose owner does not exist yet.

Derivation Owner Guard Scan surface Status
Milestone windowing (§7.1) roadmap-parser.cts lint-milestone-window-drift.cjs src/ enforced
Milestone identity (§7.2) — (Phase 6) same guard, token set widened by Phase 6 src/ contract only
Phase enumeration (§7.3) phase-locator.cts lint-phase-enumeration-drift.cjs src/ enforced for the four named consumers; buildStateFrontmatter/syncStateFrontmatter unowned
Phase completion (§7.4) verification.cts (Phase 4) Phase 4 src/ blocked on #2957
Live-plan counting (§7.5) plan-scan.cts lint-plan-count-drift.cjs src/ enforced
Live-plan counting, prompt layer (§7.5) — (Phase 8) lint-planning-prompt-drift.cjs gsd-core/workflows, commands, agents, skills ratcheted, 7 sites
Completion ratio (§7.6) phase-lifecycle.cts lint-completion-ratio-drift.cjs src/ arithmetic + rule 3 enforced; rule 4 is Phase 7
State field extraction (§7.7) state-document.cts (Phase 5) Phase 5 src/ contract only

Amendments

Amendment 1 — Phase 1 (#3183) validation: the boundary between Phases 1 and 3 was mis-cut

Decision 2 marked the contract provisional and required amendment before Phase 2 rather than a workaround in code. Phase 1 exercised it and the contract itself held — SCOPE needed no change. What did not hold was the phase boundary.

What Phase 1 found. Building the Decision 4(a) whole-repo guard — the one that may not use a file allowlist — turned up 26 live-plan re-derivations across 9 files. The epic scoped this derivation at 3 copies. Per-site triage classified them 21 true re-derivations, 2 asking a genuinely different question, 2 dead.

Why the boundary was wrong. commands.cts's cmdProgressRender re-derives both enumeration (assigned to Phase 3) and plan counting (Phase 1), on adjacent lines. So DW4's "no caller re-derives it from filenames" was unsatisfiable within Phase 1's original file scope — Phase 1 would have shipped failing its own acceptance criterion while Phase 3 inherited half a derivation.

Amended scope (maintainer decision). Phase 1 owns every live-plan-counting re-derivation repo-wide. Phase 3 narrows to milestone-window + sentinel-filter enumeration only; its files are already plan-count-clean when it starts, and its own drift guard inherits a green baseline.

Two consequential changes to Decision 1's owner surface:

  1. scanPhasePlans gains allPlanFiles (every plan on disk, pre-supersession) alongside planFiles (the live set). verify.cts conflated two questions in one loop — numbering-gap detection legitimately wants every file on disk, pairing wants the live set. The fix is for the owner to answer both explicitly, not to exempt the caller. Single ownership is preserved; the owner simply stopped under-serving. Additive — no existing field changed.
  2. findOrphanSummaries joins findUnsummarizedPlans in core-utils, sharing the same summaryCandidates rule. verify.cts needed the inverse question (summaries with no plan) and had no canonical primitive, so it had hand-rolled one — a third pairing rule of exactly the kind Decision 1 exists to prevent.

Exemptions are by documented reason, never by file allowlist (Decision 4(a)). Two sites are exempt, each carrying an inline comment stating the question it actually asks: audit.cts scanQuickTasks checks one quick task's own directory for a single completion record, and gsd2-import.cts readTasksDir reads a foreign GSD-2 tasks/ layout during a one-time import. Neither is a .planning/ phase directory.

Decision 3's Tier-2 table is re-derived for Phase 1, per its own contingency clause. Beyond the superseded-plan change, the migration also corrects: phases on the post-#3139 nested plans/ layout (previously counted as zero by every migrated site), loosely-named plan files, and stray summaries that inflated completion. The sharpest is phase.cts's cmdPhasePlanIndex, which feeds execute-phase wave scheduling — it was scheduling status: superseded plans into waves and reporting zero plans for nested-layout phases.

Regression caught during migration, recorded because it is a trap for Phases 2–5: passing the superseded-filtered planFiles into describeNonCanonicalPlans made a superseded-but-correctly-named plan report as a naming violation — the diagnostic reads non-membership as a defect. It takes allPlanFiles. The general rule: a diagnostic about file naming wants the physical set; only a question about outstanding work wants the live set. Later phases must make that choice explicitly per call site rather than swapping in planFiles mechanically.

Amendment 2 — Phase 2 (#3184) validation: the contract held; the copy count was low again

Decision 2's contract needed no change for its second consumer: SCOPE's four values covered every row of the windowing derivation's behavior table, including the two rows the epic's text does not distinguish (a free-form legacy ROADMAP with no versioned milestones is COMPLETE, not UNSCOPED — whole-document genuinely is the milestone there). Phase 2 adds no member and changes no semantics. src/planning-scope.cts needed no edit, so Phase 2 carries no .cts six-gate ripple.

What Phase 2 found. The epic and this ADR both scope milestone windowing at three copies, all inside roadmap-parser.cts. Building the Decision 4(a) whole-repo guard found two more, in a different module and one function down: state.cts buildStateFrontmatter and syncStateFrontmatter each hand-roll ^#{1,3}\s+(?!Phase\s+\S).*${escapeRegex(version)} to answer "is this milestone bounded to a versioned ROADMAP heading" — the heading-location half of the derivation, byte-identical to each other. This is Phase 1's finding repeating with a different derivation: the epic's copy counts are a lower bound derived from the reported issues, and the whole-repo guard is what makes them real. Both sites now call the owner's isMilestoneBoundedInRoadmap, which is a straight consolidation of the two identical state.cts regexes onto locateMilestoneHeadings with no behavior change — which is all it should ever have been.

A boundary tightening was tried and reverted. A first pass at locateMilestoneHeadings swapped its \b version-token boundary for the stricter (?![\w.-]) used by isMilestoneShippedInRoadmap (#2562), reasoning that v2.0 should not match inside v2.0.1 anywhere windowing happens. That broke extractCurrentMilestoneScoped's #730 contract: a milestone STATE of v8.0 legitimately selects the ## v8.0-B … active sub-milestone heading over a closed v8.0-A sibling (0 is a word character, - is not, so \b matches; (?![\w.-]) does not, because - is in its excluded set). \b is restored in locateMilestoneHeadings; the stricter boundary stays local to isMilestoneShippedInRoadmap and to the #730 detailsVersionBoundary, which answer a narrower question ("is exactly this milestone shipped" / "which Phase Details section is exactly this one's version token's") than "which heading does this milestone STATE select." The consolidation itself (three roadmap-parser.cts copies plus the two state.cts copies onto one owner) is behavior-preserving.

A composition-level re-derivation, caught in review of this phase's own diff. Decision 4(c) anticipated a consumer post-filtering an owner's result. The shape that actually appeared is its mirror: two sites re-assembling a window out of the owner's primitives — locateMilestoneHeadings → pick a heading → computeMilestoneSectionEnd → slice — in getMilestonePhaseFilter's versionOverride branch and in milestone.cts's unstarted-phase guard. Both call the canonical owner at every step, so the drift guard and an owner-level identity test are both green, and the two compositions had already diverged on whether to skip a closed milestone heading. Decision 4(c) is therefore read to cover assembly as well as post-processing: where a derivation has a composition, the composition is itself an owner. Added as sliceMilestoneWindow; both sites route through it.

Decision 3's Tier-2 table, re-derived for Phase 2 per its own contingency clause. The row this ADR predicted lands as written, plus two the prediction did not contain:

Command surface Output change
roadmap analyze gains a scope field. phase_count: 0 is still emitted verbatim — what changes is that a sibling field now says whether that zero is an answer. Stated precisely because the first draft of this row claimed the count itself changed, which is not what shipped
/gsd:progress --next Route 0 gsd-core/workflows/next.md treats a non-complete scope as scan-failed (warn + fall through to the prior-phase check) instead of looping a phase list the scan could not populate. Without this the new field would be a diagnostic no consumer reads, and #3165's actual symptom — the resume invariant reporting clean because it could not run — would still reproduce
milestone complete refuses (unless --force) when the window's scope is TRUNCATED — the milestone heading was found but its section closes before reaching any phase entries, even though the ROADMAP has phase entries elsewhere — instead of pass-all archiving every phase directory on disk (#3166). UNREADABLE and UNSCOPED are pre-existing, legitimately-handled states (missingExplicitVersion errors where that matters; a missing ROADMAP.md has its own documented graceful path) and are not refused here.
milestone complete unstarted-phase guard — not predicted the guard scoped its window by STATE.md's milestone: field while the filter beside it scoped by the version argument; the two could disagree, and the guard under-detected unstarted phases on the destructive path. Both now use the version argument.

Scope note. Phase 3 (enumeration) inherits a window layer that is now single-owner and scope-carrying; its own guard starts from a green windowing baseline, exactly as Phase 1 left plan counting clean for Phase 3.

Amendment 3 — Phase 3 (#3185) validation: the contract held; the load-bearing bug was upstream of enumeration itself

Decision 2's contract held for its third consumer: listMilestonePhaseDirs returns ScopedResult<string[]> unchanged, and SCOPE needed no new member — every case Phase 3 hit (a genuinely empty milestone, a truncated window, an unscoped/legacy ROADMAP, an unreadable ROADMAP) was already one of the four frozen values.

Declared deviation from Decision 1's provisional signature. Decision 1 locked listMilestonePhaseDirs(roadmapContent: string, phasesDir: string, deps?): ScopedResult<string[]>. That signature cannot work: the milestone window needs cwd (to read STATE.md for the active milestone version) and ws (workstream scoping), and getMilestonePhaseFilter — the post-#3184 canonical owner of "read the ROADMAP and resolve the window" — reads ROADMAP.md itself rather than accepting its content as an argument. Threading a pre-read roadmapContent string past that owner would reintroduce a second ROADMAP-reading path beside it, which is exactly the divergence class this epic removes. Shipped signature: listMilestonePhaseDirs(phasesDir, { cwd, ws, versionOverride, phaseIdConvention }). This is a signature change, not a contract change — ScopedResult<T> and SCOPE are unaffected, so it does not require re-litigating Decision 2.

The copy count was a lower bound, a third consecutive time. The epic scoped enumeration at 4 copies. The Decision 4(a) whole-repo guard (scripts/lint-phase-enumeration-drift.cjs) found 54 violations: 23 sentinel re-derivations across 8 modules, in three regex variants plus four integer-comparison forms, now consolidated onto the canonical isSentinelPhaseId (SENTINEL_RANGES [0, 999]); plus 31 unscoped phasesDir reads. Most of the 23 sentinel re-derivations tested only 999, so Phase 0 previously slipped through every one of them.

The load-bearing finding: the sentinel exclusion lived on the wrong set. The pre-existing sentinel exclusion was applied to the ROADMAP HEADING set (### Phase N: entries), not to phase-directory names — but getMilestonePhaseFilter degrades to a literal pass-all () => true when that heading set is empty, per Decision 3's documented "over-inclusive, never under-inclusive" promise. The exclusion was therefore unreachable exactly when it was needed: a backlog or pre-milestone directory has no corresponding ROADMAP heading to exclude by, so the filter that was supposed to keep it out degraded to accepting everything instead. This is the same path #3167 named, and it is why cmdStats already called getMilestonePhaseFilter and still listed backlog directories — calling the filter was not enough while the filter's own pass-all degrade could not distinguish "no phases in this milestone" from "no heading to test this directory against." The fix applies the sentinel test to directory names, unconditionally, after the window filter runs rather than folding it into the window filter's heading-matching logic. The narrowing is sentinel-only: pass-all still stands for every non-sentinel directory the window filter cannot place, so Decision 3's promise is narrowed minimally, not revoked (Decision 3 / Hyrum's Law).

#3161 is subsumed alongside #3167, as the Tier-2 table predicted. #3161 ("aggregate percent reports 100 while plans are outstanding") shared the same upstream cause: cmdStats's and cmdProgressRender's totalPlans/totalSummaries accumulation now iterates the single owner's scoped, sentinel-filtered dirs set (listMilestonePhaseDirs's value) instead of an unscoped readdirSync of the phases directory, so a 999.*/0-* directory with its own already-summarized plans can no longer inflate totalSummaries (or totalPlans) against a milestone that has not actually finished — the same backlog-dir listing bug row 3 named, manifesting in the percent aggregate rather than the phase list.

Two destructive-path defects the sweep exposed. phases clear carried the fifth sentinel copy and its third regex variant (/^999(?:\.|$)/) — it excluded 999 but not 0, so a 0-* pre-milestone directory was deleted (or, pre-#1871, hard-removed) on this irreversible path. milestone complete's phase-archival move had no sentinel filter at all on its stats/dry-run/move paths — only the milestone window — so a sentinel directory sitting inside the window's phase range could be archived alongside the milestone's own phases. Both now route through isSentinelPhaseId directly (not through listMilestonePhaseDirs, since both need every non-sentinel directory regardless of milestone window — see the generalized rule below).

An unadvertised but correct Tier-2 change. Phase 0 directories now drop out of progress/stats/phases list alongside Phase 999, because the canonical predicate treats both sentinels alike and the engine-wide convention (#1580) already declares both as sentinel ranges, while roadmap analyze (Phase 2) already honored it. This was not separately predicted by Decision 3's table — it falls out of routing every reader through one predicate that was already correct.

The guard's own false positive, worth recording. The first version of scripts/lint-phase-enumeration-drift.cjs flagged JSDoc comments and inline comments that merely documented that the code below already called the canonical owner — matching sentinel-shaped regex literals inside prose, not code. It is now comment-aware (skips block/line comments before matching). Recorded because a guard that reports prose as drift trains its readers to reflexively exempt documentation, which is the opposite of Decision 4(a)'s whole-repo, no-allowlist intent.

The generalized exemption rule, restated from Amendment 1's file-naming case in this derivation's terms: a lookup, diagnostic, archival, or mutation pass wants the physical directory set — every non-sentinel directory on disk, regardless of milestone window. Only "which phases belong to the current milestone" wants the scoped set listMilestonePhaseDirs returns. phases list --phase N and --include-archived (lookup/archive) and phases clear / milestone complete (destructive mutation) take the first; progress, stats, and the bare phases list take the second.

Decision 3's Tier-2 table, re-derived for Phase 3 per its own contingency clause:

Command surface Output change
query progress, stats, bare phases list 999.* backlog and 0-* pre-milestone directories are no longer listed or counted as current-milestone phases; the aggregate completion percentage stops reading 100 while phases in the active window are still outstanding
phases clear no longer deletes/archives a 0-* pre-milestone directory — previously excluded 999 but not 0 on this irreversible path
milestone complete its phase-archival move no longer sweeps sentinel directories into .planning/milestones/<version>-phases/ alongside the milestone's own phases
stats P0.0 plan-count correction — not predicted isDirInMilestone could not match a #1324 letter-prefixed-decimal directory (P0.0-foundation) to its own ### Phase P0.0: ROADMAP heading, so stats reported that phase with plans: 0 while its directory held real plan files. Fixed inline as part of the same sweep; not a Decision-1 owner change, but a defect the whole-repo guard's investigation surfaced in the same code path

A single owner is not always a single RULE — the 0.x split. The sharpest finding of this phase, and a correction to how Decision 1 reads. An isolated security review observed that isSentinelPhaseId classifies 0.1 / 00.1 as sentinel milestone 0 (its /^0*(\d+)/ backtracks to capture 0), and that this looked wrong against #2554. Changing the canonical predicate to exempt 0.x made the suite fail six tests, because two PINNED contracts disagree — and both are right, because they ask different questions:

Contract Question Verdict on 0.x
#2554 (roadmap-parser.test.cjs) is this directory part of the current milestone's phase SET? count it — a 00.1-<slug> dir declared as ### Phase 00.1: is a real phase
#2949 (issue-2949-phase-complete-stage3-sentinel.test.cjs) must this phase COMPLETE before the milestone can close? sentinel — a 0.x must not block is_last_phase

No single global predicate answers both. The resolution is layered, not unified: isSentinelPhaseId keeps its semantics (0.x IS a sentinel, satisfying #2949), and the milestone-WINDOW layer keeps a narrower 999-only rule (satisfying #2554), carried as a function-scoped guard exemption with a written reason rather than a second silent copy.

The lesson for Phases 4 and 5: "one owner per derivation" governs who computes an answer, not how many questions share it. Before folding a call site onto a canonical predicate, establish which question that site asks — an over-broad canonical rule is as much a defect as a divergent copy, and it fails in a worse way, because it looks like consolidation. Note also that the security review's data-completeness concern here was inference about intent, while #2949 is pinned intent; where the two conflict, the pinned contract wins and the review finding is recorded as adjudicated rather than fixed.

Scope note. Phase 3 is the last consumer of the enumeration/window layer; Phases 4 and 5 build on the completion and state-extraction derivations respectively and do not depend on listMilestonePhaseDirs.

Amendment 4 — the 2026-08-08 coverage audit: two more derivation families, a fifth enumeration copy, and a guard that was looking at half the surface

Written before Phase 3 merged, reconciled against it after. This amendment and Amendment 3 were authored concurrently and reached the same conclusion independently — the copy count is always a lower bound and only the whole-surface guard makes it real (Amendment 3 found 54 violations where the epic scoped 4). Amendment 3 states that rule; this one does not restate it. Three of this amendment's original claims were superseded by what Phase 3 actually shipped, and each is corrected in place below rather than left standing.

Source: the coverage audit posted to #3180 on 2026-08-08, which tested every open non-PR'd bug on the tracker against one question — would executing this epic's stated work, by itself, make the reported symptom stop? Four issues passed and were closed into the epic (#3164 and #3166 as already discharged by Phases 1 and 2; #3167 and #3168 as covered by open Phases 3 and 4). Seven did not. This amendment is what the seven change.

What the audit added to the copy count. A fifth enumeration copy, a third completion predicate, and two derivation families the epic never named at all. Amendment 3 states the standing rule this is the fourth instance of, and states it from a stronger position — 54 violations against a scoped 4 — so it is not restated here.

Scope changes.

# Change Why the existing phases do not cover it
1 Phase 3 widens to include buildStateFrontmatter and syncStateFrontmatter — superseded: Phase 3 merged without them. The fifth enumeration copy and #3204 are now unowned and need a phase of their own Phase 3's Done-when named only cmdProgressRender, cmdStats and cmdPhasesList, and #3222 shipped exactly that. buildStateFrontmatter/syncStateFrontmatter appear nowhere in Amendment 3, and #3204 appears nowhere in this ADR outside this row. Routing alone would not have fixed #3204 anyway — its defect is the "is the ROADMAP count trustworthy" discriminator one layer above enumeration, which is §7.1's isMilestoneBoundedInRoadmap
2 Phase 4 blocks on #2957 and its guard must fail while a third predicate exists The audit found buildStateFrontmatter computing completed phases from plan scanning alone, ignoring the ROADMAP checkbox cmdRoadmapAnalyze honors. Checkbox-override vs disk-strict is an undecided product question, not a consolidation
3 Phase 6 — milestone identity (§7.2): bind getMilestoneInfo to locateMilestoneHeadings, widen the windowing guard's token set in the same change (#3171, #3197) A sixth derivation family. Phase 2 consolidated windowing; getMilestoneInfo hand-rolls its own heading regexes to answer a different question — which milestone is this and what is it called — and no named phase touches it
4 Phase 7 — completion-ratio scoping (§7.6 rules 3–4). Corrected: #3161 is NOT fixed here — Amendment 3 subsumed it. Phase 7 keeps rule 4 and whatever consumers Phase 3 did not reach A seventh derivation family. This amendment ships the arithmetic half. Its original claim — that enumeration consolidation "changes nothing here" — was wrong, and Phase 3 proved it: routing cmdStats/cmdProgressRender's totalPlans/totalSummaries accumulation through listMilestonePhaseDirs's scoped set is §7.6 rule 3 for those two consumers, and it is what closed #3161
5 Phase 8 — the prompt layer: give the workflow layer a CLI surface to ask for plan and phase counts, and burn the ratchet baseline to zero The re-derivation lives in shell inside gsd-core/workflows/*.md. No .cts-scoped guard can see it, and no import can route it — it needs a command to call
6 Plan-lifecycle terminal states (§7.5) are declared a closed frontmatter vocabulary, with the gap recorded rather than papered over Consolidation cannot fix #1762's prose-retired plans; the canonical owner counts them live and is correct to, under the contract as written

Ordering. Phase 6 is independent of Phases 3–5 and may ship at any point. Phase 7 follows Phase 3 (its scoping is what makes rules 3–4 expressible). Phase 8 follows whichever phase first exposes the CLI surface it calls. Decision 5's locked 1→2→3→4→5 order is unchanged.

What shipped in this amendment's own change.

  • Decision 4(d) — scan surface is every authored surface, and an owner file is no longer exempt, only its named canonical functions. lint-milestone-window-drift.cjs exempted src/roadmap-parser.cts wholesale, which is constraint (a)'s forbidden allowlist pointed at the file most likely to grow the next copy — and it had: getMilestoneInfo sits inside it, invisible.
  • Decision 4(e) — the ratchet mechanism, so a surface that cannot be consolidated today is watched today.
  • Decision 7 — the behavior contract, this ADR's normative core.
  • Completion ratio consolidated: clampPercentFromFraction added beside clampPercent; six inline copies across roadmap.cts, state.cts, commands.cts (×2), workstream-inventory-builder.cts, gsd2-import.cts and state-document.cts migrated onto the owner; scripts/lint-completion-ratio-drift.cjs added, reporting zero re-derivations with no file-level exemption.
  • Prompt layer made visible: scripts/lint-planning-prompt-drift.cjs added with a shrink-only baseline covering the 7 sites across progress.md, execute-plan.md, plan-phase.md and plan-review-convergence.md. Phase 8 (#3218) owns their removal, and every baseline entry names it — Decision 4(e) requires the acknowledgment to point at the issue that removes it, not at the epic.

Decision 3's Tier-2 table, re-derived for this amendment: no rows. Every percent migration is behavior-identical — clampPercent's first line is the total > 0 ? … : 0 ternary each copy carried. The single deliberate difference is gsd2-import's pct, which gains a 100 ceiling it did not have; donePhases is a subset count of totalPhases, so the ceiling is unreachable and no emitted value changes.

Explicitly NOT absorbed, and left open on their own issues. #3165's answer remains unrepaired — Phase 2 made the truncated window decidable (SCOPE.TRUNCATED) but phase_count is still 0 and current_phase/next_phase still null, so its first acceptance criterion is unmet and the underlying document-layout ambiguity is untouched by design. #3163 belongs to #2143's sectionizer layer and is not enumerated there yet; #3169 and #3170 are standalone parser/extraction defects with no epic home. Recording them here as not covered rather than leaving them to be re-tested by the next audit.