Files
msd-core/src/phase-lifecycle.cts
Tom Boucher b9f51836e6 refactor(#3180): ADR-3180 behavior contract + cross-surface drift guardrails (#3223)
* refactor(#3180): one owner for completion ratio, a prompt-layer drift guard, and a written behavior contract

The 2026-08-08 coverage audit on #3180 found the epic's copy counts were a
lower bound for the third consecutive time, and that two derivation families
had never been named at all.

ADR-3180 gains Decision 7 — a normative behavior contract that says what the
right answer IS for each derivation, not merely who owns it. A reviewer with
no written rule can only ask "does this look like the others", which is how a
fifth copy passes review. Decision 4 gains (d) scan surface is every authored
surface and an owner FILE is never exempt, only its named functions; and (e)
a surface that cannot be consolidated today ships ratcheted, never unguarded.

Completion ratio: `clampPercent` sat exported and unused beside six hand-inlined
copies of its own body across five modules. All six now route through it;
`clampPercentFromFraction` is added for the one caller that already held a
fraction. Every migration is behaviour-identical — clampPercent's first line IS
the `total > 0 ? … : 0` ternary each copy carried. Guarded by
lint-completion-ratio-drift.cjs, which reports zero re-derivations with no
file-level exemption.

Prompt layer: workflow markdown re-derives live-plan counting in raw shell
(#1762), invisible to every `src/`-scoped guard. lint-planning-prompt-drift.cjs
scans it with a shrink-only baseline of the 7 sites that exist today — new
sites fail, and a baseline entry that stops firing fails too, so an
acknowledgment can never outlive the thing it describes.

lint-milestone-window-drift.cjs stops exempting its owner file wholesale; only
the four named canonical functions are exempt now. The blanket exemption was
pointed at the one file most likely to grow the next copy, and it had.

Refs #3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3180): link Phases 6-8 sub-issues (#3216, #3217, #3218) from ADR-3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3180): address orthogonal review — consumer-output identity tests, count-keyed ratchet, property coverage

Five findings from the two orthogonal review passes, all fixed.

Decision 4(c) breach: the completion-ratio identity test asserted at the
OWNER, which is exactly the bypass that decision exists to close — a consumer
can call clampPercent and then post-process locally, leaving both the lint and
an owner-level test green. It now drives `roadmap analyze`, `query progress`
and `stats` and asserts on their own output, over a fixture containing a
`status: superseded` plan so a consumer that re-counted raw files would report
60 where the owner reports 75.

Decision 4(e) breach: ratchet entries named the epic (#3180) rather than the
issue that removes them. They name Phase 8 (#3218) now.

The ratchet keyed on (file, text) alone, so plan-phase.md's two byte-identical
sites were one indistinguishable key and migrating either would have left the
guard green with the other alive. Entries carry an occurrence count; fewer than
acknowledged fails as a partial migration, more fails as a new copy.

Adds the missing MAX_REGEX_LITERAL_LEN boundary coverage the sibling guard's
test already had, and the fast-check property tests CONTRIBUTING requires for
clamp/budget-limit functions.

Refs #3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: stop wrapping a nested double-spawn in a 15s wall-clock budget (bug #641 probes)

`tests/ci-test-scope.test.cjs`'s `bug #641` block spawned `run-tests.cjs`
under PROBE_TIMEOUT_MS=15000; that child then spawned a nested `node --test`.
A fixed wall-clock budget around a double spawn, running inside a container
that is concurrently executing the full ~31k-test suite, fails by construction
under load.

Confirmed against three full matrix runs. Every failure was shaped
`null !== 0` — the child was KILLED, never an assertion about the thing under
test. One captured probe had already printed the correct resolution
(`suite="all" files=2: a.test.cjs b.test.cjs`) and was killed anyway. It
reproduces on `next` alone: 5 failures on linux-node22, 0 on linux-node24. The
victim subset varies by run and by lane.

What these tests are actually about is suite-token RESOLUTION — `unit` as a
bare token in --files/--files-from. Executing the seeded trivial files is
incidental and is the entire timeout surface, so the assertions move
in-process against the same functions `main()` calls, in the same order.
`parseArgs`, `selectExplicitFiles`, `selectFiles` and `walkTestFiles` are
exported for that; no behavior, signature or logic changed.

No coverage lost: `tests/run-tests-harness.test.cjs` already spawns the
harness for real and asserts exit codes end to end, on a 120s budget.

Pre-existing on `next`, fixed here rather than deferred.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: delete the three elapsed-time assertions

CLAUDE.md forbids asserting on wall-clock time. Three assertions did, and all
three are load-sensitive: on a saturated bench each can fail while the code
under test is correct. In every case the load-bearing assertion sits on the
line above and the timing line adds no discrimination.

run-with-timeout: the stated worry — "was this 124 the cap firing or the 30s
harness backstop?" — is already answered by the assertion above it. A backstop
kills by signal, which surfaces as status null, never 124. Observed directly
this session: three matrix runs produced exactly that null shape from killed
children.

normalize-test-command and context-predicates: both bounded a ReDoS check.
A threshold only ever separates "fast" from "slightly slow", which is bench
load, not correctness — catastrophic backtracking on 800 KB of input does not
take 251ms, it does not finish at all. A real regression therefore shows up as
the suite being killed on that test, which is louder and more reliable than a
number. The structural assertions (returned unchanged; cleanly rejected) are
what actually carry those tests, and they stay.

The sweep now reports zero elapsed-time assertions in tests/. The remaining
Date.now() uses are unique-path suffixes, barrier deadlines, fixture
timestamps and fake mtimes — none of them assertions.

Pre-existing on `next`, fixed here rather than deferred.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3180): backfill changeset PR number (#3223)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3180): key the prompt-drift ratchet on POSIX paths so it works on Windows

The baseline keys on (file, trimmed text). `file` came from scanTree's
`path.relative()`, which uses NATIVE separators, while the committed baseline
stores POSIX. On Windows every violation was therefore unmatched — reported as
FRESH — and every baseline entry matched nothing — reported as STALE. The guard
failed 100% of the time there, on both CI shards:

  ✖ scanRepo(repoRoot) matches the baseline exactly: zero fresh AND zero stale
    + { file: 'gsd-core\\workflows\\execute-plan.md', ... }

The remote runner this repo gates on is Linux-only and cannot see this class at
all; the GitHub Actions Windows lane is what caught it.

Normalization is unconditional — never gated on process.platform. A
platform-conditional normalizer makes the POSIX path the special case and
leaves the Windows branch unexercised on every other OS, which is the same
blind spot in a different place. It is applied at one seam inside
findPromptDrift, which builds `file` on every returned violation, so the
baseline key, the --update writer, the stderr report and the tests all consume
one normalized value.

The regression tests drive a Windows-shaped relPath directly and run on every
OS rather than skipping off-Windows — a test that only runs on the platform
where the bug lives is why this escaped. They include a sanity check that
un-normalized input does NOT match, so the assertion cannot pass vacuously.

Audited the three sibling guards: none keys against a committed cross-platform
baseline, and their exemption keys are path.join-built, so producer and
consumer share the native convention. Left correct code alone rather than
making them look alike. scripts/lib/drift-scan.cjs is untouched — normalizing
there would break those three on Windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 16:05:17 -04:00

145 lines
6.7 KiB
TypeScript

/**
* Phase Lifecycle Pure Helpers — pure-computation functions extracted from
* the phase-lifecycle SDK handler (ADR-457 build-at-publish: the hand-written
* bin/lib/phase-lifecycle.cjs collapsed to a TypeScript source of truth).
* Behaviour is preserved byte-for-behaviour from the prior hand-written .cjs;
* only types are added.
*
* I/O adapter pattern (ADR-3524 Section 4): each side supplies its own I/O
* (sync readFileSync for CJS, async readFile for SDK); the pure computation
* logic is shared via this generated artifact.
*
* Scope:
* - deriveProgressFromRoadmap(roadmapContent): count Complete rows => idempotent
* - clampPercent(completed, total): percent with 100 ceiling
*
* These two functions are the root-cause fix for issue #4.
*
* References:
* - ADR-3524 (docs/adr/3524-cjs-sdk-hard-seam.md)
* - Issue #4 (open-gsd/gsd-core)
*/
import { findTableWithColumns } from './markdown-table.cjs';
// eslint-disable-next-line @typescript-eslint/no-require-imports -- phase-id.cjs is an export= CommonJS module
import phaseIdMod = require('./phase-id.cjs');
const { isSentinelPhaseId } = phaseIdMod;
/** Result of deriveProgressFromRoadmap. */
export interface RoadmapProgress {
completedPhases: number | null;
totalPhases: number | null;
totalPlans: number | null;
}
/**
* Derive completed_phases, total_phases, and total_plans from ROADMAP content.
* Root cause fix for issue #4 — see gen-phase-lifecycle.mjs for full documentation.
*
* ADR-2143 §3 ("addressed by NAME, never ordinal"): the Progress table is
* located via the markdown-table seam's `findTableWithColumns`, which is
* column-NAME/order/count-invariant — it matches the first table whose header
* is a SUPERSET of the canonical `Phase` / `Plans Complete` / `Status` /
* `Completed` names, in any order, tolerating extra/injected unrelated
* columns (#2137's fast-check property test shuffles headers and injects
* columns and asserts the derived counts never change). This supersedes the
* earlier `findTableBySchema` exact-schema lookup, which required an exact
* canonical column SET+ORDER and returned all-null on any reordering or
* injection.
*
* Scoped to the `## Progress` section when the document has one (#2012 decoy
* avoidance — a differently-headed table sharing the same column names must
* not be picked up instead); a headingless milestone slice (#1445) falls back
* to scanning the whole input, preserving the "Progress table not under a
* `## Progress` heading, or not the first table in the document, still
* resolves" behaviour.
*
* Cells are read by column NAME (`r['Status']`, `r['Plans Complete']`,
* `r['Phase']`), fixing #2137 (the old position-based regex assumed "Status"
* was always the 3rd cell and "Plans Complete" the 2nd, which broke for the
* 5-column milestone-grouped variant that inserts a `Milestone` column ahead
* of them).
*/
export function deriveProgressFromRoadmap(roadmapContent: string): RoadmapProgress {
let completedPhases: number | null = null;
let totalPhases: number | null = null;
let totalPlans: number | null = null;
// ADR-2143 §5 (fail-loud, no null-swallow): this used to be wrapped in a
// try/catch that silently fell through to the existing (null) values on any
// thrown error. `findTableWithColumns`/`parseMarkdownTable` never throw —
// an unparseable or absent table resolves to `null` /
// `{ ok: false, reason }`, not an exception — so the catch was masking
// nothing but dead code paths. Removed per ADR-2143 §5; the public
// `RoadmapProgress` contract (nulls = absent) is unchanged.
//
// ADR-2143 §3: read the Progress table by column NAME (order/injection-invariant),
// via the markdown-table seam. Scope to the `## Progress` section when present
// (#2012 decoy avoidance); a headingless milestone slice (#1445) falls back to the
// whole input. Requires the canonical Phase/Plans Complete/Status/Completed columns
// in any order (extra columns ignored) — supersedes findTableBySchema's exact-schema lookup.
const progressMatch = roadmapContent.match(/^##[ \t]+Progress\b/im);
let scoped = roadmapContent;
if (progressMatch && progressMatch.index !== undefined) {
const afterHeading = roadmapContent.slice(progressMatch.index);
const nextHeading = afterHeading.search(/\n#{1,2}[ \t]/);
scoped = nextHeading >= 0 ? afterHeading.slice(0, nextHeading) : afterHeading;
}
const table = findTableWithColumns(scoped, ['Phase', 'Plans Complete', 'Status', 'Completed']);
if (table) {
const allRows = table.rows;
const completed = allRows.filter((r) => /^complete$/i.test((r['Status'] ?? '').trim())).length;
completedPhases = completed > 0 ? completed : null;
// Data rows only (exclude sentinel phases 0 and 999.x).
// #3185: canonical sentinel predicate (SENTINEL_RANGES [0,999]) — this was a local 999-only literal that admitted Phase 0.
const dataRows = allRows.filter((r) => {
const phase = (r['Phase'] ?? '').trim();
return /^\d/.test(phase) && !isSentinelPhaseId(phase);
});
totalPhases = dataRows.length > 0 ? dataRows.length : null;
let totalPlansSum = 0;
for (const r of allRows) {
const cell = (r['Plans Complete'] ?? '').trim();
const m = /(\d+)\s*\/\s*(\d+)/.exec(cell);
if (m) totalPlansSum += parseInt(m[2], 10);
}
totalPlans = totalPlansSum > 0 ? totalPlansSum : null;
}
return { completedPhases, totalPhases, totalPlans };
}
/**
* Compute progress percent clamped to 100 from an already-computed FRACTION.
*
* ADR-3180 Decision 7 (#3180): the completion-RATIO derivation has exactly one
* owner, and this is its kernel — the single place the `fraction -> integer
* percent` rounding and the 100 ceiling are expressed. `clampPercent` below is
* the count-shaped entry point and delegates here; a caller that already holds a
* fraction (rather than a completed/total pair) calls this directly instead of
* re-deriving `Math.min(100, Math.round(f * 100))` locally.
*
* Enforced by `scripts/lint-completion-ratio-drift.cjs`.
*/
export function clampPercentFromFraction(fraction: number): number {
return Math.min(100, Math.round(fraction * 100));
}
/**
* Compute progress percent clamped to 100.
* Root cause fix for issue #4 — see gen-phase-lifecycle.mjs for full documentation.
*
* A non-positive (or absent) denominator yields `0` — "nothing to complete" is
* reported as 0%, never as 100%. Every `.planning/` completion percentage in this
* codebase routes through here (ADR-3180 Decision 7); the `total > 0 ? ... : 0`
* ternary that used to precede each inline copy IS this function's first line.
*/
export function clampPercent(completed: number, total: number): number {
if (!total || total <= 0) return 0;
return clampPercentFromFraction(completed / total);
}