* feat(#2249): bracket phase-id core grammar — parse/render/toDir + READING-B + guards PR-1 of epic #612 (ADR-612, in-tree at docs/adr/612-bracket-phase-id-convention.md). Adds the bracket-convention grammar INSIDE src/phase-id.cts — the ADR-2121 single canonical owner — as a pure, additive extension. The 17 locked exports and PHASE_NUMBER_TOKEN_SOURCE are untouched, and normalizePhaseName is byte-identical, so the PR-0 collision anchor (tests/adr-612-collision-characterization.test.cjs) stays green. New pure round-trippable model (ADR Decision 4): - PhaseId { project, milestone, phase, subphase?, plan? }. - parsePhaseId(input): accepts display `[GSD.02] 05.03-01`, dir/token `GSD.02-05.03-slug`, or bare `GSD.02-05`; rejects ambiguous non-bracket tokens (`02-04`, `05`) rather than guessing. The rejection lives ONLY in this new parser — normalizePhaseName and every legacy reader keep accepting those tokens unchanged (conservative default; no existing path gains a throw). - renderPhaseId(id) -> `[GSD.02] 05.03-01`; toDir(id, slug) -> `GSD.02-05.03-slug` with a slug guard that sanitizes path-traversal input. - getMilestoneFromPhaseId(phaseId, convention?): READING-B derives the milestone from the `[PROJECT.MM]` prefix, gated on convention === 'bracket' and returning the `vN.0` form (parity with READING-A). The optional parameter keeps the helper pure (no config read) and byte-compatible — every existing single-arg caller resolves to the unchanged READING-A body (ADR Decision 6). - extractPhaseToken(dirName, convention?): bracket dir branch GATED on convention === 'bracket'. A bracket dir `{CODE}.{MM}-{PP}` is string-indistinguishable from the legacy #2043/#1324 letter-prefixed-decimal family (`P0.3-2`, `P0.12-34`) whenever the code ends in a digit, so no string-only discriminator is complete — an ungated auto-detect silently reinterpreted legacy reads on this CRITICAL 6-caller helper. The explicit convention signal keeps every existing convention-less call site byte-identical (pinned by a #2043 numeric-tail characterization in tests/phase-id.test.cjs). - comparator: no new code — comparePhaseNum already orders the dot-decimal `PP[.SS]` tokens extractPhaseToken yields; milestone-qualified ordering is a PR-2 resolution concern (bracketQualifiedKey), not core grammar. - SENTINEL_RANGES / isSentinelPhaseId(phaseId, convention?): {0, 999} non-milestone guard; the bracket-prefix reading is gated the same way (an ungated read called `P0.0-foundation` a sentinel), legacy leading-int form unchanged. - BRACKET_PHASE_TOKEN_SOURCE (dot-or-dash `[.-]` sub-separator; deliberately more permissive than parsePhaseId — a read-tolerance source for PR-2, not the emit grammar) and PHASE_HEADING_PREFIX_SRC exported from the drift-guard-exempt owner so PR-2 builds every bracket read regex from the canonical source and check:phase-id-drift stays green stack-wide. The bracket project code follows the repo's config-validated `[A-Z][A-Z0-9_]*` grammar (not the ADR §1 illustration's `[A-Z]{1,6}`), so every project_code the config permits parses. parsePhaseId has no live callers in PR-1, so this grammar choice is forward-facing for PR-2 with zero PR-1 behavior impact. Tests: tests/adr-612-bracket-grammar.test.cjs (28) — ADR §3 example round-trips, full 5-tuple parse, READING-B (+ legacy-unchanged and sentinel cases), extractPhaseToken bracket ON/OFF, comparator ordering of extracted tokens, sentinel + slug guards, bare-token rejection, exported-source behavioral assertions, and two generative fast-check properties: render∘parse identity over well-formed displays, and the toDir/disk↔display bijection. Plus a #2043 numeric-tail characterization (single- AND multi-digit rows) in tests/phase-id.test.cjs pinning the convention-less reading byte-identical. The compiled gsd-core/bin/lib/phase-id.cjs is gitignored (ADR-457 build-at-publish) and rebuilt by CI, so it is intentionally not committed. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(#2249): changeset fragment for PR #2258 (docs-exempt: internal grammar behind flag) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2249): reject non-canonical phase-id input + harden toDir (review B1/M1-M3) PR-1 CHANGES_REQUESTED follow-up (epic #612, ADR-612 Decision 4). B1 (blocker): parsePhaseId accepted non-canonical input (unpadded numbers, over-padded numbers, multi-space separators, stray whitespace), so render(parse(x)) === x did not hold for every well-formed x as ADR-612 Decision 4 requires. Both branches now enforce canonicality by construction: parse permissively, rebuild the canonical string via the same emit path (renderPhaseId for display, a hand-rebuilt token for dir/token), and throw "parsePhaseId: not canonical" on any mismatch. The .trim() at the parser's entry is removed — the match anchors now reject leading/trailing whitespace outright, folding into the existing "not a bracket phase id" rejection. M1 (major): toDir only ever guarded the slug; project/milestone/phase/ subphase were interpolated unsanitized, so a hand-built PhaseId (a structural, not nominal, type) could smuggle a path-traversal segment onto disk. Every field is now validated against the exact shape parsePhaseId itself would produce before use. M2 (major): a slug that sanitized to empty (e.g. '!!!') left a dangling trailing hyphen in the emitted dir name. toDir now throws in that case. M3 (major): an all-digit slug (e.g. '2026') was string-indistinguishable from the dir-branch's plan tail, so it silently broke the disk<->identity bijection on read-back. toDir now rejects all-digit slugs. Nits: toDir now rejects a non-string slug instead of coercing it to the literal token 'undefined'/'null'; sentinel boundary tests added for milestones 1/998/1000 (SENTINEL_RANGES is the two discrete values {0, 999}, not an inclusive range — these were already correct, now locked by test). Test-first: every new assertion (concrete examples + fast-check mutation property for B1; concrete cases for M1-M3 and the nits) was written and confirmed red before the implementation changes, per repo TDD convention. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(#2249): reformat changeset body to house convention (review Mi2) The fragment added in ab26190a was a plain paragraph — no bold headline, no trailing issue reference. Reformat to the repo's `**Bold headline** — symptom/explanation. (#issue)` body shape (see e.g. .changeset/agile-pandas-dance.md, .changeset/fierce-pumas-gather.md). Uses (#2249), the issue every commit on this branch references, not the PR number already carried in frontmatter (`pr: 2258`) — the changelog serializer appends `(#{pr})` unconditionally, so a body also ending in `(#2258)` would double-render as `(#2258) (#2258)`. Verified the rendered bullet directly via parseFragment + serializeChangelog: it now reads `... (#2249) (#2258)`, matching the dominant convention across the other fragments (frontmatter pr = merged PR, body reference = originating issue). Also moved the docs-exempt marker back before the paragraph -> after it (matching the file's original order): the marker sits on its own line and is stripped before the body is used, but placing it first left a leading blank line in front of the bold headline once reformatted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(#2249): widen property generators — 3+-digit numerics + subphase-pad mutation (re-review Minor 1/2) PR-1 re-review follow-up (epic #612, ADR-612 Decision 4). Test-only: closes two property-generator coverage gaps the reviewer flagged; no source change (src/phase-id.cts and gsd-core/bin/lib/phase-id.cjs are byte-unchanged). Minor 1 (3+-digit numerics never exercised): numArb capped at 99, so no property fed a 3+-digit milestone/phase/subphase/plan through parse/render/ toDir despite CANONICAL_NUMERIC_RE's dedicated `[1-9]\d{2,}` branch. Widen numArb to 1–999 so the round-trip and disk↔display bijection properties both span 3-digit widths (pad2 passes ≥3-digit values through un-truncated with no leading zero, so canonicality still holds). Add a concrete regression pinning the reviewer's hand-traced example: '[GSD.100] 05' round-trips, renders, and toDirs to 'GSD.100-05-feature' without truncation. Minor 2 (no subphase-pad mutation): the B1 mutation-rejection property covered milestone/phase pad + whitespace mutations but never a subphase pad. Add unpad-subphase / overpad-subphase to the mutation set and a generated `includeSub` boolean that decides whether the canonical carries a `.SS` (forced in for the subphase mutations so there is always a `.SS` to mutate); non-subphase mutations keep their original no-subphase coverage. Non-vacuity verified against the compiled lib by temporarily probing each widened/new property and confirming it fails: round-trip counterexample ["A",100,1,…] and bijection counterexample ["A",1,100,…,"a"] prove 3-digit tokens are genuinely generated and reach the body; a no-op unpad-subphase mutation trips the mutated===canonical guard (counterexample ["A",1,1,1,false,"unpad-subphase"]), proving the subphase branch is reached with a subphase present. Probes reverted; numRuns unchanged. Gates: tests/adr-612-bracket-grammar.test.cjs 44 pass / 0 fail; `npm run test:unit` 1079 pass / 0 fail; `npm run lint:ci` exit 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2249): consume the #2232 continuation seam at the bracket token's slug-adjacent position (review Major) BRACKET_PHASE_TOKEN_SOURCE was a sixth continuation-recognition site that re-derived the grammar as an unbounded `\d+` literal instead of consuming PHASE_CONTINUATION_SEGMENT_SOURCE, re-opening the #2232 bug class on the bracket path: a PR-2 reader interpolating it over dir `PROJ.01-14-2026-photos-…` (a slug whose first word is a year) over-collected the token as `01-14-2026` instead of `01-14`. Interpolating the cap verbatim at every position was rejected on evidence: the bracket run is `MM-PP[.SS][-LL]` and only the LAST position is slug-adjacent. The exactly-2 cap at the others would under-collect ids toDir itself emits — `PROJ.02-105-slug` (3-digit phase) reads as `02`, `[GSD.02] 05.100` (3-digit sub-phase) as `05` — because CANONICAL_NUMERIC_RE admits `[1-9]\d{2,}` and `[GSD.100] 05` is a pinned regression. Those positions are delimiter- disambiguated (a required field separator; a dot a slug can never contain), not heuristically recognized, so they have no year collision to defend against. Upstream draws the same line for the same reason: core-utils/phase cap the paired PLAN component while the leading phase component stays unbounded. So the run is now positional rather than a free `(?:[.-]\d+)*` repetition, and each position takes the width its delimiter affords: leading unbounded, dash-1 and dot canonical, and the slug-adjacent dash-2 interpolating the single-owner seam. The accepted trade-off is #2232's policy verbatim: a PLAN ≥100 is out of the token grammar. Also derives CANONICAL_NUMERIC_RE from the new BRACKET_CANONICAL_NUMERIC_SOURCE instead of re-spelling it as a literal, so the emit-side gate and the read-side token source are one rule — the same single-owner discipline this fix is about. Behaviour-identical (the anchors make the source's `(?!\d)` guard redundant). Refs #2249 * test(#2249): pin the bracket/#2232 reconciliation — parity surface 6 + divergence gate + property (review Major) The comment block alone cannot hold the divergence: src/phase-id.cts is exempt from the #2128 drift guard by construction, so lint-phase-id-drift.cjs would not catch the bracket token source drifting from the seam. Per the Generative Fix Divergence rule, the divergence is pinned behaviorally instead. Surface 6 joins the existing #2232 parity gate rather than starting a rival one: the review named the bracket token source "a sixth continuation-recognition site", and continuation-grammar-parity.test.cjs is already the invariant-named home where the five #2043 sites agree with the owner on a shared width corpus. Surface 6 asserts the same contract at the bracket run's slug-adjacent position (`01-14-<seg>-photos-…`, mirroring surface 1 with the extra milestone level), so the bracket path now fails the same gate the other five do. A second block pins the DELIBERATE half — the wider canonical width at the delimiter-disambiguated positions, plus the accepted bound (a plan >=100 is out of the grammar). Without it, "unifying" bracket onto the exactly-2 cap would look like a cleanup rather than a regression. The generative property ties the READ side to the EMIT side metamorphically: for every id toDir can produce, BRACKET_PHASE_TOKEN_SOURCE must collect exactly that id's numeric run — no more, no less. It needed a new arbitrary: the existing slugArb generates one [a-z0-9] word and so can never produce the number-leading slug the collision requires. Probe-falsified, both directions (probes reverted): - reverting the source to the old unbounded `\d+` fails 8: the parity gate reports `"01-14-2026-photos-performance" collected "01-14-2026"` — the review's scenario verbatim — and the property shrinks to ["A",1,1,undefined,"100-a"]. - interpolating the seam at EVERY position (the rejected verbatim option) leaves the repro and parity green but fails the divergence gate `'02' !== '02-105'` and the property at ["A",1,1,100,"100-a"] (3-digit sub-phase), which is the evidence that a verbatim cap under-collects ids toDir emits. Width 2 stays green under both probes — the corpus agrees with the owner exactly where the old and new rules coincide, so the gate discriminates rather than merely mirroring the regex. Refs #2249 * docs(#2249): add the new phase-id exports to the CONTEXT.md glossary bullet (round-4 Major) * test(#2249): pin deterministic grammar boundary cases (re-review m1) PR-1 re-review follow-up (epic #612, ADR-612 Decision 4). Test-only: closes the m1 proof gap — the grammar's bounds were exercised only incidentally through the fast-check domain (1-999, [a-z0-9] slugs). No source change (src/phase-id.cts and gsd-core/bin/lib/phase-id.cjs byte-unchanged). Adds a deterministic boundary block (7 describe groups, +22 tests) pinning the compiled lib's CURRENT behavior — a proof gap, not a behavior gap: - m1.1 numeric-width 99/100/101 at milestone/phase/subphase/plan: parse (display + dir) -> render/toDir round-trip byte-equality. The plan position is identity-symmetric (parse/render accept 99/100/101) but toDir drops it (filename-surface dimension only). - m1.2 read-token width is POSITIONAL: BRACKET_PHASE_TOKEN_SOURCE absorbs 99/100/101 at milestone/phase/subphase (delimiter-disambiguated) but caps the slug-adjacent plan (dash-2) at exactly 2 digits — plan >=100 is out of the token grammar (#2232 seam). Pinned as asymmetry, NOT symmetry. - m1.3 leading-zero 007 -> not-canonical rejection at every position/form. - m1.4 slug abuse: parse DROPS a null-byte/control/unicode/emoji trailing slug (never stored, never mis-read as a plan) and rejects a line terminator; toDir's allow-list sanitizer collapses each to a safe [a-z0-9-] token or rejects sanitize-to-empty. - m1.5 absolute-path slug sanitizes (next to the ../../etc traversal test); an absolute-path project on a hand-built id is rejected by PROJECT_ID_RE; an abs-path string is not a bracket id; an abs-path dir slug is dropped to a clean tuple. - m1.6 whitespace-only -> not-a-bracket-phase-id. - m1.7 very-long input (10k) resolves promptly (ReDoS smoke, behavioral): garbage/partial-prefix throw; a 10k-char slug parses (dropped)/sanitizes. No accept-not-reject case is a src bug: parse never STORES an abusive slug (dropped from the identity tuple) and toDir independently re-sanitizes on emit, so the only slug reaching disk is allow-listed. Plan >=100 accepted by parse is the documented positional design (toDir drops the plan; the read-token caps it) — divergence pinned, not papered over. Probe-falsify: corrupted one assertion in each of the 7 groups (m1.4 both its parse-side and emit-side), ran -> 8 distinct named failures, reverted -> 66/66 green. Confirms every new group executes and can fail. Gates: tests/adr-612-bracket-grammar.test.cjs 66 pass / 0 fail; grammar + continuation-grammar-parity + collision-characterization + phase-id family 175 pass / 0 fail; `npm run lint:ci` exit 0. `npm run test:unit` is green except one pre-existing, unrelated env failure (npm-integrity-gate: a live npm-audit advisory in the production dep tree — reproduces with this change stashed; no package.json/lock change here). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
690 lines
34 KiB
TypeScript
690 lines
34 KiB
TypeScript
/**
|
|
* Pure phase-id parsing/matching helpers — normalize, token match,
|
|
* milestone/phase-dir id parsing, phase-markdown regex builders.
|
|
*
|
|
* Extracted from core.cts (ADR-857 rollout phase 2a / issue #865).
|
|
* The hand-written bodies are preserved byte-for-behaviour; only the module
|
|
* boundary moved. The core.cjs re-export spine was retired in epic #1267;
|
|
* callers import phase-id helpers from phase-id.cjs directly.
|
|
*
|
|
* Dependencies: none (pure string/regex, no Node built-ins required).
|
|
*/
|
|
|
|
// ─── Phase-id helpers ─────────────────────────────────────────────────────────
|
|
|
|
function escapeRegex(value: unknown): string {
|
|
return String(value).replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
|
|
}
|
|
|
|
// project_code values start with an uppercase letter (e.g. PROJ, APP_CODE);
|
|
// leading underscores are not valid project codes per .planning/config.json.
|
|
const PROJECT_CODE_PREFIX_STRIP_RE = /^[A-Z][A-Z0-9_]*-(?=\d)/;
|
|
const PROJECT_CODE_PREFIX_STRIP_RE_I = /^[A-Z][A-Z0-9_]*-(?=\d)/i;
|
|
const PROJECT_CODE_PREFIX_CAPTURE_RE_I = /^([A-Z][A-Z0-9_]*)-(\d.*)/i;
|
|
const OPTIONAL_PROJECT_CODE_PREFIX_SOURCE = '(?:[A-Z][A-Z0-9_]*-)?';
|
|
|
|
// #1729: phase headers may carry a parenthetical tag between the number and the
|
|
// colon, e.g. `### Phase 26 (Cluster B): Title`. This optional, non-capturing
|
|
// fragment is injected at every phase-header regex call site (immediately after
|
|
// the phase-number token, before the colon/space delimiter) so the resolver
|
|
// tolerates the tag — mirroring how `[...]` is already tolerated before `Phase`.
|
|
// `[^)\n]*` keeps the match single-line (headers are one line) to avoid
|
|
// over-consuming across a malformed multi-line document. Injected at the call
|
|
// site (not baked into phaseMarkdownRegexSource) so it applies uniformly to
|
|
// both the numeric and project-code-exact escaped sources, and so the decimal
|
|
// sub-phase patterns can place it after the `.N` segment.
|
|
//
|
|
// Enumeration/parse call sites that read phase headers from a regex *literal*
|
|
// (rather than a `new RegExp` built from an interpolated phase number) cannot
|
|
// reference this constant; they inline its literal-regex mirror instead —
|
|
// `(?:\s*\([^)\n]{0,200}\))?` — kept character-for-character equivalent to this
|
|
// source. Both forms must change together; see the #1729 regression test.
|
|
const OPTIONAL_PHASE_TAG_SOURCE = '(?:\\s*\\([^)\\n]{0,200}\\))?';
|
|
|
|
// #2128: the canonical phase-NUMBER-TOKEN grammar — a phase number with an
|
|
// optional single-letter variant suffix and optional dotted sub-phases
|
|
// (1, 01, 12A, 12.1, 3.2.1). This is the ENUMERATION/scan counterpart to
|
|
// phaseMarkdownRegexSource: use phaseMarkdownRegexSource(n) to build a source
|
|
// for ONE KNOWN number; reference this constant when a call site must match ANY
|
|
// phase and capture its token. Enumeration/parse sites inline this into a
|
|
// `new RegExp(...)` instead of re-deriving the grammar as a literal, so every
|
|
// phase-token producer shares one owner. The anti-divergence guard
|
|
// (scripts/lint-phase-id-drift.cjs) fails CI if a literal re-derivation is
|
|
// introduced outside this module without a `// phase-id-owner:` justification.
|
|
const PHASE_NUMBER_TOKEN_SOURCE = '\\d+[A-Z]?(?:\\.\\d+)*';
|
|
|
|
// #2232: the canonical CONTINUATION-segment grammar — a dash-separated segment
|
|
// that extends a phase token (a zero-padded sub-phase or plan number, e.g. the
|
|
// "01" in "02-01-setup"). getPhaseDirFromPhaseId writes these zero-padded to
|
|
// exactly 2 digits, so the digit RUN of a genuine continuation is exactly 2:
|
|
// #2043's `\d{2,}` (2-or-more) over-collected a slug word that merely leads
|
|
// with ≥2 digits (a year: "14-2026-photos-…" yielded token "14-2026", so every
|
|
// phase-locating verb reported the phase as missing). The `(?!\d)` guard caps
|
|
// the run at 2 without anchoring what may follow, so call sites keep their own
|
|
// trailing grammar (letter suffixes, dotted sub-phases, segment boundaries).
|
|
// POLICY (locked by boundary tests): sub-phase/plan numbers ≥100 are out of the
|
|
// dir-token grammar — the LEADING phase number stays unbounded (`\d+`), only
|
|
// continuation segments are width-capped. Shared from here so the five #2043
|
|
// call sites cannot drift independently (see scripts/lint-phase-id-drift.cjs).
|
|
const PHASE_CONTINUATION_SEGMENT_SOURCE = '\\d{2}(?!\\d)';
|
|
const PHASE_CONTINUATION_SEGMENT_PREFIX_RE = new RegExp(`^${PHASE_CONTINUATION_SEGMENT_SOURCE}`);
|
|
function isPhaseContinuationSegment(seg: string): boolean {
|
|
return PHASE_CONTINUATION_SEGMENT_PREFIX_RE.test(seg);
|
|
}
|
|
|
|
// #612 (PR-1): bracket-convention token/heading sources, kept next to the M-NN
|
|
// PHASE_NUMBER_TOKEN_SOURCE so this owner file stays the single origin of every
|
|
// phase-token grammar. `src/phase-id.cts` is exempt from the #2128 drift guard
|
|
// (scripts/lint-phase-id-drift.cjs) by construction, and that guard fails any
|
|
// literal re-derivation of the token grammar elsewhere — so the downstream
|
|
// bracket readers (PR-2: roadmap/validate/verify) must build their regexes by
|
|
// interpolating these exports, never by copying the literal.
|
|
//
|
|
// The canonical numeric WIDTH of a bracket identity field, mirroring pad2()'s
|
|
// output: exactly 2 digits, or 3+ with no leading zero. Owned here as a SOURCE
|
|
// so the read side (BRACKET_PHASE_TOKEN_SOURCE, below) and the emit-side
|
|
// validator (CANONICAL_NUMERIC_RE, which toDir enforces) are one rule rather
|
|
// than two literals that agree today and drift tomorrow.
|
|
const BRACKET_CANONICAL_NUMERIC_SOURCE = '(?:[1-9]\\d{2,}|\\d{2})';
|
|
|
|
// BRACKET_PHASE_TOKEN_SOURCE differs from PHASE_NUMBER_TOKEN_SOURCE by a
|
|
// dot-OR-dash sub-separator: a bracket dir/heading numeric run is `MM-PP[.SS]`
|
|
// (a hyphen joins milestone↔phase, a dot joins phase↔sub-phase), whereas M-NN
|
|
// sub-phases are dot-only.
|
|
//
|
|
// The run is POSITIONAL, not a free repetition — `MM-PP[.SS][-LL]` — and each
|
|
// position gets the width its DELIMITER can actually afford:
|
|
//
|
|
// MM leading unbounded — delimited by the `{CODE}.` prefix
|
|
// -PP dash-1 canonical — the grammar REQUIRES this dash, so it is a field
|
|
// separator, not a continuation heuristic
|
|
// .SS dot canonical — a slug carries no dot (toDir sanitizes them
|
|
// away), so this position cannot collide
|
|
// -LL dash-2 #2232 cap — the ONLY slug-adjacent position, and therefore
|
|
// the only one a slug word can collide with
|
|
//
|
|
// #2232 reconciliation: the slug-adjacent position interpolates the single-owner
|
|
// PHASE_CONTINUATION_SEGMENT_SOURCE, so the #2232 bug class cannot reopen on the
|
|
// bracket path — dir `PROJ.01-14-2026-photos-…` (a slug leading with a year)
|
|
// yields `01-14`, never `01-14-2026`.
|
|
//
|
|
// DELIBERATE DIVERGENCE from the M-NN dir-token path (pinned by the parity gate
|
|
// in tests/continuation-grammar-parity.test.cjs, which fails if these two rules
|
|
// drift for a reason nobody intended): the non-slug-adjacent positions stay
|
|
// WIDER than #2232's cap. Bracket admits 3+-digit milestone/phase/sub-phase
|
|
// (CANONICAL_NUMERIC_RE — `[GSD.100] 05` is a pinned regression), and unlike the
|
|
// M-NN continuations those positions are delimiter-disambiguated rather than
|
|
// heuristically recognized, so there is no year collision to defend against.
|
|
// Interpolating the cap verbatim at every position would only under-collect ids
|
|
// that toDir itself emits: `PROJ.02-105-slug` (3-digit phase) would read as
|
|
// `02`, and `[GSD.02] 05.100` (3-digit sub-phase) as `05`. Upstream draws this
|
|
// same line for the same reason — core-utils/phase cap the paired PLAN component
|
|
// while the leading phase component stays unbounded (phase numbers ≥100 are
|
|
// legitimate). The trade-off this accepts is #2232's policy verbatim: a PLAN
|
|
// ≥100 is out of the token grammar.
|
|
//
|
|
// Still deliberately MORE PERMISSIVE than parsePhaseId's strict grammar (it
|
|
// admits a letter-suffixed and unpadded leading token that the parser rejects):
|
|
// this is a READ-TOLERANCE source for the PR-2 readers, which must recognize a
|
|
// bracket-shaped token before deciding what to do with it — it is not the
|
|
// emit/identity grammar. parsePhaseId stays the arbiter of well-formedness.
|
|
const BRACKET_PHASE_TOKEN_SOURCE =
|
|
`\\d+[A-Z]?` +
|
|
`(?:-${BRACKET_CANONICAL_NUMERIC_SOURCE}(?!\\d))?` +
|
|
`(?:\\.${BRACKET_CANONICAL_NUMERIC_SOURCE}(?!\\d))?` +
|
|
`(?:-${PHASE_CONTINUATION_SEGMENT_SOURCE})?`;
|
|
|
|
// A phase HEADING intro under bracket is either a `[...]` bracket (optionally
|
|
// followed by a `Phase ` label) or a bare `Phase ` label; a bare number is NOT
|
|
// a phase-heading intro. The `[^\]]{1,200}` bound mirrors the existing
|
|
// roadmap-parser heading regexes (ReDoS-safe: a header is one short line).
|
|
const PHASE_HEADING_PREFIX_SRC = '(?:\\[[^\\]]{1,200}\\]\\s*(?:Phase\\s+)?|Phase\\s+)';
|
|
|
|
function stripProjectCodePrefix(value: unknown, caseInsensitive = true): string {
|
|
const input = String(value);
|
|
const re = caseInsensitive ? PROJECT_CODE_PREFIX_STRIP_RE_I : PROJECT_CODE_PREFIX_STRIP_RE;
|
|
return input.replace(re, '');
|
|
}
|
|
|
|
function hasProjectCodePrefix(value: unknown): boolean {
|
|
return PROJECT_CODE_PREFIX_STRIP_RE_I.test(String(value));
|
|
}
|
|
|
|
function normalizePhaseName(phase: unknown): string {
|
|
const str = String(phase);
|
|
// Strip optional project_code prefix (e.g., 'CK-01' → '01')
|
|
const stripped = stripProjectCodePrefix(str, false);
|
|
// Milestone-prefixed phase IDs: M-NN or M-N-N (deep decomposition).
|
|
const milestoneMatch = stripped.match(/^(\d+)((?:-\d+)+)([A-Z]?(?:\.\d+)*)$/i);
|
|
if (milestoneMatch) {
|
|
const major = milestoneMatch[1].padStart(2, '0');
|
|
const subSegments = milestoneMatch[2].slice(1).split('-').map(s => s.padStart(2, '0'));
|
|
const suffix = milestoneMatch[3] || '';
|
|
return `${major}-${subSegments.join('-')}${suffix}`;
|
|
}
|
|
// Standard numeric phases: 1, 01, 12A, 12.1
|
|
const match = stripped.match(/^(\d+)([A-Z])?((?:\.\d+)*)/i);
|
|
if (match) {
|
|
const padded = match[1].padStart(2, '0');
|
|
// Preserve original case of letter suffix (#1962).
|
|
const letter = match[2] || '';
|
|
const decimal = match[3] || '';
|
|
return padded + letter + decimal;
|
|
}
|
|
// Custom phase IDs (e.g. PROJ-42, AUTH-101): return as-is
|
|
return str;
|
|
}
|
|
|
|
function getMilestoneFromPhaseId(phaseId: unknown, convention?: string): string | null {
|
|
// READING-B (#612): under the bracket convention the milestone comes from the
|
|
// `[PROJECT.MM]` / `{CODE}.{MM}-` prefix, never the phase-token leading
|
|
// integer (ADR-612 Decision 6). Gated on 'bracket' so the `null` and
|
|
// 'milestone-prefixed' (M-NN) paths keep the legacy leading-int rule
|
|
// (READING-A) below, byte-untouched. The optional parameter keeps this helper
|
|
// pure (no config read) and backward-compatible: every existing single-arg
|
|
// caller resolves to the unchanged READING-A body.
|
|
if (convention === 'bracket') {
|
|
const b = String(phaseId).match(/^([A-Z][A-Z0-9_]*)\.(\d+)/);
|
|
if (!b) return null;
|
|
const mm = parseInt(b[2], 10);
|
|
if (SENTINEL_RANGES.includes(mm)) return null; // sentinel milestones have no real milestone
|
|
return `v${mm}.0`;
|
|
}
|
|
const stripped = stripProjectCodePrefix(phaseId);
|
|
const m = stripped.match(/^0*(\d+)-\d/);
|
|
if (!m) return null;
|
|
const major = parseInt(m[1], 10);
|
|
if (major === 0 || major === 999) return null;
|
|
return `v${major}.0`;
|
|
}
|
|
|
|
function getPhaseDirFromPhaseId(phaseId: unknown, phaseName: string | null | undefined, projectCode: string | null | undefined): string | null {
|
|
const stripped = stripProjectCodePrefix(phaseId);
|
|
const m = stripped.match(/^0*(\d+)-(0*(\d+(?:-\d+)*))$/);
|
|
if (!m) return null;
|
|
const milestone = String(parseInt(m[1], 10)).padStart(2, '0');
|
|
const subParts = m[2].split('-').map(p => String(parseInt(p, 10)).padStart(2, '0'));
|
|
const sub = subParts.join('-');
|
|
const slug = phaseName
|
|
? phaseName.toLowerCase().replace(/[^a-z0-9]+/g, '-').replace(/^-+|-+$/g, '')
|
|
: '';
|
|
const parts = [milestone, sub, slug].filter(Boolean);
|
|
const base = parts.join('-');
|
|
return projectCode ? `${projectCode}-${base}` : base;
|
|
}
|
|
|
|
// ─── Bracket phase-ID grammar (#612, PR-1) ──────────────────────────────────
|
|
// One pure round-trippable model (ADR-612 §3 / Decision 4). parsePhaseId
|
|
// accepts the display form `[PROJECT.MM] PP[.SS][-LL]` or the on-disk/token form
|
|
// `{PROJECT}.{MM}-{PP}[.{SS}][-{LL|slug}]`; renderPhaseId / toDir are its two
|
|
// emitters. READING-B: the milestone lives in the `[PROJECT.MM]` prefix, so no
|
|
// token dimension is ever overloaded (the M-NN collapse pinned in
|
|
// tests/adr-612-collision-characterization.test.cjs cannot occur on this path).
|
|
// `plan` is a filename-surface dimension only — renderPhaseId emits it; toDir
|
|
// drops it (directories carry a slug, not a plan). The project code follows the
|
|
// repo's established `[A-Z][A-Z0-9_]*` grammar (the config-validated
|
|
// project_code shape shared with OPTIONAL_PROJECT_CODE_PREFIX_SOURCE), not the
|
|
// ADR §1 illustration's `[A-Z]{1,6}`, so that every project_code the config
|
|
// permits (digits / underscore / >6 chars) parses.
|
|
//
|
|
// Strict-reject posture (ADR-612 Decision 4's `render(parse(x)) === x`
|
|
// contract, held exactly): parsePhaseId accepts ONLY the canonical form of
|
|
// each branch — unpadded numbers, over-padded numbers, and multi-space or
|
|
// stray leading/trailing whitespace are all rejected rather than silently
|
|
// normalized, so two distinct input strings can never parse to the same
|
|
// tuple while one of them fails to round-trip. toDir mirrors this on the
|
|
// write side: every interpolated PhaseId field is validated (PhaseId is a
|
|
// structural type — nothing forces callers through parsePhaseId, so a hand-
|
|
// built id must not be able to smuggle a path-traversal segment onto disk),
|
|
// and the slug must sanitize to a non-empty, non-all-digit token (an empty
|
|
// slug would leave a dangling trailing hyphen; an all-digit slug is
|
|
// string-indistinguishable from the plan grammar's trailing tail and would
|
|
// silently break the disk↔identity bijection on read-back).
|
|
type PhaseId = {
|
|
project: string; // 'GSD'
|
|
milestone: string; // '02' (zero-padded, from the bracket/dir prefix)
|
|
phase: string; // '05' (zero-padded)
|
|
subphase?: string; // '03' (optional)
|
|
plan?: string; // '01' (filename surface only)
|
|
};
|
|
|
|
const pad2 = (n: string): string => String(parseInt(n, 10)).padStart(2, '0');
|
|
|
|
function parsePhaseId(input: string): PhaseId {
|
|
// No .trim(): the match anchors (`^`...`$`) then reject leading/trailing
|
|
// whitespace outright, folding that case into the same "not a bracket
|
|
// phase id" rejection below rather than needing its own check.
|
|
const str = String(input);
|
|
|
|
// Display form: [PROJECT.MM] PP[.SS][-LL]. The match itself stays
|
|
// permissive on purpose (it will happily match an unpadded number or a
|
|
// multi-space run) — canonicality is enforced UNIFORMLY below via the
|
|
// render round-trip (ADR-612 Decision 4) rather than by hand-tuning every
|
|
// numeric / whitespace sub-pattern, so a field added later inherits the
|
|
// check for free instead of needing its own regex micro-surgery.
|
|
const disp = str.match(/^\[([A-Z][A-Z0-9_]*)\.(\d+)\]\s+(\d+)(?:\.(\d+))?(?:-(\d+))?$/);
|
|
if (disp) {
|
|
const id: PhaseId = { project: disp[1], milestone: pad2(disp[2]), phase: pad2(disp[3]) };
|
|
if (disp[4] !== undefined) id.subphase = pad2(disp[4]);
|
|
if (disp[5] !== undefined) id.plan = pad2(disp[5]);
|
|
// Canonicality by construction: re-render the parsed id and require
|
|
// byte-equality with the input. This rejects unpadded ('[GSD.5] 5'),
|
|
// over-padded ('[GSD.005] 05'), and multi-space-separated ('[GSD.02] 05')
|
|
// variants uniformly, without special-casing any one of them — the emit
|
|
// path (renderPhaseId) is the single source of truth for "canonical".
|
|
if (renderPhaseId(id) !== str) {
|
|
throw new Error(`parsePhaseId: not canonical: ${JSON.stringify(input)}`);
|
|
}
|
|
return id;
|
|
}
|
|
|
|
// Dir / token form: {PROJECT}.{MM}-{PP}[.{SS}][-{plan|slug}]
|
|
const dir = str.match(/^([A-Z][A-Z0-9_]*)\.(\d+)-(\d+)(?:\.(\d+))?(?:-(.+))?$/);
|
|
if (dir) {
|
|
const id: PhaseId = { project: dir[1], milestone: pad2(dir[2]), phase: pad2(dir[3]) };
|
|
if (dir[4] !== undefined) id.subphase = pad2(dir[4]);
|
|
// Trailing segment: a pure-integer tail is the plan; anything else is a
|
|
// slug (dropped from the tuple — it is not an identity dimension). The
|
|
// plan tail participates in the canonicality check below; the slug tail
|
|
// is read-tolerant pass-through (a slug is not an identity dimension) and
|
|
// is exempt from it.
|
|
const tail = dir[5];
|
|
const tailIsPlan = tail !== undefined && /^\d+$/.test(tail);
|
|
if (tailIsPlan) id.plan = pad2(tail);
|
|
|
|
// Canonicality by construction, mirroring the display branch: rebuild the
|
|
// exact dir/token string this id would emit and require it match the
|
|
// input verbatim. Rejects unpadded milestone/phase ('GSD.2-5') and
|
|
// unpadded plan tails ('GSD.02-05-1') without special-casing either.
|
|
const sub = id.subphase ? `.${id.subphase}` : '';
|
|
const tailOut = tail === undefined ? '' : tailIsPlan ? `-${pad2(tail)}` : `-${tail}`;
|
|
const canonical = `${id.project}.${id.milestone}-${id.phase}${sub}${tailOut}`;
|
|
if (canonical !== str) {
|
|
throw new Error(`parsePhaseId: not canonical: ${JSON.stringify(input)}`);
|
|
}
|
|
return id;
|
|
}
|
|
|
|
// Ambiguous / bare tokens (e.g. `02-04`, `05`, `2-01`) match neither branch,
|
|
// as does a display/dir form carrying leading/trailing whitespace (the
|
|
// anchors never match it): reject rather than guess a tuple (ADR-612
|
|
// conservative default). The rejection lives ONLY in this new parser —
|
|
// normalizePhaseName and every other legacy reader keep accepting those
|
|
// tokens unchanged.
|
|
throw new Error(`parsePhaseId: not a bracket phase id: ${JSON.stringify(input)}`);
|
|
}
|
|
|
|
function renderPhaseId(id: PhaseId): string {
|
|
const sub = id.subphase ? `.${id.subphase}` : '';
|
|
const plan = id.plan ? `-${id.plan}` : '';
|
|
return `[${id.project}.${id.milestone}] ${id.phase}${sub}${plan}`;
|
|
}
|
|
|
|
// PhaseId is a structural type: nothing forces a caller through parsePhaseId,
|
|
// so toDir cannot trust project/milestone/phase/subphase are already
|
|
// canonical — each is validated below against the exact shape parsePhaseId
|
|
// itself would ever produce, closing off a hand-built id as a path-traversal
|
|
// vector. PROJECT_ID_RE mirrors the parser's `[A-Z][A-Z0-9_]*` grammar;
|
|
// CANONICAL_NUMERIC_RE mirrors pad2()'s output shape — exactly 2 digits, or
|
|
// 3+ digits with no leading zero. It is BUILT from
|
|
// BRACKET_CANONICAL_NUMERIC_SOURCE rather than re-spelled as a literal, so this
|
|
// emit-side gate and the read-side token source cannot disagree about what
|
|
// "canonical width" means (the anchors here make the source's trailing `(?!\d)`
|
|
// guard, which the unanchored read side needs, redundant).
|
|
const PROJECT_ID_RE = /^[A-Z][A-Z0-9_]*$/;
|
|
const CANONICAL_NUMERIC_RE = new RegExp(`^${BRACKET_CANONICAL_NUMERIC_SOURCE}$`);
|
|
|
|
function toDir(id: PhaseId, slug: string): string {
|
|
if (!PROJECT_ID_RE.test(id.project)) {
|
|
throw new Error(`toDir: invalid project: ${JSON.stringify(id.project)}`);
|
|
}
|
|
if (!CANONICAL_NUMERIC_RE.test(id.milestone)) {
|
|
throw new Error(`toDir: invalid milestone: ${JSON.stringify(id.milestone)}`);
|
|
}
|
|
if (!CANONICAL_NUMERIC_RE.test(id.phase)) {
|
|
throw new Error(`toDir: invalid phase: ${JSON.stringify(id.phase)}`);
|
|
}
|
|
if (id.subphase !== undefined && !CANONICAL_NUMERIC_RE.test(id.subphase)) {
|
|
throw new Error(`toDir: invalid subphase: ${JSON.stringify(id.subphase)}`);
|
|
}
|
|
// A non-string slug (e.g. an omitted second argument) must not be silently
|
|
// coerced by String(...) into the literal token 'undefined'/'null' on disk.
|
|
if (typeof slug !== 'string') {
|
|
throw new Error(`toDir: slug must be a string: ${JSON.stringify(slug)}`);
|
|
}
|
|
|
|
const sub = id.subphase ? `.${id.subphase}` : '';
|
|
// Slug guard: the slug becomes an on-disk path segment, so collapse it to a
|
|
// safe lowercase token — never a path separator or `..` traversal.
|
|
const safeSlug = slug.toLowerCase().replace(/[^a-z0-9]+/g, '-').replace(/^-+|-+$/g, '');
|
|
// A slug that sanitizes to nothing (e.g. '!!!') would otherwise emit a
|
|
// dangling trailing hyphen.
|
|
if (!safeSlug) {
|
|
throw new Error(`toDir: slug sanitizes to empty: ${JSON.stringify(slug)}`);
|
|
}
|
|
// An all-digit slug (e.g. '2026') is string-indistinguishable from the
|
|
// parsePhaseId dir branch's plan tail, so it would re-parse as a plan, not
|
|
// a slug — silently breaking the disk↔identity bijection on read-back.
|
|
if (/^\d+$/.test(safeSlug)) {
|
|
throw new Error(`toDir: slug must not be all-digit: ${JSON.stringify(slug)}`);
|
|
}
|
|
return `${id.project}.${id.milestone}-${id.phase}${sub}-${safeSlug}`;
|
|
}
|
|
|
|
// Milestone integers reserved as non-milestone sentinels (0.x backlog / 999.x
|
|
// icebox); a phase id in these ranges has no real milestone.
|
|
const SENTINEL_RANGES: readonly number[] = Object.freeze([0, 999]);
|
|
|
|
function isSentinelPhaseId(phaseId: unknown, convention?: string): boolean {
|
|
const s = String(phaseId);
|
|
// Bracket milestone lives in the `{CODE}.{MM}` prefix. GATED on
|
|
// convention === 'bracket' for the same reason as extractPhaseToken below and
|
|
// getMilestoneFromPhaseId above: that prefix is string-indistinguishable from
|
|
// the legacy #1324 letter-prefixed-decimal family (`P0.0-foundation` is a real
|
|
// phase, NOT sentinel milestone 0) whenever the code ends in a digit. A
|
|
// convention-less caller uses the legacy/bare leading-int rule below, so no
|
|
// existing reader gains a false positive; the bracket reading is opt-in.
|
|
if (convention === 'bracket') {
|
|
const bracket = s.match(/^[A-Z][A-Z0-9_]*\.(\d+)/); // bracket: milestone in the prefix
|
|
if (bracket) return SENTINEL_RANGES.includes(parseInt(bracket[1], 10));
|
|
}
|
|
const legacy = stripProjectCodePrefix(s).match(/^0*(\d+)/); // legacy/bare: leading int
|
|
if (!legacy) return false;
|
|
return SENTINEL_RANGES.includes(parseInt(legacy[1], 10));
|
|
}
|
|
|
|
/**
|
|
* Render a regex source fragment matching a phase number against ROADMAP/STATE
|
|
* prose regardless of zero-padding on either side.
|
|
*/
|
|
function phaseMarkdownRegexSource(phaseNum: unknown): string {
|
|
const stripped = stripProjectCodePrefix(phaseNum);
|
|
|
|
// Milestone-prefixed IDs: M-NN or M-N-N (deep).
|
|
const milestoneSegments = stripped.match(/^(\d+)((?:-\d+)*)([A-Z]?(?:\.\d+)*)$/i);
|
|
if (milestoneSegments && milestoneSegments[2]) {
|
|
const majorUnpadded = milestoneSegments[1].replace(/^0+/, '') || '0';
|
|
const subParts = milestoneSegments[2].slice(1).split('-');
|
|
const subFragments = subParts.map(s => {
|
|
const unpadded = s.replace(/^0+/, '') || '0';
|
|
return `0*${escapeRegex(unpadded)}`;
|
|
});
|
|
const suffix = milestoneSegments[3] || '';
|
|
const suffixFragment = suffix ? escapeRegex(suffix) : '';
|
|
return `0*${escapeRegex(majorUnpadded)}-${subFragments.join('-')}${suffixFragment}`;
|
|
}
|
|
|
|
// Plain numeric phase: 1, 01, 12A, 12.1
|
|
const match = stripped.match(/^0*(\d+)([A-Z])?((?:\.\d+)*)$/i);
|
|
if (!match) return escapeRegex(phaseNum);
|
|
|
|
const integer = match[1].replace(/^0+/, '') || '0';
|
|
const letter = match[2] ? escapeRegex(match[2]) : '';
|
|
const decimal = match[3] ? escapeRegex(match[3]) : '';
|
|
return `0*${escapeRegex(integer)}${letter}${decimal}`;
|
|
}
|
|
|
|
/**
|
|
* #3599: when the caller passed a project-code-prefixed ID like `PROJ-42`,
|
|
* return the exact-escaped form.
|
|
*/
|
|
function phaseMarkdownRegexSourceExact(phaseNum: unknown): string | null {
|
|
const raw = String(phaseNum);
|
|
if (!hasProjectCodePrefix(raw)) return null;
|
|
return escapeRegex(raw);
|
|
}
|
|
|
|
function comparePhaseNum(a: unknown, b: unknown): number {
|
|
// Strip optional project_code prefix before comparing
|
|
const sa = stripProjectCodePrefix(a);
|
|
const sb = stripProjectCodePrefix(b);
|
|
|
|
const milestoneA = sa.match(/^(\d+)((?:-\d+)+)([A-Z]?(?:\.\d+)*)$/i);
|
|
const milestoneB = sb.match(/^(\d+)((?:-\d+)+)([A-Z]?(?:\.\d+)*)$/i);
|
|
|
|
if (milestoneA && milestoneB) {
|
|
const segsA = [parseInt(milestoneA[1], 10), ...milestoneA[2].slice(1).split('-').map(s => parseInt(s, 10))];
|
|
const segsB = [parseInt(milestoneB[1], 10), ...milestoneB[2].slice(1).split('-').map(s => parseInt(s, 10))];
|
|
const maxSegs = Math.max(segsA.length, segsB.length);
|
|
for (let i = 0; i < maxSegs; i++) {
|
|
const av = segsA[i] !== undefined ? segsA[i] : 0;
|
|
const bv = segsB[i] !== undefined ? segsB[i] : 0;
|
|
if (av !== bv) return av - bv;
|
|
}
|
|
const sufA = milestoneA[3] || '';
|
|
const sufB = milestoneB[3] || '';
|
|
if (sufA !== sufB) return sufA < sufB ? -1 : 1;
|
|
return 0;
|
|
}
|
|
|
|
if (milestoneA || milestoneB) return String(a).localeCompare(String(b));
|
|
|
|
const pa = sa.match(/^(\d+)([A-Z])?((?:\.\d+)*)/i);
|
|
const pb = sb.match(/^(\d+)([A-Z])?((?:\.\d+)*)/i);
|
|
if (!pa || !pb) return String(a).localeCompare(String(b));
|
|
const intDiff = parseInt(pa[1], 10) - parseInt(pb[1], 10);
|
|
if (intDiff !== 0) return intDiff;
|
|
const la = (pa[2] || '').toUpperCase();
|
|
const lb = (pb[2] || '').toUpperCase();
|
|
if (la !== lb) {
|
|
if (!la) return -1;
|
|
if (!lb) return 1;
|
|
return la < lb ? -1 : 1;
|
|
}
|
|
const aDecParts = pa[3] ? pa[3].slice(1).split('.').map(p => parseInt(p, 10)) : [];
|
|
const bDecParts = pb[3] ? pb[3].slice(1).split('.').map(p => parseInt(p, 10)) : [];
|
|
const maxLen = Math.max(aDecParts.length, bDecParts.length);
|
|
if (aDecParts.length === 0 && bDecParts.length > 0) return -1;
|
|
if (bDecParts.length === 0 && aDecParts.length > 0) return 1;
|
|
for (let i = 0; i < maxLen; i++) {
|
|
const av = Number.isFinite(aDecParts[i]) ? aDecParts[i] : 0;
|
|
const bv = Number.isFinite(bDecParts[i]) ? bDecParts[i] : 0;
|
|
if (av !== bv) return av - bv;
|
|
}
|
|
return 0;
|
|
}
|
|
|
|
/**
|
|
* Extract the phase token from a directory name.
|
|
*/
|
|
function extractPhaseToken(dirName: string, convention?: string): string {
|
|
// #612 bracket dir form `{CODE}.{MM}-{PP}[.{SS}]-slug` → phase token `PP[.SS]`.
|
|
// GATED on convention === 'bracket' (mirrors getMilestoneFromPhaseId's READING-B
|
|
// decision above). A bracket dir `{CODE}.{MM}-{PP}` is string-INDISTINGUISHABLE
|
|
// from the legacy #2043/#1324 letter-prefixed-decimal family (`P0.3-2`,
|
|
// `P0.12-34`) whenever the project code ends in a digit, so NO string-only
|
|
// discriminator can separate the two conventions — auto-detecting here silently
|
|
// reinterpreted `P0.3-2` → `2` (was `P0.3-2`), a byte-identical-read regression
|
|
// on this CRITICAL 6-caller helper (ADR-2121). Requiring an explicit convention
|
|
// signal keeps every existing (convention-less) call site byte-identical to
|
|
// prior behaviour — see the #2043 numeric-tail characterization in
|
|
// tests/phase-id.test.cjs — while keeping the helper pure (optional param, no
|
|
// config read). The captured token is dot-only (`PP[.SS]`); the milestone↔phase
|
|
// hyphen and any trailing plan/slug are excluded.
|
|
if (convention === 'bracket') {
|
|
const bracketDir = dirName.match(/^[A-Z][A-Z0-9_]*\.\d+-(\d+(?:\.\d+)?)/);
|
|
if (bracketDir) return bracketDir[1];
|
|
}
|
|
|
|
const codePrefixMatch = dirName.match(PROJECT_CODE_PREFIX_CAPTURE_RE_I);
|
|
let prefix = '';
|
|
let rest = dirName;
|
|
if (codePrefixMatch) {
|
|
prefix = codePrefixMatch[1] + '-';
|
|
rest = codePrefixMatch[2];
|
|
}
|
|
|
|
const segments = rest.split('-');
|
|
const tokenSegments: string[] = [];
|
|
// #2043: distinguish a real (zero-padded) phase/sub-phase segment from a
|
|
// single-digit slug word. A pure-numeric leading segment ("46") only
|
|
// continues with exactly-2-digit segments (#2232: a ≥3-digit run is a slug
|
|
// word such as a year — "14-2026-photos-…" yields "14", not "14-2026"), so
|
|
// "46-6-rs-…" yields "46" (the "6" is the
|
|
// slug's first word), not "46-6". Milestone-prefixed ids like "M1-2" reach here
|
|
// with "M1-" already stripped as a project-code prefix (see
|
|
// PROJECT_CODE_PREFIX_CAPTURE_RE_I), so "2" is the leading segment and the same
|
|
// pure-numeric rule applies (M1-46-6-rs → "M1-46"). The firstLetterPrefixed
|
|
// carve-out covers letter+digit leading segments that survive prefix stripping
|
|
// because of punctuation (e.g. "P0.3-2"), whose single-digit continuation is
|
|
// intentionally preserved (unchanged from prior behaviour).
|
|
let firstLetterPrefixed = false;
|
|
for (let i = 0; i < segments.length; i++) {
|
|
const seg = segments[i];
|
|
if (i === 0) {
|
|
if (/^\d/.test(seg)) {
|
|
tokenSegments.push(seg);
|
|
} else if (/^[A-Za-z]{1,3}\d/.test(seg)) {
|
|
tokenSegments.push(seg);
|
|
firstLetterPrefixed = true;
|
|
} else {
|
|
break;
|
|
}
|
|
} else if (isPhaseContinuationSegment(seg) || (firstLetterPrefixed && /^\d/.test(seg))) {
|
|
tokenSegments.push(seg);
|
|
} else {
|
|
break;
|
|
}
|
|
}
|
|
|
|
if (tokenSegments.length === 0) {
|
|
return dirName;
|
|
}
|
|
|
|
return prefix + tokenSegments.join('-');
|
|
}
|
|
|
|
/**
|
|
* Check if a directory name's phase token matches the normalized phase exactly.
|
|
*/
|
|
function phaseTokenMatches(dirName: string, normalized: string): boolean {
|
|
const token = extractPhaseToken(dirName);
|
|
if (token.toUpperCase() === normalized.toUpperCase()) return true;
|
|
const stripped = stripProjectCodePrefix(dirName);
|
|
if (stripped !== dirName) {
|
|
const strippedToken = extractPhaseToken(stripped);
|
|
if (strippedToken.toUpperCase() === normalized.toUpperCase()) return true;
|
|
}
|
|
return false;
|
|
}
|
|
|
|
// ─── #2121 canonical surface (ADR-2121) ──────────────────────────────────────
|
|
|
|
/**
|
|
* Parse a phase identifier from a STATE.md `Phase:` prose field VALUE — the text
|
|
* after the `Phase:` label (e.g. `"3 of 4 (Delta)"`, `"3A — Delta (executing)"`,
|
|
* or `"Milestone v0.5 complete"`).
|
|
*
|
|
* The token is anchored to the START of the value (after an optional literal
|
|
* `Phase ` label and an optional project-code prefix) so a phase is only
|
|
* returned when the value actually begins with one. This is the #2111 fix: the
|
|
* prior unanchored `/\b(\d+[A-Z]?(?:\.\d+)*)\b/i` mined the first numeral
|
|
* anywhere, so `"Milestone v0.5 complete"` collapsed to `"5"` (the minor-version
|
|
* digit) and `"v1.0"` to `"0"` (a reserved sentinel). Here both yield
|
|
* `{ phase: null }` because they do not begin with a phase token. The name
|
|
* extraction (parenthetical or em-dash tail, minus status words) is unchanged.
|
|
*/
|
|
function parsePhaseFromProse(value: string | null): { phase: string | null; name: string | null } {
|
|
if (!value) return { phase: null, name: null };
|
|
// Coerce defensively so a non-string caller cannot throw on this canonical
|
|
// surface (mirrors the sibling #2121 functions' String(...) handling).
|
|
const str = String(value);
|
|
const phaseMatch = str.match(/^\s*(?:Phase\s+)?(?:[A-Z][A-Z0-9_]*-)?(\d+[A-Z]?(?:\.\d+)*)\b/i);
|
|
// The name-extraction quantifiers are length-bounded so a crafted long
|
|
// unterminated run (many `(` or `—`) in an untrusted STATE.md field value
|
|
// cannot drive O(n^2) regex backtracking (CPU-exhaustion DoS). A real phase
|
|
// name is far shorter than the cap.
|
|
const parenName = str.match(/\(([^)]{1,200})\)/);
|
|
const dashName = str.match(/—\s*([^(\n]{1,200}?)(?:\s*\(|$)/);
|
|
const rawName = parenName?.[1] ?? dashName?.[1] ?? null;
|
|
const name = rawName && !/^(?:complete|executing|not started)$/i.test(rawName.trim())
|
|
? rawName.trim()
|
|
: null;
|
|
return {
|
|
phase: phaseMatch ? phaseMatch[1] : null,
|
|
name,
|
|
};
|
|
}
|
|
|
|
/**
|
|
* Config-AWARE project-code prefix strip. Unlike the config-blind
|
|
* `stripProjectCodePrefix` (which strips ANY `<CODE>-` shape), this strips the
|
|
* leading `<CODE>-` ONLY when `<CODE>` case-insensitively equals the configured
|
|
* `projectCode`. A foreign prefix (`MEM-01` when the configured code is `LKML`)
|
|
* or an absent/empty `projectCode` is preserved verbatim — this is the #2104
|
|
* fix: a foreign-prefixed id must not collapse to a bare numeric phase and
|
|
* collide with a real one.
|
|
*/
|
|
function stripConfiguredProjectCodePrefix(value: unknown, projectCode: string | null | undefined): string {
|
|
const input = String(value);
|
|
const configured = typeof projectCode === 'string' ? projectCode.trim() : '';
|
|
if (!configured) return input;
|
|
const m = input.match(PROJECT_CODE_PREFIX_CAPTURE_RE_I);
|
|
if (!m) return input;
|
|
if (m[1].toUpperCase() !== configured.toUpperCase()) return input;
|
|
return m[2];
|
|
}
|
|
|
|
/**
|
|
* True when `phase` carries a project-code prefix that is NOT the configured
|
|
* `projectCode` (or when no `projectCode` is configured). The canonical
|
|
* predicate the init-command foreign-prefix guard (#2056 / PR #2105) delegates
|
|
* to, so every call site shares one foreign-prefix rule.
|
|
*/
|
|
function isForeignPrefixedPhaseQuery(phase: unknown, projectCode: unknown): boolean {
|
|
const m = String(phase).match(PROJECT_CODE_PREFIX_CAPTURE_RE_I);
|
|
if (!m) return false;
|
|
const configured = typeof projectCode === 'string' ? projectCode.trim() : '';
|
|
return !configured || m[1].toUpperCase() !== configured.toUpperCase();
|
|
}
|
|
|
|
/**
|
|
* Canonical ROADMAP heading lookup-source list (moved here from
|
|
* roadmap-parser.cts so phase-id.cts is the single owner of the ordering).
|
|
* Sources are tried in a fixed, deduplicated order: exact (only when the query
|
|
* itself is project-code-prefixed) → bare numeric / padding-tolerant →
|
|
* prefix-tolerant fallback. The bare numeric source precedes the prefix-tolerant
|
|
* form so a canonical heading (`### Phase 117:`) is preferred over a drifted
|
|
* prefixed one (`### Phase MANIFOLD-117:`) when both exist in one ROADMAP.
|
|
*/
|
|
function roadmapPhaseLookupSources(phaseNum: unknown): string[] {
|
|
const sources: string[] = [];
|
|
const exactSource = phaseMarkdownRegexSourceExact(phaseNum);
|
|
if (exactSource) sources.push(exactSource);
|
|
|
|
const numericSource = phaseMarkdownRegexSource(phaseNum);
|
|
sources.push(numericSource);
|
|
sources.push(`${OPTIONAL_PROJECT_CODE_PREFIX_SOURCE}${numericSource}`);
|
|
|
|
return [...new Set(sources)];
|
|
}
|
|
|
|
export = {
|
|
escapeRegex,
|
|
OPTIONAL_PROJECT_CODE_PREFIX_SOURCE,
|
|
OPTIONAL_PHASE_TAG_SOURCE,
|
|
PHASE_NUMBER_TOKEN_SOURCE,
|
|
PHASE_CONTINUATION_SEGMENT_SOURCE,
|
|
isPhaseContinuationSegment,
|
|
BRACKET_PHASE_TOKEN_SOURCE,
|
|
PHASE_HEADING_PREFIX_SRC,
|
|
stripProjectCodePrefix,
|
|
normalizePhaseName,
|
|
getMilestoneFromPhaseId,
|
|
getPhaseDirFromPhaseId,
|
|
parsePhaseId,
|
|
renderPhaseId,
|
|
toDir,
|
|
SENTINEL_RANGES,
|
|
isSentinelPhaseId,
|
|
phaseMarkdownRegexSource,
|
|
phaseMarkdownRegexSourceExact,
|
|
comparePhaseNum,
|
|
extractPhaseToken,
|
|
phaseTokenMatches,
|
|
parsePhaseFromProse,
|
|
stripConfiguredProjectCodePrefix,
|
|
isForeignPrefixedPhaseQuery,
|
|
roadmapPhaseLookupSources,
|
|
};
|