* fix(#2528): resolve digit-slug phase dirs by bare number — tokenizer rewind, shared bare-integer fallback, resolution-path parity gate
extractPhaseToken welded 2-digit slug words onto the phase token (phase 10
named "24/7 Autonomy" -> dir 10-24-7 -> token 10-24), making digit-prefixed
phase names unresolvable by bare number across every phase verb.
- phase-id: continuation segments must be the PURE 2-digit zero-padded form
the write side emits; a 1-digit terminator rewinds the absorbed run
(10-24-7 -> 10) while >=2-digit terminators keep the locked #2232
round-trip (14-06-2026-photos -> 14-06).
- phase-id: new matchPhaseDirs owner — primary exact-token match plus a
bare-integer leading-digit-run fallback for shapes the tokenizer cannot
rewind (05-80-20-cleanup); collisions stay #2237-loud.
- locator/find-phase/phase-plan-index all delegate selection to the owner;
plan-index gains the previously missing multi-match guard.
- tests: #2528 unit + fast-check metamorphic blocks; new 9-scenario
resolution-path parity gate across all three paths.
Fixes #2528
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(#2528): add changeset for PR #2559
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(#2528): align validation token grammar
* fix: address phase token review
* fix: restore phase grammar parity for numeric slugs
* docs: document digit-leading phase resolution
* docs: clarify ambiguous phase resolution behavior
* docs: register canonical phase directory selectors
* fix: align prefixed deep phase token parsing
* fix(#2528): route the fourth resolution site through matchPhaseDirs
Review BLOCKER. smart-entry.cts::detectVerifyFailed resolved the current
phase's directory with its own `.find(phaseTokenMatches)` and never
reached the shared selection, so the bare-integer-fallback family the
issue names — `05-80-20-cleanup`, `30-12-factor-refactor` — resolved
nowhere. The miss is silent by construction: an unresolved phase reports
"not failed", which is byte-identical to a healthy one, so a failed
verification simply never surfaced in /gsd or /gsd:progress.
`entries` is already sorted and matchPhaseDirs filters without
reordering, so matches[0] reproduces the previous selection exactly
wherever the old code resolved at all.
Wiring it into phase-resolution-parity.test.cjs as a fourth path then
exposed a second, older defect in the same function: phaseTokenFromDirName
shape-probed the UNSTRIPPED token, so a project-code-prefixed directory
(`MEM-05-…`, tokenizing to `MEM-05-80-20`) failed the leading-digit test
and was dropped before any resolution ran — every phase in a
project-coded plan was invisible to this check. The probe now runs on the
stripped token; the returned value is unchanged, so the comparePhaseNum
sort is untouched.
Path 4 has no JSON surface to compare, so the gate observes selection
indirectly: plant the failing artifact in exactly one directory and a
passing one everywhere else, then read the boolean. Reverting either fix
turns 5 of the 10 corpus scenarios red.
* refactor(#2528): collapse the duplicated extractCanonicalPlanId
Review MAJOR. The function existed as two independent, byte-identical
copies — src/core-utils.cts and src/phase.cts — and this PR had to patch
BOTH with the same single-digit-slug rewind rule. That is the generative
fix divergence CLAUDE.md names, and only the core-utils copy was under
test, so a future one-sided patch would have silently split plan-id
canonicalization between the plan listing and everything else.
Removed rather than parity-tested: core-utils was already the leaf owner
and already exported it, and phase.cts already imported that module, so
there is no second surface left for a parity test to police.
* test(#2528): pin matchPhaseDirs at the digit-width boundaries
Review MAJOR. The bare-integer fallback's correctness rests entirely on
capturing each directory's whole leading digit run before the zero-strip
compare; a regex that stopped short would turn every query into a prefix
match, and "1" would claim 10, 100, and 12 alike. The existing coverage
was example-based and never touched that boundary.
Adds the explicit 9/10 and 1/10/100 cases — including the forms where
only the wider directories exist, so an exact-width neighbour cannot
satisfy the assertion — plus a fast-check property over arbitrary
distinct leading runs. The property is stated as an invariant on the
result (every returned directory's leading run IS the query) rather than
an expected list, so it covers primary and fallback matches alike and
cannot be satisfied by reimplementing the selection in the test.
Both fail when the fallback regex is degraded to a prefix match.
* fix(#2528): route the remaining eight consumers through matchPhaseDirs
phaseTokenMatches had eight consumers left that each rebuilt the directory
selection around it by hand: phases-list, next-decimal, phase-remove, the
W021 milestone-consistency check, schema-drift, the init-manager overview,
milestone-complete's disk check, and roadmap analyze. Every one of them
reproduced the reported symptom in full after the tokenizer was fixed.
None of them derives a displayed phase number from the matched directory,
so none needs phaseNumberForMatch; the change at each site is the
selection and nothing else. matchPhaseDirs filters without reordering, so
matches[0] reproduces the prior .find() choice wherever the old code
resolved at all.
phaseTokenMatches now has no call sites outside phase-id.cts. It stays
exported as the primitive matchPhaseDirs is built from and as a pinned
canonical surface, but no consumer reaches past the owner to it.
* test(#2528): extend the parity gate to the migrated consumers
Each of the eight is observed through the surface a user sees, not
through the matcher, with a no-directory control so the assertions cannot
be satisfied by a consumer that resolves unconditionally. init-manager
and roadmap-analyze are additionally asserted to agree with each other.
* refactor(#2528): own the case-flexible phase grammar and the leading-digit-run fragment
validate.cts derived its case-flexible regex sources by running
`replaceAll('A-Z', 'A-Za-z')` over two constants exported by phase-id.cts.
That passes lint-phase-id-drift.cjs — there is no literal copy of the
grammar — but it depends on the owner rendering that exact substring. The
day phase-id.cts expresses the same class any other way the replaceAll
silently no-ops and validate.cts narrows to uppercase-only. The failure
mode is a NON-match, so nothing throws and no uppercase-only fixture
notices. Both variants are now derived once, beside the sources they
widen, and imported.
Also names the leading digit run the bare-integer fallback selects on.
It was spelled `/^(\d+)(?:-|$)/` where the fallback filters and `/^\d+/`
where phaseNumberForMatch reads the number back off the winner; selecting
on one run and displaying another would resolve a directory and then label
it with a number that never matched it.
* fix(#2528): refuse to remove a phase when two directories claim its number
cmdPhaseRemove was the only migrated site taking matches[0] with no
multi-match guard. Every sibling resolution path returns ambiguous_matches
and refuses to choose; this one is the DESTRUCTIVE path, so choosing
silently is strictly worse than anywhere else. With 05-80-20-a and
05-90-till-late on disk, `phase remove 5 --force` deleted one of them and
renumbered every phase after it — where the base resolved nothing, deleted
nothing, and the corpus in tests/phase-resolution-parity.test.cjs already
declared that exact input ambiguous.
The refusal is emitted before any file is touched and carries both
candidates. CONSUMER_SCENARIOS could not express the case — every row is
binary, resolving to one directory or to none — so the gate gains a
dedicated ambiguous test. It asserts on the filesystem, not only on the
reported directory_deleted: a null printed after an rmSync would satisfy
every other check.
* fix(#2528): pair digit-leading phase directories with their roadmap phase in validate health
W006/W007 are the ninth site of this bug class and the one a
`phaseTokenMatches` grep could never surface: they resolve roadmap↔disk by
intersecting TOKEN SETS, which is a dir→token labelling rather than the
query→dir selection matchPhaseDirs owns. On the canonical fixture the
label is wrong in both directions at once, so `validate health` reported
"Phase 5 in ROADMAP.md but no directory on disk" AND "Phase 05-80-20
exists on disk but not in ROADMAP.md" for the same directory.
collectDiskPhases now keeps the directory names behind each token, so
W006 can ask the canonical matcher whether a roadmap phase resolves to a
real directory, and W007 — which iterates directories and therefore has no
query to resolve — gets the inverse mapping it never had: a directory is
claimed when some roadmap phase resolves to it.
Both checks are additive: the token intersection still decides every shape
it already decided, and the resolution can only REMOVE a warning. The
regression test carries controls in the other direction — a roadmap phase
with no directory must still raise W006, an unclaimed directory must still
raise W007 — so it cannot be satisfied by a check that stopped reporting.
* docs(#2528): state and pin the directory-side scope of the bare-integer fallback
The matchPhaseDirs docblock claimed deep-decomposition lookups were
untouched. That is true of the QUERY side only — no non-bare query enters
the fallback — but the DIRECTORY side is what changed classification: a
bare `5` now reaches a lone `05-01-auth` and resolves it (phase_number
"05", phase_name "01-auth") where the base found nothing.
The widening is irreducible from directory names alone. `05-01-auth`
(sub-phase 5.1) and `30-12-factor-refactor` (phase 30 named "12-Factor
Refactor") are the same `NN-NN-<slug>` shape, and the discriminator that
would separate them — "is the second segment a valid decimal sub-phase" —
accepts `5.1` and `30.12` equally. Any rule strong enough to exclude the
first excludes the second, which is the defect #2528 exists to fix. So the
tie is broken in favour of resolving, the docblock now says so, and the
consequence is bounded where it matters: two such directories are two
matches, and every caller (including phase remove) refuses to choose.
Pins both directions, since nothing observed the directory side before.
* fix(#2528): count surviving phases by identity in phase remove's STATE resync
#2640 landed on `next` while this branch was open. Its STATE.md phase-count
resync re-derives "which directory was removed" from the query with
`phaseTokenMatches`, which is the tenth site of this issue's defect: the
bare-integer fallback resolves `05-80-20-cleanup` for query `5`, but the
token predicate does not, so the just-deleted directory is counted as still
present and the written `Total Phases` is one too high.
`targetDir` already IS the directory that was removed, and the block is gated
on it being non-null, so identity answers the question exactly — which is also
what the comment above the filter already claimed it did. This keeps
`phaseTokenMatches` out of `phase.cts` rather than re-importing it to satisfy
one call site: the module's public surface should not grow for a question that
does not need re-derivation.
Pinned in the parity gate with a control on a directory the tokenizer reads
correctly, so the assertion is about the digit-leading shape and not about the
counting rule changing for everything.
* fix(#2528): let the resolution layer own the digit-leading slug family alone
The tokenizer rewind this fix carried — pop the last absorbed continuation when
the segment that stopped the scan is a bare single digit — reads
"10-24-7-autonomy" (phase 10 named "24/7 Autonomy") correctly and silently
re-reads "10-24-7-zip" (sub-phase 10.24 named "7-Zip Integration") from "10-24"
to "10". The two names are string-identical in shape, so no local signal
separates them; the rule traded the reported ambiguity for the symmetric one a
level down, on a 15-caller chokepoint whose output also feeds query-less
derivations (STATE.md phase counts, W007, the #2562 key surface). A well-formed
sub-phase directory became unresolvable by its own id — the very symptom #2528
was filed about.
It also bought nothing. The bare-integer fallback in matchPhaseDirs already
resolves "10-24-7-autonomy" for query "10" whatever the token is: no primary
match, bare query, leading digit run "10". The reported case was covered twice,
by two rules, and the two disagreed about the case nobody reported.
So the rewind is removed rather than narrowed, in the tokenizer and in the five
surfaces kept in lockstep with it (BRACKET_PHASE_TOKEN_SOURCE,
PHASE_TOKEN_FROM_DIR_RE, canonicalPlanStem's pair grammar and its collision
branch, roadmap-parser's numericRe, extractCanonicalPlanId), together with the
SINGLE_DIGIT_RUN_SEGMENT_SOURCE owner constant they shared. Disambiguation now
lives only where a QUERY exists to disambiguate against, which is the same
bounded mechanism the "05-80-20-cleanup" shape already used.
Measured, not argued: over 29800 generated directory names, extractPhaseToken is
byte-identical to `next` on every input except the lowercase-continuation class
("01-20a", "05-80-20-25abc") — a rule about the segment itself, not a guess about
its neighbour.
Both readings now stay reachable by their own ids:
matchPhaseDirs(['10-24-7-autonomy'], '10') -> the dir (fallback)
matchPhaseDirs(['10-24-7-zip'], '10') -> the dir (fallback)
matchPhaseDirs(['10-24-7-zip'], '10-24') -> the dir (primary)
* test(#2528): pin the one-continuation boundary the rewind had no coverage for
The regressing shape was invisible to the suite by construction, not by luck:
the deep-rewind property built its cases from `continuationArb` with
`minLength: 2`, so it never exercised the single-continuation case — exactly one
genuine sub-phase level before a digit-leading slug — and every hand-written
fixture used the ambiguous shape only where "phase-plus-slug" was the intended
reading.
`continuationArb` is now `minLength: 1` and the property states the invariant
instead of the old rule: for any prefix, phase, 1-5 continuations and any
one-digit terminator, the token equals the FULL continuation run on both the
imperative and the regex surface, and `matchPhaseDirs([dir], token)` returns that
dir. That third assertion is the one that catches the class on its own — the old
behaviour made a well-formed directory unresolvable by its own id, which is a
property, not a fixture.
Around it: "10-24-7-zip" and "10-24-3d-printer" now sit beside
"10-24-7-autonomy" everywhere the family is pinned, so the two readings can never
diverge again; the 24/7 metamorphic property asserts the RESOLUTION result rather
than the token (the token is precisely the part no surface may decide); the
end-to-end parity corpus gains "a sub-phase with a digit-leading slug resolves by
its full id" across all four resolution paths; and the milestone-scoping residual
is pinned in three directions rather than left to prose.
Mutation: re-inserting the rewind and rebuilding turns 9 tests red, the
`minLength: 1` property first, and nothing else. Build success checked separately
(build:lib reports 0 `error TS`), so the mutation reached the artifact under test.
* test(#2528): pin the #2946 guard against digit-leading phase directories
The #2946 fix makes the milestone-complete unstarted-phase guard run
unconditionally, so whether it fires now rides entirely on the
directory-resolution owner this PR replaces. Two cases, both with STATE.md
carrying no `milestone:` field so the #2946 path is the one exercised:
- ROADMAP Phase 5, disk `05-80-20-cleanup` → guard must stay silent.
RED on next (fail-closed: the guard blocks a legitimate one-way-door
operation because phaseTokenMatches resolves neither 05 nor 80 for
that directory).
- ROADMAP Phase 80, same directory → guard must still fire. Green on
both sides; it pins the fail-open direction against a future widening
of the matcher.
* fix(#3175): stop the injection scanner reading RegExp.exec as code execution
The apostrophe fix in 27aa40f6 replaced ["\x27] with a real ["'] class.
The old class never contained an apostrophe at all (POSIX bracket
expressions do not honour backslash escapes, so it was the set ", \, x,
2, 7), so only exec(" matched. Single-quoted method calls now match for
the first time, and RegExp.prototype.exec takes a subject string, not
code: any PR touching a file that tests a regex goes red. On next, six
files match the scanner's own pattern across 16 method calls.
A plain [^[:alnum:]] boundary cannot separate the two forms because . is
not alnum, so exec gets [^[:alnum:].] and the command-execution vector
moves to a dedicated member-call pattern. Bare exec('rm -rf /'),
cp.exec(...) and child_process.exec(...) all still fire.
Mutation: reverting the boundary reds 1 test and only it; removing the
member-call pattern reds the 2 non-weakening tests and only them.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(#2528): keep exec( detection receiver-blind, allowlist the two grammar suites
The left boundary [^[:alnum:].] added in 58a7b560 excluded a preceding dot,
which dropped every member-position .exec('…') from the scanner. The follow-up
receiver pattern only restored three literal spellings (child_process,
childProcess, cp), so require('child_process').exec('…') — the most common Node
spelling of the vector this pattern exists to catch — became invisible, along
with any opaque receiver (conn.exec, shelljs.exec).
Revert the pattern to its receiver-blind form and handle the RegExp.prototype
.exec false positive where the script already handles this class: per-file
ALLOWLIST entries for the two phase-token grammar suites. Mutation-checked —
removing the two entries reds exactly those two files and nothing else.
The four assertions written around the old patterns are replaced by a
table-driven set covering all six spellings, including the three the narrowed
pattern silently lost.
Co-Authored-By: Claude <noreply@anthropic.com>
* docs(#2528): pin the three undeclared grammar edges, correct the ambiguity claim
Review round 10 asked for declaration, not behavior change, on four items. All
four have zero production consumers or preserve their caller's prior rule, so
each is pinned as a test or corrected in prose rather than reverted.
- BRACKET_PHASE_TOKEN_SOURCE: the (?=-|$) terminator is what keeps the bracket
read path in step with the other surfaces, and it costs the display shapes
(`05.03: Title`, `12A: X`, `05.03]` no longer tokenize). Pinned so widening
the terminator class is a deliberate act rather than a lookahead deletion.
- canonicalPlanStem: uppercase plan suffixes still strip, lowercase and dotted
sub-plans now fall through. Dead export; pinned as a decision on record.
- getMilestonePhaseFilter: `12A-01-foo` now yields `12A-01`, matching what
`12-01-foo` has always yielded. The letter suffix was the only reason a
sub-phase directory folded into its parent phase's milestone window; the two
shapes now agree. Not named in the review — found auditing the same commit.
- matchPhaseDirs docblock claimed every caller refuses on multi-match. Four do;
five take matches[0]. Replaced the claim with the actual two-tier policy and
the honest caveat that the bare fallback makes multi-match newly reachable
for queries that previously found nothing.
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
962 lines
49 KiB
TypeScript
962 lines
49 KiB
TypeScript
/**
|
|
* Pure phase-id parsing/matching helpers — normalize, token match,
|
|
* milestone/phase-dir id parsing, phase-markdown regex builders.
|
|
*
|
|
* Extracted from core.cts (ADR-857 rollout phase 2a / issue #865).
|
|
* The hand-written bodies are preserved byte-for-behaviour; only the module
|
|
* boundary moved. The core.cjs re-export spine was retired in epic #1267;
|
|
* callers import phase-id helpers from phase-id.cjs directly.
|
|
*
|
|
* Dependencies: none (pure string/regex, no Node built-ins required).
|
|
*/
|
|
|
|
// ─── Phase-id helpers ─────────────────────────────────────────────────────────
|
|
|
|
function escapeRegex(value: unknown): string {
|
|
return String(value).replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
|
|
}
|
|
|
|
// project_code values start with an uppercase letter (e.g. PROJ, APP_CODE);
|
|
// leading underscores are not valid project codes per .planning/config.json.
|
|
const PROJECT_CODE_PREFIX_STRIP_RE = /^[A-Z][A-Z0-9_]*-(?=\d)/;
|
|
const PROJECT_CODE_PREFIX_STRIP_RE_I = /^[A-Z][A-Z0-9_]*-(?=\d)/i;
|
|
const PROJECT_CODE_PREFIX_CAPTURE_RE_I = /^([A-Z][A-Z0-9_]*)-(\d.*)/i;
|
|
const OPTIONAL_PROJECT_CODE_PREFIX_SOURCE = '(?:[A-Z][A-Z0-9_]*-)?';
|
|
|
|
// #1729: phase headers may carry a parenthetical tag between the number and the
|
|
// colon, e.g. `### Phase 26 (Cluster B): Title`. This optional, non-capturing
|
|
// fragment is injected at every phase-header regex call site (immediately after
|
|
// the phase-number token, before the colon/space delimiter) so the resolver
|
|
// tolerates the tag — mirroring how `[...]` is already tolerated before `Phase`.
|
|
// `[^)\n]*` keeps the match single-line (headers are one line) to avoid
|
|
// over-consuming across a malformed multi-line document. Injected at the call
|
|
// site (not baked into phaseMarkdownRegexSource) so it applies uniformly to
|
|
// both the numeric and project-code-exact escaped sources, and so the decimal
|
|
// sub-phase patterns can place it after the `.N` segment.
|
|
//
|
|
// Enumeration/parse call sites that read phase headers from a regex *literal*
|
|
// (rather than a `new RegExp` built from an interpolated phase number) cannot
|
|
// reference this constant; they inline its literal-regex mirror instead —
|
|
// `(?:\s*\([^)\n]{0,200}\))?` — kept character-for-character equivalent to this
|
|
// source. Both forms must change together; see the #1729 regression test.
|
|
const OPTIONAL_PHASE_TAG_SOURCE = '(?:\\s*\\([^)\\n]{0,200}\\))?';
|
|
|
|
// #2128: the canonical phase-NUMBER-TOKEN grammar — a phase number with an
|
|
// optional single-letter variant suffix and optional dotted sub-phases
|
|
// (1, 01, 12A, 12.1, 3.2.1). This is the ENUMERATION/scan counterpart to
|
|
// phaseMarkdownRegexSource: use phaseMarkdownRegexSource(n) to build a source
|
|
// for ONE KNOWN number; reference this constant when a call site must match ANY
|
|
// phase and capture its token. Enumeration/parse sites inline this into a
|
|
// `new RegExp(...)` instead of re-deriving the grammar as a literal, so every
|
|
// phase-token producer shares one owner. The anti-divergence guard
|
|
// (scripts/lint-phase-id-drift.cjs) fails CI if a literal re-derivation is
|
|
// introduced outside this module without a `// phase-id-owner:` justification.
|
|
const PHASE_NUMBER_TOKEN_SOURCE = '\\d+[A-Z]?(?:\\.\\d+)*';
|
|
|
|
// #2528 review: the CASE-FLEXIBLE renderings of the two sources above, for call
|
|
// sites that scan directory names (where a project code or a variant suffix may
|
|
// legitimately be lowercase) and therefore cannot use a case-sensitive class.
|
|
//
|
|
// They live HERE, beside the sources they widen, because the alternative in use
|
|
// was `SOURCE.replaceAll('A-Z', 'A-Za-z')` at the consuming site — a derivation
|
|
// that depends on the owner rendering that exact literal. It passes
|
|
// lint-phase-id-drift.cjs (no literal copy of the grammar), but the day this
|
|
// module expresses the same class any other way (`[[:upper:]]`, a named
|
|
// fragment, an escaped range) the replaceAll silently no-ops and the consumer
|
|
// quietly narrows to uppercase-only — the failure is a NON-match, so nothing
|
|
// throws and no test that only feeds uppercase input notices. Deriving it once,
|
|
// where the source is defined, makes that impossible: a rename here is a
|
|
// compile-visible change, not a silent behavior change three modules away.
|
|
const CASE_FLEXIBLE_PROJECT_CODE_PREFIX_SOURCE =
|
|
OPTIONAL_PROJECT_CODE_PREFIX_SOURCE.replaceAll('A-Z', 'A-Za-z');
|
|
const CASE_FLEXIBLE_PHASE_NUMBER_TOKEN_SOURCE =
|
|
PHASE_NUMBER_TOKEN_SOURCE.replaceAll('A-Z', 'A-Za-z');
|
|
|
|
// #2232: the canonical CONTINUATION-segment grammar — a dash-separated segment
|
|
// that extends a phase token (a zero-padded sub-phase or plan number, e.g. the
|
|
// "01" in "02-01-setup"). getPhaseDirFromPhaseId writes these zero-padded to
|
|
// exactly 2 digits, so the digit RUN of a genuine continuation is exactly 2:
|
|
// #2043's `\d{2,}` (2-or-more) over-collected a slug word that merely leads
|
|
// with ≥2 digits (a year: "14-2026-photos-…" yielded token "14-2026", so every
|
|
// phase-locating verb reported the phase as missing). The `(?!\d)` guard caps
|
|
// the run at 2 without anchoring what may follow, so call sites keep their own
|
|
// trailing grammar (letter suffixes, dotted sub-phases, segment boundaries).
|
|
// POLICY (locked by boundary tests): sub-phase/plan numbers ≥100 are out of the
|
|
// dir-token grammar — the LEADING phase number stays unbounded (`\d+`), only
|
|
// continuation segments begin with a two-digit run; consuming sites retain
|
|
// their established suffix and boundary grammar. Shared from here so the five #2043
|
|
// call sites cannot drift independently (see scripts/lint-phase-id-drift.cjs).
|
|
const PHASE_CONTINUATION_SEGMENT_SOURCE = '\\d{2}(?!\\d)';
|
|
const PHASE_CONTINUATION_SEGMENT_PREFIX_RE = new RegExp(`^${PHASE_CONTINUATION_SEGMENT_SOURCE}`);
|
|
function isPhaseContinuationSegment(seg: string): boolean {
|
|
return PHASE_CONTINUATION_SEGMENT_PREFIX_RE.test(seg);
|
|
}
|
|
|
|
// #612 (PR-1): bracket-convention token/heading sources, kept next to the M-NN
|
|
// PHASE_NUMBER_TOKEN_SOURCE so this owner file stays the single origin of every
|
|
// phase-token grammar. `src/phase-id.cts` is exempt from the #2128 drift guard
|
|
// (scripts/lint-phase-id-drift.cjs) by construction, and that guard fails any
|
|
// literal re-derivation of the token grammar elsewhere — so the downstream
|
|
// bracket readers (PR-2: roadmap/validate/verify) must build their regexes by
|
|
// interpolating these exports, never by copying the literal.
|
|
//
|
|
// The canonical numeric WIDTH of a bracket identity field, mirroring pad2()'s
|
|
// output: exactly 2 digits, or 3+ with no leading zero. Owned here as a SOURCE
|
|
// so the read side (BRACKET_PHASE_TOKEN_SOURCE, below) and the emit-side
|
|
// validator (CANONICAL_NUMERIC_RE, which toDir enforces) are one rule rather
|
|
// than two literals that agree today and drift tomorrow.
|
|
const BRACKET_CANONICAL_NUMERIC_SOURCE = '(?:[1-9]\\d{2,}|\\d{2})';
|
|
|
|
// BRACKET_PHASE_TOKEN_SOURCE differs from PHASE_NUMBER_TOKEN_SOURCE by a
|
|
// dot-OR-dash sub-separator: a bracket dir/heading numeric run is `MM-PP[.SS]`
|
|
// (a hyphen joins milestone↔phase, a dot joins phase↔sub-phase), whereas M-NN
|
|
// sub-phases are dot-only.
|
|
//
|
|
// The run is POSITIONAL, not a free repetition — `MM-PP[.SS][-LL]` — and each
|
|
// position gets the width its DELIMITER can actually afford:
|
|
//
|
|
// MM leading unbounded — delimited by the `{CODE}.` prefix
|
|
// -PP dash-1 canonical — the grammar REQUIRES this dash, so it is a field
|
|
// separator, not a continuation heuristic
|
|
// .SS dot canonical — a slug carries no dot (toDir sanitizes them
|
|
// away), so this position cannot collide
|
|
// -LL dash-2 #2232 cap — the ONLY slug-adjacent position, and therefore
|
|
// the only one a slug word can collide with
|
|
//
|
|
// #2232 reconciliation: the slug-adjacent position interpolates the single-owner
|
|
// PHASE_CONTINUATION_SEGMENT_SOURCE, so the #2232 bug class cannot reopen on the
|
|
// bracket path — dir `PROJ.01-14-2026-photos-…` (a slug leading with a year)
|
|
// yields `01-14`, never `01-14-2026`.
|
|
//
|
|
// DELIBERATE DIVERGENCE from the M-NN dir-token path (pinned by the parity gate
|
|
// in tests/continuation-grammar-parity.test.cjs, which fails if these two rules
|
|
// drift for a reason nobody intended): the non-slug-adjacent positions stay
|
|
// WIDER than #2232's cap. Bracket admits 3+-digit milestone/phase/sub-phase
|
|
// (CANONICAL_NUMERIC_RE — `[GSD.100] 05` is a pinned regression), and unlike the
|
|
// M-NN continuations those positions are delimiter-disambiguated rather than
|
|
// heuristically recognized, so there is no year collision to defend against.
|
|
// Interpolating the cap verbatim at every position would only under-collect ids
|
|
// that toDir itself emits: `PROJ.02-105-slug` (3-digit phase) would read as
|
|
// `02`, and `[GSD.02] 05.100` (3-digit sub-phase) as `05`. Upstream draws this
|
|
// same line for the same reason — core-utils/phase cap the paired PLAN component
|
|
// while the leading phase component stays unbounded (phase numbers ≥100 are
|
|
// legitimate). The trade-off this accepts is #2232's policy verbatim: a PLAN
|
|
// ≥100 is out of the token grammar.
|
|
//
|
|
// Still deliberately MORE PERMISSIVE than parsePhaseId's strict grammar (it
|
|
// admits a letter-suffixed and unpadded leading token that the parser rejects):
|
|
// this is a READ-TOLERANCE source for the PR-2 readers, which must recognize a
|
|
// bracket-shaped token before deciding what to do with it — it is not the
|
|
// emit/identity grammar. parsePhaseId stays the arbiter of well-formedness.
|
|
const BRACKET_PHASE_TOKEN_SOURCE =
|
|
`\\d+[A-Z]?` +
|
|
`(?:-${BRACKET_CANONICAL_NUMERIC_SOURCE}(?!\\d))?` +
|
|
`(?:\\.${BRACKET_CANONICAL_NUMERIC_SOURCE}(?!\\d))?` +
|
|
`(?:-${PHASE_CONTINUATION_SEGMENT_SOURCE})?` +
|
|
`(?=-|$)`;
|
|
|
|
// A phase HEADING intro under bracket is either a `[...]` bracket (optionally
|
|
// followed by a `Phase ` label) or a bare `Phase ` label; a bare number is NOT
|
|
// a phase-heading intro. The `[^\]]{1,200}` bound mirrors the existing
|
|
// roadmap-parser heading regexes (ReDoS-safe: a header is one short line).
|
|
const PHASE_HEADING_PREFIX_SRC = '(?:\\[[^\\]]{1,200}\\]\\s*(?:Phase\\s+)?|Phase\\s+)';
|
|
|
|
function stripProjectCodePrefix(value: unknown, caseInsensitive = true): string {
|
|
const input = String(value);
|
|
const re = caseInsensitive ? PROJECT_CODE_PREFIX_STRIP_RE_I : PROJECT_CODE_PREFIX_STRIP_RE;
|
|
return input.replace(re, '');
|
|
}
|
|
|
|
function hasProjectCodePrefix(value: unknown): boolean {
|
|
return PROJECT_CODE_PREFIX_STRIP_RE_I.test(String(value));
|
|
}
|
|
|
|
function normalizePhaseName(phase: unknown): string {
|
|
const str = String(phase);
|
|
// Strip optional project_code prefix (e.g., 'CK-01' → '01')
|
|
const stripped = stripProjectCodePrefix(str, false);
|
|
// Milestone-prefixed phase IDs: M-NN or M-N-N (deep decomposition).
|
|
const milestoneMatch = stripped.match(/^(\d+)((?:-\d+)+)([A-Z]?(?:\.\d+)*)$/i);
|
|
if (milestoneMatch) {
|
|
const major = milestoneMatch[1].padStart(2, '0');
|
|
const subSegments = milestoneMatch[2].slice(1).split('-').map(s => s.padStart(2, '0'));
|
|
const suffix = milestoneMatch[3] || '';
|
|
return `${major}-${subSegments.join('-')}${suffix}`;
|
|
}
|
|
// Standard numeric phases: 1, 01, 12A, 12.1
|
|
const match = stripped.match(/^(\d+)([A-Z])?((?:\.\d+)*)/i);
|
|
if (match) {
|
|
const padded = match[1].padStart(2, '0');
|
|
// Preserve original case of letter suffix (#1962).
|
|
const letter = match[2] || '';
|
|
const decimal = match[3] || '';
|
|
return padded + letter + decimal;
|
|
}
|
|
// Custom phase IDs (e.g. PROJ-42, AUTH-101): return as-is
|
|
return str;
|
|
}
|
|
|
|
function getMilestoneFromPhaseId(phaseId: unknown, convention?: string): string | null {
|
|
// READING-B (#612): under the bracket convention the milestone comes from the
|
|
// `[PROJECT.MM]` / `{CODE}.{MM}-` prefix, never the phase-token leading
|
|
// integer (ADR-612 Decision 6). Gated on 'bracket' so the `null` and
|
|
// 'milestone-prefixed' (M-NN) paths keep the legacy leading-int rule
|
|
// (READING-A) below, byte-untouched. The optional parameter keeps this helper
|
|
// pure (no config read) and backward-compatible: every existing single-arg
|
|
// caller resolves to the unchanged READING-A body.
|
|
if (convention === 'bracket') {
|
|
const b = String(phaseId).match(/^([A-Z][A-Z0-9_]*)\.(\d+)/);
|
|
if (!b) return null;
|
|
const mm = parseInt(b[2], 10);
|
|
if (SENTINEL_RANGES.includes(mm)) return null; // sentinel milestones have no real milestone
|
|
return `v${mm}.0`;
|
|
}
|
|
const stripped = stripProjectCodePrefix(phaseId);
|
|
const m = stripped.match(/^0*(\d+)-\d/);
|
|
if (!m) return null;
|
|
const major = parseInt(m[1], 10);
|
|
if (major === 0 || major === 999) return null;
|
|
return `v${major}.0`;
|
|
}
|
|
|
|
function getPhaseDirFromPhaseId(phaseId: unknown, phaseName: string | null | undefined, projectCode: string | null | undefined): string | null {
|
|
const stripped = stripProjectCodePrefix(phaseId);
|
|
const m = stripped.match(/^0*(\d+)-(0*(\d+(?:-\d+)*))$/);
|
|
if (!m) return null;
|
|
const milestone = String(parseInt(m[1], 10)).padStart(2, '0');
|
|
const subParts = m[2].split('-').map(p => String(parseInt(p, 10)).padStart(2, '0'));
|
|
const sub = subParts.join('-');
|
|
const slug = phaseName
|
|
? phaseName.toLowerCase().replace(/[^a-z0-9]+/g, '-').replace(/^-+|-+$/g, '')
|
|
: '';
|
|
const parts = [milestone, sub, slug].filter(Boolean);
|
|
const base = parts.join('-');
|
|
return projectCode ? `${projectCode}-${base}` : base;
|
|
}
|
|
|
|
// ─── Bracket phase-ID grammar (#612, PR-1) ──────────────────────────────────
|
|
// One pure round-trippable model (ADR-612 §3 / Decision 4). parsePhaseId
|
|
// accepts the display form `[PROJECT.MM] PP[.SS][-LL]` or the on-disk/token form
|
|
// `{PROJECT}.{MM}-{PP}[.{SS}][-{LL|slug}]`; renderPhaseId / toDir are its two
|
|
// emitters. READING-B: the milestone lives in the `[PROJECT.MM]` prefix, so no
|
|
// token dimension is ever overloaded (the M-NN collapse pinned in
|
|
// tests/adr-612-collision-characterization.test.cjs cannot occur on this path).
|
|
// `plan` is a filename-surface dimension only — renderPhaseId emits it; toDir
|
|
// drops it (directories carry a slug, not a plan). The project code follows the
|
|
// repo's established `[A-Z][A-Z0-9_]*` grammar (the config-validated
|
|
// project_code shape shared with OPTIONAL_PROJECT_CODE_PREFIX_SOURCE), not the
|
|
// ADR §1 illustration's `[A-Z]{1,6}`, so that every project_code the config
|
|
// permits (digits / underscore / >6 chars) parses.
|
|
//
|
|
// Strict-reject posture (ADR-612 Decision 4's `render(parse(x)) === x`
|
|
// contract, held exactly): parsePhaseId accepts ONLY the canonical form of
|
|
// each branch — unpadded numbers, over-padded numbers, and multi-space or
|
|
// stray leading/trailing whitespace are all rejected rather than silently
|
|
// normalized, so two distinct input strings can never parse to the same
|
|
// tuple while one of them fails to round-trip. toDir mirrors this on the
|
|
// write side: every interpolated PhaseId field is validated (PhaseId is a
|
|
// structural type — nothing forces callers through parsePhaseId, so a hand-
|
|
// built id must not be able to smuggle a path-traversal segment onto disk),
|
|
// and the slug must sanitize to a non-empty, non-all-digit token (an empty
|
|
// slug would leave a dangling trailing hyphen; an all-digit slug is
|
|
// string-indistinguishable from the plan grammar's trailing tail and would
|
|
// silently break the disk↔identity bijection on read-back).
|
|
type PhaseId = {
|
|
project: string; // 'GSD'
|
|
milestone: string; // '02' (zero-padded, from the bracket/dir prefix)
|
|
phase: string; // '05' (zero-padded)
|
|
subphase?: string; // '03' (optional)
|
|
plan?: string; // '01' (filename surface only)
|
|
};
|
|
|
|
const pad2 = (n: string): string => String(parseInt(n, 10)).padStart(2, '0');
|
|
|
|
function parsePhaseId(input: string): PhaseId {
|
|
// No .trim(): the match anchors (`^`...`$`) then reject leading/trailing
|
|
// whitespace outright, folding that case into the same "not a bracket
|
|
// phase id" rejection below rather than needing its own check.
|
|
const str = String(input);
|
|
|
|
// Display form: [PROJECT.MM] PP[.SS][-LL]. The match itself stays
|
|
// permissive on purpose (it will happily match an unpadded number or a
|
|
// multi-space run) — canonicality is enforced UNIFORMLY below via the
|
|
// render round-trip (ADR-612 Decision 4) rather than by hand-tuning every
|
|
// numeric / whitespace sub-pattern, so a field added later inherits the
|
|
// check for free instead of needing its own regex micro-surgery.
|
|
const disp = str.match(/^\[([A-Z][A-Z0-9_]*)\.(\d+)\]\s+(\d+)(?:\.(\d+))?(?:-(\d+))?$/);
|
|
if (disp) {
|
|
const id: PhaseId = { project: disp[1], milestone: pad2(disp[2]), phase: pad2(disp[3]) };
|
|
if (disp[4] !== undefined) id.subphase = pad2(disp[4]);
|
|
if (disp[5] !== undefined) id.plan = pad2(disp[5]);
|
|
// Canonicality by construction: re-render the parsed id and require
|
|
// byte-equality with the input. This rejects unpadded ('[GSD.5] 5'),
|
|
// over-padded ('[GSD.005] 05'), and multi-space-separated ('[GSD.02] 05')
|
|
// variants uniformly, without special-casing any one of them — the emit
|
|
// path (renderPhaseId) is the single source of truth for "canonical".
|
|
if (renderPhaseId(id) !== str) {
|
|
throw new Error(`parsePhaseId: not canonical: ${JSON.stringify(input)}`);
|
|
}
|
|
return id;
|
|
}
|
|
|
|
// Dir / token form: {PROJECT}.{MM}-{PP}[.{SS}][-{plan|slug}]
|
|
const dir = str.match(/^([A-Z][A-Z0-9_]*)\.(\d+)-(\d+)(?:\.(\d+))?(?:-(.+))?$/);
|
|
if (dir) {
|
|
const id: PhaseId = { project: dir[1], milestone: pad2(dir[2]), phase: pad2(dir[3]) };
|
|
if (dir[4] !== undefined) id.subphase = pad2(dir[4]);
|
|
// Trailing segment: a pure-integer tail is the plan; anything else is a
|
|
// slug (dropped from the tuple — it is not an identity dimension). The
|
|
// plan tail participates in the canonicality check below; the slug tail
|
|
// is read-tolerant pass-through (a slug is not an identity dimension) and
|
|
// is exempt from it.
|
|
const tail = dir[5];
|
|
const tailIsPlan = tail !== undefined && /^\d+$/.test(tail);
|
|
if (tailIsPlan) id.plan = pad2(tail);
|
|
|
|
// Canonicality by construction, mirroring the display branch: rebuild the
|
|
// exact dir/token string this id would emit and require it match the
|
|
// input verbatim. Rejects unpadded milestone/phase ('GSD.2-5') and
|
|
// unpadded plan tails ('GSD.02-05-1') without special-casing either.
|
|
const sub = id.subphase ? `.${id.subphase}` : '';
|
|
const tailOut = tail === undefined ? '' : tailIsPlan ? `-${pad2(tail)}` : `-${tail}`;
|
|
const canonical = `${id.project}.${id.milestone}-${id.phase}${sub}${tailOut}`;
|
|
if (canonical !== str) {
|
|
throw new Error(`parsePhaseId: not canonical: ${JSON.stringify(input)}`);
|
|
}
|
|
return id;
|
|
}
|
|
|
|
// Ambiguous / bare tokens (e.g. `02-04`, `05`, `2-01`) match neither branch,
|
|
// as does a display/dir form carrying leading/trailing whitespace (the
|
|
// anchors never match it): reject rather than guess a tuple (ADR-612
|
|
// conservative default). The rejection lives ONLY in this new parser —
|
|
// normalizePhaseName and every other legacy reader keep accepting those
|
|
// tokens unchanged.
|
|
throw new Error(`parsePhaseId: not a bracket phase id: ${JSON.stringify(input)}`);
|
|
}
|
|
|
|
function renderPhaseId(id: PhaseId): string {
|
|
const sub = id.subphase ? `.${id.subphase}` : '';
|
|
const plan = id.plan ? `-${id.plan}` : '';
|
|
return `[${id.project}.${id.milestone}] ${id.phase}${sub}${plan}`;
|
|
}
|
|
|
|
// PhaseId is a structural type: nothing forces a caller through parsePhaseId,
|
|
// so toDir cannot trust project/milestone/phase/subphase are already
|
|
// canonical — each is validated below against the exact shape parsePhaseId
|
|
// itself would ever produce, closing off a hand-built id as a path-traversal
|
|
// vector. PROJECT_ID_RE mirrors the parser's `[A-Z][A-Z0-9_]*` grammar;
|
|
// CANONICAL_NUMERIC_RE mirrors pad2()'s output shape — exactly 2 digits, or
|
|
// 3+ digits with no leading zero. It is BUILT from
|
|
// BRACKET_CANONICAL_NUMERIC_SOURCE rather than re-spelled as a literal, so this
|
|
// emit-side gate and the read-side token source cannot disagree about what
|
|
// "canonical width" means (the anchors here make the source's trailing `(?!\d)`
|
|
// guard, which the unanchored read side needs, redundant).
|
|
const PROJECT_ID_RE = /^[A-Z][A-Z0-9_]*$/;
|
|
const CANONICAL_NUMERIC_RE = new RegExp(`^${BRACKET_CANONICAL_NUMERIC_SOURCE}$`);
|
|
|
|
function toDir(id: PhaseId, slug: string): string {
|
|
if (!PROJECT_ID_RE.test(id.project)) {
|
|
throw new Error(`toDir: invalid project: ${JSON.stringify(id.project)}`);
|
|
}
|
|
if (!CANONICAL_NUMERIC_RE.test(id.milestone)) {
|
|
throw new Error(`toDir: invalid milestone: ${JSON.stringify(id.milestone)}`);
|
|
}
|
|
if (!CANONICAL_NUMERIC_RE.test(id.phase)) {
|
|
throw new Error(`toDir: invalid phase: ${JSON.stringify(id.phase)}`);
|
|
}
|
|
if (id.subphase !== undefined && !CANONICAL_NUMERIC_RE.test(id.subphase)) {
|
|
throw new Error(`toDir: invalid subphase: ${JSON.stringify(id.subphase)}`);
|
|
}
|
|
// A non-string slug (e.g. an omitted second argument) must not be silently
|
|
// coerced by String(...) into the literal token 'undefined'/'null' on disk.
|
|
if (typeof slug !== 'string') {
|
|
throw new Error(`toDir: slug must be a string: ${JSON.stringify(slug)}`);
|
|
}
|
|
|
|
const sub = id.subphase ? `.${id.subphase}` : '';
|
|
// Slug guard: the slug becomes an on-disk path segment, so collapse it to a
|
|
// safe lowercase token — never a path separator or `..` traversal.
|
|
const safeSlug = slug.toLowerCase().replace(/[^a-z0-9]+/g, '-').replace(/^-+|-+$/g, '');
|
|
// A slug that sanitizes to nothing (e.g. '!!!') would otherwise emit a
|
|
// dangling trailing hyphen.
|
|
if (!safeSlug) {
|
|
throw new Error(`toDir: slug sanitizes to empty: ${JSON.stringify(slug)}`);
|
|
}
|
|
// An all-digit slug (e.g. '2026') is string-indistinguishable from the
|
|
// parsePhaseId dir branch's plan tail, so it would re-parse as a plan, not
|
|
// a slug — silently breaking the disk↔identity bijection on read-back.
|
|
if (/^\d+$/.test(safeSlug)) {
|
|
throw new Error(`toDir: slug must not be all-digit: ${JSON.stringify(slug)}`);
|
|
}
|
|
return `${id.project}.${id.milestone}-${id.phase}${sub}-${safeSlug}`;
|
|
}
|
|
|
|
// Milestone integers reserved as non-milestone sentinels (0.x backlog / 999.x
|
|
// icebox); a phase id in these ranges has no real milestone.
|
|
const SENTINEL_RANGES: readonly number[] = Object.freeze([0, 999]);
|
|
|
|
function isSentinelPhaseId(phaseId: unknown, convention?: string): boolean {
|
|
const s = String(phaseId);
|
|
// Bracket milestone lives in the `{CODE}.{MM}` prefix. GATED on
|
|
// convention === 'bracket' for the same reason as extractPhaseToken below and
|
|
// getMilestoneFromPhaseId above: that prefix is string-indistinguishable from
|
|
// the legacy #1324 letter-prefixed-decimal family (`P0.0-foundation` is a real
|
|
// phase, NOT sentinel milestone 0) whenever the code ends in a digit. A
|
|
// convention-less caller uses the legacy/bare leading-int rule below, so no
|
|
// existing reader gains a false positive; the bracket reading is opt-in.
|
|
if (convention === 'bracket') {
|
|
const bracket = s.match(/^[A-Z][A-Z0-9_]*\.(\d+)/); // bracket: milestone in the prefix
|
|
if (bracket) return SENTINEL_RANGES.includes(parseInt(bracket[1], 10));
|
|
}
|
|
const legacy = stripProjectCodePrefix(s).match(/^0*(\d+)/); // legacy/bare: leading int
|
|
if (!legacy) return false;
|
|
return SENTINEL_RANGES.includes(parseInt(legacy[1], 10));
|
|
}
|
|
|
|
/**
|
|
* Render a regex source fragment matching a phase number against ROADMAP/STATE
|
|
* prose regardless of zero-padding on either side.
|
|
*/
|
|
function phaseMarkdownRegexSource(phaseNum: unknown): string {
|
|
const stripped = stripProjectCodePrefix(phaseNum);
|
|
|
|
// Milestone-prefixed IDs: M-NN or M-N-N (deep).
|
|
const milestoneSegments = stripped.match(/^(\d+)((?:-\d+)*)([A-Z]?(?:\.\d+)*)$/i);
|
|
if (milestoneSegments && milestoneSegments[2]) {
|
|
const majorUnpadded = milestoneSegments[1].replace(/^0+/, '') || '0';
|
|
const subParts = milestoneSegments[2].slice(1).split('-');
|
|
const subFragments = subParts.map(s => {
|
|
const unpadded = s.replace(/^0+/, '') || '0';
|
|
return `0*${escapeRegex(unpadded)}`;
|
|
});
|
|
const suffix = milestoneSegments[3] || '';
|
|
const suffixFragment = suffix ? escapeRegex(suffix) : '';
|
|
return `0*${escapeRegex(majorUnpadded)}-${subFragments.join('-')}${suffixFragment}`;
|
|
}
|
|
|
|
// Plain numeric phase: 1, 01, 12A, 12.1
|
|
const match = stripped.match(/^0*(\d+)([A-Z])?((?:\.\d+)*)$/i);
|
|
if (!match) return escapeRegex(phaseNum);
|
|
|
|
const integer = match[1].replace(/^0+/, '') || '0';
|
|
const letter = match[2] ? escapeRegex(match[2]) : '';
|
|
const decimal = match[3] ? escapeRegex(match[3]) : '';
|
|
return `0*${escapeRegex(integer)}${letter}${decimal}`;
|
|
}
|
|
|
|
/**
|
|
* #3599: when the caller passed a project-code-prefixed ID like `PROJ-42`,
|
|
* return the exact-escaped form.
|
|
*/
|
|
function phaseMarkdownRegexSourceExact(phaseNum: unknown): string | null {
|
|
const raw = String(phaseNum);
|
|
if (!hasProjectCodePrefix(raw)) return null;
|
|
return escapeRegex(raw);
|
|
}
|
|
|
|
function comparePhaseNum(a: unknown, b: unknown): number {
|
|
// Strip optional project_code prefix before comparing
|
|
const sa = stripProjectCodePrefix(a);
|
|
const sb = stripProjectCodePrefix(b);
|
|
|
|
const milestoneA = sa.match(/^(\d+)((?:-\d+)+)([A-Z]?(?:\.\d+)*)$/i);
|
|
const milestoneB = sb.match(/^(\d+)((?:-\d+)+)([A-Z]?(?:\.\d+)*)$/i);
|
|
|
|
if (milestoneA && milestoneB) {
|
|
const segsA = [parseInt(milestoneA[1], 10), ...milestoneA[2].slice(1).split('-').map(s => parseInt(s, 10))];
|
|
const segsB = [parseInt(milestoneB[1], 10), ...milestoneB[2].slice(1).split('-').map(s => parseInt(s, 10))];
|
|
const maxSegs = Math.max(segsA.length, segsB.length);
|
|
for (let i = 0; i < maxSegs; i++) {
|
|
const av = segsA[i] !== undefined ? segsA[i] : 0;
|
|
const bv = segsB[i] !== undefined ? segsB[i] : 0;
|
|
if (av !== bv) return av - bv;
|
|
}
|
|
const sufA = milestoneA[3] || '';
|
|
const sufB = milestoneB[3] || '';
|
|
if (sufA !== sufB) return sufA < sufB ? -1 : 1;
|
|
return 0;
|
|
}
|
|
|
|
if (milestoneA || milestoneB) return String(a).localeCompare(String(b));
|
|
|
|
const pa = sa.match(/^(\d+)([A-Z])?((?:\.\d+)*)/i);
|
|
const pb = sb.match(/^(\d+)([A-Z])?((?:\.\d+)*)/i);
|
|
if (!pa || !pb) return String(a).localeCompare(String(b));
|
|
const intDiff = parseInt(pa[1], 10) - parseInt(pb[1], 10);
|
|
if (intDiff !== 0) return intDiff;
|
|
const la = (pa[2] || '').toUpperCase();
|
|
const lb = (pb[2] || '').toUpperCase();
|
|
if (la !== lb) {
|
|
if (!la) return -1;
|
|
if (!lb) return 1;
|
|
return la < lb ? -1 : 1;
|
|
}
|
|
const aDecParts = pa[3] ? pa[3].slice(1).split('.').map(p => parseInt(p, 10)) : [];
|
|
const bDecParts = pb[3] ? pb[3].slice(1).split('.').map(p => parseInt(p, 10)) : [];
|
|
const maxLen = Math.max(aDecParts.length, bDecParts.length);
|
|
if (aDecParts.length === 0 && bDecParts.length > 0) return -1;
|
|
if (bDecParts.length === 0 && aDecParts.length > 0) return 1;
|
|
for (let i = 0; i < maxLen; i++) {
|
|
const av = Number.isFinite(aDecParts[i]) ? aDecParts[i] : 0;
|
|
const bv = Number.isFinite(bDecParts[i]) ? bDecParts[i] : 0;
|
|
if (av !== bv) return av - bv;
|
|
}
|
|
return 0;
|
|
}
|
|
|
|
/**
|
|
* Extract the phase token from a directory name.
|
|
*/
|
|
function extractPhaseToken(dirName: string, convention?: string): string {
|
|
// #612 bracket dir form `{CODE}.{MM}-{PP}[.{SS}]-slug` → phase token `PP[.SS]`.
|
|
// GATED on convention === 'bracket' (mirrors getMilestoneFromPhaseId's READING-B
|
|
// decision above). A bracket dir `{CODE}.{MM}-{PP}` is string-INDISTINGUISHABLE
|
|
// from the legacy #2043/#1324 letter-prefixed-decimal family (`P0.3-2`,
|
|
// `P0.12-34`) whenever the project code ends in a digit, so NO string-only
|
|
// discriminator can separate the two conventions — auto-detecting here silently
|
|
// reinterpreted `P0.3-2` → `2` (was `P0.3-2`), a byte-identical-read regression
|
|
// on this CRITICAL 6-caller helper (ADR-2121). Requiring an explicit convention
|
|
// signal keeps every existing (convention-less) call site byte-identical to
|
|
// prior behaviour — see the #2043 numeric-tail characterization in
|
|
// tests/phase-id.test.cjs — while keeping the helper pure (optional param, no
|
|
// config read). The captured token is dot-only (`PP[.SS]`); the milestone↔phase
|
|
// hyphen and any trailing plan/slug are excluded.
|
|
if (convention === 'bracket') {
|
|
const bracketDir = dirName.match(/^[A-Z][A-Z0-9_]*\.\d+-(\d+(?:\.\d+)?)/);
|
|
if (bracketDir) return bracketDir[1];
|
|
}
|
|
|
|
const codePrefixMatch = dirName.match(PROJECT_CODE_PREFIX_CAPTURE_RE_I);
|
|
let prefix = '';
|
|
let rest = dirName;
|
|
if (codePrefixMatch) {
|
|
prefix = codePrefixMatch[1] + '-';
|
|
rest = codePrefixMatch[2];
|
|
}
|
|
|
|
const segments = rest.split('-');
|
|
const tokenSegments: string[] = [];
|
|
// #2043: distinguish a real (zero-padded) phase/sub-phase segment from a
|
|
// single-digit slug word. A pure-numeric leading segment ("46") only
|
|
// continues with exactly-2-digit segments (#2232: a ≥3-digit run is a slug
|
|
// word such as a year — "14-2026-photos-…" yields "14", not "14-2026"), so
|
|
// "46-6-rs-…" yields "46" (the "6" is the
|
|
// slug's first word), not "46-6". Milestone-prefixed ids like "M1-2" reach here
|
|
// with "M1-" already stripped as a project-code prefix (see
|
|
// PROJECT_CODE_PREFIX_CAPTURE_RE_I), so "2" is the leading segment and the same
|
|
// pure-numeric rule applies (M1-46-6-rs → "M1-46"). The firstLetterPrefixed
|
|
// carve-out covers letter+digit leading segments that survive prefix stripping
|
|
// because of punctuation (e.g. "P0.3-2"), whose single-digit continuation is
|
|
// intentionally preserved (unchanged from prior behaviour).
|
|
let firstLetterPrefixed = false;
|
|
for (let i = 0; i < segments.length; i++) {
|
|
const seg = segments[i];
|
|
if (i === 0) {
|
|
if (/^\d/.test(seg)) {
|
|
tokenSegments.push(seg);
|
|
} else if (/^[A-Za-z]{1,3}\d/.test(seg)) {
|
|
tokenSegments.push(seg);
|
|
firstLetterPrefixed = true;
|
|
} else {
|
|
break;
|
|
}
|
|
} else if (
|
|
(firstLetterPrefixed && /^\d/.test(seg)) ||
|
|
(!firstLetterPrefixed && isPhaseContinuationSegment(seg))
|
|
) {
|
|
tokenSegments.push(seg);
|
|
} else {
|
|
break;
|
|
}
|
|
}
|
|
|
|
if (tokenSegments.length === 0) {
|
|
return dirName;
|
|
}
|
|
|
|
// #2528 (re-review): the tokenizer deliberately does NOT try to tell a 2-digit
|
|
// slug word ("24" of "24/7 Autonomy") from a genuine zero-padded continuation
|
|
// ("24" of sub-phase 10.24) — by width alone they are the same string, the gap
|
|
// between #2043's 1-digit and #2232's ≥3-digit guards, and no LOCAL signal
|
|
// separates them. An earlier revision of this fix rewound the token when the
|
|
// segment that stopped the scan was a 1-digit word, which reads
|
|
// "10-24-7-autonomy" correctly but silently re-tokenizes the equally real
|
|
// "10-24-7-zip" (sub-phase 10.24 named "7-Zip …") from "10-24" to "10" — it
|
|
// trades the reported ambiguity for the symmetric one a level down, on a
|
|
// CRITICAL 15-caller chokepoint whose output feeds query-less derivations
|
|
// (STATE.md phase counts, W007, the #2562 key surface).
|
|
//
|
|
// So the token stays the LITERAL reading of the name, and disambiguation lives
|
|
// ONE layer up, in matchPhaseDirs, where a QUERY exists to disambiguate
|
|
// against: a bare-integer lookup falls back to the directory's leading digit
|
|
// run and resolves "10-24-7-autonomy" for "10" without touching what the
|
|
// directory's own token is. That is the same bounded mechanism the
|
|
// "05-80-20-cleanup" shape already uses — one rule for the whole
|
|
// digit-leading-slug family instead of two overlapping ones.
|
|
//
|
|
// A generated slug is lowercase. If the owner admitted a two-digit prefix
|
|
// from a digit+letter slug segment ("10x", "25abc"), remove only that final
|
|
// segment. Uppercase suffixes remain available to the established plan-ID
|
|
// grammar, and dotted continuations remain intact.
|
|
if (
|
|
!firstLetterPrefixed &&
|
|
tokenSegments.length > 1 &&
|
|
/^\d{2}[a-z][a-z0-9]*$/.test(tokenSegments[tokenSegments.length - 1])
|
|
) {
|
|
tokenSegments.pop();
|
|
}
|
|
|
|
return prefix + tokenSegments.join('-');
|
|
}
|
|
|
|
/**
|
|
* Check if a directory name's phase token matches the normalized phase exactly.
|
|
*/
|
|
function phaseTokenMatches(dirName: string, normalized: string): boolean {
|
|
const token = extractPhaseToken(dirName);
|
|
if (token.toUpperCase() === normalized.toUpperCase()) return true;
|
|
const stripped = stripProjectCodePrefix(dirName);
|
|
if (stripped !== dirName) {
|
|
const strippedToken = extractPhaseToken(stripped);
|
|
if (strippedToken.toUpperCase() === normalized.toUpperCase()) return true;
|
|
}
|
|
return false;
|
|
}
|
|
|
|
/**
|
|
* #2528: the LEADING DIGIT RUN of a directory name — the fragment the
|
|
* bare-integer fallback selects on, and the one `phaseNumberForMatch` then
|
|
* displays. Named (per this module's convention of naming grammar fragments
|
|
* rather than inlining them) because the two sites must not drift: selecting on
|
|
* one run and displaying another would resolve a directory and then label it
|
|
* with a number that never matched.
|
|
*
|
|
* `LEADING_DIGIT_RUN_RE` anchors a trailing `-`-or-end so the run is a whole
|
|
* segment; `_PREFIX` is the same run without that boundary, for reading the run
|
|
* back off a name already known to match.
|
|
*/
|
|
const LEADING_DIGIT_RUN_SOURCE = '\\d+';
|
|
const LEADING_DIGIT_RUN_RE = new RegExp(`^(${LEADING_DIGIT_RUN_SOURCE})(?:-|$)`);
|
|
const LEADING_DIGIT_RUN_PREFIX_RE = new RegExp(`^${LEADING_DIGIT_RUN_SOURCE}`);
|
|
const BARE_INTEGER_RE = new RegExp(`^${LEADING_DIGIT_RUN_SOURCE}$`);
|
|
|
|
/** Strip leading zeros for numeric-equality compare, keeping a lone "0". */
|
|
const unpad = (digits: string): string => digits.replace(/^0+(?=\d)/, '');
|
|
|
|
/**
|
|
* #2528: the CANONICAL phase-directory match selection — the one rule every
|
|
* directory-resolution path (the shared locator plus the `find-phase` and
|
|
* `phase-plan-index` command scans) applies to a candidate dir list. Extracted
|
|
* here because the surrounding scan/ambiguity/shaping code exists per site and
|
|
* had already diverged; the selection itself must not.
|
|
*
|
|
* Two passes:
|
|
* 1. PRIMARY — exact token match (`phaseTokenMatches`), unchanged behavior.
|
|
* 2. BARE-INTEGER FALLBACK — only when the primary pass matched NOTHING and
|
|
* the query is a bare integer, re-filter by each directory's own LEADING
|
|
* digit run (zero-padded compare). This catches digit-leading slug shapes
|
|
* the tokenizer cannot disambiguate from genuine sub-phase segments
|
|
* (e.g. "05-80-20-cleanup", phase 5 named "80/20 Cleanup", whose token
|
|
* "05-80-20" is byte-identical in shape to a real deep-decomposition dir).
|
|
* The fallback can only turn a silent not-found into a resolution or into
|
|
* a surfaced ambiguity (callers keep their #2237 multi-match guards) —
|
|
* never override a primary match.
|
|
*
|
|
* SCOPE, precisely (#2528 re-review). Non-bare QUERIES ("46-6", "12A",
|
|
* "PROJ-42") never enter the fallback, so nothing changes about how a
|
|
* deep-decomposition or letter-suffix lookup is asked. What DOES change is the
|
|
* DIRECTORY side: a bare query now reaches directories the tokenizer classified
|
|
* as multi-segment, and a genuine sub-phase directory has exactly that shape.
|
|
* So `5` against a lone `05-01-auth` resolves (phase_number "05", phase_name
|
|
* "01-auth") where it previously found nothing.
|
|
*
|
|
* That widening is DELIBERATE and it is irreducible from directory names alone.
|
|
* `05-01-auth` (sub-phase 5.1) and `30-12-factor-refactor` (phase 30 named
|
|
* "12-Factor Refactor") are the same string shape — `NN-NN-<slug>` — and the
|
|
* discriminator that would separate them, "is the second segment a valid decimal
|
|
* sub-phase", accepts both (`5.1` and `30.12` are equally well-formed). Any rule
|
|
* strong enough to exclude `05-01-auth` also excludes `30-12-factor-refactor`,
|
|
* which is the defect #2528 exists to fix. The tie is therefore broken in favour
|
|
* of resolving, and the consequence is bounded on the side that matters: when
|
|
* BOTH readings have a directory (`05-01-auth` + `05-02-api`) the result is two
|
|
* matches. `tests/phase-resolution-parity.test.cjs` pins both directions: the
|
|
* lone-directory resolution and the two-directory refusal.
|
|
*
|
|
* WHAT IS SHARED IS SELECTION, NOT AMBIGUITY POLICY. This function is the one
|
|
* owner of "which directories does this query name". What a caller does with
|
|
* two of them stays the caller's own decision, and the callers split in two
|
|
* tiers on purpose:
|
|
*
|
|
* REFUSE on `matches.length > 1` — `searchPhaseInDir`, `cmdFindPhase`,
|
|
* `cmdPhasePlanIndex`, `cmdPhaseRemove`. These either act destructively or
|
|
* answer "which phase is this", so guessing is worse than reporting the
|
|
* candidates (#2237).
|
|
*
|
|
* TAKE `matches[0]` — `cmdPhasesList`, `cmdInitManager`, `cmdRoadmapAnalyze`,
|
|
* `cmdVerifySchemaDrift`, `detectVerifyFailed`. Each read a directory to
|
|
* DECORATE a row they are already emitting; each used `.find()` before this
|
|
* PR, so first-match is their prior behavior preserved verbatim, and each is
|
|
* order-stable because the directory list is sorted and this function filters
|
|
* without reordering.
|
|
*
|
|
* The honest caveat on that second tier: the bare-number fallback makes
|
|
* multi-match newly REACHABLE for inputs that previously found nothing, so those
|
|
* five can now silently pick one of several candidates where they used to report
|
|
* not-found. That is a widening of an existing first-match rule, not a new rule
|
|
* — but it is a widening, and promoting any of them to refusal is a UX decision
|
|
* about their own output, not a change to selection, so it does not belong here.
|
|
*
|
|
* `usedBareFallback` tells callers to derive the displayed phase number from
|
|
* the directory's leading digit run instead of `extractPhaseToken` (whose
|
|
* token for these dirs is the mis-absorbed multi-segment form).
|
|
*/
|
|
function matchPhaseDirs(dirs: string[], normalized: string): { matches: string[]; usedBareFallback: boolean } {
|
|
const primary = dirs.filter(d => phaseTokenMatches(d, normalized));
|
|
if (primary.length > 0) return { matches: primary, usedBareFallback: false };
|
|
|
|
const bare = String(normalized);
|
|
if (!BARE_INTEGER_RE.test(bare)) return { matches: primary, usedBareFallback: false };
|
|
const want = unpad(bare);
|
|
|
|
const fallback = dirs.filter(d => {
|
|
const m = stripProjectCodePrefix(d).match(LEADING_DIGIT_RUN_RE);
|
|
return m !== null && unpad(m[1]) === want;
|
|
});
|
|
return { matches: fallback, usedBareFallback: fallback.length > 0 };
|
|
}
|
|
|
|
/**
|
|
* #2528: the display phase number for a directory selected by matchPhaseDirs.
|
|
* Primary matches keep the extracted token; bare-fallback matches use the
|
|
* directory's leading digit run (the whole point of the fallback is that the
|
|
* extracted token is wrong for these dirs).
|
|
*/
|
|
function phaseNumberForMatch(dirName: string, usedBareFallback: boolean): string {
|
|
if (!usedBareFallback) return extractPhaseToken(dirName);
|
|
const stripped = stripProjectCodePrefix(dirName);
|
|
const prefix = dirName.slice(0, dirName.length - stripped.length);
|
|
const m = stripped.match(LEADING_DIGIT_RUN_PREFIX_RE);
|
|
return m ? prefix + m[0] : extractPhaseToken(dirName);
|
|
}
|
|
|
|
// ─── Canonical phase KEY surface (#2562) ─────────────────────────────────────
|
|
//
|
|
// A phase "key" is the padding-, case- and project-code-insensitive identity of
|
|
// a phase, for use as a Map/Set key when two independently-derived phase
|
|
// references (a ROADMAP table cell and a phase directory name, say) must be
|
|
// compared. Promoted here from a local pair in state.cts (#2445) so every
|
|
// consumer derives BOTH sides of a comparison from the SAME function — deriving
|
|
// one side with a bespoke regex is the #2562 defect class (a `01` table cell
|
|
// never matching a `1-slug` directory, silently zeroing a rollup).
|
|
|
|
/**
|
|
* Canonical key for an already-extracted phase TOKEN (`"5"`, `"05"`, `"005"`,
|
|
* `"12A"`, `"30.1"`, `"PROJ-05"`). Padding- and case-insensitive: every
|
|
* spelling of a number collapses to one key.
|
|
*
|
|
* Leading zeros are stripped per hyphen-separated segment BEFORE
|
|
* `normalizePhaseName` pads to the 2-digit convention. Padding alone is not a
|
|
* normalisation — `padStart(2)` is a no-op once the input is already ≥2
|
|
* characters, so `5` yielded `05` while `005` stayed `005` and the two never
|
|
* compared equal. The strip is deliberately confined to this key surface:
|
|
* `normalizePhaseName` itself is a RENDERING function whose verbatim treatment
|
|
* of wide IDs (`001.10`) is relied on by plan-ID capture and wave assignment.
|
|
* Arithmetic is avoided (`parseInt` would lose precision on a long digit run).
|
|
*/
|
|
function phaseKeyFromToken(token: unknown): string {
|
|
const stripped = String(token)
|
|
.split('-')
|
|
.map(segment => segment.replace(/^0+(?=\d)/, ''))
|
|
.join('-');
|
|
return normalizePhaseName(stripped).toUpperCase();
|
|
}
|
|
|
|
/**
|
|
* Canonical key for a phase DIRECTORY name (`"05-schedule-8"` → `"05"`,
|
|
* `"PROJ-5-x"` → `"05"`, `"30.1-follow-up"` → `"30.1"`).
|
|
*/
|
|
function phaseKeyFromDir(dirName: string): string {
|
|
return phaseKeyFromToken(extractPhaseToken(dirName));
|
|
}
|
|
|
|
/**
|
|
* Canonical key for a phase referenced in PROSE — a ROADMAP `## Progress` table
|
|
* cell (`"30. Schedule 8 rollout"`, `"**05.1 Follow-up**"`) or a STATE.md
|
|
* `Phase:` value. Markdown emphasis is stripped first so a bolded cell is not
|
|
* mistaken for a non-phase. Returns null when the value does not BEGIN with a
|
|
* phase token (`parsePhaseFromProse` anchoring, #2111).
|
|
*/
|
|
function phaseKeyFromProse(value: string | null | undefined): string | null {
|
|
if (value == null) return null;
|
|
const { phase } = parsePhaseFromProse(String(value).replace(/[*_`~]/g, ''));
|
|
return phase === null ? null : phaseKeyFromToken(phase);
|
|
}
|
|
|
|
/**
|
|
* The PARENT phase key of a sub-phase key (`"30.1"` → `"30"`), or null for a
|
|
* top-level phase. A sub-phase directory inserted mid-milestone frequently has
|
|
* no ROADMAP row of its own and inherits its parent's milestone (#2562).
|
|
*/
|
|
function parentPhaseKey(key: string): string | null {
|
|
const dot = key.indexOf('.');
|
|
return dot === -1 ? null : key.slice(0, dot);
|
|
}
|
|
|
|
// ─── #2121 canonical surface (ADR-2121) ──────────────────────────────────────
|
|
|
|
/**
|
|
* Parse a phase identifier from a STATE.md `Phase:` prose field VALUE — the text
|
|
* after the `Phase:` label (e.g. `"3 of 4 (Delta)"`, `"3A — Delta (executing)"`,
|
|
* or `"Milestone v0.5 complete"`).
|
|
*
|
|
* The token is anchored to the START of the value (after an optional literal
|
|
* `Phase ` label and an optional project-code prefix) so a phase is only
|
|
* returned when the value actually begins with one. This is the #2111 fix: the
|
|
* prior unanchored `/\b(\d+[A-Z]?(?:\.\d+)*)\b/i` mined the first numeral
|
|
* anywhere, so `"Milestone v0.5 complete"` collapsed to `"5"` (the minor-version
|
|
* digit) and `"v1.0"` to `"0"` (a reserved sentinel). Here both yield
|
|
* `{ phase: null }` because they do not begin with a phase token. The name
|
|
* extraction (parenthetical or em-dash tail, minus status words) is unchanged.
|
|
*/
|
|
function parsePhaseFromProse(value: string | null): { phase: string | null; name: string | null } {
|
|
if (!value) return { phase: null, name: null };
|
|
// Coerce defensively so a non-string caller cannot throw on this canonical
|
|
// surface (mirrors the sibling #2121 functions' String(...) handling).
|
|
const str = String(value);
|
|
const phaseMatch = str.match(/^\s*(?:Phase\s+)?(?:[A-Z][A-Z0-9_]*-)?(\d+[A-Z]?(?:\.\d+)*)\b/i);
|
|
// The name-extraction quantifiers are length-bounded so a crafted long
|
|
// unterminated run (many `(` or `—`) in an untrusted STATE.md field value
|
|
// cannot drive O(n^2) regex backtracking (CPU-exhaustion DoS). A real phase
|
|
// name is far shorter than the cap.
|
|
const parenName = str.match(/\(([^)]{1,200})\)/);
|
|
// #2736 (the #1695 AC #3 residual): status-keyword-aware precedence. The
|
|
// first-party writer shapes are `N — Name (aside)` (completePhaseCore),
|
|
// `N (Name) — EXECUTING` (beginPhaseCore), `N — COMPLETE`, and the
|
|
// gsd2-import `N (slug) — Milestone: Title`. A blind paren-first read
|
|
// harvests the aside as the name on the first shape; a blind dash-first
|
|
// read harvests the status keyword on the others. Prefer the em-dash name
|
|
// when it is a genuine name, else fall back to the parenthetical. Still
|
|
// lossy for names that themselves contain a parenthetical — transitions
|
|
// that hold the exact name bypass this parser entirely via the
|
|
// syncStateFrontmatter authoritative override.
|
|
//
|
|
// The em-dash separator is searched on a paren-stripped copy, so an em-dash
|
|
// INSIDE a parenthetical name (`16 (Native — Global Hotkey) — EXECUTING`)
|
|
// can never be mistaken for the name separator.
|
|
const strNoParens = str.replace(/\([^)\n]{0,200}\)/g, ' ');
|
|
const dashName = strNoParens.match(/—\s*([^(\n]{1,200}?)\s*$/);
|
|
// The precedence-decision vocabulary is deliberately broader than the final
|
|
// name-nulling filter below: a dash tail that merely LOOKS like a status
|
|
// annotation should lose to a parenthetical name, without changing which
|
|
// extracted names are nulled (that set stays the long-standing three).
|
|
const STATUS_WORD_RE = /^(?:complete|executing|not started)$/i;
|
|
const STATUSY_TAIL_RE = /^(?:completed?|executing|not started|planning|planned|ready(?:\s+to\s+\S.{0,50})?|done|in progress|blocked|paused|verifying)$/i;
|
|
const dashRaw = dashName?.[1]?.trim() ?? null;
|
|
const dashIsName = dashRaw !== null && dashRaw.length > 0
|
|
&& !STATUSY_TAIL_RE.test(dashRaw)
|
|
&& !/^milestone\s*:/i.test(dashRaw)
|
|
// A lone ALL-CAPS token after the dash reads as a status marker whenever a
|
|
// parenthetical name exists to prefer (the beginPhase writer's systematic
|
|
// `(Name) — STATUS` shape); with no parenthetical it stays the best guess.
|
|
&& !(parenName && /^[A-Z][A-Z0-9_-]*$/.test(dashRaw));
|
|
const rawName = dashIsName ? dashRaw : (parenName?.[1] ?? dashRaw ?? null);
|
|
const name = rawName && !STATUS_WORD_RE.test(rawName.trim())
|
|
? rawName.trim()
|
|
: null;
|
|
return {
|
|
phase: phaseMatch ? phaseMatch[1] : null,
|
|
name,
|
|
};
|
|
}
|
|
|
|
/**
|
|
* Config-AWARE project-code prefix strip. Unlike the config-blind
|
|
* `stripProjectCodePrefix` (which strips ANY `<CODE>-` shape), this strips the
|
|
* leading `<CODE>-` ONLY when `<CODE>` case-insensitively equals the configured
|
|
* `projectCode`. A foreign prefix (`MEM-01` when the configured code is `LKML`)
|
|
* or an absent/empty `projectCode` is preserved verbatim — this is the #2104
|
|
* fix: a foreign-prefixed id must not collapse to a bare numeric phase and
|
|
* collide with a real one.
|
|
*/
|
|
function stripConfiguredProjectCodePrefix(value: unknown, projectCode: string | null | undefined): string {
|
|
const input = String(value);
|
|
const configured = typeof projectCode === 'string' ? projectCode.trim() : '';
|
|
if (!configured) return input;
|
|
const m = input.match(PROJECT_CODE_PREFIX_CAPTURE_RE_I);
|
|
if (!m) return input;
|
|
if (m[1].toUpperCase() !== configured.toUpperCase()) return input;
|
|
return m[2];
|
|
}
|
|
|
|
/**
|
|
* True when `phase` carries a project-code prefix that is NOT the configured
|
|
* `projectCode` (or when no `projectCode` is configured). The canonical
|
|
* predicate the init-command foreign-prefix guard (#2056 / PR #2105) delegates
|
|
* to, so every call site shares one foreign-prefix rule.
|
|
*/
|
|
function isForeignPrefixedPhaseQuery(phase: unknown, projectCode: unknown): boolean {
|
|
const m = String(phase).match(PROJECT_CODE_PREFIX_CAPTURE_RE_I);
|
|
if (!m) return false;
|
|
const configured = typeof projectCode === 'string' ? projectCode.trim() : '';
|
|
return !configured || m[1].toUpperCase() !== configured.toUpperCase();
|
|
}
|
|
|
|
/**
|
|
* Canonical ROADMAP heading lookup-source list (moved here from
|
|
* roadmap-parser.cts so phase-id.cts is the single owner of the ordering).
|
|
* Sources are tried in a fixed, deduplicated order: exact (only when the query
|
|
* itself is project-code-prefixed) → bare numeric / padding-tolerant →
|
|
* prefix-tolerant fallback. The bare numeric source precedes the prefix-tolerant
|
|
* form so a canonical heading (`### Phase 117:`) is preferred over a drifted
|
|
* prefixed one (`### Phase MANIFOLD-117:`) when both exist in one ROADMAP.
|
|
*/
|
|
function roadmapPhaseLookupSources(phaseNum: unknown): string[] {
|
|
const sources: string[] = [];
|
|
const exactSource = phaseMarkdownRegexSourceExact(phaseNum);
|
|
if (exactSource) sources.push(exactSource);
|
|
|
|
const numericSource = phaseMarkdownRegexSource(phaseNum);
|
|
sources.push(numericSource);
|
|
sources.push(`${OPTIONAL_PROJECT_CODE_PREFIX_SOURCE}${numericSource}`);
|
|
|
|
return [...new Set(sources)];
|
|
}
|
|
|
|
export = {
|
|
escapeRegex,
|
|
OPTIONAL_PROJECT_CODE_PREFIX_SOURCE,
|
|
OPTIONAL_PHASE_TAG_SOURCE,
|
|
PHASE_NUMBER_TOKEN_SOURCE,
|
|
CASE_FLEXIBLE_PROJECT_CODE_PREFIX_SOURCE,
|
|
CASE_FLEXIBLE_PHASE_NUMBER_TOKEN_SOURCE,
|
|
PHASE_CONTINUATION_SEGMENT_SOURCE,
|
|
isPhaseContinuationSegment,
|
|
BRACKET_PHASE_TOKEN_SOURCE,
|
|
PHASE_HEADING_PREFIX_SRC,
|
|
stripProjectCodePrefix,
|
|
normalizePhaseName,
|
|
getMilestoneFromPhaseId,
|
|
getPhaseDirFromPhaseId,
|
|
parsePhaseId,
|
|
renderPhaseId,
|
|
toDir,
|
|
SENTINEL_RANGES,
|
|
isSentinelPhaseId,
|
|
phaseMarkdownRegexSource,
|
|
phaseMarkdownRegexSourceExact,
|
|
comparePhaseNum,
|
|
extractPhaseToken,
|
|
phaseTokenMatches,
|
|
matchPhaseDirs,
|
|
phaseNumberForMatch,
|
|
phaseKeyFromToken,
|
|
phaseKeyFromDir,
|
|
phaseKeyFromProse,
|
|
parentPhaseKey,
|
|
parsePhaseFromProse,
|
|
stripConfiguredProjectCodePrefix,
|
|
isForeignPrefixedPhaseQuery,
|
|
roadmapPhaseLookupSources,
|
|
};
|