* test(#3185): failing-first phase-enumeration single-owner suite Covers the enumeration rows with direct code evidence: 999.* backlog dirs listed by progress/stats, the phase-0 sentinel divergence, the #1324 letter-prefixed-decimal negative space, and the destructive-path find — cmdPhasesClear carries a fifth sentinel copy (/^999(?:\.|$)/) that excludes 999 but not 0, so a 0-* directory roadmap.analyze preserves is deleted there. Also covers the pass-all degrade, which is where the defect actually lives: when the milestone window declares no phases the filter becomes a literal () => true and its heading-side sentinel exclusion is unreachable. A fixture carrying phase headings keeps the filter active and never reaches that path. Named for the derivation, not a module: the suite drives commands, phase, milestone, workstream-inventory and state, and both the phase and phase-locator buckets are already at the per-module test-file cap. Committed alone so the remote runner records the failure before the fix. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): phase enumeration has one owner and a decidable scope Adds phase-locator.cts::listMilestonePhaseDirs as the single canonical owner of "which phase directories belong to the current milestone". It applies the milestone window AND the sentinel filter and returns a ScopedResult, so a caller can tell a genuinely-empty milestone from an enumeration that could not be scoped. The sentinel test now runs against DIRECTORY NAMES and is unconditional. getMilestonePhaseFilter excludes sentinels from its ROADMAP heading set, but degrades to a literal () => true pass-all predicate when that set is empty -- at which point the heading set is never consulted and its sentinel exclusion is unreachable exactly when it is needed. That degrade is the #3167 path, and it is why stats already used the filter and still listed backlog directories. The narrowing is sentinel-only: pass-all stays over-inclusive otherwise. Sentinel copies deleted, canonical isSentinelPhaseId adopted: - cmdRoadmapAnalyze's local closure (parseInt === 0 || === 999), 2 call sites - cmdPhasesClear's /^999(?:\.|$)/ -- the DESTRUCTIVE path, which excluded 999 but not 0, so a 0-* directory roadmap.analyze preserves was deleted cmdStats also seeded rows from ROADMAP headings with no sentinel filter, so a 999 heading produced a row with no directory; that seed is filtered now. cmdPhasesList routes only its ENUMERATION. --phase lookup searches the physical set (scoping it would report an out-of-window phase as not found) and --include-archived still merges archived dirs (they are by definition from other milestones). Both exempt by documented reason, never a file allowlist. Fixed inline, found while building: isDirInMilestone could not match a #1324 letter-prefixed-decimal directory (P0.0-foundation) to its own Phase P0.0 heading, so stats reported the phase with plans: 0 while its directory held plan files. Defers to phase-id's extractPhaseToken rather than widening a fourth bespoke regex; additive, so it can only admit directories. Refs #3180. Closes #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): route the last two enumeration re-derivations workstream-inventory countRoadmapPhases counted every `Phase` heading across the whole ROADMAP -- no window, no sentinel filter -- so it counted 999.* backlog and Phase 0 and spanned every milestone the document ever had. Its own caller already resolved a currentVersion and passed it to getMilestonePhaseFilter elsewhere in the same file; this was the sibling copy that never got the fix. state.cts phaseInventoryProvider enumerated phase dirs with its own /^(\d+)-(.+)$/ convention regex and neither filter, so a rebuilt STATE.md inventory carried backlog and sentinel directories as current-milestone phases. A non-COMPLETE enumeration scope now throws to the outer catch as a real scan failure rather than reporting a confident undercount, mirroring the per-phase scanPhasePlans contract beside it. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): consolidate 23 sentinel re-derivations onto one predicate The whole-repo drift guard (ADR-3180 Decision 4a, no file allowlist) found the sentinel rule re-implemented 23 times across 8 modules, in three regex variants plus four integer-comparison forms. Most tested 999 only, so Phase 0 slipped through them while roadmap.analyze and the engine-wide convention (#1580) both treat 0 and 999 alike. That disagreement is the defect class this epic removes. All 23 now call phase-id's isSentinelPhaseId (SENTINEL_RANGES [0,999]). Sites: init recommended-actions and backlog counts, milestone phase scan, the phase-lifecycle progress table, phase.cts used-number collection and the four renumber-on-remove guards, roadmap-parser's heading and bullet milestone counts, roadmap get-phase fallbacks, and state's heading denominator. Excluding Phase 0 at these sites is a deliberate behavior change and the point of the consolidation — several carried comments already saying 0 should be excluded while the literal beside them caught only 999. Adds scripts/lint-phase-enumeration-drift.cjs, wired into lint:ci. It scans the whole src/ tree with no file allowlist and reports both shapes: an independent phases-dir enumeration, and an independent sentinel literal. Exemptions are function-scoped with a written reason. The guard is comment-aware — its first pass flagged JSDoc and a comment documenting that the code below uses the canonical owner, which would have trained readers to exempt prose. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): resolve every phases-dir enumeration; drift guard reports zero Per-site triage of the 31 remaining whole-repo guard hits, applying the rule generalized from #3183's Amendment 1: a LOOKUP, DIAGNOSTIC, ARCHIVAL or MUTATION pass wants the physical set; only "which phases belong to this milestone" wants the scoped set. Routed (10): init new-milestone phase_dir_count, init milestone-op fallback count, init manager, init progress, milestone complete stats/dry-run/archive move, phase complete's next-phase scan, state update-progress, state frontmatter stats, and uat audit's active set. Exempt with a written function-scoped reason (never a file allowlist): the audit/UAT/verification sweeps that deliberately scan every directory to report gaps, phase create/insert/rename/renumber mutations, single-phase lookups, roadmap-upgrade's cross-milestone migration, cmdPhasesClear's whole-tree destructive pass, and the reads that list a phase dir's FILES rather than enumerating the phases dir at all. Latent defects fixed by the routing: sentinel directories leaked into cmdInitNewMilestone's phase_dir_count, cmdMilestoneComplete's stats, dry-run AND ARCHIVE MOVE, cmdStateUpdateProgress, buildStateFrontmatter and cmdAuditUat's active set — every one of those hand-rolled an isDirInMilestone filter with no sentinel exclusion, so `milestone complete` was archiving backlog directories. scripts/lint-phase-enumeration-drift.cjs now reports 0 re-derivations and npm run lint:ci is green. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * docs(#3185): document milestone-scoped enumeration and record ADR Amendment 3 Changeset fragment (Changed), CLI-TOOLS/COMMANDS/USER-GUIDE updates for the scoped output of progress, stats, phases list, phases clear and milestone complete, the CONTEXT.md Phase Locator glossary entry naming listMilestonePhaseDirs, and ADR-3180 Amendment 3. Amendment 3 records: the SCOPE contract held unchanged; the declared deviation from Decision 1's provisional signature (the window needs cwd/ws, which the locked roadmapContent parameter cannot supply); the copy count being a lower bound for the third consecutive phase (4 scoped vs 54 found); the load-bearing finding that the sentinel exclusion sat on the heading set and was unreachable under the pass-all degrade; the two destructive-path defects; and the generalized exemption rule. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * fix(#3185): wire scope to consumers; revert two wrong routings the suite caught Review + remote runner findings, all fixed: The three consumers computed the enumeration scope and threw it away, so TRUNCATED/UNSCOPED/UNREADABLE collapsed into the same output as COMPLETE -- reproducing this epic's own output-identical-failure defect one layer up. progress, stats and phases list now emit phase_scope (null on the phases list --phase lookup path, which performs no enumeration). Two routings were wrong and the suite proved it: roadmap-parser's two milestone phase-count scans are reverted to the 999-only literal. isSentinelPhaseId is BROADER than what it replaced: its legacy branch runs /^0*(\d+)/ over "00.1", which backtracks to capture 0, so it read #2554's decimal phase ids as sentinel milestone 0 and stopped counting them. state.cts phaseInventoryProvider is reverted to the physical disk scan. `state rebuild` is a RECONCILIATION pass -- scoping it made it throw on healthy trees whose fixture resolves no window, swallowed the raw readdirSync fault message #3057 B1 requires verbatim, and stopped it dropping orphan STATE.md rows, which is the job. Both are now function-scoped guard exemptions with written reasons, not silent reverts. This is the consolidation trap named in the epic: a canonical rule can cover MORE than the copy it replaces, and only real inputs show it. Adds phases list coverage, a scope-branch test, and a drift-guard unit suite; backports comment-awareness to the milestone-window and plan-count guards so all three siblings share one false-positive profile; names #3161 alongside #3167 in Amendment 3's subsumption record. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * fix(#3185): correct isSentinelPhaseId's decimal-zero misclassification An isolated security review caught this branch committing the epic's own sin: the over-broad predicate was worked around at ONE call site and left live at the destructive ones. isSentinelPhaseId's legacy branch ran /^0*(\d+)/, which backtracks so any id whose leading digit run is all zeros before a non-digit captures 0 -- "0.1", "00.1" and "0.2554" all read as sentinel milestone 0. Two pinned contracts disagree with that: #2554 requires "00.1" to be counted as a real phase, and the 999 icebox is a whole reserved milestone so "999.1" must stay sentinel. The rule is asymmetric and now says so explicitly: 999 is sentinel with or without a decimal part; 0 is sentinel only when bare. A decimal phase under either is a real phase for 0 and reserved for 999, because 999 reserves a MILESTONE while 0 reserves a PHASE. Fixing the owner lets the earlier workaround go: getMilestonePhaseFilter's two scans route through isSentinelPhaseId again and the guard exemption that existed only to accommodate the defect is deleted. The state.cts cmdStateRebuild exemption stays -- that one is a genuine reconciliation-wants-the-physical-set case. Also corrects tests/adr-612-bracket-grammar.test.cjs, which asserted isSentinelPhaseId('0.1') === true and so had encoded the defect as expected behavior. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * fix(#3185): keep isSentinelPhaseId's semantics — 0.x is layered, not wrong Reverts the previous commit. The remote suite failed six tests proving it wrong, and the reason is the sharpest finding of this phase. An isolated security review observed that isSentinelPhaseId reads 0.1 and 00.1 as sentinel milestone 0 and judged that a defect against #2554. Correcting the canonical predicate broke #2949. Both contracts are pinned and both are right, because they ask different questions: #2554 is this dir part of the current milestone's phase SET? -> count 00.1 #2949 must this phase COMPLETE before the milestone closes? -> 0.x sentinel No single global predicate answers both. isSentinelPhaseId keeps its semantics (0.x IS a sentinel, #2949), and the milestone-window layer keeps a narrower 999-only rule (#2554) as a function-scoped guard exemption with a written reason — not a second silent copy. That corrects how Decision 1 reads: "one owner per derivation" governs who computes an answer, not how many questions share it. An over-broad canonical rule is as much a defect as a divergent copy and fails worse, because it looks like consolidation. Recorded in Amendment 3 as the lesson for Phases 4 and 5. Where a review's inference about intent conflicts with a pinned contract, the pinned contract wins; the finding is adjudicated, not fixed. The boundary tables in the enumeration suite are corrected to assert 0.x IS a sentinel, with the layering explained. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * chore(#3185): set changeset fragment pr to 3222 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
313 lines
16 KiB
JavaScript
313 lines
16 KiB
JavaScript
#!/usr/bin/env node
|
|
'use strict';
|
|
|
|
/**
|
|
* Anti-divergence drift guard for the live-plan-counting seam
|
|
* (epic #3180, issue #3183, ADR-3180 "Planning Semantic Model Single Owner").
|
|
*
|
|
* `src/plan-scan.cts`'s `scanPhasePlans` is the SINGLE canonical owner of
|
|
* live-plan/summary counting: which files on disk are a "plan", which are a
|
|
* "summary", and how the two pair up. Every other module that reads a phase
|
|
* directory and re-derives that filename grammar itself — `readdirSync(...)`
|
|
* filtered by an inline `-PLAN.md` / `PLAN.md` / `-SUMMARY.md` / `SUMMARY.md`
|
|
* pattern — is a re-derivation that can silently drift from the owner (the
|
|
* exact failure class this epic removes; see #2349, #1988).
|
|
*
|
|
* Per ADR-3180 Decision 4(a) this guard discovers call sites by SCANNING THE
|
|
* WHOLE `src/` TREE, not by consulting an allowlist of known files — an
|
|
* allowlist only measures re-derivations in files someone remembered to
|
|
* list, and a new call site added anywhere else would sail through silently.
|
|
*
|
|
* Detection is intentionally NARROW and mirrors the existing
|
|
* `lint-phase-id-drift.cjs` precedent: a small, readable per-line regex pair
|
|
* over authored TypeScript source, with a short, explicitly-named exemption
|
|
* list — not a general-purpose AST/control-flow analysis. A line counts as a
|
|
* re-derivation when it contains BOTH:
|
|
* (a) a filename-TEST operation — `.filter(`, `.test(`, `.match(`,
|
|
* `.exec(`, `.endsWith(`, `.startsWith(`, `.includes(`, `.some(`,
|
|
* `.every(`, or `===` — and
|
|
* (b) a plan/summary filename-suffix pattern, either a quoted literal
|
|
* ('-PLAN.md', 'PLAN.md', '-SUMMARY.md', 'SUMMARY.md') OR an unquoted
|
|
* regex literal that mentions PLAN or SUMMARY and `\.md` together
|
|
* (`/-PLAN\.md$/`, `/^PLAN-\d+.*\.md$/i`)
|
|
* on the same source line. #3183 originally required (a) to be specifically
|
|
* `.filter(` on the SAME line as the literal — that missed a regex-literal
|
|
* predicate (`files.filter(f => /-PLAN\.md$/.test(f))`, no quotes) and a
|
|
* predicate defined on one line and consumed by `.filter(` on another
|
|
* (`const isPlan = f => f.endsWith('-PLAN.md'); … files.filter(isPlan)`).
|
|
* Widening (a) to any filename-test operator — not just `.filter(` itself —
|
|
* catches both: the predicate's OWN line already carries a qualifying test
|
|
* operation (`.test(`/`.endsWith(`) alongside the literal, independent of
|
|
* where `.filter(` ends up.
|
|
*
|
|
* KNOWN, ACCEPTED limits of a per-line textual scan (same tradeoff the
|
|
* phase-id-drift guard documents): a re-derivation that filters via a
|
|
* hand-rolled loop with none of the listed test operators (e.g. a manual
|
|
* character-index scan), or one whose literal and test operator are split
|
|
* across two DIFFERENT lines with no single line carrying both, is not
|
|
* caught by this narrow shape. That is left to code review, not this regex.
|
|
*
|
|
* The tree-walk / root-confinement / regex-literal-tokenizer / sanitizer
|
|
* machinery below is SHARED with `scripts/lint-milestone-window-drift.cjs`
|
|
* (#3184) via `scripts/lib/drift-scan.cjs` — see that module for the
|
|
* `isInsideRoot` case-sensitivity note, the `walk` symlink-confinement
|
|
* rationale, and the `readRegexLiteralAt` tokenizer's ReDoS-avoidance
|
|
* rationale. It is deliberately NOT duplicated here a second time (ADR-3180
|
|
* Decision 4's own "Rejected" list: "let the new drift guard copy Phase 1's
|
|
* tree-walk / root-confinement / sanitizer").
|
|
*
|
|
* A regex literal longer than MAX_REGEX_LITERAL_LEN (400) characters is not
|
|
* read, and is therefore not caught. That bound is what keeps the scan
|
|
* linear; no real plan/summary filename filter approaches it. The scan is
|
|
* scoped to SCAN_DIRS (`src`) with SCAN_EXT (.cts/.ts/.mts) — 186 files and
|
|
* 4,299 lines matching FILENAME_TEST_RE within SCAN_DIRS/SCAN_EXT as of this
|
|
* commit, 43 of them holding 7 or more backslashes and one (`src/milestone.cts`)
|
|
* holding 16. (Definition used, so this is reproducible: walk SCAN_DIRS
|
|
* filtering by SCAN_EXT exactly as `walk` does, split each file on `\n`, and
|
|
* count every line for which the exported `FILENAME_TEST_RE.test(line)` is
|
|
* true — independent of whether a PLAN/SUMMARY literal is also present on
|
|
* that line.) Those are the lines the old backtracking detector had to
|
|
* survive, and the reason the detector is now a tokenizer.
|
|
*/
|
|
|
|
const path = require('node:path');
|
|
const driftScan = require('./lib/drift-scan.cjs');
|
|
const { readRegexLiteralAt, MAX_REGEX_LITERAL_LEN, isInsideRoot, sanitizeForReport, scanTree } = driftScan;
|
|
|
|
// A `.filter(` call on the line — the shape every current re-derivation uses
|
|
// to turn a directory listing into a plan-or-summary subset. Kept as its own
|
|
// export for back-compat / documentation; FILENAME_TEST_RE below is the
|
|
// broadened detector actually used (any filename-test operator, not just
|
|
// `.filter(`).
|
|
const FILTER_CALL_RE = /\.filter\(/;
|
|
|
|
// A filename-TEST operation: `.filter(`, `.test(`, `.match(`, `.exec(`,
|
|
// `.endsWith(`, `.startsWith(`, `.includes(`, `.some(`, `.every(`, or a
|
|
// strict-equality comparison. Any one of these on a line asking "is this
|
|
// filename a plan/summary" is a re-derivation, independent of whether the
|
|
// literal shows up as a `.filter(...)` predicate specifically.
|
|
const FILENAME_TEST_RE = /\.(?:filter|test|match|exec|endsWith|startsWith|includes|some|every)\(|===/;
|
|
|
|
// A quoted plan/summary filename-suffix literal: 'PLAN.md', '-PLAN.md',
|
|
// 'SUMMARY.md', or '-SUMMARY.md', single- or double-quoted (opening and
|
|
// closing quote must match).
|
|
const PLAN_SUMMARY_LITERAL_RE = /(['"])-?(?:PLAN|SUMMARY)\.md\1/;
|
|
|
|
// The two tokens that, appearing together INSIDE one regex literal, make it a
|
|
// plan/summary filename filter. `\.md` is matched as literal source text, not
|
|
// as a pattern, so there is nothing here to backtrack.
|
|
const PLAN_SUMMARY_TOKEN_RE = /PLAN|SUMMARY/i;
|
|
const ESCAPED_MD_TOKEN = '\\.md';
|
|
|
|
// Authored TypeScript source only (the generated bin/lib/*.cjs mirror it).
|
|
const SCAN_DIRS = ['src'];
|
|
const SCAN_EXT = new Set(['.cts', '.ts', '.mts']);
|
|
|
|
// The canonical owner defines the grammar; it is exempt by construction.
|
|
const OWNER_FILE = path.join('src', 'plan-scan.cts');
|
|
|
|
// core-utils.cts's canonical pairing rule (#1988/#2648): these three
|
|
// functions build/match `*-SUMMARY.md` CANDIDATE strings for a given plan —
|
|
// that IS the single pairing rule, not a re-derivation of it. Scoped to just
|
|
// these functions (not the whole file) so an unrelated re-derivation added
|
|
// elsewhere in core-utils.cts is still caught.
|
|
const CORE_UTILS_FILE = path.join('src', 'core-utils.cts');
|
|
const CORE_UTILS_EXEMPT_FUNCTIONS = new Set([
|
|
'summaryCandidates',
|
|
'countMatchedSummaries',
|
|
'findUnsummarizedPlans',
|
|
'findOrphanSummaries',
|
|
]);
|
|
|
|
// Per ADR-3180 Decision 4(a): NOT a bare file allowlist — each entry below is
|
|
// scoped to the SPECIFIC function asking a documented, different question
|
|
// (see the inline comment at each site), so an unrelated re-derivation added
|
|
// anywhere else in these same files is still caught. Mirrors the
|
|
// CORE_UTILS_EXEMPT_FUNCTIONS mechanism above, generalized per-file.
|
|
//
|
|
// - audit.cts scanQuickTasks: scans a quick task's OWN directory
|
|
// (`.planning/quick/<task>/`) for that ONE task's completion record —
|
|
// not a phase directory's live-plan/summary counting question.
|
|
// - gsd2-import.cts readTasksDir: reads a FOREIGN GSD-2 legacy project's
|
|
// `tasks/` dir convention during a one-time import, not this project's
|
|
// `.planning/phases/` layout at all.
|
|
// - estimate-cli.cts collectCalibrationSamples: pairs a PLAN.md and a
|
|
// SUMMARY.md by their identical `<stem>` to build an estimation
|
|
// CALIBRATION sample (projected vs. actual token counts) — a stem-keyed
|
|
// join for a statistics question, not a live-plan/completion count.
|
|
// It intentionally does NOT use the canonical three-candidate pairing
|
|
// rule (marker-swap / `-SUMMARY.md` / extended) or exclude superseded
|
|
// plans — an unmatched or superseded plan simply yields no sample,
|
|
// which is correct for calibration, not a live-completion determination.
|
|
// - roadmap.cts cmdRoadmapAnnotateDependencies: matches a plan-ID token
|
|
// out of an ALREADY-RENDERED ROADMAP.md checklist LINE OF TEXT
|
|
// (`- [ ] 01-01-PLAN.md — …`), not a filesystem directory listing — it
|
|
// can never diverge from scanPhasePlans's file-existence rule because it
|
|
// never tests file existence at all.
|
|
// - worktree-safety.cts defaultFindSummaryFiles: a recursive walk of the
|
|
// ENTIRE `.planning/` tree (not a single phase directory) for a
|
|
// pre-merge rescue of any `*SUMMARY.md` artifact, deliberately mirroring
|
|
// the shell fallback's own `find … -name "*SUMMARY.md"` glob (quick.md,
|
|
// #2296/#2070/#2838) byte-for-behaviour rather than the phase-scoped
|
|
// plan-scan owner's root+nested rule — "rescue every summary before a
|
|
// merge blows it away" is not a live-plan/completion count.
|
|
// - verify.cts cmdValidateConsistency: the strict `-(\d{2})-PLAN\.md$`
|
|
// match extracts a zero-padded SEQUENCE NUMBER from filenames the owner
|
|
// (`allPlanFiles`) already classified as plans — it does not re-derive
|
|
// "is this a plan", it answers a different, narrower question (does the
|
|
// canonical 2-digit numbering sequence have a gap) that the owner's
|
|
// boolean plan/summary classification cannot answer. See the extended
|
|
// inline comment at that call site for the full Question 1/2/3 split.
|
|
const FUNCTION_SCOPED_EXEMPTIONS = new Map([
|
|
[CORE_UTILS_FILE, CORE_UTILS_EXEMPT_FUNCTIONS],
|
|
[path.join('src', 'audit.cts'), new Set(['scanQuickTasks'])],
|
|
[path.join('src', 'gsd2-import.cts'), new Set(['readTasksDir'])],
|
|
[path.join('src', 'estimate-cli.cts'), new Set(['collectCalibrationSamples'])],
|
|
[path.join('src', 'roadmap.cts'), new Set(['cmdRoadmapAnnotateDependencies'])],
|
|
[path.join('src', 'worktree-safety.cts'), new Set(['defaultFindSummaryFiles'])],
|
|
[path.join('src', 'verify.cts'), new Set(['cmdValidateConsistency'])],
|
|
]);
|
|
|
|
// Optional `export ` modifier: `collectCalibrationSamples` (estimate-cli.cts)
|
|
// is declared `export function …` rather than a bare `function …`, and the
|
|
// function-boundary tracker below must still recognize it for its
|
|
// FUNCTION_SCOPED_EXEMPTIONS entry above to take effect.
|
|
const TOP_LEVEL_FUNCTION_RE = /^(?:export\s+)?function\s+([A-Za-z0-9_]+)\s*\(/;
|
|
|
|
/**
|
|
* The regex literal on `line` that mentions PLAN or SUMMARY together with an
|
|
* escaped `.md` suffix — e.g. `/-PLAN\.md$/`, `/^PLAN-\d+.*\.md$/i`,
|
|
* `/-SUMMARY-\d+.*\.md$/i` — or null if there is none. Replaces the former
|
|
* `REGEX_LITERAL_MD_RE`, which was both exponentially/cubically backtracking
|
|
* (CodeQL js/redos; this guard runs in `lint:ci` on fork PRs) and unable to
|
|
* see a `[\\/]` character class.
|
|
*/
|
|
function findRegexLiteralMdMatch(line) {
|
|
for (let i = 0; i < line.length; i++) {
|
|
if (line[i] !== '/') continue;
|
|
const literal = readRegexLiteralAt(line, i);
|
|
if (!literal) continue;
|
|
// Case-insensitive `\.md` test — the regex literal this replaced carried
|
|
// the `i` flag, so `\.MD`/`\.Md` must still match. A lowercased-copy
|
|
// `.includes()` preserves that behaviour without reintroducing a
|
|
// backtracking regex.
|
|
if (PLAN_SUMMARY_TOKEN_RE.test(literal.text) && literal.text.toLowerCase().includes(ESCAPED_MD_TOKEN)) {
|
|
return literal.text;
|
|
}
|
|
}
|
|
return null;
|
|
}
|
|
|
|
/**
|
|
* Strip comment text from a line before detection. A guard that fires on a
|
|
* COMMENT — including a comment documenting that the code below uses the
|
|
* canonical owner, or prose quoting this guard's own detector shapes — reports
|
|
* prose as drift and trains readers to add exemptions for documentation.
|
|
* Handles the three shapes that appear in this codebase: a whole-line
|
|
* block-comment continuation (`*` or `/*` leading), a `//` line comment, and
|
|
* a trailing `//` after code. Mirrors `lint-phase-enumeration-drift.cjs`'s
|
|
* own copy (not shared — each guard applies it at a slightly different point
|
|
* in its detection pipeline).
|
|
*
|
|
* Deliberately simple and conservative: it does not attempt full block-comment
|
|
* state tracking across lines (this is a per-line scan, same tradeoff the
|
|
* sibling guards document). A `//` inside a string literal would be stripped
|
|
* early — accepted, because the effect is to UNDER-report on a pathological
|
|
* line, never to over-report prose as drift.
|
|
*/
|
|
function stripComments(line) {
|
|
const trimmed = line.trim();
|
|
// Whole-line block comment or JSDoc continuation.
|
|
if (trimmed.startsWith('*') || trimmed.startsWith('/*') || trimmed.startsWith('//')) return '';
|
|
// Trailing line comment after code.
|
|
const idx = line.indexOf('//');
|
|
return idx === -1 ? line : line.slice(0, idx);
|
|
}
|
|
|
|
/**
|
|
* Pure: find every unsanctioned plan/summary-filter re-derivation in `text`.
|
|
* `relPath` is the repo-relative path, used both to report file:line and to
|
|
* apply the narrow, function-scoped core-utils.cts exemption.
|
|
* Returns [{ line, found }].
|
|
*/
|
|
function findPlanCountDrift(text, relPath) {
|
|
const out = [];
|
|
const lines = text.split('\n');
|
|
const exemptFunctions = FUNCTION_SCOPED_EXEMPTIONS.get(relPath) || null;
|
|
let currentFunction = null;
|
|
for (let i = 0; i < lines.length; i++) {
|
|
const line = lines[i];
|
|
const fnMatch = TOP_LEVEL_FUNCTION_RE.exec(line);
|
|
if (fnMatch) currentFunction = fnMatch[1];
|
|
|
|
const code = stripComments(line);
|
|
if (!code.trim()) continue;
|
|
|
|
if (!FILENAME_TEST_RE.test(code)) continue;
|
|
const quoted = PLAN_SUMMARY_LITERAL_RE.exec(code);
|
|
const found = quoted ? quoted[0] : findRegexLiteralMdMatch(code);
|
|
if (!found) continue;
|
|
|
|
if (exemptFunctions && exemptFunctions.has(currentFunction)) continue;
|
|
|
|
out.push({ line: i + 1, found });
|
|
}
|
|
return out;
|
|
}
|
|
|
|
/**
|
|
* Scan the authored source tree and return every unsanctioned re-derivation,
|
|
* each annotated with the repo-relative file path.
|
|
*/
|
|
function scanRepo(root) {
|
|
return scanTree({
|
|
root,
|
|
scanDirs: SCAN_DIRS,
|
|
scanExt: SCAN_EXT,
|
|
onFile(rel, text) {
|
|
// `rel` is already the REAL (canonical) path (scanTree resolves
|
|
// symlinks before calling onFile), so this comparison — and
|
|
// FUNCTION_SCOPED_EXEMPTIONS above, also keyed on `rel` — match
|
|
// consistently regardless of which symlink reached the file.
|
|
if (rel === OWNER_FILE) return [];
|
|
return findPlanCountDrift(text, rel).map((d) => ({ file: rel, ...d }));
|
|
},
|
|
});
|
|
}
|
|
|
|
function main() {
|
|
const root = path.join(__dirname, '..');
|
|
const violations = scanRepo(root);
|
|
if (violations.length === 0) {
|
|
process.stdout.write('ok plan-count-drift: no unsanctioned plan/summary re-derivations outside plan-scan.cts\n');
|
|
return;
|
|
}
|
|
process.stderr.write('plan-count-drift: independent re-derivation(s) of plan/summary filename filtering found.\n');
|
|
process.stderr.write('Use src/plan-scan.cjs `scanPhasePlans` (or core-utils.cjs `getPhaseFileStats`, which now\n');
|
|
process.stderr.write('sources plans/summaries from it) instead of re-deriving the -PLAN.md/-SUMMARY.md filter:\n');
|
|
for (const d of violations) {
|
|
// `d.file` is exactly as attacker-controlled as `d.found`: a repo can
|
|
// legally track a filename containing control bytes / bidi overrides,
|
|
// and it is a fork-PR-authored value reaching a CI log the same way the
|
|
// matched literal does — sanitize it at the same reporting boundary.
|
|
process.stderr.write(` ${sanitizeForReport(d.file)}:${d.line} ${sanitizeForReport(d.found)}\n`);
|
|
}
|
|
process.exitCode = 1;
|
|
}
|
|
|
|
if (require.main === module) main();
|
|
|
|
module.exports = {
|
|
findPlanCountDrift,
|
|
scanRepo,
|
|
FILTER_CALL_RE,
|
|
FILENAME_TEST_RE,
|
|
PLAN_SUMMARY_LITERAL_RE,
|
|
findRegexLiteralMdMatch,
|
|
readRegexLiteralAt,
|
|
MAX_REGEX_LITERAL_LEN,
|
|
isInsideRoot,
|
|
sanitizeForReport,
|
|
stripComments,
|
|
};
|