Files
msd-core/scripts/gen-adr-index.cjs
Tom Boucher bbdf7e8e84 chore(#4654): add local/no-unconfined-path-join and drain it to zero — Phase 4 of #4636 (#4674)
* chore(#4654): add local/no-unconfined-path-join and drain it to zero

Phase 4 of epic #4636 — the ratchet, and the phase that makes the epic hold.

THE MEASUREMENT THAT RESHAPED THE PHASE. An AST census (the repo's own parser,
not grep) found what the epic never enumerated: ADR-4650 named seven containment
implementations; `src/` alone held roughly 24 more hand-rolled gates across ~13
files, several guarding a write or an `fs.rmSync`. Two verified by reading rather
than pattern-matching — `research-store.cts` comments its own as "ensure the
resolved file path stays inside the store dir" immediately before a write, and
`capability-lifecycle.cts` gates `fs.rmSync` with one.

So the epic's Done-when "one containment predicate, used at every site" was FALSE
when Phase 3 reported it satisfied. It is true now: the rule is clean across
src/, scripts/, gsd-core/bin/ and hooks/ with an EMPTY allowlist.

WHY NOT THE RULE THE ISSUE PROPOSED. #4654 proposed flagging `path.join` whose
first argument is a managed root and whose later arguments derive from argv. That
is a taint analysis over 2046 call sites, in ESLint, without type information;
"derives from argv" is not locally decidable. Any approximation either floods or
is trivially evaded, and a rule that fires on hundreds of correct sites earns an
allowlist of hundreds — the opposite of a ratchet. What is actually duplicated is
the COMPARISON, not the join, and that has one recognizable shape.

  Arm 1  X.startsWith(Y + sep)            the hand-rolled containment idiom
  Arm 2  a containment predicate called as a bare statement, answer discarded

Arm 2 is the issue's "asserts the result was narrowed, not merely that a helper
was called". Its example `validatePath(x, root).resolved` is already
structurally impossible — Phase 3 un-exported `validatePath` — so the remaining
expressible failure is ignoring the answer, which is the defect that recurred
five times in this epic. The census found exactly one live instance
(`milestone.cts:1643`); it now returns the proven `ContainedPath` so consumers
stop re-deriving the path the comment above it was extracted to stop them
re-deriving.

The rule deliberately does NOT try to catch validate-one-path-use-another where
the answer is used but a different variable flows onward. That needs flow
analysis; the branded `ContainedPath` from Phase 3 is the defense there, and the
two are complementary.

PER-SITE FAMILY CHOICE, NOT A DEFAULT. Phase 3's lesson binds: collapsing a
lexical site onto the realpath family broke four tests and was caught only by the
matrix. Every migrated site was triaged individually. The six
installer-migrations tree-walks and the six capability-lifecycle gates take the
LEXICAL family because their operands are already realpath-resolved and they
deliberately treat the final component as a link; boundary sites take realpath.

TWO SITES WITH AN INVERTED CONTRACT, which a mechanical swap would have broken.
`installer-migrations.cts:127` and `runtime-artifact-install-plan.cts:144` REJECT
`target === root` by contract, while the canonical comparison ACCEPTS it. Swapped
naively, a migration could `rmdir` the user's config root and a third-party
descriptor could write at configHome itself. Both keep `=== root` as an explicit
additional arm alongside the predicate call — the predicate decides containment,
the call site keeps its own extra condition (ADR-4650 decision 6).

ONE DUPLICATE DELETED OUTRIGHT: `planning-inspect.cts`'s `isWithinRoot` was
byte-identical to `isContainedIn` and said so in its own docstring.
`isContainedIn` is now exported for callers that have already resolved both
operands and need only the comparison, with a doc note that a caller which has
NOT resolved them must use a full predicate instead.

THE MARKER, AND WHY IT IS NOT THE ALLOWLIST. Nine sites are justified holdouts and
carry `// allow-handrolled-containment: <reason>` with a mandatory, reviewable
reason. Two justifications: (a) not a containment decision — an ancestor-walk loop
condition, sub-repo grouping, worktree identity matching, declared-path coverage;
(b) it IS containment but the canonical predicate is unreachable —
`capability-validator.cjs` is a committed pre-build `.cjs` and the compiled
`security.cjs` is untracked build output, so requiring it would break a fresh
clone. `scripts/lib/drift-scan.cjs` runs under `lint:ci` with the same exposure.
The marker was renamed from `allow-lexical-prefix-match` mid-phase because that
name asserted only (a) and would have stated something false at the (b) sites.

A marker suppresses BEFORE the violation counter increments, so a file whose
every occurrence is marked still reports `staleAllowlistEntry` — otherwise a
drained entry lingers and silently re-permits the site later.

DEMONSTRATED RED, per #4654: a hand-rolled copy reintroduced into a real `src/`
file made `npm run lint` fail with the rule's full guidance message; removing it
returned the tree to clean. Both halves recorded — red alone proves nothing,
since a rule red for an unrelated reason looks identical.

DISCLOSED: `defaultRequireFromInstallRoot` (gsd-tools.cjs) previously carried two
distinct rejection messages and two manual realpath calls; routing it through
`tryWithinRoot` collapses them to one message, and a missing module now surfaces
as MODULE_NOT_FOUND rather than ENOENT. No test asserts either message. The
security property is preserved and slightly strengthened — the candidate is
realpathed and containment re-checked, and the dangling-symlink oracle closure
comes along with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4654): record the containment ratchet in CONTEXT.md and the security model

Both entries previously described the seam without the thing that keeps it a
seam. They now state what the rule bans, and — more usefully for whoever reads
this next — what it deliberately does NOT attempt: deciding per path.join call
whether an argument came from user input. That question is not locally
decidable, and an approximation across ~2000 join sites would earn an exemption
list of hundreds, which is the opposite of a ratchet.

Also records the marker's two legitimate justifications and that its reason is
mandatory, so the escape stays reviewable rather than becoming a mute button.

Glossary gate 270 refs exit 0; install-tree goldens and CONTEXT-INDEX.json
regenerated and confirmed byte-identical rather than assumed — which also
confirms eslint-rules/ is not a shipped path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4654): close review findings and the two matrix failures

MATRIX FAILURE 1 — a collapsed message broke a negative-proof test, and my
evidence for collapsing it was wrong. I searched tests/ for the literal string
"resolves outside its install root", found nothing, and reported that no test
asserted it. The test matches a REGEX SUBSTRING, /outside its install root/, so
the literal search missed it. What broke was "NEGATIVE PROOF: a symlinked module
pointing OUTSIDE the install root is not loaded" — the test guarding the exact
property I claimed was preserved. defaultRequireFromInstallRoot now does both
checks again with both messages byte-identical, each routed through the
canonical predicate, which is better than the original since that hand-rolled
both comparisons.

MATRIX FAILURE 2 — shipped migrations are checksum-locked, and a marker cannot
serve there. migrationChecksum hashes plan.toString(), which INCLUDES comments,
so a suppression marker inside a plan body drifts the baseline exactly as an
edit does. Measured: with markers in place, two of the four still differed from
their committed checksums. The four shipped bodies are now byte-identical to
next, and the rule's config excludes those four paths BY NAME rather than by a
directory wildcard, so a NEW migration is still covered. Six containment
comparisons stay un-ratcheted there; that gap is recorded in the rule's Known
gaps, in CONTEXT.md and in the security model rather than left implicit.
Justification (c) is removed from the marker's documented reasons, because a
marker was proven unable to express it.

ADVERSARIAL REVIEW — the sharpest finding was that the rule banned the CORRECT
shape while permitting the incorrect one: startsWith(root) with no separator is
the genuinely unsafe form, since it accepts a sibling such as root-evil, and my
own test blessed it as valid. Flagging every bare startsWith would swamp the
rule, so that stays a STATED gap rather than a silent one. Closed for real: the
template-literal spelling, which the census never saw because it only inspected
plus-concatenation — that surfaced TWELVE more sites, now triaged and migrated.
A separator reached through a const alias is now resolved via scope analysis.
And isContainedIn, exported in Phase 3, was missing from the discarded-result
set, so a bare no-op call went unflagged on the one function the epic funnels
through.

SECURITY REVIEW — the marker could over-suppress two ways: a block comment
worked identically to a line comment, and one marker silently covered every
violation sharing its line. It now requires a Line comment positioned after the
flagged node ends, so it anchors to the node it trails. Four sites had dropped
an unreachable-but-deliberate equality rejection against the root; each is
restored as the call site's own arm. eslint.config.mjs still documented the OLD
marker token, which my rename missed — it would have sent the next author in
circles.

A FALSE GREEN, recorded because it nearly stuck: lint:ci reported exit 0 from a
stale eslint cache while twelve real violations existed. Every lint check here
now clears the cache first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4654): anchor a suppression marker to the violation it actually trails

The matrix caught this; my own test caught it, on its first execution. The case
"two violations on one line: trailing marker suppresses only the one it trails"
expected 1 error and got 0 — both were suppressed.

ROOT CAUSE: the anchoring accepted any Line comment on the node's line whose
range started at or after the node's end. A trailing marker at the END of a line
sits after EVERY node on that line, so that condition held for all of them.
"After the node" does not identify WHICH node the marker trails. The fix reads
as correct and is not.

FIX: deferred reporting. Violations accumulate during traversal instead of being
reported immediately; at Program:exit each marker claims exactly ONE pending
violation — the one on its line whose end is nearest before the marker begins —
and every unclaimed violation is then counted and reported. One marker, one
suppression. An earlier violation sharing the line is still reported, which is
the property the security review asked for and the previous attempt only
appeared to deliver.

The counter now increments at flush time rather than during traversal, so a
suppressed occurrence still does not keep an allowlist entry alive.

AND A TOOL THAT SHOULD HAVE EXISTED BEFORE THE FIRST MATRIX RUN. `node --test`
is hard-blocked here, so this rule's test file could only ever be executed on
the remote matrix — which is why a broken anchoring shipped into a run. ESLint's
programmatic Linter API is not a test runner, and exercising the rule through it
verifies every case locally in seconds. All 24 now pass locally, including the
two-on-one-line case that failed remotely. That loop should have been built
before the rule was first sent to the matrix rather than after it failed twice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4654): backfill PR 4674 into the changeset and complete 70-docs.json

The phase gate requires enablementSequence and the Diataxis quadrants; 70-docs
now carries both, with the how-to quadrant skipped for a stated reason rather
than an empty field. The audience for this deliverable is a contributor who
trips the rule, and the task-oriented guidance reaches them in the ESLint
message itself — which names the correct predicate, says how to choose between
the realpath and lexical families, cites the Phase 3 regression caused by
choosing wrong, and gives the marker syntax. A docs/how-to page would be a
second, driftable copy read by nobody at the moment of failure.

enablementSequence is recorded as what it actually is: a VERIFICATION sequence,
not an enablement one. The rule is never off, so there is no off-to-on
transition to describe.

scripts/lint-docs-required.cjs now passes (ok_docs_updated) — it could not
evaluate against the mandated pr:0 placeholder.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 22:17:46 -04:00

1225 lines
52 KiB
JavaScript

#!/usr/bin/env node
'use strict';
/**
* Generates the ADR index table in docs/adr/README.md from the ADR files
* themselves, and validates the corpus' lifecycle invariants.
*
* The index is a DERIVED artifact: it is regenerated from every
* `docs/adr/<id>-<slug>.md` on disk, so it cannot silently drift out of date
* the way a hand-maintained table does. CI re-runs this with `--check` and
* fails on any diff or invariant violation.
*
* Invariants enforced (see docs/adr/README.md "Lifecycle rules"):
* 1. Every ADR declares `- **Status:** <Token>` with Token in STATUSES.
* 2. A Superseded/Retired ADR names its successor as a markdown link to the
* target file — never a bare "ADR-N", which is ambiguous (ADR-0010 and
* ADR-0011 each resolve to more than one file).
* 3. Supersession is symmetric: if A supersedes B, B records superseded-by A.
* 4. An ADR whose H1 declares an id must match its filename's id.
* 5. The committed index equals the generated index.
*
* Usage:
* node scripts/gen-adr-index.cjs # print the index to stdout
* node scripts/gen-adr-index.cjs --write # rewrite the index in README.md
* node scripts/gen-adr-index.cjs --check # exit 1 if stale or invalid
* node scripts/gen-adr-index.cjs --json # --check semantics; JSON report on stdout
*/
const fs = require('node:fs');
const path = require('node:path');
const { ExitError, runMain } = require('./lib/cli-exit.cjs');
const { escapeRegex: escapeRegExp } = require('../gsd-core/bin/lib/pattern.cjs');
const { isContainedIn } = require('../gsd-core/bin/lib/security.cjs');
const ROOT = path.resolve(__dirname, '..');
const ADR_DIR = path.join(ROOT, 'docs', 'adr');
const README_PATH = path.join(ADR_DIR, 'README.md');
const START_MARKER = '<!-- ADR-INDEX:START — generated by scripts/gen-adr-index.cjs; do not edit by hand -->';
const END_MARKER = '<!-- ADR-INDEX:END -->';
/**
* The canonical status vocabulary.
*
* `Legacy` and `Retired` are deliberately distinct from `Superseded`:
* - Superseded — a specific newer ADR replaced this decision. Names it.
* - Retired — the thing this ADR decided no longer exists at all, and no
* single ADR replaced it (e.g. a deleted package boundary).
* - Legacy — frozen historical record, kept for provenance, not a
* pattern to imitate.
*
* NOTE: `Legacy` describes a DECISION's standing, not a filename. The
* `0001-`..`0012-` sequential *naming* era is legacy, but many of those ADRs
* (e.g. 0002, 0004, 0008, 0009) are Accepted and load-bearing today. Do not
* conflate the two: grep the naming rule in README.md, not this enum.
*/
const STATUSES = ['Accepted', 'Proposed', 'Superseded', 'Legacy', 'Retired'];
/**
* Stable reason codes for every lifecycle violation this gate can emit.
* Tests assert via `assert.equal(record.reason, REASON.X)` (or `.some(...)`
* over the `--json` `violations` array) rather than regex-matching stderr
* prose — see CONTRIBUTING.md "Prohibited: Raw Text Matching on Test
* Outputs" and the worked example in `bin/verify-reapply-patches.cjs`.
*
* Adding a reason is a deliberate three-part change: a new entry here, the
* emitting `add(...)` call site, and the corpus test that locks
* `Object.keys(REASON).sort()` — so a new violation class cannot ship
* without its own typed identity.
*/
const REASON = Object.freeze({
FILENAME_INVALID: 'filename_invalid',
STATUS_MISSING: 'status_missing',
STATUS_INVALID: 'status_invalid',
STATUS_BRACKET_MISMATCH: 'status_bracket_mismatch',
ID_MISMATCH: 'id_mismatch',
SUPERSEDED_NO_SUCCESSOR: 'superseded_no_successor',
SUPERSEDED_BARE_ID: 'superseded_bare_id',
RELATION_LINK_MISSING: 'relation_link_missing',
RELATION_BARE_ID_MISSING: 'relation_bare_id_missing',
RELATION_BARE_ID_UNLINKED: 'relation_bare_id_unlinked',
RELATION_ASYMMETRIC: 'relation_asymmetric',
LINK_UNRESOLVED: 'link_unresolved',
LINK_ESCAPES_REPO: 'link_escapes_repo',
LINK_ESCAPES_REPO_SYMLINK: 'link_escapes_repo_symlink',
DIRENT_UNREADABLE: 'dirent_unreadable',
DIRENT_ESCAPES_REPO_SYMLINK: 'dirent_escapes_repo_symlink',
});
/**
* The H1 trailing-bracket vocabulary, derived from `STATUSES` — not a second
* hand-written literal. Before this PR, `parseAdr`'s title strip carried its
* own copy of these five tokens, and nothing asserted the two lists agreed:
* a textbook `DEFECT.GENERATIVE-FIX` instance (a generated surface and its
* hand-authored source drifting apart with no parity check). A 6th status
* added to `STATUSES` now covers the bracket for free, and the corpus's
* parity test iterates the real exported array rather than a copy.
*/
// Escaped for defence-in-depth, not a live-bug fix: `STATUSES` is a static
// array literal today, so nothing in it can currently carry a regex
// metacharacter. But nothing enforces that it STAYS static — if a future
// change ever derives it from external input (a config file, a corpus scan),
// an unescaped `join('|')` would let a status token break out of the
// alternation it is meant to be one branch of.
const STATUS_BRACKET_RE = new RegExp(String.raw`\s*\[(${STATUSES.map(escapeRegExp).join('|')})\]\s*$`, 'i');
/**
* Header fields that assert a lifecycle relation.
*
* Two DISTINCT relations, deliberately not conflated:
*
* supersedes — the target decision is REPLACED. The target's status becomes
* Superseded and it must name this ADR. (ADR-0174 → ADR-0005.)
*
* subsumes — the target decision still HOLDS, but a broader ADR now frames
* it; the target keeps its Accepted status and becomes a component of the
* larger decision. (ADR-1239/EoS subsumes ADR-1016 "as the declarative
* adapter" — the descriptor is still real and still correct.)
*
* Both directions are symmetry-checked, but only `supersedes` implies a status
* change on the target. Collapsing subsumption into supersession would mark
* four live, load-bearing ADRs as dead — the opposite of the truth.
*/
const RELATION_FIELDS = new Map([
['supersedes', { kind: 'supersedes', dir: 'out' }],
['superseded by', { kind: 'supersedes', dir: 'in' }],
['subsumes', { kind: 'subsumes', dir: 'out' }],
['subsumed by', { kind: 'subsumes', dir: 'in' }],
]);
/** Relation kinds and the header field a reader should add to fix each gap. */
const RELATION_SPEC = {
supersedes: { out: 'Supersedes', in: 'Superseded by' },
subsumes: { out: 'Subsumes', in: 'Subsumed by' },
};
/**
* A relation field whose value opens with "nothing"/"none"/"n/a" asserts the
* absence of the relation, whatever prose follows it.
*/
const NEGATED_RELATION_RE = /^\s*(?:nothing|none|n\/a|[—–-])\s*(?:$|[;,.]|\s)/i;
/**
* Header fields appear in two shapes across the corpus, both legitimate:
* bullet — `- **Status:** Accepted`
* table — `| **Status** | Accepted |`
* Yield [field, value] for either.
*/
function* headerFields(header) {
const bullet = /^\s*[-*]\s*\*\*([^*:]+?)(?::)?\*\*\s*(.*)$/gm;
let m;
while ((m = bullet.exec(header)) !== null) yield [m[1].trim(), m[2].trim()];
const row = /^\s*\|\s*\*\*([^*|]+?)(?::)?\*\*\s*\|\s*(.*?)\s*\|\s*$/gm;
while ((m = row.exec(header)) !== null) yield [m[1].trim(), m[2].trim()];
}
/** Numeric identity of an ADR: "0011" and "11" are the same id. */
function canonicalId(raw) {
return String(raw).replace(/^0+(?=\d)/, '');
}
/** The documented filename shape: `<issue#>-<kebab-slug>.md`. */
const ADR_FILENAME_RE = /^[0-9]+-[a-z0-9-]+\.md$/;
/**
* Whole-segment containment test: true if `abs` is NOT inside `root`.
*
* The SINGLE copy of this predicate. Before this PR it was hand-written
* inline in two places (the lexical pre-stat check in `validateLinks` and the
* post-realpath escape check in `existsCaseExact`) with no shared name; a
* third copy for `markdownFilesInAdrDir`'s own symlink check would have made
* three. All three call sites now share this one function.
*
* `rel.startsWith('..')` alone would also match an in-repo path whose first
* segment merely BEGINS with two dots (`..hidden.md`), and a false "escapes
* the repository" on a valid path is a worse failure than a miss — hence the
* whole-segment `rel === '..' || rel.startsWith('..' + sep)` form.
*
* Delegates to the shared `isContainedIn` (ADR-4650): every call site here
* already resolved both operands itself (either lexically, before any
* filesystem call, or via realpathSync after a symlink) and needs only this
* comparison step — the exact already-resolved-caller case `isContainedIn`
* documents.
*/
function escapesRoot(abs, root) {
return !isContainedIn(abs, root);
}
/**
* `fs.realpathSync(ROOT)`, tolerant of an unreadable/vanished ROOT (degrades
* to the lexical ROOT itself rather than throwing — callers still get a
* usable comparison root, just without symlink-normalization on hosts where
* ROOT itself sits under a symlinked ancestor, e.g. macOS's /var -> /private/var).
*/
function realRootOrFallback() {
try {
return fs.realpathSync(ROOT);
} catch {
return ROOT;
}
}
/**
* Whether `joined` (a `*.md` dirent directly under docs/adr/) should be
* treated as an ADR file: a regular file, or a symlink that resolves to a
* regular file WITHOUT leaving the repository. Applies the same rule to which
* FILES are read as `validateLinks` already applies to which link TARGETS
* resolve — a symlink escaping the repo is never followed and its content is
* never touched (no `statSync`/`readFileSync` past the `lstatSync`/
* `realpathSync` calls below), because `parseAdr` and `validateLinks` both
* read the FULL body of every accepted file and echo fragments into stderr.
*/
function isAcceptedAdrEntry(joined, realRoot) {
let lst;
try {
lst = fs.lstatSync(joined);
} catch {
return false; // vanished / unreadable
}
if (lst.isFile()) return true;
if (!lst.isSymbolicLink()) return false;
let real;
try {
real = fs.realpathSync(joined);
} catch {
return false; // broken symlink
}
if (escapesRoot(real, realRoot)) return false; // escapes the repository
try {
return fs.statSync(real).isFile();
} catch {
return false; // vanished between realpath and stat (TOCTOU)
}
}
/**
* Every markdown file directly under docs/adr/, README.md included.
*
* Single source of the traversal rule: `adrFiles()` is this minus README.md
* (which is the index, not an ADR), and the link-resolution pass is this
* unfiltered (the generated index can point nowhere too). Two hand-copied
* readdir filters would drift the moment either grew a rule — the exact
* DEFECT.GENERATIVE-FIX shape this gate now enforces against the corpus.
*/
function markdownFilesInAdrDir() {
let entries;
try {
entries = fs.readdirSync(ADR_DIR);
} catch {
// An unreadable docs/adr/ itself degrades to "no files" here — the caller
// (validate/validateLinks) surfaces the real problem elsewhere; this
// traversal helper must never throw a raw fs error up into `runMain`,
// which would print `err.stack` (absolute host paths) to public CI logs.
entries = [];
}
const realRoot = realRootOrFallback();
return entries
.filter((f) => f.endsWith('.md'))
.filter((f) => {
try {
return isAcceptedAdrEntry(path.join(ADR_DIR, f), realRoot);
} catch {
// A `*.md` dirent that cannot be classified — most commonly a broken
// symlink — is excluded here rather than crashing the caller. It is
// not silently dropped from the gate: `validateLinks` diffs this
// filtered list against the raw `readdirSync` listing and reports
// the exclusion as its own violation, naming the file.
return false;
}
})
.sort();
}
function adrFiles() {
return markdownFilesInAdrDir().filter((f) => f !== 'README.md');
}
/**
* Split the directory into files this tool can parse and files it cannot.
*
* A file without a numeric prefix is not merely unparseable — it is invisible
* to the index, which is the failure this gate exists to prevent. Report it as
* a violation naming the convention, rather than crashing on `match(...)[1]`
* or silently skipping it.
*/
function partitionAdrFiles() {
const conforming = [];
const nonConforming = [];
for (const f of adrFiles()) (ADR_FILENAME_RE.test(f) ? conforming : nonConforming).push(f);
return { conforming, nonConforming };
}
/** Extract the leading bullet-field header block (everything before the first `##`). */
function headerBlock(text) {
const body = text.split(/\r?\n/);
const stop = body.findIndex((l) => /^##\s/.test(l));
return (stop === -1 ? body : body.slice(0, stop)).join('\n');
}
/**
* A relation may also be declared as a whole SECTION rather than a header field.
* ADR-0174 is the exemplar: a `## Supersedes` heading over a table whose first
* column links each superseded ADR and whose remaining columns explain why.
* That is the richest form in the corpus and must count — reading only the
* header block would report the repo's best-documented supersession as missing.
*
* Returns { supersedes: [file…], subsumes: [file…] } from matching sections.
*/
const RELATION_SECTION_RE = /^##\s+(Supersedes|Subsumes)\b[^\n]*$/i;
function relationSections(text) {
const lines = text.split(/\r?\n/);
const out = { supersedes: [], subsumes: [] };
for (let i = 0; i < lines.length; i++) {
const m = lines[i].match(RELATION_SECTION_RE);
if (!m) continue;
const kind = m[1].toLowerCase() === 'supersedes' ? 'supersedes' : 'subsumes';
// Collect until the next heading of any level.
let j = i + 1;
const body = [];
for (; j < lines.length && !/^#{1,6}\s/.test(lines[j]); j++) body.push(lines[j]);
const chunk = body.join('\n');
if (NEGATED_RELATION_RE.test(chunk.trim())) continue;
out[kind].push(...linkedAdrFiles(chunk));
i = j - 1;
}
return out;
}
/** All markdown links to sibling ADR files inside a chunk of text. */
function linkedAdrFiles(text) {
const out = [];
const re = /\]\(\s*(?:\.\/)?([0-9]+-[a-z0-9-]+\.md)\s*\)/gi;
let m;
while ((m = re.exec(text)) !== null) out.push(m[1]);
return out;
}
/** Bare `ADR-123` / `ADR 123` mentions that are NOT part of a markdown link. */
function bareAdrRefs(text) {
const withoutLinks = text.replace(/\[[^\]]*\]\([^)]*\)/g, '');
const out = [];
const re = /\bADR[-\s]0*(\d+)\b/gi;
let m;
while ((m = re.exec(withoutLinks)) !== null) out.push(canonicalId(m[1]));
return out;
}
/**
* Mask code (fenced blocks and inline spans) so link resolution never reads a
* `[…](…)` sequence that markdown does not render as a link. The corpus has
* two real examples of this: `mod[entry.router]({ args, cwd, raw, error })`
* inside a ``` fence, and `` `require(module)[router]()` `` inline — both
* ordinary JavaScript, neither a link.
*
* The output is the SAME LENGTH as the input, with every masked character
* replaced by a single space and every newline left untouched — so a finding
* computed against the masked text still names the correct 1-indexed line
* (Kernighan's Law: keep the debug surface honest rather than deleting text).
*/
function maskCode(text) {
// Capturing split keeps the line terminators as their own array elements
// (even indices are line content, odd indices are the terminator that
// followed), so the rebuild below never has to guess LF vs CRLF.
const parts = String(text).split(/(\r?\n)/);
// null outside a fence; otherwise the marker char ('`' or '~') and the
// length of the run that opened it — both are load-bearing for closing:
// only the SAME char with a run length >= the opener's closes the fence.
let fence = null;
for (let i = 0; i < parts.length; i += 2) {
const line = parts[i];
if (fence) {
// Whichever way this line resolves, it is code: the closing fence line
// is still a fence delimiter, not prose.
const closeRe = fence.char === '`' ? /^ {0,3}(`{3,})\s*$/ : /^ {0,3}(~{3,})\s*$/;
const close = line.match(closeRe);
parts[i] = ' '.repeat(line.length);
if (close && close[1].length >= fence.len) fence = null;
continue;
}
const open = line.match(/^ {0,3}(`{3,}|~{3,})/);
if (open) {
fence = { char: open[1][0], len: open[1].length };
parts[i] = ' '.repeat(line.length);
continue;
}
parts[i] = maskInlineCodeSpans(line);
}
return parts.join('');
}
/**
* Mask backtick-delimited inline code spans within a single line (fences are
* handled by the caller, per-line, before this runs — a span never crosses a
* newline). CommonMark's rule: a run of N backticks opens a span, closed by
* the NEXT run of exactly N backticks; a run of any other length in between
* is part of the span's content, not a delimiter. An opening run with no
* matching close is literal text, not a span.
*
* LINEAR, not the naive per-opener rescan this replaced: the old
* implementation, for every backtick run, rescanned the entire remainder of
* the line looking for a same-length closer. A line of strictly-ascending-
* length backtick runs (nothing ever closes) forced a near-full rescan per
* run — measured ~O(n^1.6) and unbounded (34ms/50KB -> 220ms/200KB ->
* 1.76s/800KB on adversarial input). This version scans the line ONCE to
* collect every backtick run as `{start, end, len}`, then walks that run
* list left to right with a per-length cursor (`byLen`/`cursor` below) that
* only ever advances forward — so finding "the next run of equal length" is
* amortized O(1) per step and the whole pass is O(line length).
*
* Behavior is identical to the rescan version for every input: once an
* opener at run `r` is paired with the next same-length run `r'`, every run
* strictly between them is consumed as span content and is never
* reconsidered as its own delimiter — exactly what the old code did by
* jumping `i` straight to the close and never revisiting the interior.
*/
function maskInlineCodeSpans(line) {
const runs = [];
let i = 0;
while (i < line.length) {
if (line[i] !== '`') {
i += 1;
continue;
}
const start = i;
while (i < line.length && line[i] === '`') i += 1;
runs.push({ start, end: i, len: i - start });
}
if (runs.length === 0) return line;
// Every run's index, grouped by length, in left-to-right order (already
// sorted — `runs` was built in scan order).
const byLen = new Map();
for (let idx = 0; idx < runs.length; idx += 1) {
const len = runs[idx].len;
if (!byLen.has(len)) byLen.set(len, []);
byLen.get(len).push(idx);
}
const cursor = new Map(); // len -> next unexamined index into byLen.get(len)
const spans = []; // [start, end) ranges to mask, in order, non-overlapping
let r = 0;
while (r < runs.length) {
const len = runs[r].len;
const candidates = byLen.get(len);
let c = cursor.get(len) || 0;
// Skip past any candidate at or before `r`: `r` itself, or an index
// already consumed as interior content of an earlier matched span (a run
// inside a completed span is never revisited — same as the rescan
// version never re-examining a delimiter it has already masked over).
while (c < candidates.length && candidates[c] <= r) c += 1;
if (c < candidates.length) {
const closeIdx = candidates[c];
spans.push([runs[r].start, runs[closeIdx].end]);
cursor.set(len, c + 1);
r = closeIdx + 1;
} else {
cursor.set(len, c);
r += 1; // no closer of equal length anywhere ahead — literal text
}
}
let out = '';
let pos = 0;
for (const [s, e] of spans) {
out += line.slice(pos, s) + ' '.repeat(e - s);
pos = e;
}
out += line.slice(pos);
return out;
}
/**
* Every inline `[text](dest)` / `![alt](dest)` link or image in `text`, with
* code masked out first so a code-shaped bracket/paren sequence is never
* misread as a link (see `maskCode`).
*
* `[^\][\n]*` for the link-text class deliberately excludes BOTH bracket
* characters, not just `]` — so `[see [1]](x.md)` does not match (nested
* brackets are out of the inline-links-only scope this gate supports) and a
* regex character class in prose like `[A-Z][A-Z0-9_]` cannot be misread as
* one either. Reference-style links (`[text][ref]`) are correspondingly not
* supported: the corpus has zero reference definitions to resolve against.
*
* Returns `{ line, target }` per match — `line` is 1-indexed, `target` is the
* RAW parenthesized capture, untrimmed and unresolved; callers normalize.
*/
function extractLinks(text) {
const masked = maskCode(String(text));
const lines = masked.split(/\r?\n/);
const out = [];
const re = /!?\[[^\][\n]*\]\(([^()\n]*)\)/g;
for (let i = 0; i < lines.length; i += 1) {
re.lastIndex = 0;
let m;
while ((m = re.exec(lines[i])) !== null) {
out.push({ line: i + 1, target: m[1] });
}
}
return out;
}
/**
* Case-exact existence of `abs` (which MUST already be verified inside ROOT
* by the caller — LEXICALLY, via `path.relative`). Walks each path segment
* against a cached, real `readdirSync` listing of its parent rather than
* calling `fs.existsSync(abs)` directly: existsSync resolves through the OS's
* case-folding rules, which pass on macOS/Windows for a link that 404s on
* github.com and reds the Linux CI lane — every platform must agree, so
* resolution never trusts the filesystem's own case sensitivity (or lack
* of it).
*
* SYMLINK ESCAPE (the reason this function is more than a readdir loop): the
* caller's containment check is purely lexical string math on `abs` — it
* proves nothing about what is actually ON DISK at each segment. But
* `readdirSync` FOLLOWS symlinks at the OS level while walking further down a
* path. A contributor can commit `docs/adr/x -> /etc` (a symlink; Linux CI
* lanes, including fork PRs, preserve symlinks) plus an ADR linking
* `[t](x/passwd)`: the caller's lexical check sees `docs/adr/x/passwd`, which
* LOOKS repo-internal, and this walk would then list the real external
* directory — and the "Did you mean X?" hint below is built from exactly that
* listing, so a wrong-case probe (`[t](x/PASSWD)`) would echo a real filename
* from OUTSIDE the repo into PUBLIC CI LOGS on a fork PR. So every segment is
* lstat'd, and a symlink is realpath'd and re-checked against the REAL root,
* BEFORE this walk ever descends into or reads what it points at.
*
* `dirCache` is a Map<directory, Set<entryName>|null> (null = unreadable),
* built and owned by the caller so repeated links into the same directory
* cost one `readdirSync` total, not one per link.
*
* `realRoot` is `fs.realpathSync(ROOT)`, computed ONCE by the caller (never
* per-segment/per-link here) and passed in — ROOT itself may sit under a
* symlinked path (macOS's `/var` -> `/private/var`), so comparing a
* realpath'd descendant against a non-realpath'd ROOT would misclassify every
* legitimate path on such a host as an escape.
*
* Returns `{ exists, hint, escaped }`. `escaped: true` means a symlink
* resolved outside `realRoot`; in that case `exists` is `false` and `hint` is
* ALWAYS `null` — the caller must report a distinct "escapes" message and
* never fall back to the generic "does not resolve" wording or a hint, both
* of which would leak into the escape's own disclosure hazard.
*/
function existsCaseExact(abs, dirCache, realRoot) {
const rel = path.relative(ROOT, abs);
if (rel === '') return { exists: true, hint: null, escaped: false }; // ROOT itself
const segments = rel.split(path.sep);
let dir = ROOT;
for (const seg of segments) {
let entries = dirCache.get(dir);
if (entries === undefined) {
try {
entries = new Set(fs.readdirSync(dir));
} catch {
entries = null;
}
dirCache.set(dir, entries);
}
if (!entries || !entries.has(seg)) {
const hint = entries ? [...entries].find((e) => e.toLowerCase() === seg.toLowerCase()) : null;
return { exists: false, hint: hint || null, escaped: false };
}
const joined = path.join(dir, seg);
// `entries.has(seg)` above proved a directory ENTRY named `seg` exists —
// it says nothing about what that entry IS. Check before descending.
let lst;
try {
lst = fs.lstatSync(joined);
} catch {
// Vanished between readdir and lstat (TOCTOU race) — degrade to "does
// not resolve", never throw.
return { exists: false, hint: null, escaped: false };
}
if (lst.isSymbolicLink()) {
let real;
try {
real = fs.realpathSync(joined);
} catch {
// Broken symlink — degrade to "does not resolve", never throw.
return { exists: false, hint: null, escaped: false };
}
if (escapesRoot(real, realRoot)) {
// No further readdirSync down this path, and no hint: both would
// disclose facts about a directory outside the repo.
return { exists: false, hint: null, escaped: true };
}
dir = real; // resolves inside the repo — continue the walk from there.
continue;
}
dir = joined;
}
return { exists: true, hint: null, escaped: false };
}
/**
* The link-resolution pass: every inline link/image target in every `*.md`
* file directly under `docs/adr/` — INCLUDING README.md (the generated index
* can point nowhere too) and files that fail the naming convention (their
* naming violation is reported separately by `partitionAdrFiles`, but a
* reader still follows their links). Non-recursive, matching `adrFiles()`.
*
* Errors are reported through the same `add(file, msg)` channel `validate`
* uses elsewhere, keeping the `${file}: ${msg}` prefix uniform — but the
* "file" half of that prefix is `${file}:${line}` here, so the emitted line
* reads `<file>:<line>: <prose>` (a literal colon immediately before the line
* number, compiler-diagnostic style) rather than `<file>: <line>: <prose>`.
*/
function validateLinks(add) {
const files = markdownFilesInAdrDir();
// Computed ONCE per pass, never per-link/per-dirent: see existsCaseExact's
// doc comment for why comparing against the REAL root (not the lexical
// ROOT constant) is required to avoid false escapes when the repo checkout
// itself sits under a symlinked ancestor (e.g. macOS's /var -> /private/var).
const realRoot = realRootOrFallback();
// Report any `*.md` dirent that `markdownFilesInAdrDir` silently excluded —
// because it could not be stat'd (e.g. a broken symlink) OR because it IS a
// symlink that resolves outside the repository — so it surfaces as a gate
// finding instead of quietly vanishing from the index. Reading the
// directory again here (rather than threading a second return value
// through `markdownFilesInAdrDir`) keeps that function's contract
// (`string[]`) simple for its other callers. Wrapped in try/catch for the
// same reason as inside `markdownFilesInAdrDir`: an unreadable ADR_DIR
// degrades to "nothing more to report" here, never a crash.
let dirents;
try {
dirents = fs.readdirSync(ADR_DIR);
} catch {
dirents = [];
}
const included = new Set(files);
const BROKEN_MSG = 'could not be read (broken symlink?) and was excluded from the index. Remove it or fix its target.';
for (const f of dirents) {
if (!f.endsWith('.md') || included.has(f)) continue;
const joined = path.join(ADR_DIR, f);
let lst;
try {
lst = fs.lstatSync(joined);
} catch {
add(f, BROKEN_MSG, { reason: REASON.DIRENT_UNREADABLE, line: null });
continue;
}
if (!lst.isSymbolicLink()) {
// Not a symlink and still excluded — some other legitimate reason
// (e.g. it's a directory literally named `*.md`), not unreadable.
continue;
}
let real;
try {
real = fs.realpathSync(joined);
} catch {
add(f, BROKEN_MSG, { reason: REASON.DIRENT_UNREADABLE, line: null });
continue;
}
if (escapesRoot(real, realRoot)) {
// Distinct message from the broken-symlink one above, and — same
// discipline as the link-target escape below — no path or hint from
// outside the repo is ever included: only the in-repo dirent name.
add(
f,
'is a symlink that escapes the repository and was excluded from the index. Point it at a file inside docs/adr/, or remove it.',
{ reason: REASON.DIRENT_ESCAPES_REPO_SYMLINK, line: null },
);
continue;
}
// Resolves inside the repo but is not a regular file (e.g. a symlink to
// a directory) — a legitimate exclusion, not a disclosure hazard.
}
const dirCache = new Map();
for (const file of files) {
const text = fs.readFileSync(path.join(ADR_DIR, file), 'utf8');
for (const { line, target: rawTarget } of extractLinks(text)) {
let t = String(rawTarget).trim();
if (t.startsWith('<') && t.endsWith('>')) {
t = t.slice(1, -1).trim();
} else {
// A link title: `dest "Title"` / `dest 'Title'`. Drop it, keep dest.
const titled = t.match(/^(\S+)\s+(?:"[^"]*"|'[^']*')$/);
if (titled) t = titled[1];
}
if (t === '' || t.startsWith('#') || t.startsWith('//') || /^[a-z][a-z0-9+.-]*:/i.test(t)) {
continue; // empty, same-document anchor, protocol-relative, or any URI scheme — out of scope
}
t = t.split('#')[0];
if (t === '') continue; // was only a fragment
try {
t = decodeURIComponent(t);
} catch {
// Malformed escape (e.g. "%zz"): resolve the raw, non-decoded text
// rather than throwing — an unresolvable literal is still reportable.
}
const abs = t.startsWith('/') ? path.resolve(ROOT, t.slice(1)) : path.resolve(ADR_DIR, t);
// Containment BEFORE any filesystem call: never `stat` outside ROOT.
const rel = path.relative(ROOT, abs);
if (escapesRoot(abs, ROOT)) {
add(`${file}:${line}`, `link "${rawTarget}" escapes the repository. Link a path inside the repo.`, {
reason: REASON.LINK_ESCAPES_REPO,
file,
line,
target: rawTarget,
});
continue;
}
const { exists, hint, escaped } = existsCaseExact(abs, dirCache, realRoot);
if (escaped) {
// A symlink under the (lexically in-repo) target path resolves
// outside the repository. Distinct message from the generic
// "does not resolve" below, and — deliberately — no hint: the hint
// itself would be the disclosure (see existsCaseExact's doc comment).
add(`${file}:${line}`, `link "${rawTarget}" escapes the repository via a symlink. Link a path inside the repo.`, {
reason: REASON.LINK_ESCAPES_REPO_SYMLINK,
file,
line,
target: rawTarget,
});
continue;
}
if (!exists) {
const relFromRoot = rel.split(path.sep).join('/');
const hintSuffix = hint ? ` Did you mean ${hint}? — link targets are case-sensitive on github.com.` : '';
add(
`${file}:${line}`,
`link "${rawTarget}" does not resolve — no such file or directory at ${relFromRoot}.${hintSuffix}`,
{ reason: REASON.LINK_UNRESOLVED, file, line, target: rawTarget, resolved: relFromRoot },
);
}
}
}
}
function parseAdr(file) {
const full = path.join(ADR_DIR, file);
const text = fs.readFileSync(full, 'utf8');
const lines = text.split(/\r?\n/);
// `fileId` is the numeric identity used for comparison ("0011" === "11");
// `displayId` preserves the filename's prefix exactly as written, because the
// corpus and its cross-references say "ADR-0001" and "ADR-58", not "ADR-1".
const rawId = file.match(/^([0-9]+)-/)[1];
const fileId = canonicalId(rawId);
const displayId = rawId;
const h1 = (lines.find((l) => /^#\s/.test(l)) || '').replace(/^#\s+/, '').trim();
// Capture the trailing status bracket against the RAW h1, before the title
// strip below discards it. This is the value the H1-vs-Status comparison in
// `validate` checks against — the strip alone throws the information away.
const bracketMatch = h1.match(STATUS_BRACKET_RE);
const bracketStatus = bracketMatch ? bracketMatch[1] : null;
// Title as displayed: drop a leading "ADR-123 — " / "ADR-123: " prefix and a
// trailing "[Proposed]"-style status bracket, both of which the index renders
// from structured fields instead.
const title = h1
.replace(/^ADR[-\s]?0*\d+\s*(?:[—:-]\s*)?/i, '')
.replace(STATUS_BRACKET_RE, '')
.trim();
const declaredIdMatch = h1.match(/^ADR[-\s]?0*(\d+)\b/i);
const declaredId = declaredIdMatch ? canonicalId(declaredIdMatch[1]) : null;
const header = headerBlock(text);
let statusRaw = null;
// relations[kind][dir] = [{field, value, links, bare}]
const relations = { supersedes: { out: [], in: [] }, subsumes: { out: [], in: [] } };
for (const [field, value] of headerFields(header)) {
if (field.toLowerCase() === 'status') {
if (statusRaw === null) statusRaw = value;
continue;
}
// "Supersedes (generalizes)" / "Subsumes as adapters" → "supersedes" / "subsumes"
const key = field.toLowerCase().replace(/\s*\([^)]*\)\s*/g, ' ').replace(/\s+as\s+.*$/, '').trim();
const spec = RELATION_FIELDS.get(key);
if (!spec) continue;
// "Supersedes: nothing; amends the ADR-1239 harness" asserts NO relation. Such a
// field routinely name-drops other ADRs in its prose ("related", "amends", "builds
// on"); reading those as supersession claims invents links that were never made.
if (NEGATED_RELATION_RE.test(value)) continue;
relations[spec.kind][spec.dir].push({ field, value, links: linkedAdrFiles(value), bare: bareAdrRefs(value) });
}
const statusToken = statusRaw ? (statusRaw.match(/^([A-Za-z]+)/) || [])[1] : null;
// A "Superseded by X" written into the Status line itself is the relation.
if (statusRaw && /^Superseded\b/i.test(statusRaw)) {
relations.supersedes.in.push({ field: 'Status', value: statusRaw, links: linkedAdrFiles(statusRaw), bare: bareAdrRefs(statusRaw) });
}
// `## Supersedes` / `## Subsumes` sections count as OUT claims (ADR-0174's table).
const sections = relationSections(text);
for (const kind of ['supersedes', 'subsumes']) {
if (sections[kind].length === 0) continue;
relations[kind].out.push({ field: `## ${kind === 'supersedes' ? 'Supersedes' : 'Subsumes'} section`, value: '', links: sections[kind], bare: [] });
}
return { file, fileId, displayId, title, declaredId, statusRaw, statusToken, bracketStatus, relations, text };
}
function buildCorpus() {
const { conforming, nonConforming } = partitionAdrFiles();
const adrs = conforming.map(parseAdr);
const byFile = new Map(adrs.map((a) => [a.file, a]));
const byId = new Map();
for (const a of adrs) {
if (!byId.has(a.fileId)) byId.set(a.fileId, []);
byId.get(a.fileId).push(a);
}
return { adrs, byFile, byId, nonConforming };
}
function validate({ adrs, byFile, byId, nonConforming }) {
const errors = [];
const violations = [];
// `record` carries the STRUCTURED half of every violation — `reason` plus
// whatever typed fields a `--json` consumer needs (line, target, resolved,
// expected/actual, status, …) — kept alongside, never instead of, the
// human `${file}: ${msg}` line: a large pre-existing test suite asserts on
// that string verbatim and is out of scope to migrate. `record` may supply
// its own `file` (spread AFTER the outer `file`) so link-violation records
// carry the plain filename as a field distinct from the human message's
// `<file>:<line>` prefix string.
const add = (file, msg, record) => {
errors.push(`${file}: ${msg}`);
violations.push({ file, ...record });
};
for (const f of nonConforming) {
add(
f,
'filename does not match the `<issue#>-<kebab-slug>.md` convention, so it cannot appear in the index. ' +
'Rename it (see docs/adr/README.md "Naming Convention"), or move it out of docs/adr/ if it is not an ADR.',
{ reason: REASON.FILENAME_INVALID, line: null },
);
}
for (const a of adrs) {
if (!a.statusToken) {
add(a.file, 'no `- **Status:** <Token>` field found in the header block.', {
reason: REASON.STATUS_MISSING,
line: null,
});
continue;
}
if (!STATUSES.includes(a.statusToken)) {
add(a.file, `status "${a.statusToken}" is not one of ${STATUSES.join(' | ')} (full line: "${a.statusRaw}").`, {
reason: REASON.STATUS_INVALID,
line: null,
status: a.statusToken,
});
} else if (
// Only compare when the status token is itself valid — an already-invalid
// token gets its own report above, and piling a bracket-disagreement
// message on top of it would be a second complaint about the same defect.
a.bracketStatus &&
a.bracketStatus.toLowerCase() !== a.statusToken.toLowerCase()
) {
add(
a.file,
`H1 status bracket [${a.bracketStatus}] contradicts the Status field (${a.statusToken}). ` +
'Update the H1 bracket (or the Status field) so they agree — a stale bracket is the first thing a reader sees.',
{ reason: REASON.STATUS_BRACKET_MISMATCH, line: null, expected: a.statusToken, actual: a.bracketStatus },
);
}
if (a.declaredId && a.declaredId !== a.fileId) {
add(a.file, `H1 declares ADR-${a.declaredId} but the filename says ${a.fileId}. The id must match the filename.`, {
reason: REASON.ID_MISMATCH,
line: null,
expected: a.fileId,
actual: a.declaredId,
});
}
// A Superseded ADR must point at its successor by FILE LINK.
if (a.statusToken === 'Superseded') {
const links = a.relations.supersedes.in.flatMap((r) => r.links);
if (links.length === 0) {
const bare = a.relations.supersedes.in.flatMap((r) => r.bare);
if (bare.length) {
add(
a.file,
`status is Superseded and mentions ADR-${bare.join('/')} but not as a markdown link to the file. ` +
'Bare ids are ambiguous (ADR-0010 and ADR-0011 each resolve to multiple files) — link the target file.',
{ reason: REASON.SUPERSEDED_BARE_ID, line: null, bare },
);
} else {
add(a.file, 'status is Superseded but names no successor. Write `Superseded by [ADR-N](N-slug.md)`.', {
reason: REASON.SUPERSEDED_NO_SUCCESSOR,
line: null,
});
}
}
}
// Every relation link must resolve; every bare id must be linked (and exist).
for (const kind of Object.keys(RELATION_SPEC)) {
for (const dir of ['out', 'in']) {
for (const rel of a.relations[kind][dir]) {
for (const l of rel.links) {
if (!byFile.has(l)) {
add(a.file, `"${rel.field}" links "${l}", which does not exist in docs/adr/.`, {
reason: REASON.RELATION_LINK_MISSING,
line: null,
field: rel.field,
target: l,
});
}
}
// The synthetic relation lifted out of the Status line is already covered by
// the dedicated Superseded check above; reporting it again just duplicates.
if (rel.field === 'Status') continue;
// Ids already linked ANYWHERE in this field. A field legitimately
// reads "…([ADR-58](58-x.md)) — see 'Relation to ADR-58' below": the
// trailing prose repeats an id that is linked earlier, and flagging
// that would be noise. But an id that appears ONLY bare is an
// unchecked claim — and testing `rel.links.length` instead of the
// specific id silently dropped every bare claim in a field that
// happened to carry one link.
const linkedIds = new Set(rel.links.map((l) => (byFile.get(l) || {}).fileId).filter(Boolean));
for (const b of rel.bare) {
if (linkedIds.has(b)) continue;
const candidates = byId.get(b) || [];
if (candidates.length === 0) {
add(
a.file,
`"${rel.field}" names ADR-${b}, which does not exist in docs/adr/. If it is an ISSUE number, write "#${b}" — not "ADR-${b}".`,
{ reason: REASON.RELATION_BARE_ID_MISSING, line: null, field: rel.field, target: b },
);
} else {
add(
a.file,
`"${rel.field}" names ADR-${b} without a file link` +
(candidates.length > 1 ? ` (ambiguous — resolves to ${candidates.length} files: ${candidates.map((c) => c.file).join(', ')})` : '') +
'. Link the target file so the relation is checkable.',
{
reason: REASON.RELATION_BARE_ID_UNLINKED,
line: null,
field: rel.field,
target: b,
candidates: candidates.map((c) => c.file),
},
);
}
}
}
}
}
}
// Symmetry, per relation kind: A -out-> B <=> B -in-> A.
//
// Only a RATIFIED (Accepted) claimant is owed the back-link. A Proposed ADR's
// supersession claim is prospective — it has not taken effect, so stamping its
// target as superseded would assert something untrue. When such an ADR is
// ratified to Accepted, this check starts demanding the back-links at exactly
// the right moment — as ADR-857 shows: it was Proposed when this guard was
// written, was ratified to Accepted on 2026-07-17 (its claim over ADR-0011 /
// ADR-58 restated as "Subsumes", since both remain Accepted and live), and
// both targets now carry the reciprocal "Subsumed by" back-link this demands.
const OPPOSITE = { out: 'in', in: 'out' };
for (const a of adrs) {
for (const kind of Object.keys(RELATION_SPEC)) {
for (const dir of ['out', 'in']) {
// The ratification guard applies to the OUT direction only: an unratified
// ADR's claim over someone else is prospective and must not obligate the
// target. The IN direction is this ADR's statement about ITSELF ("I am
// superseded by X") and is always owed a reciprocal — guarding it too
// would skip every Superseded ADR (statusToken !== 'Accepted') and leave
// dangling one-way claims unchecked, which is the bug this gate exists
// to catch.
if (dir === 'out' && a.statusToken !== 'Accepted') continue;
for (const target of new Set(a.relations[kind][dir].flatMap((r) => r.links))) {
const b = byFile.get(target);
if (!b) continue;
const back = new Set(b.relations[kind][OPPOSITE[dir]].flatMap((r) => r.links));
if (back.has(a.file)) continue;
const needed = RELATION_SPEC[kind][OPPOSITE[dir]];
const claim = dir === 'out' ? `it ${kind} this ADR` : `it is ${kind === 'supersedes' ? 'superseded' : 'subsumed'} by this ADR`;
add(
target,
`${a.file} declares ${claim}, but this ADR does not record it. ` +
`Add \`- **${needed}:** [ADR-${a.displayId}](${a.file})\` so a reader of THIS file learns the decision moved on.`,
{ reason: REASON.RELATION_ASYMMETRIC, line: null, source: a.file, kind, neededField: needed },
);
}
}
}
}
// Link resolution reads the directory directly rather than the parsed
// `adrs` list — it must ALSO cover README.md and naming-violation files,
// neither of which is in `adrs` (see `validateLinks`'s own doc comment).
validateLinks(add);
return { errors, violations };
}
const GROUPS = [
{
heading: 'Active decisions',
blurb: 'These govern the system as it stands. Cite these.',
match: (a) => a.statusToken === 'Accepted',
},
{
heading: 'Proposed',
blurb: 'Decided in principle, not yet ratified. Do not cite as settled architecture.',
match: (a) => a.statusToken === 'Proposed',
},
{
heading: 'Superseded, Retired, and Legacy',
blurb: 'Historical record. **Do not follow these** — each names what replaced it, or why it was retired.',
match: (a) => ['Superseded', 'Retired', 'Legacy'].includes(a.statusToken),
},
];
/**
* Render ADR-authored text (a title) into a markdown table cell.
*
* Three hazards, all from text this script does not control:
* - `|` would split the cell and corrupt the row.
* - An HTML comment would be emitted verbatim into README.md. A title
* containing the END marker relocates it, so the NEXT `--write` splices
* against the wrong boundary and silently eats the rest of the file.
* Escaping `<`/`>` makes a comment sequence unformable, which also blocks
* any other HTML injected through a title.
* - A backslash is markdown's own escape character, so it MUST be escaped
* first. Escaping `|` → `\|` without it turns the input `\|` into `\\|`,
* which markdown reads as a literal backslash followed by an UNESCAPED
* pipe — re-opening the cell break the pipe escape exists to prevent.
* Order is load-bearing: backslash first, then everything that emits one.
*/
function cellText(text) {
return String(text)
.replace(/\\/g, '\\\\')
.replace(/\|/g, '\\|')
.replace(/</g, '&lt;')
.replace(/>/g, '&gt;')
.replace(/\r?\n/g, ' ')
.trim();
}
function linkCell(files, byFile) {
if (files.length === 0) return '—';
return files.map((l) => `[ADR-${(byFile.get(l) || {}).displayId || '?'}](${l})`).join(', ');
}
function renderIndex(corpus) {
const { byFile } = corpus;
const out = [START_MARKER, ''];
for (const g of GROUPS) {
const rows = corpus.adrs.filter(g.match).sort((x, y) => Number(x.fileId) - Number(y.fileId));
if (rows.length === 0) continue;
// No row count in the heading: it is a numeric cell inside the generated
// region, shared by every ADR-adding PR. Two PRs that add different ADRs
// touch different table rows and merge cleanly — but both rewrite this
// same count line, so whichever lands second gets a stale local --check
// pass and a red CI --check against the merged tree (#3251). Same failure
// mode CHANGELOG.md and drift-acks already solved by moving to per-PR
// fragment files (.changeset/, tests/emitted-drift-acks/); here the fix is
// simpler still — the count carries no verification value (--check
// regenerates and diffs the whole region regardless) and is trivially
// derivable by counting the table rows. Do not add it back.
out.push(`### ${g.heading}`, '', g.blurb, '');
const isHistorical = g.heading.startsWith('Superseded');
// "Read first" points at the broader ADR that now frames this one. It is how a
// reader of a still-Accepted component decision (e.g. the runtime descriptor)
// discovers the wider decision that reframed it (e.g. EoS) instead of assuming
// the component IS the architecture.
out.push(
isHistorical ? '| ADR | Title | Status | Replaced by |' : '| ADR | Title | Status | Read first |',
isHistorical ? '|-----|-------|--------|-------------|' : '|-----|-------|--------|------------|',
);
for (const a of rows) {
const cells = [`[ADR-${a.displayId}](${a.file})`, cellText(a.title), a.statusToken];
cells.push(
isHistorical
? linkCell([...new Set(a.relations.supersedes.in.flatMap((r) => r.links))], byFile)
: linkCell([...new Set(a.relations.subsumes.in.flatMap((r) => r.links))], byFile),
);
out.push(`| ${cells.join(' | ')} |`);
}
out.push('');
}
// No total ADR count either, for the same reason as the per-group heading
// count above: it is a second shared mutable cell in the generated region
// that every ADR-adding PR would rewrite, guaranteeing the identical merge
// race (#3251). Leave it out; the count is derivable by reading the table.
out.push(
`_Generated by \`scripts/gen-adr-index.cjs\` — run \`--write\` after adding or restatusing an ADR._`,
'',
END_MARKER,
);
return out.join('\n');
}
function spliceIntoReadme(readme, index) {
const start = readme.indexOf(START_MARKER);
const end = readme.indexOf(END_MARKER);
if (start === -1 || end === -1) {
throw new ExitError(
1,
`docs/adr/README.md is missing the index markers.\nExpected:\n ${START_MARKER}\n ${END_MARKER}\n`,
);
}
return readme.slice(0, start) + index + readme.slice(end + END_MARKER.length);
}
/**
* Parse CLI flags from `argv` (already sliced to just the flags, i.e.
* `process.argv.slice(2)`). Supports `--write`, `--check`, `--json` in any
* order. FAIL-CLOSED on an unrecognized flag: silently falling through to
* the no-flags "print the index" behavior would mask a typo (e.g.
* `--jsno`) as a clean run instead of failing loudly, so any argument that
* is not one of the three recognized flags throws `ExitError(1, …)` naming
* the offender rather than being ignored.
*/
function parseArgs(argv) {
const opts = { write: false, check: false, json: false };
for (const arg of argv) {
if (arg === '--write') opts.write = true;
else if (arg === '--check') opts.check = true;
else if (arg === '--json') opts.json = true;
else throw new ExitError(1, `unknown flag: ${arg}\nRecognized flags: --write, --check, --json.`);
}
return opts;
}
function main() {
const { write, check, json } = parseArgs(process.argv.slice(2));
const corpus = buildCorpus();
const { errors, violations } = validate(corpus);
const index = renderIndex(corpus);
if (json) {
// `--json` implies `--check` semantics (`--check --json` is identical to
// `--json` alone) but emits a single JSON document to stdout instead of
// the human stderr report, and writes nothing to stderr at all. Unlike
// the human `--check` path below — which short-circuits on lifecycle
// violations and never even reads README.md to check staleness — the
// JSON report always computes BOTH facts (`violations` and
// `indexStale`) independently, since a consumer parsing the document
// needs the complete picture in one shot rather than one violation
// class masking the other.
const readme = fs.readFileSync(README_PATH, 'utf8');
const expected = spliceIntoReadme(readme, index);
const indexStale = expected !== readme;
const ok = violations.length === 0 && !indexStale;
process.stdout.write(JSON.stringify({ ok, adrCount: corpus.adrs.length, indexStale, violations }) + '\n');
return ok ? 0 : 1;
}
if (errors.length > 0 && !write) {
process.stderr.write(
`docs/adr/ has ${errors.length} lifecycle violation(s).\n` +
'See docs/adr/README.md "Lifecycle rules" for the contract.\n\n',
);
for (const e of errors) process.stderr.write(` ✗ ${e}\n`);
process.stderr.write('\n');
throw new ExitError(1);
}
// `--write` takes precedence over a co-supplied `--check`: neither
// combination is part of this CLI's documented contract (the flags exist
// to be used one at a time, or as `--check --json`), so this is an
// arbitrary-but-safe tiebreak rather than a specified behavior.
if (write) {
const readme = fs.readFileSync(README_PATH, 'utf8');
fs.writeFileSync(README_PATH, spliceIntoReadme(readme, index));
process.stdout.write(`Wrote ADR index into ${README_PATH} (${corpus.adrs.length} ADRs).\n`);
if (errors.length > 0) {
process.stderr.write(`\n${errors.length} lifecycle violation(s) remain — --check will fail:\n\n`);
for (const e of errors) process.stderr.write(` ✗ ${e}\n`);
}
} else if (check) {
const readme = fs.readFileSync(README_PATH, 'utf8');
const expected = spliceIntoReadme(readme, index);
if (expected !== readme) {
process.stderr.write(
'docs/adr/README.md index is stale. Run:\n node scripts/gen-adr-index.cjs --write\n\n',
);
throw new ExitError(1);
}
process.stdout.write(`docs/adr/README.md index is up to date (${corpus.adrs.length} ADRs).\n`);
} else {
process.stdout.write(index + '\n');
}
}
// Guarded: `require`-ing this module (the test suite imports STATUSES and
// the pure scanner directly) must not also run the generator as a side
// effect of loading it.
if (require.main === module) runMain(main);
module.exports = { STATUSES, REASON, extractLinks, maskCode };