* feat(#3987): guard slug re-derivation, and record why the swallow shape cannot be guarded Epic #3473's Decision 1 requires the wrong call site be UNREPRESENTABLE. #3984 measured that two of the nine §8 rules had no guard at all and recorded both as "Shipped - test-covered". This closes one of them, proves the other cannot be closed the same way, and corrects two false claims I merged yesterday. 1. §8.3 - scripts/lint-slug-derivation-drift.cjs. generateSlugInternal (src/core-utils.cts) is the canonical owner; #3883 removed 11 inline copies. Nothing prevented a twelfth: no slug guard existed in scripts/ or eslint-rules/. The detector is STATEMENT-scoped and matches the shape the real copies took - one statement carrying BOTH .replace(<negated class>, '-') and .replace(/^-+|-+$/, ''). Statement scoping is what buys the precision: the loose LINE-level form yields 18 hits with 7 unrelated, a material false-positive rate. Measured on the tree: 5 flags, 2 TRUE, 3 SANCTIONED, 0 FALSE. The three sanctioned sites are allowlisted with a reason each, following lint-phase-enumeration-drift's form rather than a bare denylist. The owner itself is listed explicitly even though it escapes by construction - an implicit escape is a latent bug, and the next person to touch line 192 would not know the guard depended on it. 2. Both TRUE positives were live defects, not style. scripts/qa-smell-ratchet.cjs reproduced the canonical formula including the 60-cap but trimmed BEFORE truncating - the #2849 bug - and never transliterated. The divergence is total, not cosmetic: canonical "privet-mir-privet-mir-privet-mir-privet-mir-privet-mir-prive" inline "tail" Cyrillic collapsed to nothing and only the ASCII remainder survived, so the ratchet was keying on wrong identifiers for any non-ASCII input. tests/planning-inspect.test.cjs carried a helper whose comment claimed parity with getPhaseDirFromPhaseId. That function now transliterates; the helper did not, so the test asserted against a stale formula while looking correct. Both now route through the seam. 3. §8.5 - measured, and deliberately NOT shipped. A candidate detector (swallowing catch + errno-retry-set test in the same function) gives 26 flags across 11 functions: 0 TRUE, 26 FALSE. Every one is best-effort unlink/rm/close cleanup, lost-rename-race backoff, or a deliberate null fallback. The file-scoped variant is worse at 71. Worse than the noise: the only known true instance was removed by #3885, so there is NO POSITIVE CONTROL - the guard cannot be shown capable of failing, which this repo requires of every drift guard. Shipping it would add a guard nobody can trust and nobody can test. The ADR now records the measurement and the reason, keeps §8.5 at "Shipped - test-covered", and points at the #1884 regression test as what actually enforces it. An honest "not detectable at acceptable precision" beats a guard that only ever passes. 4. Two claims I merged into the ADR yesterday were wrong. §8.9 said 17 of 19 subsumed children have a test citing their issue number, and that #3364 and #3812 have none. Both halves are false, and the claim came from a NUMBER-GREP - inside an amendment whose own subject is that a text match is not a fact. #3364 IS cited: tests/runtime-marker-resolution.test.cjs:107, T3 installMarkerResolvesWhenEnvAndConfigAbsent_3897 (#3364), asserting at :115-119. #3812 IS covered: tests/gen-state-md-docs.test.cjs:374, asserting at :382. Corrected to 19 of 19. #3812 does carry a real finding, though a different one: it is PARTIALLY DELIVERED on a CLOSED issue. The shipped fix declares cardinality for frontmatter keys, but #3812's stated acceptance was about the ## Current Position BODY section, and docs/reference/state-md.md:196-208 still has no normative single-valued/overwrite sentence and no pointer to ## Performance Metrics for history. Recorded in the ADR and left for #3812 to re-open - fixing it here would bury a scope question inside an unrelated PR. Note on B6: this ADDS a guard, and B6 said the net count must fall. #3951 already amended that clause - a guard ledger is a claim about COVERAGE, not count - which is what makes adding this one honest rather than contradictory. Verified: the guard flags 0 on the fixed tree, and PROVES IT CAN FAIL - a fresh inline copy planted in src/ makes it exit 1 naming the exact statement. All three sanctioned sites were confirmed exempt BY the allowlist, not by accident of the pattern, by re-attributing each to a non-exempt path and watching it flag. build:lib, lint and lint:ci all exit 0. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3987): add the changeset fragment Doc-only, so it carries forward from the verified sha rather than costing a second matrix run. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): §8.5 IS guardable — I was wrong, and the guard found a live defect Two orthogonal reviews. The correctness review overturned my central judgment, and it was right. 1. I concluded §8.5 was "not detectable at acceptable precision" and recorded that in the ADR. False. My evidence was 26 flags / 0 TRUE / 26 FALSE. The reviewer pointed out what I had not: all 26 false positives are CLEANUP verbs - rmSync 54, unlinkSync 43, closeSync 17, chmodSync 12 - and the obvious narrower predicate was never tried. A swallowed cleanup is legitimate best-effort. A swallowed CREATION is a precondition silently lost, which is exactly the #1884 shape. Measured properly, in three stages: swallowing catch 911 + try-block calls a CREATION verb 24 + enclosing function references a *_ERRNOS set 0 0 flags, 0 false positives. The `*_ERRNOS` naming key is empirically total - all 10 retry/tolerate sets in src/ follow it. My second claim was worse. I wrote that no positive control exists because #3885 removed the only true instance, so the guard "cannot be shown capable of failing". That is self-refuting: this very PR's slug guard proves-it-can-fail on a synthetic tree, and the pre-#3885 blob is available as exactly such a fixture. It is now the control, and it works in both directions - the rule flags 0c43d853e^:src/planning-workspace.cts at line 210, the line the fix commit's own message cites, and reports zero on the post-fix code. I stopped at the first negative result on the option that meant less work. Shipped as eslint-rules/no-swallowed-precondition.cjs, wired into the existing src/**/*.cts ESLint block rather than a scripts/lint-*-drift.cjs: no script in scripts/ requires typescript/espree/acorn, and scripts/ ships to consumers, so a .cts-parsing standalone guard would add a devDep at consumer runtime. The ESLint block already parses .cts for free. 2. The guard immediately found a live defect of the same class. src/capability-lock.cts swallowed a mkdirSync on the lock directory, then acquireLock classified the follow-on failure as `code !== 'EEXIST' → return null`. A real EACCES/EROFS makes openSync(lockPath,'wx') fail ENOENT, which is not EEXIST - so a fatal filesystem error was laundered into "lock unavailable". Same defect as #1884, different laundering target. Fixed the way #3885 fixed #1884: the creation failure propagates. Regression test proven fail-first by hand - with the fix stashed, EACCES was laundered to null; restored, it throws. The strict rule does NOT catch this shape (its errno classification is an inline literal, not a named set). The rule is deliberately left strict: the broadened form had 2 false positives - capability-lock.cts:408, the deliberate EEXIST steal protocol, and commonjs-marker.cts:131, which returns a distinct documented outcome. The gap is noted in code rather than papered over with a noisy predicate. 3. The security review found the slug guard's exemption FAILED OPEN. currentFunction was never reset, and only a column-0 `function` declaration updated it, so exemption bled from an allowlisted declaration to the next one. generateSlugInternal exempted 50 lines for an 11-line function. A re-derivation planted anywhere in that window was silently exempt - the same fail-open shape that produced a blocker in #3897, and an allowlist is a SUBTRACTION so a mismatch fails open by construction. Extent is now tracked by real brace depth, and a test plants a violation after each allowlisted function's real closing brace and asserts it IS flagged. 4. Also from the security review: the guard was a CI-DoS and narrower than I claimed. Its unbounded [^\]]* was re-scanned from every `.replace(/[^` start: 54.3s on a 1.28MB line. It imported MAX_REGEX_LITERAL_LEN and never called readRegexLiteralAt - the bounded tokenizer that exists for exactly this. Now routed through it with a 2MB file cap: ~200ms. 15 of 25 genuine re-derivations evaded. Widened to catch replaceAll, {1,}, \s*-wrapped classes, escaped ], literal new RegExp(...), five trim spellings, .split().join(), and multi-line .replace( args - still 0 false positives. Two forms still evade and are documented as deliberate gaps with negative tests: the two-statement/temp-var form and new RegExp built from a variable. Both need data flow, and guessing at it is how a guard becomes noisy. Also fixed: // inside a string truncated the line, a ; inside the collapse regex split the statement (a one-character bypass), and SCAN_EXT omitted .mjs/.tsx/.jsx. 5. A regression I introduced, caught by the same review. qa-smell-ratchet.cjs top-level-required a build output that is not git-tracked, so the script hard-failed MODULE_NOT_FOUND before build:lib - including for --help, which previously had no build dependency. The require is now lazy at the point of use. 6. Four of my own tests were vacuous or weak. T9's input yielded an identical string under the buggy formula, so it passed on the implementation it was meant to catch. T12 compared maxLen null vs 60 on an 18-char name, where they agree trivially. T9-T12 all asserted generateSlugInternal directly, so they would pass unchanged if both call-site fixes were reverted. And prove-it-can-fail was scoped to scanRepo, never the CLI - dropping main()'s exit-code line would have kept every row green. All rewritten with discriminating inputs, per-call-site rows that red when the fix is reverted, 59/60/61 boundaries, an entirely-non-alphanumeric row, and a CLI row asserting the real subprocess exit code and both sanitizeForReport sites. Verified: both guards flag 0 on the tree and both prove they can fail. The swallow rule's control is confirmed in both directions - pre-#1884 shape flagged, post-#3885 shape clean. build:lib, lint and lint:ci all exit 0. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3987): record that §8.5 IS guardable, and correct a correction that made a ledger worse Three ADR corrections, two of them to text this branch wrote hours ago. §8.5 advances to Enforced. Its previous entry said the rule was not detectable at acceptable precision. That was wrong twice: the 26 false positives were uniformly CLEANUP verbs, which is a reason to narrow the predicate rather than abandon it, and the claim that no positive control exists was self-refuting - the pre-#3885 blob is available as a fixture and this repo's own guards prove-it-can- fail on synthetic trees. Narrowed to creation verbs plus a *_ERRNOS reference: 911 -> 24 -> 0 flags, 0 false positives, control confirmed in both directions. The entry keeps the wrong reasoning visible, because a high false-positive count being evidence the predicate is wrong - not evidence the rule is unguardable - is the transferable part, and the first negative result is most seductive when it is also the answer that means less work. §8.9's correction is itself corrected. The original 17-of-19 claim was CORRECT for the predicate it stated; this branch silently swapped cited -> covered and declared 19 of 19. #3812 appears in zero test files. Changing what a word means to make a ledger read better is a worse failure than the miscount it claimed to repair. Both predicates are now reported separately - 18 of 19 cited, 19 of 19 covered - because §8.9 asks for a test NAMING each child, so 18 is the number that answers it. #3812 is also re-opened for real, rather than the first draft's promise that it could be. §8.3 stays Shipped - test-covered rather than advancing. The slug guard catches the copy-paste class and a dozen variants, but two forms still evade by decision (temp-var split, new RegExp from a variable) because both need data flow. Naming them keeps the status honest: the wrong call site is much harder to write, not unrepresentable. Closes #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3987): backfill changeset pr number Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): replace my own wall-clock assertion, and close the guard that let me write it CI went red on ubuntu shard 2/3. The failing test was mine, and the failure was the test, not the code. a 1.28MB line ... scans in well under a second (was 54.3s pre-fix) 7368ms It asserted ELAPSED TIME. ~200ms locally, 7.4s on a shared CI runner. The bound introduced for the MAJOR-2 DoS fix works - 7.4s against a 54.3s pre-fix baseline is the fix doing its job - but an absolute wall-clock threshold on shared hardware is a race, not an assertion. CLAUDE.md says so directly: "Clock Seams: Do not assert on wall-clock time." I wrote the anti-pattern the project bans, in a PR about guards. Raising the threshold would only move the flake. The row now asserts a DETERMINISTIC bound instead: an instrumentation seam on drift-scan.cjs counts readRegexLiteralAt calls and characters examined, and the test asserts charsExamined stays under an absolute ceiling. Measured on the same 1.28MB fixture: 120,000 calls, 48,000,000 chars - two orders under the ceiling. The pathological fixture is kept; only the thing being asserted changed. Proven to still discriminate: with MAX_REGEX_LITERAL_LEN raised to simulate the unbounded pre-fix behavior, the same fixture does not complete in 120 seconds, versus ~0.3s bounded. It is a real regression test, not a tautology. Then the second half, which is the same defect class as the rest of this PR. eslint-rules/no-elapsed-assertion.cjs matched only the EXACT identifiers ^(elapsed|duration|took|ms)$. I used `elapsedMs`. It evaded the rule entirely. tookMs, durationMs, elapsedTime and msElapsed evade the same way. A guard that cannot see the violation it exists to catch is exactly what this PR is about - it just happened to be an existing rule rather than one of the two I came here for, and it was found because I committed the violation it should have blocked. Widened to /^(?:elapsed|duration|took|ms)(?:[A-Z]\w*)?$/ plus a narrow start/endMs delta pair. Deliberately NOT a blanket *Ms suffix: a first draft did that and produced 2 false positives on `timeoutMs` in plan-phase-stall-detection, which is a configured timeout and not a measurement. Verified negative on params, items, forms, terms, dirnames, timeoutMs, cacheTtlMs and staleAfterMs. Measured over the five files carrying camelCase timing identifiers: 0 true positives beyond my own, so nothing else needed rewriting. The rule's own test file gains a row asserting `elapsedMs` flags, proven to fail against the pre-widening rule - the same prove-it-can-fail standard both new guards in this PR are held to. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): a comment I added leaked a Claude reference into every runtime install The runner went red with 4 failures in tests/install.test.cjs: Leaking: .hermes/scripts/lib/drift-scan.cjs Leaking: .qwen/scripts/lib/drift-scan.cjs The instrumentation seam added for the deterministic bound carried a comment naming CLAUDE.md as the source of the no-wall-clock-assertions rule. scripts/ SHIPS to consumers, so that comment was installed verbatim into hermes and qwen trees, and the install suite scans for exactly this - a Claude-specific reference reaching a non-Claude runtime. The rule is real and worth citing; the filename is not portable. The comment now says "this repo's test rules" and states the rule inline, which is what a reader of an installed tree actually needs anyway. Worth noting what caught it: not lint, and not the two guards this PR adds - the install suite's full-tree scan, which exists precisely because a shipped file is read by runtimes that have never heard of CLAUDE.md. Same lesson as the rest of this PR from the other direction: the check that matters is the one that can see the surface where the defect actually lands. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): a test fixture swallowed 46 git exit codes and produced a silent false negative CI red on ubuntu shard 3/3: tests/health-validation.test.cjs:2029 expected exactly one W024, got [{"code":"W006", ...}] Not caused by this branch, and the evidence is decisive rather than a hunch: the SIBLING test at :2039 builds the IDENTICAL fixture with the identical commitsAhead and asserts the same thing, and it PASSED in the same process, same file, same run. Same input, both outcomes - which rules out logic, ordering, sharding and environment, and leaves a per-invocation nondeterministic failure inside one fixture build. The mechanism is an unchecked exit code, 46 times over. The W024 fixture performs ~46 runGit spawns and never checks a single one. runGit returns failures as DATA and never throws, so one silently-failed `git commit` yields 19 commits instead of 20, or a silently-empty `git rev-parse HEAD` yields a blank state_head. Either drops readStateHeadFreshness below the advisory threshold, W024 never fires, and only W006 remains. Reproduced exactly: 20 commits -> ["W006","W024"]; 19 -> ["W006"]; blank state_head -> ["W006"] - byte-identical to the CI assertion dump. The arithmetic is what hid it. At threshold-1 and threshold+1 a lost commit still produces the asserted answer; only the exactly-at-threshold cases sit one commit from a false negative. Two of the seven tests are in that position, and CI hit one. That is why it had never been seen before, and why it surfaced now: this branch adds three test files, which reshuffles the cost-weighted shard partition and moved this file into a chunk where the latent flake fired. My files were checked as suspects first and cleared: all fixtures mkdtemp-unique, no process.chdir, no .planning/ writes, no git spawns, and node --test gives per-file process isolation regardless. Fixed at the cause, not the symptom. A mustGit wrapper throws on a non-zero exit with the command, exit code and stderr, and all nine call sites route through it. The fixture now asserts its OWN preconditions before the assertion under test runs - the seed head is non-empty, and `git rev-list --count <seed>..HEAD` equals the requested commitsAhead - so a fixture that did not build what it claims fails loudly as a FIXTURE ERROR naming got-versus-asked, instead of quietly handing a weaker input to the assertion. Proven: dropping one commit now raises FIXTURE ERROR: requested commitsAhead=19 but git rev-list --count reports 18 where it previously produced a silent ["W006"] pass-for-the-wrong-reason. 64/64 tests in that block pass unperturbed. Deliberately NOT done: no threshold change, no retry, no loosened assertion, no skip. The assertion was correct; the input was silently wrong. Worth naming, because it is the same shape from the other side: this PR ships eslint-rules/no-swallowed-precondition.cjs, whose entire subject is a swallowed precondition failure being laundered into a plausible downstream outcome. This fixture is that defect in test code - the swallowed git failure was laundered into a legitimate-looking "W024 did not fire". The rule does not cover test fixtures, so the connection is noted at the fix site rather than enforced. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3987): two tests wrote to committed files; the shard packing decided when that mattered CI red on windows-latest shard 1/3 only: "gen-exit-code-registry: CLI" > "a --write run redirected to a tmpdir leaves every committed artifact untouched" AssertionError: hooks artifact must be untouched The Linux runner passed the same sha at 40425/40425. It is Linux-only, so a Windows-scheduling defect is structurally invisible to it. Root cause, established by measurement rather than inference. tests/cli-exit.test.cjs appended a corruption marker to the REAL COMMITTED hooks/lib/exit-code-registry.js, held it corrupted across a full subprocess, and restored it in a finally. tests/exit-code-registry.test.cjs reads that same real file before and after its own subprocess and asserts byte equality. If it samples while the other test holds the file corrupted, it fails. The landmine is pre-existing, from2ea5efc15(#3911). What this branch changed is WHEN the two run together. scripts/run-tests.cjs shards by cost-weighted LPT over the sorted unit list, so adding three test files repacks the bins: merge-basec3e667df3(838 files): cli-exit -> shard 1, exit-code-registry -> shard 3 HEAD 03b342601 (841 files): BOTH in shard 1, same argv chunk, one node --test process, concurrent Co-location is necessary but not sufficient - Linux shard 1/3 also had both and passed. Windows loses because TEST_CONCURRENCY defaults to 2 there against 4 elsewhere, spawn cost is ~10x, and the sibling corruptor holds one of only two slots through a ~90s tsc compile. That turns a sub-second overlap into seconds. Not a path-separator or case-sensitivity issue, and not CRLF - .gitattributes pins * text=auto eol=lf. Redirection was not at fault either: ensureScriptsOut derives all five --out flags correctly and gen-exit-code-registry.cjs honours them with no __dirname escape. Fixed at the cause: no test writes to a committed file any more. Both corruptors now copy to a mkdtempSync tmpdir, corrupt the COPY, and point the generator at it. Repinning or reordering the shards would have turned CI green while leaving the landmine armed for the next reshuffle. That required closing an inconsistency between two sibling generators. gen-exit-code-registry.cjs already accepts --out/--scripts-out/--hooks-out/--dts-out/--sh-out and honours them under --check; gen-hooks-cli-exit.cjs hardcoded OUTPUT_PATH and had no flag surface at all, so its corruptor could not be redirected anywhere. It now takes --out in the same style, honoured by both --write and --check, and is a no-op when absent - verified: a bare --check on the default path still exits 0. ensureScriptsOut moved to tests/helpers/exit-code-artifact-flags.cjs and both test files import it. Hand-rolling a second copy of the flag derivation would have been a re-derivation of exactly the kind this PR ships a guard against. Verified: both tests still detect corruption (proven by defeating the check and watching them red, with a positive control showing an uncorrupted copy exits 0); SHA-256 of hooks/lib/exit-code-registry.js and hooks/lib/cli-exit.js identical before and after running both rewritten bodies, and git reports nothing under hooks/ modified - that is the property that was violated. A repo-wide search for the corrupt-then-restore-in-finally shape against hooks/ found no other instances. One detail worth recording: the tmpdir test keeps --declaration pointed at the real committed declaration rather than copying it, because the generated banner embeds path.relative(REPO_ROOT, declarationPath) - copying it would produce a false drift unrelated to the injected corruption. The declaration is read-only on that path and never written. Refs #3987 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
572 lines
29 KiB
TypeScript
572 lines
29 KiB
TypeScript
/**
|
|
* Shared cross-process mutual-exclusion lock primitive — #1459 finding 4 + #1462 finding 1.
|
|
*
|
|
* A SINGLE hardened lockfile protocol shared by BOTH capability-lifecycle (the `.gsd/capabilities/.lock`
|
|
* mutation lock) and capability-consent (the consent-store `.consent.lock`). Before this extraction the
|
|
* two locks had DIFFERENT, divergent steal policies: the lifecycle lock was hardened (#1462 — pid +
|
|
* process-start-time identity + hard deadman, never steals a verified-live same-host holder), while the
|
|
* consent lock used a naive mtime-only 60 s steal that would STEAL A LIVE WRITER (a slow/paused holder
|
|
* past 60 s is reclaimed; the original writer then resumes and overwrites — a lost update). Sharing one
|
|
* primitive makes the consent lock as safe as the lifecycle lock (single source of truth — mirrors the
|
|
* shared-validator / shared bounded-reader lessons).
|
|
*
|
|
* STEAL PROTOCOL (never deadlocks AND never steals a verified-live SAME-host holder). The age is bound
|
|
* to the BODY instance the acquirer acts on — `age = now - body.ts` for a JSON body (a fresh replacement
|
|
* body carries a fresh ts), falling back to `now - mtime` for a legacy/no-`ts` body — and the
|
|
* (dev, ino, ts) identity is re-confirmed immediately before the atomic rename-steal:
|
|
* - age <= LOCK_STALE_MS → FRESH: never stolen (genuinely held → blocked).
|
|
* - age > LOCK_STALE_MS:
|
|
* · SAME host: VERIFIED-LIVE (pid alive AND recorded startTime present AND observed startTime ===
|
|
* recorded) → NEVER steal (even past the deadman). NOT verified-live (dead pid, start-time
|
|
* MISMATCH = pid-reuse, or start-time unobtainable) → STEAL (fast local recovery).
|
|
* · DIFFERENT host / no parseable pid (legacy/oversized/garbage body) → liveness unverifiable →
|
|
* steal ONLY after age > LOCK_DEADMAN_MS (the deadman fallback).
|
|
*
|
|
* The lockfile body is UNTRUSTED: it is read via the shared fd-based bounded reader
|
|
* (ledgerMod.readSmallRegularFile) so a FIFO/device/oversized body cannot block or read unbounded.
|
|
*
|
|
* Test seam: _setLockProbes / _resetLockProbes inject deterministic isPidAlive / getProcessStartTime so
|
|
* the start-time liveness branches are exercised without depending on real OS pids beyond the current
|
|
* process. capability-lifecycle re-exports these so its existing #1462 lock tests keep driving them.
|
|
*
|
|
* Imports: node:fs, node:path, node:os, node:crypto, and the ledger's shared bounded readSmallRegularFile
|
|
* + execTool (for the rare start-time shell-out on win32/macOS).
|
|
*/
|
|
|
|
import fs from 'node:fs';
|
|
import path from 'node:path';
|
|
import os from 'node:os';
|
|
import crypto from 'node:crypto';
|
|
|
|
/* eslint-disable @typescript-eslint/no-require-imports */
|
|
const ledgerMod = require('./capability-ledger.cjs') as {
|
|
readSmallRegularFile: (filePath: string, maxBytes: number) => string | null;
|
|
};
|
|
const { execTool, retryRenameSync } = require('./shell-command-projection.cjs') as {
|
|
execTool: (
|
|
program: string,
|
|
args: string[],
|
|
opts?: { cwd?: string; env?: Record<string, string>; timeout?: number },
|
|
) => { exitCode: number; stdout: string; stderr: string; signal: NodeJS.Signals | null; error: Error | null };
|
|
retryRenameSync: (fromPath: string, toPath: string) => void;
|
|
};
|
|
/* eslint-enable @typescript-eslint/no-require-imports */
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Constants
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/**
|
|
* A lock older than this is a CANDIDATE for stealing (the holder may have crashed). A same-host
|
|
* lock past this age whose recorded pid is DEAD is stolen immediately (fast local recovery).
|
|
*/
|
|
const LOCK_STALE_MS = 60_000;
|
|
/**
|
|
* HARD deadman timeout. A lock older than this is stolen REGARDLESS of pid liveness or host. This is
|
|
* the only thing that can break a permanent deadlock caused by:
|
|
* - PID REUSE: a crashed holder's pid reused by an unrelated long-lived process makes isPidAlive
|
|
* return true forever, so the dead-pid fast-recovery branch never fires.
|
|
* - CROSS-HOST (NFS): a remote holder's pid is meaningless to local process.kill(pid,0), so liveness
|
|
* cannot be judged at all — only the deadman can reclaim such a lock.
|
|
* Much larger than LOCK_STALE_MS so a genuinely slow-but-live SAME-host holder is given a wide grace
|
|
* window (it is protected by the same-host liveness check until then); 10 minutes is far longer than
|
|
* any real sub-second capability fs critical section.
|
|
*/
|
|
const LOCK_DEADMAN_MS = 600_000;
|
|
/**
|
|
* The lockfile body is UNTRUSTED content. A well-formed lock body is a tiny JSON object. The body is
|
|
* read via the shared fd-based bounded reader (open → fstat → require a REGULAR file → enforce this
|
|
* size cap → read exactly size). A non-regular/oversized body is treated as UNPARSEABLE → routed to the
|
|
* deadman policy (cannot verify liveness → steal only after the deadman). 64 KiB is orders of magnitude
|
|
* larger than any legitimate lock body.
|
|
*/
|
|
const LOCK_MAX_BODY_BYTES = 64 * 1024;
|
|
/**
|
|
* DEFAULT bounded steal/retry attempts so a pathological never-acquirable lock cannot recurse forever.
|
|
* A caller may raise it (the consent store passes a larger budget — two genuinely-racing same-machine
|
|
* consent writers must SERIALIZE, not fail, before the lock-acquire-failure throw kicks in #1459
|
|
* finding 3). The lifecycle's sub-second critical section is happy with the small default.
|
|
*/
|
|
const LOCK_MAX_ATTEMPTS = 8;
|
|
const LOCK_RETRY_BACKOFF_MS = 25;
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Types
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/**
|
|
* A held lock: the lockfile path, the unique OWNER TOKEN we wrote into it, and the (dev, ino) of the
|
|
* lockfile inode captured at acquire. releaseLock re-confirms BOTH the token AND the captured dev/ino
|
|
* still match the path on disk immediately before rmSync, so a successor lock that replaced ours at the
|
|
* same path (different inode) is never deleted. dev/ino are null when the post-create stat could not be
|
|
* taken (best-effort) — then release falls back to the token check alone.
|
|
*/
|
|
interface LockHandle { path: string; token: string; dev: number | null; ino: number | null; }
|
|
|
|
/**
|
|
* Parsed view of a lockfile body. `hostname` is null for a legacy lock (no hostname recorded) — treated
|
|
* as SAME-host (conservative, backward compatible). `startTime` is the holder process's recorded
|
|
* start-time; null for a legacy lock or one whose body did not record it — a null recorded start-time
|
|
* cannot be matched, so liveness cannot be verified and the holder is treated as NOT verified-live.
|
|
*/
|
|
interface ParsedLock { pid: number | null; hostname: string | null; startTime: string | null; ts: number | null; }
|
|
|
|
/** Per-body IDENTITY used to confirm the lock being stolen is still the same instance just before steal. */
|
|
interface LockIdentity { dev: number | null; ino: number | null; ts: number | null; }
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Tokens + backoff
|
|
// ---------------------------------------------------------------------------
|
|
|
|
let _lockSeq = 0;
|
|
/**
|
|
* A per-acquire unique token so release is owner-safe (never deletes a successor's lock). The FIRST
|
|
* `-`-delimited segment is the holder PID — acquireLock parses it back out to check liveness before
|
|
* stealing a stale lock.
|
|
*/
|
|
function newLockToken(): string {
|
|
return `${process.pid}-${Date.now()}-${++_lockSeq}`;
|
|
}
|
|
|
|
let _lockSleepBuf: Int32Array | null = null;
|
|
function lockBackoff(): void {
|
|
// Small jittered backoff between steal attempts (yields the thread via Atomics.wait).
|
|
if (_lockSleepBuf === null) _lockSleepBuf = new Int32Array(new SharedArrayBuffer(4));
|
|
const jitter = Math.floor(Math.random() * LOCK_RETRY_BACKOFF_MS);
|
|
Atomics.wait(_lockSleepBuf, 0, 0, LOCK_RETRY_BACKOFF_MS + jitter);
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Body parse / age / host
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/**
|
|
* Parse the holder PID from a legacy plain-token lockfile body (the first `-`-delimited segment).
|
|
* Returns null when the body has no numeric leading segment (e.g. JSON content, or legacy no-pid).
|
|
*/
|
|
function lockHolderPid(body: string): number | null {
|
|
const seg = body.split('-')[0];
|
|
if (!/^\d+$/.test(seg)) return null;
|
|
const pid = Number(seg);
|
|
return Number.isInteger(pid) && pid > 0 ? pid : null;
|
|
}
|
|
|
|
/**
|
|
* Parse a lockfile body into { pid, hostname, startTime, ts }. The new format is JSON
|
|
* `{ token, pid, hostname, startTime, ts }`; a legacy body is a plain `pid-ts-seq` token (or
|
|
* non-numeric junk). Never throws — unparseable content yields all-null.
|
|
*
|
|
* `ts` is the body's OWN recorded timestamp. The age decision is bound to `now - ts` (a FRESH
|
|
* replacement body carries a FRESH ts → small age → not stolen), NOT to the file `mtime`. A legacy/
|
|
* no-`ts` body yields ts:null and the caller falls back to the file `mtime` age.
|
|
*/
|
|
function parseLockBody(body: string): ParsedLock {
|
|
const trimmed = body.trim();
|
|
if (trimmed.startsWith('{')) {
|
|
try {
|
|
const parsed: unknown = JSON.parse(trimmed);
|
|
if (parsed && typeof parsed === 'object' && !Array.isArray(parsed)) {
|
|
const p = parsed as Record<string, unknown>;
|
|
const pidVal = p['pid'];
|
|
const pid = typeof pidVal === 'number' && Number.isInteger(pidVal) && pidVal > 0 ? pidVal : null;
|
|
const hostVal = p['hostname'];
|
|
const hostname = typeof hostVal === 'string' && hostVal ? hostVal : null;
|
|
const stVal = p['startTime'];
|
|
const startTime = typeof stVal === 'string' && stVal ? stVal : null;
|
|
const tsVal = p['ts'];
|
|
const ts = typeof tsVal === 'number' && Number.isFinite(tsVal) ? tsVal : null;
|
|
return { pid, hostname, startTime, ts };
|
|
}
|
|
} catch { /* fall through to legacy parse */ }
|
|
}
|
|
// Legacy plain-token body: hostname/startTime/ts were never recorded → null.
|
|
return { pid: lockHolderPid(trimmed), hostname: null, startTime: null, ts: null };
|
|
}
|
|
|
|
/**
|
|
* Derive the lock AGE (ms) from the body's own `ts` when trustworthy, else fall back to the file
|
|
* `mtime`. A `ts` is distrusted when it is in the FUTURE (planted body / clock-skewed writer): a
|
|
* trusted future `ts` would keep age <= LOCK_STALE_MS forever → permanent block. A future `mtime` is
|
|
* likewise distrusted past a half-stale-window jitter tolerance → MAX_SAFE_INTEGER so the lock routes
|
|
* into the normal steal decision tree (verified-live holders are still protected there).
|
|
*/
|
|
function lockAgeMs(ts: number | null, mtimeMs: number): number {
|
|
if (ts !== null) {
|
|
const age = Date.now() - ts;
|
|
if (age >= 0 && age <= Number.MAX_SAFE_INTEGER) return age;
|
|
}
|
|
const mtimeAge = Date.now() - mtimeMs;
|
|
if (mtimeAge >= 0) return mtimeAge;
|
|
return mtimeAge >= -(LOCK_STALE_MS / 2) ? 0 : Number.MAX_SAFE_INTEGER;
|
|
}
|
|
|
|
/** Is the parsed lock from THIS host? A null (legacy) hostname is treated as same-host. */
|
|
function isSameHost(parsed: ParsedLock): boolean {
|
|
return parsed.hostname === null || parsed.hostname === os.hostname();
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Process start-time (the pid-reuse discriminator)
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/**
|
|
* Best-effort process start-time for `pid`, as an OPAQUE platform-specific string used ONLY for
|
|
* equality comparison (never parsed as a date). The pair (pid, startTime) uniquely identifies a
|
|
* process instance: even if a crashed holder's pid is REUSED, the new process's start-time differs.
|
|
* Returns null on ANY error / unobtainable value (liveness cannot be VERIFIED → steal-eligible past
|
|
* the deadman). The shell-outs only run on the rare STEAL-decision path, never the happy path.
|
|
*/
|
|
function getProcessStartTime(pid: number): string | null {
|
|
if (!Number.isInteger(pid) || pid <= 0) return null;
|
|
try {
|
|
if (process.platform === 'linux') {
|
|
const stat = fs.readFileSync(`/proc/${pid}/stat`, 'utf8');
|
|
const rparen = stat.lastIndexOf(')');
|
|
if (rparen === -1) return null;
|
|
const rest = stat.slice(rparen + 1).trim().split(/\s+/);
|
|
const starttime = rest[19]; // overall field 22 → index 19 after comm.
|
|
return typeof starttime === 'string' && /^\d+$/.test(starttime) ? starttime : null;
|
|
}
|
|
if (process.platform === 'win32') {
|
|
const res = execTool(
|
|
'powershell',
|
|
['-NoProfile', '-NonInteractive', '-Command', `(Get-Process -Id ${pid}).StartTime.Ticks`],
|
|
{ timeout: 5_000 },
|
|
);
|
|
if (res.exitCode !== 0 || res.error) return null;
|
|
const out = res.stdout.trim();
|
|
return /^\d+$/.test(out) ? out : null;
|
|
}
|
|
const res = execTool('ps', ['-p', String(pid), '-o', 'lstart='], { timeout: 5_000 });
|
|
if (res.exitCode !== 0 || res.error) return null;
|
|
const out = res.stdout.trim();
|
|
return out ? out : null;
|
|
} catch {
|
|
return null;
|
|
}
|
|
}
|
|
|
|
/** THIS process's start-time, captured ONCE at module load so we never re-shell on every lock write. */
|
|
const _selfStartTime: string | null = getProcessStartTime(process.pid);
|
|
|
|
/** Serialize the lockfile body: JSON carrying the owner token, pid, hostname, cached start-time, ts. */
|
|
function lockFileBody(token: string): string {
|
|
return JSON.stringify({ token, pid: process.pid, hostname: os.hostname(), startTime: _selfStartTime, ts: Date.now() });
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Liveness probes (test seam)
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/** Is `pid` a live process? process.kill(pid, 0) succeeds for a live (signalable) process. */
|
|
function _realIsPidAlive(pid: number): boolean {
|
|
try {
|
|
process.kill(pid, 0);
|
|
return true; // signalable → alive
|
|
} catch (err) {
|
|
// EPERM means the process exists but we cannot signal it (still ALIVE). ESRCH means it's gone.
|
|
return (err as NodeJS.ErrnoException).code === 'EPERM';
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Test seams: the steal-decision path goes through these indirections so unit tests can mock liveness +
|
|
* process start-time DETERMINISTICALLY. The defaults are the real implementations.
|
|
*/
|
|
const _lockProbes: {
|
|
isPidAlive: (pid: number) => boolean;
|
|
getProcessStartTime: (pid: number) => string | null;
|
|
} = { isPidAlive: _realIsPidAlive, getProcessStartTime };
|
|
|
|
function isPidAlive(pid: number): boolean {
|
|
return _lockProbes.isPidAlive(pid);
|
|
}
|
|
|
|
/**
|
|
* Is the recorded SAME-host holder VERIFIED-LIVE? True ONLY when ALL hold: the pid signals alive AND
|
|
* the lock recorded a non-null start-time AND the pid's CURRENT observed start-time matches that
|
|
* recorded value. Any failure — dead pid, no recorded start-time, unobtainable current start-time, or a
|
|
* MISMATCH (= pid-reuse) — means NOT verified-live, so the holder may be stolen. This defeats pid-reuse
|
|
* WITHOUT ever stealing a genuinely-live holder.
|
|
*/
|
|
function holderVerifiedLive(parsed: ParsedLock): boolean {
|
|
if (parsed.pid === null) return false;
|
|
if (!isPidAlive(parsed.pid)) return false;
|
|
if (parsed.startTime === null) return false;
|
|
const observed = _lockProbes.getProcessStartTime(parsed.pid);
|
|
if (observed === null) return false;
|
|
return observed === parsed.startTime;
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Bounded body read + identity recheck
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/**
|
|
* Parse the lockfile body via the SHARED fd-based bounded reader. The body is untrusted: a FIFO/device/
|
|
* oversized/garbage body returns all-null (routed to the deadman policy). Never throws.
|
|
*/
|
|
function readParsedLockBounded(lockPath: string): ParsedLock {
|
|
const allNull: ParsedLock = { pid: null, hostname: null, startTime: null, ts: null };
|
|
try {
|
|
const body = ledgerMod.readSmallRegularFile(lockPath, LOCK_MAX_BODY_BYTES);
|
|
if (body === null) return allNull; // vanished/missing — cannot verify anything.
|
|
return parseLockBody(body);
|
|
} catch {
|
|
return allNull; // non-regular / oversized / unreadable untrusted body → unparseable.
|
|
}
|
|
}
|
|
|
|
/**
|
|
* The per-body IDENTITY used to confirm, immediately before the atomic rename-steal, that the lock the
|
|
* acquirer decided to steal is STILL the same body instance. Binds (dev, ino) from a fresh stat AND the
|
|
* body's own `ts` (when JSON). A null on any field means we could not read it → caller treats it as
|
|
* "changed" and retries rather than stealing. Never throws.
|
|
*/
|
|
function lockIdentity(lockPath: string): LockIdentity {
|
|
let dev: number | null = null;
|
|
let ino: number | null = null;
|
|
try {
|
|
const st = fs.statSync(lockPath);
|
|
dev = typeof st.dev === 'number' ? st.dev : null;
|
|
ino = typeof st.ino === 'number' ? st.ino : null;
|
|
} catch {
|
|
return { dev: null, ino: null, ts: null }; // vanished/unstatable — treat as changed.
|
|
}
|
|
const ts = readParsedLockBounded(lockPath).ts;
|
|
return { dev, ino, ts };
|
|
}
|
|
|
|
/**
|
|
* Two lock identities refer to the SAME body instance only when dev AND ino match AND the `ts` is
|
|
* unchanged. A null dev/ino on EITHER side is a CHANGE (fail-safe: do not steal). If the DECISION body
|
|
* had a non-null JSON `ts`, the recheck body MUST carry the SAME non-null `ts` (a disappearing ts is a
|
|
* CHANGE → do not steal, retry).
|
|
*/
|
|
function sameLockInstance(a: LockIdentity, b: LockIdentity): boolean {
|
|
if (a.dev === null || a.ino === null || b.dev === null || b.ino === null) return false;
|
|
if (a.dev !== b.dev || a.ino !== b.ino) return false;
|
|
if (a.ts !== null && a.ts !== b.ts) return false;
|
|
return true;
|
|
}
|
|
|
|
/** Extract the owner token from a lockfile body (JSON `token` field), or null if not JSON/absent. */
|
|
function lockBodyToken(body: string): string | null {
|
|
const trimmed = body.trim();
|
|
if (!trimmed.startsWith('{')) return null;
|
|
try {
|
|
const parsed: unknown = JSON.parse(trimmed);
|
|
if (parsed && typeof parsed === 'object' && !Array.isArray(parsed)) {
|
|
const t = (parsed as Record<string, unknown>)['token'];
|
|
return typeof t === 'string' ? t : null;
|
|
}
|
|
} catch { /* not JSON */ }
|
|
return null;
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Acquire / release
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/**
|
|
* Acquire an exclusive lock at `lockPath` (a single lockfile created with O_EXCL), stamping a JSON body
|
|
* that records a unique owner token, our PID, our HOSTNAME, our process START-TIME, and a timestamp.
|
|
* The containing directory is mkdir'd (recursive, best-effort). Returns a LockHandle on success, or
|
|
* null if another LIVE operation holds it / the attempt budget is exhausted.
|
|
*
|
|
* `opts.maxAttempts` raises the bounded steal/retry budget (default LOCK_MAX_ATTEMPTS) so a caller with
|
|
* legitimately-contended writers (the consent store) can SERIALIZE rather than fail under brief
|
|
* contention. The budget is always bounded — no unbounded recursion.
|
|
*
|
|
* `opts.waitForFresh` (consent store) changes the BLOCKED-held disposition: when a held lock is NOT
|
|
* steal-eligible (fresh under the stale window, a verified-live same-host holder, or an unverifiable
|
|
* holder under the deadman), the DEFAULT (lifecycle) returns null IMMEDIATELY (fail-fast — the caller
|
|
* does not retry). With waitForFresh the acquirer instead BACKS OFF AND RETRIES (within the bounded
|
|
* budget) so two genuinely-racing same-machine writers SERIALIZE — the loser waits for the holder to
|
|
* release its sub-ms critical section and then wins the O_EXCL create. It still returns null once the
|
|
* budget is exhausted (then #1459 finding 3 turns that into a throw rather than an unlocked write). This
|
|
* NEVER steals a non-steal-eligible holder — it only WAITS for it; the steal protocol is unchanged.
|
|
*
|
|
* The steal itself is atomic (rename-then-recreate, so only ONE racing process can rename the inode),
|
|
* and the whole thing is a BOUNDED iterative loop.
|
|
*/
|
|
function acquireLock(lockPath: string, opts?: { maxAttempts?: number; waitForFresh?: boolean }): LockHandle | null {
|
|
// A genuine failure here (EACCES/ENOSPC/EROFS) MUST surface immediately, matching the #1884
|
|
// fix (0c43d853e / PR #3472) for withPlanningLock's identical shape. The prior
|
|
// `catch { /* best-effort */ }` swallowed it, and the subsequent `fs.openSync(lockPath, 'wx')`
|
|
// below then failed with ENOENT (parent dir missing) — which is NOT 'EEXIST', so the
|
|
// `if (code !== 'EEXIST') return null;` branch laundered a fatal filesystem error into an
|
|
// ordinary "lock unavailable" (null) result, indistinguishable from another live process
|
|
// legitimately holding the lock (#3987). `mkdirSync(recursive:true)` does not throw when the
|
|
// directory already exists, so the normal path (dir already present) is unaffected; only real
|
|
// creation failures propagate.
|
|
fs.mkdirSync(path.dirname(lockPath), { recursive: true });
|
|
const maxAttempts = (opts && Number.isInteger(opts.maxAttempts) && (opts.maxAttempts as number) > 0)
|
|
? (opts.maxAttempts as number)
|
|
: LOCK_MAX_ATTEMPTS;
|
|
const waitForFresh = !!(opts && opts.waitForFresh);
|
|
// A held lock that is NOT steal-eligible: fail-fast (return null) by default, or BACK OFF + RETRY
|
|
// (continue) when waitForFresh and a retry budget remains — so a contended consent writer serializes.
|
|
const blocked = (attempt: number): LockHandle | null | 'retry' => {
|
|
if (waitForFresh && attempt + 1 < maxAttempts) { lockBackoff(); return 'retry'; }
|
|
return null;
|
|
};
|
|
|
|
for (let attempt = 0; attempt < maxAttempts; attempt++) {
|
|
const token = newLockToken();
|
|
try {
|
|
const fd = fs.openSync(lockPath, 'wx'); // exclusive create — fails if held
|
|
// Once the exclusive create SUCCEEDS, a writeSync/closeSync failure must NOT leave the empty
|
|
// lockfile behind — an orphan body self-blocks every later acquirer until the deadman. On any
|
|
// write/close error, best-effort unlink the file we just created and bail. fs.writeFileSync(fd, …)
|
|
// flushes the WHOLE buffer (no short-write) unlike a bare fs.writeSync.
|
|
try {
|
|
fs.writeFileSync(fd, lockFileBody(token));
|
|
} catch (writeErr) {
|
|
try { fs.closeSync(fd); } catch { /* best-effort */ }
|
|
try { fs.unlinkSync(lockPath); } catch { /* best-effort — no orphan */ }
|
|
throw writeErr;
|
|
}
|
|
try {
|
|
fs.closeSync(fd);
|
|
} catch (closeErr) {
|
|
try { fs.unlinkSync(lockPath); } catch { /* best-effort — no orphan */ }
|
|
throw closeErr;
|
|
}
|
|
// Capture the lock inode's (dev, ino) so releaseLock can confirm, immediately before rmSync, that
|
|
// the path still holds OUR inode. Best-effort: a null dev/ino just falls back to the token check.
|
|
let dev: number | null = null;
|
|
let ino: number | null = null;
|
|
try {
|
|
const lst = fs.statSync(lockPath);
|
|
dev = typeof lst.dev === 'number' ? lst.dev : null;
|
|
ino = typeof lst.ino === 'number' ? lst.ino : null;
|
|
} catch { /* best-effort — release falls back to the token check alone */ }
|
|
return { path: lockPath, token, dev, ino };
|
|
} catch (err) {
|
|
// EEXIST → held (fall through to the steal decision). Any other error here is the create failing
|
|
// for a real reason OR a write/close failure we already cleaned up → bail out.
|
|
if ((err as NodeJS.ErrnoException).code !== 'EEXIST') return null;
|
|
}
|
|
// Held — decide whether to steal.
|
|
let st: fs.Stats;
|
|
try {
|
|
st = fs.statSync(lockPath);
|
|
} catch {
|
|
continue; // lock vanished between open and stat — retry the create immediately.
|
|
}
|
|
|
|
// Bind the age decision to the SAME body instance we act on. Parse the (bounded) body ONCE; derive
|
|
// age from the body's own `ts` for a JSON body so a FRESH replacement (fresh ts) is seen as fresh
|
|
// even if the file `mtime` is stale-old. A legacy/garbage/no-`ts` body — and a FUTURE/implausible
|
|
// `ts` — falls back to the file `mtime` age so a planted/clock-skewed future ts can never deadlock.
|
|
const parsed = readParsedLockBounded(lockPath);
|
|
const age = lockAgeMs(parsed.ts, st.mtimeMs);
|
|
if (age <= LOCK_STALE_MS) { // genuinely held (fresh) — blocked.
|
|
const b = blocked(attempt);
|
|
if (b === 'retry') continue;
|
|
return b;
|
|
}
|
|
|
|
const decisionIdentity: LockIdentity = {
|
|
dev: typeof st.dev === 'number' ? st.dev : null,
|
|
ino: typeof st.ino === 'number' ? st.ino : null,
|
|
ts: parsed.ts,
|
|
};
|
|
|
|
if (isSameHost(parsed) && parsed.pid !== null) {
|
|
// SAME host with a parseable pid → we CAN verify liveness via the (pid, start-time) pair. A
|
|
// VERIFIED-LIVE holder is NEVER stolen — even past the deadman. Otherwise → steal.
|
|
if (holderVerifiedLive(parsed)) { // provably-live same-host holder — blocked.
|
|
const b = blocked(attempt);
|
|
if (b === 'retry') continue;
|
|
return b;
|
|
}
|
|
// else fall through to the atomic steal.
|
|
} else {
|
|
// DIFFERENT host, or no parseable pid → liveness cannot be verified locally. Only the deadman can
|
|
// reclaim it; under the deadman, leave it (blocked).
|
|
if (age <= LOCK_DEADMAN_MS) {
|
|
const b = blocked(attempt);
|
|
if (b === 'retry') continue;
|
|
return b;
|
|
}
|
|
// else (age > deadman) → fall through to the atomic steal.
|
|
}
|
|
|
|
// Re-stat + re-read the body IMMEDIATELY before the rename and confirm it is the SAME instance
|
|
// (dev/ino unchanged AND, for a JSON body, ts unchanged). If a racer stole+recreated a FRESH lock
|
|
// between our decision and now, the identity differs → do NOT steal; RETRY the bounded loop.
|
|
if (!sameLockInstance(decisionIdentity, lockIdentity(lockPath))) {
|
|
if (attempt + 1 < maxAttempts) lockBackoff();
|
|
continue; // the body changed under us — re-evaluate from scratch rather than steal a replacement.
|
|
}
|
|
|
|
// Steal atomically (only one racer can rename the inode).
|
|
const stolen = `${lockPath}.stale-${process.pid}-${Date.now()}-${crypto.randomBytes(4).toString('hex')}`;
|
|
try { retryRenameSync(lockPath, stolen); } catch { return null; } // another process won the steal
|
|
try { fs.rmSync(stolen, { force: true }); } catch { /* best-effort */ }
|
|
if (attempt + 1 < maxAttempts) lockBackoff();
|
|
}
|
|
return null; // attempt budget exhausted (pathological contention) — never throws/recurses.
|
|
}
|
|
|
|
/**
|
|
* Release a lock only if it still carries our owner token (PRIMARY discriminator) — and, as a best-
|
|
* effort SECONDARY check, if its inode still matches the (dev, ino) we captured at acquire, so the
|
|
* common path never deletes a lock that was stale-stolen out from under us.
|
|
*
|
|
* The TOKEN re-check is the load-bearing protection: a real successor wrote a DIFFERENT token, so we
|
|
* read a non-matching token and refuse to delete on every filesystem. The dev/ino recheck is best-
|
|
* effort secondary hardening (may be defeated by inode reuse on some filesystems). The body is read via
|
|
* the bounded reader so a FIFO/oversized body at the path is never read or deleted by us.
|
|
*/
|
|
function releaseLock(handle: LockHandle | null): void {
|
|
if (!handle) return;
|
|
try {
|
|
let body: string | null;
|
|
try {
|
|
body = ledgerMod.readSmallRegularFile(handle.path, LOCK_MAX_BODY_BYTES);
|
|
} catch {
|
|
return; // non-regular / oversized / unreadable → not ours; do not read or delete.
|
|
}
|
|
if (body === null) return; // gone / missing — nothing of ours to release.
|
|
// The body is JSON `{ token, … }`; release only if the recorded token is still OURS. A legacy
|
|
// plain-token body (whole body === token) is also honored.
|
|
if (lockBodyToken(body) !== handle.token && body !== handle.token) return; // not our token (PRIMARY).
|
|
if (handle.dev !== null && handle.ino !== null) {
|
|
let cur: fs.Stats;
|
|
try {
|
|
cur = fs.statSync(handle.path);
|
|
} catch {
|
|
return; // vanished/unstatable between read and rmSync → nothing of ours to release.
|
|
}
|
|
if (cur.dev !== handle.dev || cur.ino !== handle.ino) return; // successor inode — not ours.
|
|
}
|
|
fs.rmSync(handle.path, { force: true });
|
|
} catch { /* already gone / stale-stolen / unreadable — nothing of ours to release */ }
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Exports
|
|
// ---------------------------------------------------------------------------
|
|
|
|
export = {
|
|
acquireLock,
|
|
releaseLock,
|
|
getProcessStartTime,
|
|
LOCK_STALE_MS,
|
|
LOCK_DEADMAN_MS,
|
|
LOCK_MAX_BODY_BYTES,
|
|
// Test seams (shared by capability-lifecycle's #1462 lock tests via re-export): inject deterministic
|
|
// isPidAlive / getProcessStartTime so the start-time liveness branches are exercised without real pids.
|
|
_setLockProbes(probes: Partial<{ isPidAlive: (pid: number) => boolean; getProcessStartTime: (pid: number) => string | null }>): void {
|
|
if (typeof probes.isPidAlive === 'function') _lockProbes.isPidAlive = probes.isPidAlive;
|
|
if (typeof probes.getProcessStartTime === 'function') _lockProbes.getProcessStartTime = probes.getProcessStartTime;
|
|
},
|
|
_resetLockProbes(): void {
|
|
_lockProbes.isPidAlive = _realIsPidAlive;
|
|
_lockProbes.getProcessStartTime = getProcessStartTime;
|
|
},
|
|
};
|