Files
msd-core/tests
Dave 510f37661a fix(core): acquireStateLock must not leak fd + orphan lock on write error (M9)
Root cause: in acquireStateLock, once openSync(O_CREAT|O_EXCL) created the
lock file, the subsequent writeSync(pid)/closeSync were unguarded. A
recoverable errno (EAGAIN etc., in ACQUIRE_LOCK_RETRY_ERRNOS) made the catch
do checkBudgetAndSleep + continue WITHOUT closing the fd or unlinking the
just-created empty lock — leaking a descriptor every occurrence and stranding
a content-less lock (the #500/#905/#1230 STATE.md write-corruption family).

Fix: wrap writeSync/closeSync in an inner try that guardedly closeSync(fd) +
unlinkSync(lockPath) then re-throws to the existing outer catch (DRY errno
classification). Recoverable errno retries from a clean slate; a FATAL errno
(e.g. ENOSPC, not recoverable) still propagates after cleanup — not masked.
Mirrors the already-shipped capability-lock.cts:415-425 pattern.

Extends the M8 test seam with a one-shot simulateWriteError errno + an
onLoopIteration snapshot hook so the orphan-before-retry is deterministically
observable. New tests prove RED (orphan stranded / fatal leaves orphan)
before the cleanup and GREEN after.

Source of truth src/state.cts (ADR-457); bin/lib/state.cjs is generated.

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz
2026-06-21 13:11:22 -04:00
..