Commit Graph

3 Commits

Author SHA1 Message Date
Tom Boucher
e22596be04 fix(#471): make perf-407 lock-buffer-alloc test deterministic via clock-seam; remove real-worker race (#472)
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 13:20:32 -04:00
Tom Boucher
199251c0f6 test(#432): make perf-316/perf-407 lock-race tests deterministic (#450)
The perf-316 and perf-407 regression tests had two race-driven failure modes:

1. False-fail: setTimeout(80) before assert.ok(fs.existsSync(lockPath)) could
   fire before Worker A finished writing the lock file under CI load. The
   handoff noted this fired on origin/next's coverage job.

2. False-pass (latent): if Worker B raced past Worker A's release before
   retrying, sabCount === 1 from the no-retry success path matched what
   the PRE-FIX buggy code produced (one SAB for the single successful
   open/write). The existing test had no witness for retry-path coverage,
   violating test-rigor Contract 4 (exercise the path you claim to cover).

Fix (test-only):

- Replace setTimeout(80) with await-{pid}-message synchronization. Worker A
  posts {pid} AFTER fs.writeFileSync/openSync returns (single-thread source
  order within the worker), so by the time the parent receives it the lock
  file exists on disk. The MessagePort buffers messages posted before the
  parent attaches its listener, so there is no listener-race.
  Ref: https://nodejs.org/api/worker_threads.html#event-message_1

- Add a 5000ms safety timeout on the lock-written signal so a hung Holder
  worker surfaces as a specific error, not a generic test-timeout.

- Add a fs.openSync/fs.writeFileSync stub in Worker B that counts atomic-
  create attempts (O_CREAT|O_EXCL for perf-316 / { flag: 'wx' } for perf-407).
  Assert lockAttempts >= 2 to prove the SUT entered the retry loop. This
  closes the Contract-4 hole: sabCount === 1 now discriminates pre-fix
  (sabCount === lockAttempts) from post-fix (sabCount === 1, hoisted).

- Bump holdMs from 400ms to 1000ms to guarantee >=4 retries (perf-316,
  200ms+jitter delay) or >=9 retries (perf-407, 100ms delay) on the
  slowest CI worker. Test wall time grows ~600ms; well under the existing
  8000ms timeout.

Closes #432

Co-authored-by: CI Rebase Check <ci@open-gsd.dev>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-28 20:10:08 -04:00
Tom Boucher
ec0a32ea07 fix(#407): hoist sleep SharedArrayBuffer out of withPlanningLock retry loop (#418) 2026-05-27 22:48:31 -04:00