Files
msd-core/tests/external-job.test.cjs
Tom Boucher 53ea8e0664 fix(#3057): make a guard's failure distinguishable from its benign result — Wave 1 (#3088)
* fix(#3057): refuse the write when the duplicate scan cannot complete

writeManifest documents itself as a fail-closed duplicate guard: if any
existing manifest shares plan_id with a different, non-terminal job_id it must
refuse, because dispatching again would duplicate the external job.

It could not honour that. The scan reads every sibling manifest looking for the
duplicate, and an unreadable or unparseable sibling was `continue`d past. If
the corrupt file was the one holding the live duplicate, the scan found nothing
and a duplicate external job dispatched.

The asymmetry is what gives it away: a malformed TARGET refused with
malformed_existing because clobbering is unacceptable, while a malformed
SIBLING was skipped — yet siblings are the only thing the duplicate check
reads.

Adds a scan_incomplete verdict that refuses and names the offending file, so an
operator can quarantine or repair it. Fail-closed alone would let one stale
corrupt manifest wedge every dispatch for that planning dir permanently; naming
the file is what makes refusing survivable. malformed_existing is untouched, so
the target/sibling distinction stays visible. The docstring is updated — it
previously stated a rule the function did not keep.

memFs() gains an optional failReads map so these branches are reachable at all;
they had zero coverage because the fake could not express a per-file read
fault. The signature is additive and every existing caller is unchanged.

The regression is proved by a pair, not a single test. A control writes a
readable sibling holding a genuine non-terminal duplicate and asserts
duplicate_plan_id, establishing the scenario is real; the regression then makes
that same path unreadable and asserts scan_incomplete. A first draft of this
test used a corrupt-JSON fixture containing no plan_id at all while its comment
claimed otherwise — it duplicated the unparseable-sibling case and proved
nothing, which is the defect class this phase exists to remove.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3057): make a guard's failure distinguishable from its benign result

Wave 1 of the negative-space backfill: the branches where a guard that could
not verify something reported the same value it reports when everything is
fine. That indistinguishability is the defect; every fix here makes the two
states tellable apart, and every test proves it with a pair — one for the
failure, one for the benign case. A single test cannot establish that two
states are distinguishable, which is the whole property being fixed.

state.cts phaseInventoryProvider returned null for both a real disk-scan
failure and a genuinely empty phases dir, so `state rebuild` could report
success while phase-table reconciliation never ran. It now returns a
discriminated result and the CLI surfaces phase_inventory_scan_failed plus a
reason. The reason field turned out never to have been wired into the emitted
JSON at all — it existed only as an internal variable — so a test could only
assert on the operator-facing note. It is a real field now.

state.cts treated an unreadable lock body the same as an empty one, applying
the 1-second stealable floor. A lock we cannot read is not a lock we know is
stale; an unreadable body is now held to the deadman ceiling like a live
holder.

verification.cts findStaleVerificationSummary returned null on any fs, scan or
clock failure — meaning "not stale". It now returns a discriminated
StaleCheckResult and the caller records that the check was indeterminate.

git-base-branch resolveBaseBranch returned 'main' both when no candidate branch
existed and when every git tier timed out. A diagnostics variant now reports
whether the answer was verified, and the CLI writes an unverified-fallback note
to stderr. The stdout contract five workflows parse is untouched.

worktree-safety snapshotWorktreeInventory left exists:true when statSync threw,
so a guard that could not check reported the worktree present; exists is now
tri-state and a stat failure surfaces as an 'unverified' finding.
planWorktreePrune reported 'no_worktrees' for a parse failure, which is not the
same as an empty list — and it drives a prune. It now reports 'parse_failed'.

Fixing the inventory change exposed a second fail-open in verify.cts: the
validate-health consumer silently dropped findings whose kind it did not
recognise, so the new kind would have vanished. That is closed too — worth
noting that the survey enumerated producers of degraded verdicts, not consumers
that discard them.

worktree-base-ref and state-transition gain the distinguishing signal without
changing what they do: headAbsenceVerified, and a phase-inventory scan meta.
Whether those guards should ACT differently is a product question this change
does not answer, and both are flagged rather than quietly settled.

rescueSummaryArtifacts is left alone: rescuing on an uncertain cat-file is
deliberate per #2556. It now has tests proving it, and a recorded negative
finding — git cat-file -e returns 128 for both "absent from HEAD" and a fatal
error, so "uncertain" and "certain-and-fine" are not separable at the git
level.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3057): assert typed values, not rendered text

Ten assertions in the rebuild CLI suite matched substrings of produced output —
STATE.md body fields, a markdown table row, an audit-log heading, and JSON keys
read as text. CONTRIBUTING prohibits that: if the code under test produces
text, the test asserts on its structured surface instead.

No production surface had to be built. Every one already existed and was
already compiled into bin/lib: stateExtractField for body fields,
parseMarkdownTable for the phase table, collectSection for the audit-log
section, and result.data.log — already a typed RebuildLogEntry[]. The tests
were matching rendered text sitting next to the structured data.

One of those assertions was passing for the wrong reason. `stdout.includes
('rebuilt')` matched the JSON KEY name, not a value: the dry-run path emits
`mutated` and the real path emits `rebuilt`, so it would have passed whether
the value was true or false. It now asserts the value.

external-job's refusal already had to name the offending file — that naming is
why the fail-closed variant is survivable rather than a permanent wedge — but
the tests proved it by substring of a prose message. The failure result now
carries offendingPath as its own field and the tests assert it by value. The
human message is unchanged; operators read it.

Array membership is left alone. `phaseIds.includes('99')` and
`result.updated.includes('Completed Phases')` are membership checks on real
arrays, not text matching, and converting them would weaken nothing and clarify
nothing.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3057): execute acquireStateLock instead of grepping its source

The non-EEXIST lock test asserted on the TEXT of the built .cjs and never
called acquireStateLock. It carried an allow-test-rule: architectural-invariant
exemption to permit that. A source grep proves a literal is present in a file,
not that the behaviour works — it is weaker than a liveness test, which at
least runs the code, and it was the only coverage the fatal-errno path had.

Replaced with tests that inject the errno through fs and assert what actually
happens: a fatal EACCES propagates out of acquireStateLock with zero backoff
sleeps, while EAGAIN/EINTR/EINVAL/EIO/ENOENT/ESTALE/EPERM/EBUSY retry once and
succeed. The exemption is removed and its allowlist entry with it.

One old assertion is deliberately not carried over: it checked the retryable
errnos were expressed as a Set rather than an inline literal. That is a shape
check with no runtime signature; the behavioural tests fail if the code reverts
to the old inline check, which is the regression it was really guarding.

The #3057 lock-body tests move into that same file rather than a new one, which
is what lint-test-file-count asks for and puts every acquireStateLock test in
one place.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3057): surface an indeterminate staleness check to its callers

An isolated review caught an inconsistency inside this wave. Two of the three
"add the distinguishing signal" fixes wire through to something a user sees:
git base-branch writes an unverified-fallback diagnostic to stderr, and an
unverifiable worktree surfaces as a W020 finding. The third set
staleCheckIndeterminate on readVerificationStatus's result and nothing read it.

A signal nobody consumes leaves the fail-open exactly as silent as before: the
staleness check could fail and the operator saw precisely what they would see
if the answer were genuinely "not stale". That is the defect this issue exists
to remove, so it is not defensible as scaffolding when its two siblings in the
same change already wire through.

All five callers now surface it, each through the channel it already had rather
than a mechanism imposed uniformly: phase complete adds it to its existing
warnings array and, on the blocked path, as an additive note on the error text;
init and roadmap carry it as a field on output they already emit; the UAT
report carries it without ever gating passed/blockers; workstream inventory
takes an injectable writeDiagnostic mirroring the git base-branch idiom,
because its return shape had nowhere to hang a per-phase field without
rippling the builder's types.

The routing decision is unchanged everywhere. What changes is only that a
caller and an operator can now tell a failed check from a completed one.

That diagnostic carries structured meta rather than being asserted by regex —
the default still writes only the human message to stderr, but tests assert
phaseDir and reason by value. Two earlier assertions in this branch were
converted the same way; this was the last raw-text assertion left.

Also records a scope correction: the completePhaseCore guards now compare
stateReplaceField's result to the body instead of testing truthiness, so a
field whose substitution produced identical text no longer reports as updated.
That is a real behaviour fix, not the signal-only change this file was
described as carrying, and its tests cover both the changed and unchanged
cases.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3057): bound two heavy subprocesses for a loaded bench, not an idle one

The remote matrix surfaced three failures unrelated to this branch's changes.
All were bad tests, and a re-run would have hidden every one of them.

The reviewer-flags parse block bounded bash -> node -> a full gsd-tools cold
start at 5 seconds. On a bench running thirty thousand tests in parallel that
is not a hang, it is a busy machine. Raised to 30s, matching the convention
sibling suites already use for script invocations, with a comment saying what
the budget covers so nobody tightens it back. Two further copies of the same
5-second spawn in the same file had the identical defect and are raised too —
they were not in the failure report, but they will be next time.

The fragment-propagation test bounded npm run regen:derived — a full build plus
eight generators, the heaviest subprocess in the suite — at five minutes, and
node22 was killed near the end. The captured output proves it: every generator
had written its files and gen:install-tree had emitted all fifteen runtimes
before the kill. Raised to fifteen minutes.

That failure read as `null !== 0`, which says nothing. status null means killed,
not a non-zero exit, and the two want different responses: one is a timeout to
size correctly, the other is a real build break. The assertion now distinguishes
them and names the signal.

Neither test's assertions were weakened and no retry was added. A retry here
would suppress exactly the signal the timeout exists to produce.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3057): capture fd 1 through the mock tracker, not a raw reassignment

The phase suite reported zero test results on both lanes while running for five
and a half minutes and exiting 1. No assertion text, no stderr, four events for
the whole file: enqueue, start, dequeue, complete. That shape is not a failing
assertion — it is the runner being unable to read the child at all, because it
parses its event stream from the child's stdout.

The cause was the capture helper reassigning fs.writeSync directly. Proven
rather than assumed: a standalone probe patched fs.writeSync and called
process.stdout.write, and the interception fired only when fd 1 resolved to a
FILE, not when it was a pipe. The remote runner captures the event stream to a
file, so a helper that was invisible against a pipe swallowed the reporter's own
output on the bench. That is also why the two sibling suites wired the same way
in this change pass cleanly — they use the mock tracker, the seam io.test.cjs
established for this exact function.

The helper now uses t.mock.method with an explicit restore after each call, so
teardown belongs to node:test rather than a second hand-rolled implementation,
and the interception cannot outlive the one synchronous call it wraps even if
that call throws. Ten call sites thread the test context through; three test
callbacks gained the parameter they lacked.

The three B3 tests are untouched — same assertions, same fault injection. Only
how the context reaches the helper changed.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3057): capture phase-complete output from a subprocess, not fd 1

Two attempts to make in-process fd-1 interception safe both failed on the
bench. The suite reported zero test results on either lane while exiting 1 —
four events for the whole file — because the runner parses its event stream
from the child's stdout, and process.stdout.write routes through fs.writeSync
whenever fd 1 resolves to a file, which is how the runner captures. Patching
that seam anywhere in a file can therefore destroy the file's own reporting,
and tightening the window only moved the runtime from 326s to 125s without
recovering a single event.

So the interception is gone rather than tuned. The helper now spawns gsd-tools
as a real subprocess and reads stdout the way the OS already gives it to us,
which is what the rest of the suite does. It asserts the command succeeded
before parsing, so a genuine failure can no longer present as a JSON parse
error.

The two fault-injecting tests could not survive that move as written: a
subprocess cannot see a mock installed in the parent. Instead of reinstating
the interception they now produce the fault on disk — the summary artifact is
created as a dangling symlink, so the staleness check's real statSync throws
inside the child. That is a more honest fixture than a mock in any case, since
it is a condition a user's tree can actually be in. Skipped on Windows, matching
the existing symlink precedent in the write-guard suite.

Three further call sites turned out to depend on parent-process writeFileSync
mocks the subprocess could not see. Those call the CJS function directly, which
is what they always wanted — they never needed stdout at all.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3057): one name for one signal, one encoding for one distinction

Standards review found four things this branch introduced, all of them
inconsistencies with itself rather than with the repo.

One upstream bit reached its consumers under three names —
verification_stale_check_indeterminate in two modules, the same value with
"stale" dropped in a third, and stderr only in the fourth. Standardised on the
long name wherever it is a field. The workstream inventory keeps its stderr
channel, since its return shape has nowhere to hang a per-phase field without
rippling the builder's types, but it now says the same word for the same thing.

worktree-safety encoded one three-way distinction two ways in a single file: a
named union for a finding's kind, and boolean|null for an inventory entry's
existence. The second is now a named union too.

Two assertions matched human prose because the blocked and non-blocked
completion paths carried no typed field for the signal. Both now assert typed
values. The first round of this fix added the field but left the regex beside
it, which is the banned pattern sitting next to its own replacement; the second
removed it and added an assertion on the reason enum so nothing was lost.

The remaining two were reasoned away before being fixed, and both reasons were
bad. "No typed surface exists" is the condition CONTRIBUTING says to fix by
adding one — it took three lines. "The file already does this dozens of times"
is not licence to add instance number thirty-one; a convention that violates a
documented rule is debt, not precedent.

Vocabulary differing across DIFFERENT modules is left alone: CONTEXT.md rejects
a single shared result envelope, so per-module shapes are precedented, and a
baseline smell does not outrank a documented standard.

A census of every line this branch adds to a test file now finds no regex or
substring assertion on produced prose: 87 strictEqual, 25 ok (all non-empty or
shape guards), 12 equal, 3 throws (all typed err.code predicates), 3
deepStrictEqual, 2 notStrictEqual.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3057): backfill changeset pr number to 3088

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 16:00:52 -04:00

439 lines
19 KiB
JavaScript

'use strict';
// Producer-half of the async external-job contract (#1164 / #1105).
// The CORE consumer half (external_job_waiting) is covered by
// external-job-waiting.test.cjs. These tests assert the scheduler-adapter
// Capability's pure module: SLURM state mapping, manifest build/validate,
// sbatch/squeue/sacct parsers, and the fail-closed manifest writer that
// mirrors the consumer's duplicate-execution guard
// (docs/reference/planning-artifacts.md).
const { test } = require('node:test');
const assert = require('node:assert');
const path = require('node:path');
const fc = require('fast-check');
const m = require('../gsd-core/bin/lib/external-job.cjs');
const {
MANIFEST_VERSION,
MANIFEST_STATUS,
NON_TERMINAL_STATUSES,
TERMINAL_FAILURE_STATUSES,
mapSlurmState,
buildManifest,
validateManifest,
parseSbatchParsable,
parseSqueueLine,
parseSacctRow,
writeManifest,
manifestPath,
} = m;
// ─── Closed status enum (Hyrum's Law: stability contract) ─────────────────────
test('MANIFEST_STATUS is the closed scheduler-agnostic enum from the contract', () => {
assert.deepStrictEqual([...MANIFEST_STATUS].sort(), [
'cancelled',
'completed-unverified',
'failed',
'running',
'submitted',
'timeout',
]);
assert.strictEqual(MANIFEST_VERSION, '1.0');
});
test('NON_TERMINAL / TERMINAL_FAILURE partition the enum without overlap', () => {
for (const s of MANIFEST_STATUS) {
const inNon = NON_TERMINAL_STATUSES.includes(s);
const inTerm = TERMINAL_FAILURE_STATUSES.includes(s);
// completed-unverified is neither non-terminal nor a failure — its own bucket.
if (s === 'completed-unverified') {
assert.ok(!inNon && !inTerm, 'completed-unverified is its own bucket');
} else {
assert.ok(inNon !== inTerm, `${s} must sit in exactly one partition`);
}
}
});
// ─── SLURM state mapping ──────────────────────────────────────────────────────
test('mapSlurmState maps every documented SLURM state to the closed enum', () => {
const cases = {
PENDING: 'submitted',
CONFIGURING: 'submitted',
RUNNING: 'running',
COMPLETED: 'completed-unverified',
COMPLETING: 'running',
FAILED: 'failed',
CANCELLED: 'cancelled',
TIMEOUT: 'timeout',
OUT_OF_MEMORY: 'failed',
BOOT_FAIL: 'failed',
NODE_FAIL: 'failed',
PREEMPTED: 'failed',
};
for (const [slurm, expected] of Object.entries(cases)) {
assert.strictEqual(mapSlurmState(slurm), expected, `${slurm} -> ${expected}`);
}
});
test('mapSlurmState is case-insensitive and trims whitespace', () => {
assert.strictEqual(mapSlurmState('running'), 'running');
assert.strictEqual(mapSlurmState(' PENDING '), 'submitted');
assert.strictEqual(mapSlurmState('Cancelled'), 'cancelled');
});
test('mapSlurmState returns null for unknown states (no guessing)', () => {
// Boundary: unknown must not collapse to a terminal failure silently.
assert.strictEqual(mapSlurmState('NO_SUCH_STATE'), null);
assert.strictEqual(mapSlurmState(''), null);
assert.strictEqual(mapSlurmState('COMPLETED2'), null);
});
// ─── Manifest build ───────────────────────────────────────────────────────────
function baseInput() {
return {
plan_id: '3.1',
phase: '3',
job_id: '12345',
backend: 'slurm',
submit_command: 'sbatch --parsable ./run.sh',
status: 'submitted',
expected_artifacts: ['Artifacts/jobs/12345/result.h5'],
verification_command: 'python -m verify.py 12345',
resume_command: '/gsd:execute-phase 3',
};
}
test('buildManifest stamps version and submitted_at via the clock seam', () => {
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const out = buildManifest(baseInput(), { clock });
assert.strictEqual(out.version, '1.0');
assert.strictEqual(out.submitted_at, '2020-06-15T12:00:00.000Z');
assert.strictEqual(out.terminal_details, null, 'non-terminal -> null terminal_details');
assert.strictEqual(out.plan_id, '3.1');
});
test('buildManifest rejects missing required fields', () => {
for (const key of ['plan_id', 'phase', 'job_id', 'backend', 'submit_command', 'status', 'expected_artifacts', 'verification_command', 'resume_command']) {
const bad = baseInput();
delete bad[key];
assert.throws(() => buildManifest(bad), { message: new RegExp(key) }, `missing ${key} must throw`);
}
});
test('buildManifest rejects an out-of-enum status', () => {
const bad = baseInput();
bad.status = 'done';
assert.throws(() => buildManifest(bad), /status/i);
});
test('buildManifest sets terminal_details when status is a terminal failure', () => {
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
for (const status of TERMINAL_FAILURE_STATUSES) {
const input = { ...baseInput(), status, terminal_details: { reason: 'oom', exit_code: 137 } };
const out = buildManifest(input, { clock });
assert.deepStrictEqual(out.terminal_details, { reason: 'oom', exit_code: 137 }, `${status} carries terminal_details`);
}
});
// ─── validateManifest (producer-side mirror of the trust boundary) ────────────
test('validateManifest accepts a well-formed manifest', () => {
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const res = validateManifest(buildManifest(baseInput(), { clock }));
assert.strictEqual(res.ok, true);
});
test('validateManifest rejects bad version, status, and missing plan_id', () => {
const good = buildManifest(baseInput(), { clock: { nowIso: () => '2020-06-15T12:00:00.000Z' } });
const badVersion = { ...good, version: '9.9' };
assert.strictEqual(validateManifest(badVersion).ok, false);
const badStatus = { ...good, status: 'finished' };
assert.strictEqual(validateManifest(badStatus).ok, false);
const noPlan = { ...good };
delete noPlan.plan_id;
assert.strictEqual(validateManifest(noPlan).ok, false);
});
// ─── Parsers ──────────────────────────────────────────────────────────────────
test('parseSbatchParsable parses a bare number and a number;cluster form', () => {
assert.deepStrictEqual(parseSbatchParsable('12345'), { ok: true, job_id: '12345' });
assert.deepStrictEqual(parseSbatchParsable('12345;mycluster\n'), { ok: true, job_id: '12345' });
});
test('parseSbatchParsable fails closed on empty or non-numeric output', () => {
assert.strictEqual(parseSbatchParsable('').ok, false);
assert.strictEqual(parseSbatchParsable('Submitted batch job 12345').ok, false, 'non-parsable prose rejected');
assert.strictEqual(parseSbatchParsable('abc;cluster').ok, false);
});
test('parseSqueueLine parses "<jobid> <state>" and returns null for malformed', () => {
assert.deepStrictEqual(parseSqueueLine('12345 RUNNING'), { job_id: '12345', state: 'RUNNING' });
assert.strictEqual(parseSqueueLine('header'), null);
assert.strictEqual(parseSqueueLine(''), null);
});
test('parseSacctRow parses [jobid, state] columns', () => {
assert.deepStrictEqual(parseSacctRow(['12345', 'COMPLETED']), { job_id: '12345', state: 'COMPLETED' });
assert.strictEqual(parseSacctRow(['x']), null);
assert.strictEqual(parseSacctRow([], ), null);
});
// ─── writeManifest (fail-closed duplicate guard + fs injection) ────────────────
function memFs(files = {}, opts = {}) {
const store = new Map(Object.entries(files));
const failReads = new Map(Object.entries(opts.failReads || {}));
return {
mkdirSync: () => undefined,
readdirSync: (d) => {
const set = store.get(d);
return Array.isArray(set) ? set : [];
},
readFileSync: (p) => {
if (failReads.has(p)) {
throw new Error(failReads.get(p));
}
if (!store.has(p)) { const e = new Error('enoent'); e.code = 'ENOENT'; throw e; }
return store.get(p);
},
writeFileSync: (p, c) => { store.set(p, c); },
existsSync: (p) => store.has(p),
};
}
test('manifestPath projects to .planning/async-jobs/<job>.json', () => {
assert.strictEqual(
manifestPath('.planning', '12345'),
path.join('.planning', 'async-jobs', '12345.json'),
);
});
test('writeManifest writes a new manifest and returns its path', () => {
const fs = memFs({ [path.join('.planning', 'async-jobs')]: [] });
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const manifest = buildManifest(baseInput(), { clock });
const res = writeManifest(manifest, '.planning', { fs, clock });
assert.strictEqual(res.ok, true);
assert.ok(res.path.endsWith(path.join('async-jobs', '12345.json')));
const written = JSON.parse(fs.readFileSync(res.path));
assert.strictEqual(written.plan_id, '3.1');
});
test('writeManifest allows updating the SAME job_id (status progression)', () => {
const dir = path.join('.planning', 'async-jobs');
const existingPath = path.join(dir, '12345.json');
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const submitted = buildManifest(baseInput(), { clock });
const existing = { ...submitted };
const fs = memFs({ [dir]: ['12345.json'], [existingPath]: JSON.stringify(existing) });
const running = buildManifest({ ...baseInput(), status: 'running' }, { clock });
const res = writeManifest(running, '.planning', { fs, clock });
assert.strictEqual(res.ok, true, 'same job_id progression must be allowed');
});
test('writeManifest FAILS CLOSED when a different non-terminal job exists for the same plan_id', () => {
// Duplicate-execution guard: a second dispatch for the same plan would
// duplicate the external job. Mirror of planning-artifacts.md fail-closed.
const dir = path.join('.planning', 'async-jobs');
const otherPath = path.join(dir, '99999.json');
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const other = buildManifest({ ...baseInput(), job_id: '99999' }, { clock });
const fs = memFs({ [dir]: ['99999.json'], [otherPath]: JSON.stringify(other) });
const second = buildManifest({ ...baseInput(), job_id: '12345' }, { clock });
const res = writeManifest(second, '.planning', { fs, clock });
assert.strictEqual(res.ok, false);
assert.strictEqual(res.kind, 'duplicate_plan_id');
});
test('writeManifest allows a NEW job once the prior plan_id job is terminal', () => {
const dir = path.join('.planning', 'async-jobs');
const otherPath = path.join(dir, '99999.json');
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const dead = buildManifest({ ...baseInput(), job_id: '99999', status: 'failed', terminal_details: { code: 1 } }, { clock });
const fs = memFs({ [dir]: ['99999.json'], [otherPath]: JSON.stringify(dead) });
const next = buildManifest({ ...baseInput(), job_id: '12345' }, { clock });
const res = writeManifest(next, '.planning', { fs, clock });
assert.strictEqual(res.ok, true, 'terminal prior job must not block a new dispatch');
});
test('writeManifest fails closed on a malformed existing manifest', () => {
const dir = path.join('.planning', 'async-jobs');
const brokenPath = path.join(dir, '12345.json');
const fs = memFs({ [dir]: ['12345.json'], [brokenPath]: '{not json' });
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const manifest = buildManifest(baseInput(), { clock });
const res = writeManifest(manifest, '.planning', { fs, clock });
assert.strictEqual(res.ok, false);
assert.strictEqual(res.kind, 'malformed_existing');
});
// ─── writeManifest — negative-space: an incomplete duplicate scan (#3057) ─────
// A corrupt/unreadable SIBLING manifest must not be silently skipped: the
// duplicate scan cannot be trusted to have inspected it, so the write must
// refuse rather than risk dispatching a duplicate external job.
test('an unreadable sibling manifest refuses the write', () => {
// A1
const dir = path.join('.planning', 'async-jobs');
const siblingPath = path.join(dir, '99999.json');
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const fs = memFs(
{ [dir]: ['99999.json'] },
{ failReads: { [siblingPath]: 'eio: input/output error' } },
);
const manifest = buildManifest(baseInput(), { clock });
const res = writeManifest(manifest, '.planning', { fs, clock });
assert.strictEqual(res.ok, false);
assert.strictEqual(res.kind, 'scan_incomplete');
assert.strictEqual(res.offendingPath, siblingPath, 'offendingPath must name the unreadable file');
});
test('an unparseable sibling manifest refuses the write', () => {
// A2
const dir = path.join('.planning', 'async-jobs');
const siblingPath = path.join(dir, '99999.json');
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const fs = memFs({ [dir]: ['99999.json'], [siblingPath]: '{not json' });
const manifest = buildManifest(baseInput(), { clock });
const res = writeManifest(manifest, '.planning', { fs, clock });
assert.strictEqual(res.ok, false);
assert.strictEqual(res.kind, 'scan_incomplete');
assert.strictEqual(res.offendingPath, siblingPath, 'offendingPath must name the unparseable file');
});
test('target and sibling failures stay distinct kinds', () => {
// A3
const dir = path.join('.planning', 'async-jobs');
const targetPath = path.join(dir, '12345.json');
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const fs = memFs({ [dir]: ['12345.json'], [targetPath]: '{not json' });
const manifest = buildManifest(baseInput(), { clock });
const res = writeManifest(manifest, '.planning', { fs, clock });
assert.strictEqual(res.ok, false);
assert.strictEqual(res.kind, 'malformed_existing', 'the TARGET failure must stay malformed_existing, not scan_incomplete');
});
// A4 (paired): prove the sibling really holds a duplicate (control), then
// prove that making that SAME content unreadable still refuses (regression).
test('control: a readable sibling with a non-terminal duplicate plan_id is duplicate_plan_id', () => {
// A4a — control. Same construction as the "FAILS CLOSED on duplicate
// plan_id" test above: valid JSON, same plan_id, different job_id,
// non-terminal status. This establishes the sibling content genuinely
// triggers the duplicate guard when it CAN be read.
const dir = path.join('.planning', 'async-jobs');
const siblingPath = path.join(dir, '99999.json');
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const duplicate = buildManifest({ ...baseInput(), job_id: '99999' }, { clock });
const fs = memFs({ [dir]: ['99999.json'], [siblingPath]: JSON.stringify(duplicate) });
const manifest = buildManifest(baseInput(), { clock });
const res = writeManifest(manifest, '.planning', { fs, clock });
assert.strictEqual(res.ok, false);
assert.strictEqual(res.kind, 'duplicate_plan_id', 'the sibling content must be a real duplicate trigger when readable');
});
test('a corrupt sibling can no longer hide a duplicate dispatch', () => {
// A4b — the bug, as a test. Same sibling PATH as the control above, which
// that test proves can hold a genuine non-terminal duplicate for this
// plan_id. Here the read itself faults (opts.failReads), so no content is
// reachable at all — and that is the point: the guard now refuses on ANY
// unreadable sibling precisely because it cannot know whether that file
// was the duplicate. The control supplies the "it could have been"; this
// supplies the "and we no longer gamble on it".
//
// Before the fix, an unreadable sibling was `continue`d past, the
// duplicate was never seen, and this returned ok:true — dispatching a
// duplicate external job.
const dir = path.join('.planning', 'async-jobs');
const siblingPath = path.join(dir, '99999.json');
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const fs = memFs(
{ [dir]: ['99999.json'] },
{ failReads: { [siblingPath]: 'eio: input/output error' } },
);
const manifest = buildManifest(baseInput(), { clock });
const res = writeManifest(manifest, '.planning', { fs, clock });
assert.strictEqual(res.ok, false, 'a corrupt sibling must refuse, not silently permit a possible duplicate dispatch');
assert.strictEqual(res.kind, 'scan_incomplete');
});
test('happy path unaffected', () => {
// A5
const dir = path.join('.planning', 'async-jobs');
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const fs = memFs({ [dir]: [] });
const manifest = buildManifest(baseInput(), { clock });
const res = writeManifest(manifest, '.planning', { fs, clock });
assert.strictEqual(res.ok, true);
});
test('a terminal prior job is still allowed', () => {
// A6
const dir = path.join('.planning', 'async-jobs');
const otherPath = path.join(dir, '99999.json');
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const dead = buildManifest({ ...baseInput(), job_id: '99999', status: 'failed', terminal_details: { code: 1 } }, { clock });
const fs = memFs({ [dir]: ['99999.json'], [otherPath]: JSON.stringify(dead) });
const manifest = buildManifest(baseInput(), { clock });
const res = writeManifest(manifest, '.planning', { fs, clock });
assert.strictEqual(res.ok, true, 'the duplicate guard only protects against re-dispatching live work');
});
test('a directory read failure is io_error, not scan_incomplete', () => {
// A7
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const fs = memFs({}, { failReads: {} });
fs.readdirSync = () => { throw new Error('eio: input/output error'); };
const manifest = buildManifest(baseInput(), { clock });
const res = writeManifest(manifest, '.planning', { fs, clock });
assert.strictEqual(res.ok, false);
assert.strictEqual(res.kind, 'io_error');
});
// ─── Property-based (CLAUDE.md: parsers/contracts need a fast-check test) ─────
test('property: mapSlurmState is total and idempotent over the known alphabet', () => {
fc.assert(
fc.property(fc.constantFrom(
'PENDING', 'CONFIGURING', 'RUNNING', 'COMPLETING', 'COMPLETED',
'FAILED', 'CANCELLED', 'TIMEOUT', 'OUT_OF_MEMORY', 'BOOT_FAIL', 'NODE_FAIL', 'PREEMPTED',
), (state) => {
const a = mapSlurmState(state);
const b = mapSlurmState(state);
return a !== null && a === b && MANIFEST_STATUS.includes(a);
}),
{ numRuns: 200 },
);
});
test('property: buildManifest -> validateManifest round-trips for valid generated input', () => {
fc.assert(
fc.property(
fc.record({
plan_id: fc.stringMatching(/^[0-9]+\.[0-9]+$/),
phase: fc.stringMatching(/^[0-9]+$/),
job_id: fc.stringMatching(/^[0-9]{1,8}$/),
backend: fc.constantFrom('slurm'),
submit_command: fc.constantFrom('sbatch --parsable ./run.sh'),
status: fc.constantFrom(...MANIFEST_STATUS),
expected_artifacts: fc.array(fc.constantFrom('Artifacts/jobs/x/out.h5'), { minLength: 1 }),
verification_command: fc.constantFrom('python -m verify.py'),
resume_command: fc.constantFrom('/gsd:execute-phase 3'),
terminal_details: fc.oneof(fc.constant(null), fc.record({ code: fc.integer() })),
}),
(input) => {
const clock = { nowIso: () => '2020-06-15T12:00:00.000Z' };
const td = input.status === 'completed-unverified' ? null : input.terminal_details;
const built = buildManifest({ ...input, terminal_details: td }, { clock });
return validateManifest(built).ok === true;
},
),
{ numRuns: 100 },
);
});