* test(#3871): failing-first regressions for the dropped curated progress block Pins ADR-3473 §8.6 / #3756 at the consumer's output: state record-session and state add-decision on an archived-milestone project drop the curated progress frontmatter entirely, exit 0, and report nothing. Reproduced against the real CLI before writing the tests, not inferred from the issue text. Also adds the unit-level probe that applyStatePreservation's preserve-always row is inert on a resyncing write, and an over-preservation guard that an empty project is never inflated. Refs #3871 * feat(#3871): make the STATE.md pre-write snapshot mandatory via open()/rebuild() ADR-3473 §8.6. StatePreservationInput's nullable preFm and the always-present preFmSnapshot were the same extractFrontmatter call, one of them nulled on resync — a policy flag baked into a snapshot. Both collapse into a single StateTransaction whose snapshot cannot be absent: openStateTransaction() applies preservation, rebuildStateTransaction() does not, and both carry the snapshot because the reporting phase needs it either way. An absent snapshot is now a construction failure; an empty one stays legal, because that is what a document with no parseable frontmatter honestly has. writeStateMd requires a rebuild transaction, which types ADR-3408 §8.3's closed exception list at both call sites (state sync, health --repair) instead of matching them as strings in a ratcheted baseline. Fixes the dropped curated progress block: an all-zero or absent derived total set is an unmeasured scan, not a measurement, so the curated block stands. Also fixes two defects surfaced while building — preserve-always reported a mutation even when it restored an identical value, and it re-entered the curated object by reference, which would alias the snapshot the next phase diffs against. Refs #3871 * fix(#3871): close the three remaining subsumed defects and restore the arm the type does not replace Review of the first two commits found four things. The guard shrink deleted the seam-bypass axis whole, but only its writeStateMd( arm became redundant. Its other arm catches a call site re-assembling syncStateFrontmatter + applyPostSyncPreservation instead of the owned composition, which the transaction type does not make unrepresentable and which #3469 found live. Restored as findCompositionBypasses, terminal rather than ratcheted. Three of the four issues this phase claims were untouched. All three are the epic's own shape and are fixed at the seam: current_phase_name is reasserted from the curated value when the caller names none, and cmdStateJson stops carrying a hand-maintained list parallel to FIELD_CLASSIFICATION and projects it instead. The construction failure that is the point of this phase had no test. Every enumerated matrix row now has one, including the measured-versus-unmeasured coercion boundary and a seeded property that no curated key is ever dropped. ADR-3473 §8.6 said the guard 'keeps only its raw-write check'. Verified against next: there was no raw-write check, and four other checks it does not name. Amended in place with the evidence. ARCHITECTURE.md separately advertised a preservation policy the code had deleted. Refs #3871 * fix(#3871): do not let the unmeasured-scan rule block an explicitly-requested resync The remote matrix caught over-preservation, the failure this phase's own negative space says must not happen. state update Progress re-derives the block from the body the caller just rewrote; on a project with no phase dirs the derivation yields zero totals, the unmeasured rule read that as 'the scan measured nothing', and the stale curated percent was restored over the resync the user asked for. preserve-always already said what the missing condition was: never overwrite unless the caller explicitly names this field. explicitProgressField carries it and is derived from shouldResyncStateProgress, not set by hand at a call site, so it cannot drift from what the caller asked for. Two defects found in the same mechanism and fixed with it. readModifyWriteStateMd enumerates its option keys, so a new option was silently dropped rather than rejected. And the raw-write axis captured its first argument up to the first comma, which lands inside a nested path.join, so a write to a STATE.md literal was invisible to it — the prove-it-can-fail test caught that one immediately. No test assertion was weakened; all three frontmatter rows encode #3242, #1969 B3 and #1972 and stand unchanged. Refs #3871 * docs(#3871): record why the raw-write check is kept, not why it was named The amendment justified findRawStateWrites as 'written because §8.6 requires it to exist', which is cargo-culting the contract and would have been the wrong reason to keep anything. The real reason is that writeStateMd acquires the STATE.md lockfile and a raw fs.writeFileSync acquires nothing, so this is a lock bypass and lost-update is the #500/#905/#1230 family — and after this phase it is the one reachable path into the file that nothing else covers. Also records why ADR-3408 §8.6's deletion of the 'clear' policy is not the precedent it looks like: 'clear' was dead vocabulary in a closed enum, this is coverage of a reachable path. Refs #3871 * chore(#3871): backfill changeset PR number --------- Co-authored-by: sim <sim@local>
137 lines
6.9 KiB
JavaScript
137 lines
6.9 KiB
JavaScript
'use strict';
|
|
// allow-test-rule: architectural-invariant (see #1531)
|
|
// writeStateMd's "scan happens INSIDE the lock" property is a concurrency invariant.
|
|
// A single-threaded test cannot observe the difference between scan-before-lock and
|
|
// scan-after-lock unless something mutates the disk in the window between the two.
|
|
// The afterAcquire test hook (fired inside writeStateMd right after the lock is
|
|
// taken) is the deterministic seam that simulates a concurrent writer landing in
|
|
// exactly that window — the only level at which the TOCTOU is observable.
|
|
|
|
/**
|
|
* M8 — writeStateMd scans the disk (syncStateFrontmatter / PLAN-SUMMARY count)
|
|
* BEFORE taking the lock, so a concurrent writer that commits a new PLAN/SUMMARY
|
|
* between our scan and our lock acquisition makes writeStateMd stamp STALE
|
|
* progress counts (a lost-update of the frontmatter progress block).
|
|
* readModifyWriteStateMd (the atomic variant) correctly scans INSIDE its lock —
|
|
* this non-atomic variant was the outlier.
|
|
*
|
|
* Deterministic repro (no wall-clock, no threads): the afterAcquire test hook
|
|
* fires inside writeStateMd immediately after the lock is acquired and adds a
|
|
* second PLAN file to the phase dir — simulating a concurrent writer who landed
|
|
* in the scan→lock window. The written frontmatter's progress.total_plans then
|
|
* reveals whether the scan ran before the hook (stale: 1) or after it (fresh: 2).
|
|
*
|
|
* RED (pre-fix): scan runs BEFORE acquire → before the hook → total_plans = 1.
|
|
* GREEN (post-fix): scan runs AFTER acquire → after the hook → total_plans = 2.
|
|
*
|
|
* Recurring closed family this guards: #500 / #905 / #1230 (STATE.md write
|
|
* corruption). #453 deleted the flaky race tests in favor of seams, so this exact
|
|
* path was under-tested — the hook restores deterministic coverage.
|
|
*/
|
|
|
|
const { test, describe, beforeEach, afterEach } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
const os = require('node:os');
|
|
|
|
const stateMod = require('../gsd-core/bin/lib/state.cjs');
|
|
const { writeStateMd } = stateMod;
|
|
const { rebuildStateTransaction } = require('../gsd-core/bin/lib/state-transition.cjs');
|
|
const { extractFrontmatter } = require('../gsd-core/bin/lib/frontmatter.cjs');
|
|
const { cleanup } = require('./helpers.cjs');
|
|
|
|
// ADR-3473 §8.6: writeStateMd's third argument is now a transaction. Both
|
|
// tests below mirror MINIMAL_STATE_MD's own pre-write content back to itself
|
|
// (no frontmatter on disk yet, so the snapshot is legitimately {}) — this is
|
|
// the M8 concurrency invariant under test, not a preservation scenario, so
|
|
// a `rebuild()` transaction (no preservation applied) is correct here.
|
|
function rebuildTransactionFor(content) {
|
|
return rebuildStateTransaction({ snapshot: extractFrontmatter(content) });
|
|
}
|
|
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
// Helpers
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
const MINIMAL_STATE_MD = [
|
|
'# Project State',
|
|
'',
|
|
'**Status:** Planning',
|
|
'**Current Phase:** 01',
|
|
].join('\n') + '\n';
|
|
|
|
/** Parse progress.total_plans out of the STATE.md frontmatter block. */
|
|
function readTotalPlans(statePath) {
|
|
const written = fs.readFileSync(statePath, 'utf-8');
|
|
// eslint-disable-next-line local/no-unbounded-quantifier -- parses STATE.md this test just wrote via a fixture, fixed-size test-controlled content
|
|
const fmMatch = written.match(/^---\r?\n([\s\S]*?)\r?\n---/);
|
|
assert.ok(fmMatch, 'STATE.md must have a frontmatter block after writeStateMd');
|
|
const m = fmMatch[1].match(/total_plans:\s*(\d+)/);
|
|
assert.ok(m, 'frontmatter must carry a progress.total_plans line');
|
|
return parseInt(m[1], 10);
|
|
}
|
|
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
// M8 — afterAcquire hook proves the scan runs INSIDE the lock
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
describe('M8: writeStateMd scans disk AFTER acquiring the lock (scan-in-lock)', () => {
|
|
let tmpDir;
|
|
let statePath;
|
|
let phaseDir;
|
|
|
|
beforeEach(() => {
|
|
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-m8-'));
|
|
const planningDir = path.join(tmpDir, '.planning');
|
|
phaseDir = path.join(planningDir, 'phases', '01-init');
|
|
fs.mkdirSync(phaseDir, { recursive: true });
|
|
// Start with exactly ONE plan file on disk.
|
|
fs.writeFileSync(path.join(phaseDir, '01-PLAN.md'), '# Plan 01\n');
|
|
statePath = path.join(planningDir, 'STATE.md');
|
|
fs.writeFileSync(statePath, MINIMAL_STATE_MD);
|
|
});
|
|
|
|
afterEach(() => {
|
|
stateMod._resetStateLockTestHooks();
|
|
try { fs.unlinkSync(statePath + '.lock'); } catch { /* ok */ }
|
|
cleanup(tmpDir);
|
|
});
|
|
|
|
test('a PLAN added in the post-acquire window is reflected in the written progress count', () => {
|
|
// The hook simulates a concurrent writer who commits a second PLAN file in the
|
|
// window between scan and lock. It MUST be observed only if the scan runs after
|
|
// the lock (and therefore after this hook fires).
|
|
let fired = 0;
|
|
stateMod._setStateLockTestHooks({
|
|
afterAcquire() {
|
|
fired++;
|
|
fs.writeFileSync(path.join(phaseDir, '02-PLAN.md'), '# Plan 02\n');
|
|
},
|
|
});
|
|
|
|
writeStateMd(statePath, MINIMAL_STATE_MD, rebuildTransactionFor(MINIMAL_STATE_MD), tmpDir);
|
|
|
|
assert.equal(fired, 1, 'afterAcquire hook must fire exactly once inside writeStateMd');
|
|
|
|
const totalPlans = readTotalPlans(statePath);
|
|
// RED pre-fix: scan ran before the hook → counts only 01-PLAN.md → 1.
|
|
// GREEN post-fix: scan ran after the hook → counts both PLANs → 2.
|
|
assert.equal(
|
|
totalPlans, 2,
|
|
'writeStateMd must scan the disk INSIDE the lock (after the concurrent ' +
|
|
'writer landed), stamping total_plans=2 — not the stale pre-lock count of 1'
|
|
);
|
|
});
|
|
|
|
test('single-threaded callers (no hook) are byte-for-behaviour unchanged: count = 1', () => {
|
|
// Regression guard: with no concurrent writer (hook unset), the count must be
|
|
// exactly the on-disk truth — the fix must NOT change the uncontended result.
|
|
writeStateMd(statePath, MINIMAL_STATE_MD, rebuildTransactionFor(MINIMAL_STATE_MD), tmpDir);
|
|
assert.equal(
|
|
readTotalPlans(statePath), 1,
|
|
'uncontended writeStateMd must stamp the real on-disk plan count (1)'
|
|
);
|
|
});
|
|
});
|