Files
msd-core/tests/estimate-calibrate.test.cjs
Tom Boucher 15af0f5536 enhance(#3951): B6+B7 — widen two unreachable lint rules and make the guard ledger true (#3965)
* fix(#3951): two lint rules that could not reach the code they govern

B6 names two widenings. Measuring them first turned up a defect the criterion did
not know about, and refuted the reason it gave for one of them.

1. no-adhoc-markdown-parsing self-gates on its own filename.

   Lines 107-110 short-circuit create() to {} unless the path matches
   /(?:^|\/)src\/[^/]+\.cts$/. B6 says to widen the files: glob in
   eslint.config.mjs - but doing only that ships an INERT rule, because the gate
   still returns {} for every new path. Both halves have to change, and the gate
   is the load-bearing one.

   That same regex hides a live hole: [^/]+ is FLAT-ONLY, so it requires the file
   to sit directly in src/. The registered glob is src/**/*.cts, which includes
   subdirectories. 28 .cts files - health-diagnostic-rules/ (10),
   installer-migrations/ (11), observability/ (3), host-integration-adapters/ (2),
   vendor/ (2) - are inside the registered glob and silently skipped.

   Measured with the gate neutralized: 0 violations there today. The hole is
   hiding nothing right now, and is fixed anyway, because "no violations today" is
   not a property that keeps holding.

   The fix is not invented: require-subprocess-timeout.cjs:196 already carries the
   correct form of this guard, /(?:^|\/)src\/.*\.cts$/ with .*, one directory over.
   Checked the other 21 rules for the same bug - no-adhoc-regex-escape and
   no-private-binary-resolution short-circuit only to exempt their own seam file,
   which is the right shape, and no-crlf-fragile-split has no filename gate at
   all. This bug is unique to the one rule.

2. no-adhoc-regex-escape could not see the shape that actually occurs.

   Line 396 gated the whole UNSAFE-NEW-REGEXP arm on arg.type === 'Identifier'.
   Every check below it - the _SOURCE provenance check, the
   isSoleReturnOfOwnParameter shape - lives inside that branch, so
   new RegExp(obj['key']) and new RegExp(cfg.pattern) were never examined at all.
   Runtime data arrives as a property access far more often than as a bare
   identifier, which is exactly why this rule never fired on the #3477 ReDoS.

   Widened to MemberExpression, measured by AST walk across all five registered
   blocks rather than by grep. 27 sites, zero TSAsExpression:

     18  safe new RegExp(X.source, flags)  -> exempted, keyed strictly on the
         PROPERTY being `source`, never on the object. Keying on the object would
         wave through X.anything and buy nothing. B6 estimated ~10; that was an
         undercount.
      3  _SOURCE-suffixed constants reached through a required module namespace
         (phaseId.BRACKET_PHASE_TOKEN_SOURCE) -> the same provenance-exempt class
         the rule already recognizes for bare identifiers, extended to reach them.
         Without this the widening produces 3 false flags.
      6  real findings -> marked, each a test extracting a pattern from a shipped
         file at test time, where the runtime contract IS the product.

   Deliberately the NARROW MemberExpression form. The rule's own
   isSoleReturnOfOwnParameter doc comment records that an earlier broad
   "any non-literal identifier" heuristic produced ~25 false positives and was
   rejected; a re-run of the census after this change flags exactly the 6 above
   and nothing else.

Verified by execution, not by reading: the gate now accepts src/<subdir>/x.cts,
still accepts flat src/x.cts, and still exempts paths outside src/ - each pinned
by a test proven to fail against the old regex. build:lib, lint and lint:ci all
exit 0.

Refs #3951

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3951): give no-adhoc-markdown-parsing its reach, and fix the 80 parses it finds

The rule self-gates on filename AND is registered on one glob, so widening either
half alone is inert. Both move here: the gate now accepts tests/**/*.cjs and
scripts/**/*.cjs alongside src/**/*.cts, and eslint.config.mjs registers it on the
same two.

A test pins that the gate and the registration AGREE, in both directions. The
original defect was a gate narrower than its registration; the failure mode of
this fix is a gate wider than its registration. Both are silent, so the test
asserts the pair rather than either half.

80 violations across 43 files, all in tests/, zero in scripts/. 70 are routed
through the existing seams - scanFencedBlocks, collectSection, stripFencedCode,
tokenizeHeadings from markdown-sectionizer; splitTableRow, parseMarkdownTable,
findTableWithColumns from markdown-table. Headerless STATE.md tables use
splitTableRow per line, because parseMarkdownTable needs a real delimiter row.

10 are suppressed, 12.5%, well under the third that would have meant the rule is
mis-scoped for tests/ rather than the tests carrying debt. Each names its reason:
three regression guards (#3873 / bug-#21) are deliberately independent of the
generator's own fence handling, and routing them through the seam would have them
test the generator against itself; one is a negative-text probe that extracts
nothing; six are a shell-pipe-to-jq detector whose regex coincidentally matches the
table fingerprint and is not markdown parsing at all.

All ten sit in tests whose subject is .md content, which is normally a reason to
prefer the seam. The marker used is allow-adhoc-markdown, distinct from
no-source-grep's allow-test-rule, and lint:ci's lint-allow-test-rule-refs reports
the same 280/280 unverified count as before - checked rather than assumed, because
those two markers are easy to conflate.

The widening earned its keep immediately: it found a test that passed for the
wrong reason.

  tests/config-field-docs.test.cjs asserted notEqual(<cell>, '600') against the
  TYPE column instead of the DEFAULT column. notEqual('number', '600') is true
  forever, so the guard against workflow.subagent_timeout regressing to the old
  seconds default could never fire. docs/CONFIGURATION.md:434 is
  `| workflow.subagent_timeout | number | 300000 | ... |`, so the default is cell
  index 2; the assertion is now row-scoped through splitTableRow and reads 300000.

That is the argument for the widening in one case: the violation was invisible to
lint, the suite was green, and the assertion was vacuous. A rule that cannot reach
a file cannot tell you the file is lying.

Not fixed here, and recorded rather than assumed: #3426/#3239 are NOT reachable by
this widening. tests/package-legitimacy-gate.test.cjs yields zero violations even
with the gate bypassed - its hand-rolled scans are real, but built from line
filters and split('|') rather than the regex-literal fingerprints this rule
detects. They need new detectors. The epic assumed a wider glob would catch them.

build:lib, lint and lint:ci all exit 0; the post-fix census across tests/** and
scripts/** is 0 violations.

Refs #3951

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3951): B7 — and #3356's defects were still live in the code

B7 asks that each closed child be driven fail-first with a behavioral identity
test at the CONSUMER's output. Four of eleven children had no test citing their
issue number. Auditing them by BEHAVIOR rather than by number-grep changed the
answer for three of the four.

#3364 and #2540 — traceability only. Both were implemented by #3941 and their
consumer-output tests exist and were shown failing-first; neither cited its
originating issue, so an audit that greps for the number reports them uncovered.
Tagged the specific asserting test in each file, following the citation form those
files already use.

#3372 — covered, but only at helper level, and the triage narrowed it. Of the four
commands the issue names, only estimate-cli's collectCalibrationSamples actually
enumerates phase dirs from disk; smart-entry, audit and roadmap-upgrade derive from
ROADMAP/body text and never reach the sentinel path, so they are benign by
construction and were left alone rather than "fixed" into churn. The existing #3882
rows asserted the helper's return value. Added a consumer-output test driving
`query estimate-calibrate` and asserting sample_count and the persisted document.
RED proof: reverted collectCalibrationSamples to a raw readdirSync and ran the real
CLI - sample_count 3, sentinel leaked; restored - sample_count 2.

#3356 — NOT covered, and BOTH halves of the defect were still live in source. The
issue is closed; the bug was not fixed. Fixed here rather than writing tests that
document a bug as correct.

  Defect 1, the contradicted row. quick.md:627 claimed
  `quick-tasks-append` performs "the equivalent write" to the Step 7c row. It did
  not: the `#` cell was a positional ordinal and `Directory` read `—`, because the
  route had no way to receive a quick id or task directory. Added OPTIONAL
  `--quick-id` / `--slug` / `--directory`. A caller with neither - fast.md, the
  original #2133 caller - omits them and gets the byte-identical prior row, so
  nothing existing changes. A caller that HAS a real id and directory now gets the
  canonical row quick.md:632 renders. The false-equivalence sentence itself is
  corrected rather than left to mislead the next reader.

  Defect 2, the forced re-derive. The route called readModifyWriteStateMd with no
  options, so a body-only append to the Quick Tasks table triggered a full
  re-derive of the disk-derived progress.* frontmatter. Every other body-only
  writer passes { resync: false } - src/state.cts's own docstring prescribes it -
  and this route was the lone outlier. RED proof: reverted the option, seeded a
  project with 2 real phase dirs and a curated total_phases of 25, ran
  quick-tasks-append; total_phases collapsed to 2. Restored; it stayed 25.

That second one is the shape this epic exists to close: a silent write that
replaces curated state with a re-derivation nobody asked for, exit 0 throughout.

build:lib, lint and lint:ci all exit 0.

Refs #3951

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3951): amend B6's ledger to what was measured, and document the new flags

The ADR gains a ledger amendment in its own correction style - the sixth wrong
premise it records, found the same way as the other five, by measuring before
building.

B6 says the net guard count must fall. It rose: 62 -> 69, +7, measured from the
epic's filing commit to origin/next. The attribution is the point, though. Five of
the seven came from PRs unrelated to this epic, one was added by a phase of it, and
the epic did retire something sub-file - #3884 removed a detector with an explicit
"net: -1 detector, 0 added" ledger. Every named casualty is load-bearing, two
already carry retractions in this same document, and a sweep of all 22 rules plus
every scripts/lint-* found no provably dead guard. There is no honest way to make
the count fall; forcing it would trade coverage for a number, which is the Goodhart
outcome Decision 6 exists to prevent.

The amendment also records that B6's own prescribed fix for one widening was inert.
no-adhoc-markdown-parsing self-gates on its filename, so widening only the files:
glob - which is what the criterion says to do - ships a rule that still returns {}
for every new path. And #3426/#3239 are not reachable by that widening at all;
their scans use line filters and split('|'), not the regex fingerprints the rule
detects. The roster row tracked them against the wrong mechanism.

Three roster rows updated from aspiration to fact: the two widenings are DONE with
their measured counts, and lint-phase-enumeration-drift is marked RETAINED rather
than "expected casualty - verify before retiring", because Phase 5 verified it and
kept it.

The rule Decision 6 should carry forward is stated plainly: a guard ledger is a
claim about COVERAGE, not about COUNT. "Net count must fall" is measurable and
wrong. "Every guard is reachable, and each retirement names what makes its defect
unrepresentable" is the property that was actually wanted.

CLI-TOOLS.md documents the optional --quick-id/--slug/--directory flags and says
plainly that omitting them keeps the pre-#3356 row byte-identical, plus that the
append no longer re-derives progress frontmatter.

New features fragment (id 3951); FEATURES.md regenerated rather than hand-edited.
Changeset is Changed, pr:0 pending backfill.

Refs #3951

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3951): correct four rows that pinned the lint rule's old narrow reach

The remote suite came back RED with 5 failures, all in tests/eslint-rules.test.cjs.
They are stale tests, not a regression: four rows assert that
no-adhoc-markdown-parsing is inert outside src/*.cts, which is exactly the
contract this deliverable changes.

Confirmed by reading rather than inferred from the names - the row at :1981 used
filename: 'tests/some.test.cjs' and filename: 'scripts/helper.cjs', the two roots
the rule now covers on purpose.

Worth recording WHY local gates missed this. npm run lint and lint:ci were green,
and the touched test files passed standalone. Lint only reports violations in real
files; these rows assert the rule's REACH using synthetic RuleTester filenames, so
nothing but the full suite could see them. Local green on a rule change says
nothing about the rule's own tests.

Each row is rewritten with BOTH halves rather than flipped from valid to invalid:

  - the same fingerprint under tests/ or scripts/ is now flagged, with the right
    messageId
  - the negative space is preserved - the same fingerprint under a path outside
    all three roots (gsd-core/bin/lib/foo.cjs) is still NOT flagged

The second half is the one that matters. Without it the rule has no boundary and
nothing would catch an over-wide gate later, which is the mirror image of the bug
this deliverable just fixed.

Each row is renamed to state the current contract; the old names said
"non-src/*.cts ... is not flagged" and would have been actively misleading once
the bodies changed.

Proven to test the widening rather than restate it: every flagged half was run
against HEAD~2's pre-widening rule and does NOT fire there, then against the
current rule and does. 12/12 on that probe; the full file is 178/178.

Swept for the same staleness elsewhere and found none.
require-subprocess-timeout's own "inert outside src/*.cts" row is untouched -
that rule's gate was not widened here - and no-adhoc-regex-escape's test file
already carries correctly-targeted rows.

Refs #3951

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3951): acknowledge the quick.md growth the attribution guard reported

The full suite came back RED with one failure, and it is mine:

  1 file(s) grew without an acknowledgment:
    quick.md grew 364 bytes

gsd-core/workflows/quick.md is runtime-loaded emitted content, so correcting
its false 'performs the equivalent write' claim trips emitted-attribution by
construction. This is the acknowledgment, not a workaround - there is nothing
to regenerate.

The fragment names ONE path, which is the only one the guard reported. The four
spent acknowledgments it also listed (audit-uat, plan-phase, progress, review)
belong to other fragments whose ripple the base already absorbs; they are inert,
not failures, and are deliberately NOT copied here - naming paths I did not
change would make this record false in the other direction.

Byte figure corrected before committing: the guard reported 37220 -> 37584
(+364), but origin/next has since moved and quick.md is 37232 there now, so the
measured delta is +352. The reason text says so and names the base as a moving
figure rather than pinning a number that is already stale.

Refs #3951

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3951): move the quick.md growth ack to a trailer, delete the obsolete fragment

The acknowledgment mechanism changed under this branch. Merging next brought in
the redesign - it also deleted .github/workflows/ack-fragment-sweep.yml, which
was in the merge status and which I did not register at the time - and the guard
now says so directly:

  Add a trailer to a commit in this PR (never a new file).
    Emitted-Drift-Ack-Growth: quick.md - <why this growth is deliberate>

So tests/emitted-drift-acks/3951-quick-append-equivalence.json is obsolete on
arrival. A fragment file is no longer read by anything, and leaving it would be a
dead record that looks like an active one. It is deleted here rather than kept
"just in case".

The byte figure moved again with the merge: 37232 -> 37596, +364. The earlier
fragment said +352, measured before the merge auto-merged quick.md itself. The
trailer carries no number, which is the better design - the figure was stale
twice in two attempts.

Refs #3951

Emitted-Drift-Ack-Growth: quick.md — #3356/#3951 replaces a false claim with an accurate one. Line 627 said the `quick-tasks-append` shortcut "performs the equivalent write" to the Step 7c row rendered above it; it did not, and that was the documented half of #3356 — with no quick id or task directory the route emitted a positional ordinal in `#` and an em-dash in `Directory`, a visibly different row. The corrected sentence has to carry three facts the original elided: what the shortcut actually writes when it has neither input, that this is honest behavior for its real caller (`fast.md`, which has neither), and how a caller with both now gets the byte-identical canonical row via the new optional `--quick-id`/`--slug`/`--directory` flags. Prose is the product here — an executing agent reads this line to decide whether the shortcut is safe for its case, and a shorter correction would either drop the flags (leaving the reader unable to act on the fix) or drop the limitation (recreating the false claim in gentler words).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3951): backfill changeset pr number

Refs #3951

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 23:10:49 -04:00

654 lines
30 KiB
JavaScript

// docs-guard-exempt: docPath is a .planning/estimation-calibration.json tmp fixture; the docs/adr and docs/reference citations are comment-only.
/**
* estimate-calibrate — build the calibration document from completed phases.
*
* Epic #1952 Phase 3 (#2632). Design lock: docs/adr/2629-phase-effort-estimation-calibration.md.
*
* This is the verb that makes AC4 real. Phase 1 shipped the calibration MATH;
* Phase 2 made the planner emit an estimate. Neither closes the loop, because
* nothing pairs a plan's `estimate` with its summary's `actuals` and writes the
* result. Leaving that to agent prose would make "estimates improve over time"
* unverifiable — so the pairing and the write are deterministic here, and
* extract-learnings just invokes them.
*
* The headline test is `a consistently-underestimated project produces an
* upward correction`: that is epic acceptance criterion AC4 stated as an
* executable claim.
*/
const { describe, test } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const { createTempProject, cleanup, runGsdTools } = require('./helpers.cjs');
const est = require('../gsd-core/bin/lib/phase-estimation.cjs');
const estimateCli = require('../gsd-core/bin/lib/estimate-cli.cjs');
// io.cjs owns error()/ERROR_REASON/JSON-error-mode — driven directly here so
// H1 below can assert a typed `reason`, mirroring tests/config-get-default.test.cjs's
// established in-process CLI-error-path pattern.
const io = require('../gsd-core/bin/lib/io.cjs');
/** Write a phase dir containing a PLAN with an estimate and a SUMMARY with actuals. */
function writePhase(tmpDir, phaseDir, { estTokens, actTokens, tasks = 3, commits = 4 }) {
const dir = path.join(tmpDir, '.planning', 'phases', phaseDir);
fs.mkdirSync(dir, { recursive: true });
if (estTokens !== null) {
fs.writeFileSync(path.join(dir, '01-PLAN.md'), [
'---',
'phase: ' + phaseDir,
'plan: 01',
'estimate:',
` tokens: ${estTokens}`,
` tasks: ${tasks}`,
' confidence: low',
'must_haves:',
' truths: []',
'---',
'<objective>x</objective>',
'',
].join('\n'));
}
if (actTokens !== null) {
fs.writeFileSync(path.join(dir, '01-SUMMARY.md'), [
'---',
'phase: ' + phaseDir,
'plan: 01',
'actuals:',
` tokens: ${actTokens}`,
` tasks: ${tasks}`,
` commits: ${commits}`,
'---',
'## What shipped',
'',
].join('\n'));
}
return dir;
}
describe('estimate-calibrate', () => {
test('AC4: a consistently-underestimated project produces an upward correction', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
// Three phases that each cost ~2x their estimate.
writePhase(tmpDir, '01-alpha', { estTokens: 50000, actTokens: 98000 });
writePhase(tmpDir, '02-beta', { estTokens: 60000, actTokens: 121000 });
writePhase(tmpDir, '03-gamma', { estTokens: 40000, actTokens: 82000 });
const r = runGsdTools('query estimate-calibrate', tmpDir);
assert.ok(r.success, `estimate-calibrate should succeed: ${r.error}`);
const out = JSON.parse(r.output);
assert.equal(out.sample_count, 3, 'all three phases pair up');
assert.equal(out.applied, true);
assert.ok(out.factor > 1, `expected an upward correction, got ${out.factor}`);
// The document must be persisted where estimate-calibration reads it.
const docPath = path.join(tmpDir, '.planning', 'estimation-calibration.json');
assert.ok(fs.existsSync(docPath), 'calibration document must be written');
assert.deepEqual(
est.parseCalibrationDocument(fs.readFileSync(docPath, 'utf8')).length, 3,
'persisted document must carry all three samples',
);
// And the read verb must now agree — this is the loop actually closing.
const readBack = JSON.parse(runGsdTools('query estimate-calibration', tmpDir).output);
assert.equal(readBack.factor, out.factor, 'estimate-calibration must see what estimate-calibrate wrote');
assert.equal(readBack.applied, true);
// A subsequent estimate is therefore larger than the raw projection.
const check = JSON.parse(runGsdTools('query estimate-check --tokens 50000', tmpDir).output);
assert.ok(check.calibrated_tokens > 50000,
`a later estimate must be corrected upward, got ${check.calibrated_tokens}`);
});
test('a consistently-overestimated project produces a downward correction', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
writePhase(tmpDir, '01-a', { estTokens: 100000, actTokens: 60000 });
writePhase(tmpDir, '02-b', { estTokens: 80000, actTokens: 48000 });
writePhase(tmpDir, '03-c', { estTokens: 90000, actTokens: 54000 });
const out = JSON.parse(runGsdTools('query estimate-calibrate', tmpDir).output);
assert.ok(out.factor < 1, `expected a downward correction, got ${out.factor}`);
});
test('boundary: inert below the minimum sample count, applied at it', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
writePhase(tmpDir, '01-a', { estTokens: 100, actTokens: 200 });
writePhase(tmpDir, '02-b', { estTokens: 100, actTokens: 200 });
let out = JSON.parse(runGsdTools('query estimate-calibrate', tmpDir).output);
assert.equal(out.sample_count, 2);
assert.equal(out.applied, false, '2 samples must not apply a correction');
assert.equal(out.factor, 1);
writePhase(tmpDir, '03-c', { estTokens: 100, actTokens: 200 });
out = JSON.parse(runGsdTools('query estimate-calibrate', tmpDir).output);
assert.equal(out.sample_count, 3);
assert.equal(out.applied, true, '3 samples must apply');
assert.equal(out.factor, 2);
});
test('phases missing either side are skipped, not guessed', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
writePhase(tmpDir, '01-paired', { estTokens: 100, actTokens: 200 });
writePhase(tmpDir, '02-plan-only', { estTokens: 100, actTokens: null });
writePhase(tmpDir, '03-summary-only', { estTokens: null, actTokens: 200 });
const out = JSON.parse(runGsdTools('query estimate-calibrate', tmpDir).output);
assert.equal(out.sample_count, 1, 'only the fully-paired phase counts');
});
test('a phase whose PLAN has no estimate block contributes nothing', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
const dir = path.join(tmpDir, '.planning', 'phases', '01-noest');
fs.mkdirSync(dir, { recursive: true });
fs.writeFileSync(path.join(dir, '01-PLAN.md'), '---\nphase: 01-noest\nplan: 01\n---\nbody\n');
fs.writeFileSync(path.join(dir, '01-SUMMARY.md'), '---\nphase: 01-noest\nactuals:\n tokens: 5\n tasks: 1\n commits: 1\n---\nx\n');
const out = JSON.parse(runGsdTools('query estimate-calibrate', tmpDir).output);
assert.equal(out.sample_count, 0);
assert.equal(out.applied, false);
});
test('no phases at all is a clean no-op, not an error', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
const r = runGsdTools('query estimate-calibrate', tmpDir);
assert.ok(r.success, 'must not fail on an empty project');
const out = JSON.parse(r.output);
assert.equal(out.sample_count, 0);
assert.equal(out.factor, 1);
});
test('re-running is idempotent — it rebuilds, never appends duplicates', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
writePhase(tmpDir, '01-a', { estTokens: 100, actTokens: 200 });
writePhase(tmpDir, '02-b', { estTokens: 100, actTokens: 200 });
writePhase(tmpDir, '03-c', { estTokens: 100, actTokens: 200 });
const first = JSON.parse(runGsdTools('query estimate-calibrate', tmpDir).output);
const second = JSON.parse(runGsdTools('query estimate-calibrate', tmpDir).output);
assert.deepEqual(second, first, 'a second run must produce an identical result');
const doc = est.parseCalibrationDocument(
fs.readFileSync(path.join(tmpDir, '.planning', 'estimation-calibration.json'), 'utf8'),
);
assert.equal(doc.length, 3, 'samples must not accumulate across runs');
});
test('a corrupt pre-existing document is replaced, not merged', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
fs.writeFileSync(path.join(tmpDir, '.planning', 'estimation-calibration.json'), '{ not json');
writePhase(tmpDir, '01-a', { estTokens: 100, actTokens: 200 });
const r = runGsdTools('query estimate-calibrate', tmpDir);
assert.ok(r.success, 'a corrupt prior document must not fail the rebuild');
const doc = est.parseCalibrationDocument(
fs.readFileSync(path.join(tmpDir, '.planning', 'estimation-calibration.json'), 'utf8'),
);
assert.equal(doc.length, 1);
});
test('the written document round-trips through the parser', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
writePhase(tmpDir, '01-a', { estTokens: 12345, actTokens: 23456 });
runGsdTools('query estimate-calibrate', tmpDir);
const raw = fs.readFileSync(path.join(tmpDir, '.planning', 'estimation-calibration.json'), 'utf8');
const parsed = est.parseCalibrationDocument(raw);
assert.deepEqual(parsed, [{ estimateTokens: 12345, actualTokens: 23456 }]);
assert.equal(JSON.parse(raw).schema_version, est.CALIBRATION_SCHEMA_VERSION,
'must stamp the current schema version so a future reader can refuse it');
});
});
// ─── convergence guard (#2632) ─────────────────────────────────────────────
describe('calibration converges instead of oscillating', () => {
// The loop must measure actual/RAW, not actual/calibrated. Measuring against
// the already-corrected figure is self-defeating: once the correction works
// the observed ratio approaches 1, dragging the median back toward 1, which
// un-corrects the next estimate. This test pins convergence over enough
// phases for that oscillation to show up — it fails at ~1.41 if the basis
// regresses to the calibrated value.
const RAW = 50000;
const TRUE_COST = 100000; // the planner is consistently 2x low
const simulate = (useRawBasis) => {
const samples = [];
for (let phase = 0; phase < 10; phase += 1) {
const cal = est.computeCalibration(samples);
const emitted = est.applyCalibration(RAW, cal.factor);
const estimate = { tokens: emitted, rawTokens: RAW, tasks: 3, confidence: cal.confidence };
samples.push({
estimateTokens: useRawBasis ? est.calibrationBasis(estimate) : estimate.tokens,
actualTokens: TRUE_COST,
});
}
return est.computeCalibration(samples).factor;
};
test('measuring against the raw projection converges on the true ratio', () => {
assert.ok(Math.abs(simulate(true) - 2) < 1e-9,
`expected convergence on 2.0, got ${simulate(true)}`);
});
test('measuring against the calibrated figure does NOT converge', () => {
// Negative proof that the basis choice is load-bearing, not incidental.
assert.ok(simulate(false) < 1.9,
'if this passes at ~2.0 the two bases are equivalent and this guard is vacuous');
});
test('calibrationBasis prefers raw_tokens and falls back for older plans', () => {
assert.equal(est.calibrationBasis({ tokens: 100000, rawTokens: 50000, tasks: 3, confidence: 'med' }), 50000);
assert.equal(est.calibrationBasis({ tokens: 60000, tasks: 3, confidence: 'low' }), 60000,
'a pre-#2632 plan with no raw_tokens must still contribute a sample');
});
test('estimate-calibrate uses raw_tokens from the plan when present', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
// tokens=100000 (calibrated) but raw_tokens=50000; actual=100000.
// Ratio must be 100000/50000 = 2, NOT 100000/100000 = 1.
for (const phase of ['01-a', '02-b', '03-c']) {
const dir = path.join(tmpDir, '.planning', 'phases', phase);
fs.mkdirSync(dir, { recursive: true });
fs.writeFileSync(path.join(dir, '01-PLAN.md'),
`---\nphase: ${phase}\nestimate:\n tokens: 100000\n raw_tokens: 50000\n tasks: 3\n confidence: med\nmust_haves:\n---\nx\n`);
fs.writeFileSync(path.join(dir, '01-SUMMARY.md'),
`---\nphase: ${phase}\nactuals:\n tokens: 100000\n tasks: 3\n commits: 5\n---\nx\n`);
}
const out = JSON.parse(runGsdTools('query estimate-calibrate', tmpDir).output);
assert.equal(out.sample_count, 3);
assert.equal(out.factor, 2,
'ratio must be actual/raw (2.0), not actual/calibrated (1.0)');
});
});
// ─── multi-plan pairing (#2632 review BLOCKER) ─────────────────────────────
describe('multi-plan phases pair per plan, not per phase', () => {
// A phase routinely holds several plans (`<NN>-<PP>-PLAN.md`, one per plan —
// docs/reference/planning-artifacts.md). An earlier implementation took the
// first PLAN carrying an estimate and the first SUMMARY carrying actuals
// INDEPENDENTLY, which cross-paired one plan's projection with another plan's
// cost and discarded every later plan. The whole suite passed because its
// helper only ever wrote `01-PLAN.md`.
/** Write one plan/summary pair inside a phase, using the real `<NN>-<PP>` naming. */
const writePlan = (tmpDir, phase, pp, { estTokens, actTokens }) => {
const dir = path.join(tmpDir, '.planning', 'phases', phase);
fs.mkdirSync(dir, { recursive: true });
const nn = phase.slice(0, 2);
if (estTokens !== null) {
fs.writeFileSync(path.join(dir, `${nn}-${pp}-PLAN.md`),
`---\nphase: ${phase}\nplan: ${pp}\nestimate:\n tokens: ${estTokens}\n`
+ ` raw_tokens: ${estTokens}\n tasks: 3\n confidence: low\nmust_haves:\n---\nx\n`);
} else {
fs.writeFileSync(path.join(dir, `${nn}-${pp}-PLAN.md`), `---\nphase: ${phase}\nplan: ${pp}\n---\nx\n`);
}
fs.writeFileSync(path.join(dir, `${nn}-${pp}-SUMMARY.md`),
`---\nphase: ${phase}\nplan: ${pp}\nactuals:\n tokens: ${actTokens}\n tasks: 3\n commits: 4\n---\nx\n`);
};
test('never cross-pairs one plan\'s estimate with another plan\'s actuals', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
// Plan 01 has NO estimate but cheap actuals; plan 02 has both (true 2.5x).
writePlan(tmpDir, '04-multi', '01', { estTokens: null, actTokens: 30000 });
writePlan(tmpDir, '04-multi', '02', { estTokens: 80000, actTokens: 200000 });
runGsdTools('query estimate-calibrate', tmpDir);
const doc = est.parseCalibrationDocument(
fs.readFileSync(path.join(tmpDir, '.planning', 'estimation-calibration.json'), 'utf8'),
);
assert.deepEqual(doc, [{ estimateTokens: 80000, actualTokens: 200000 }],
'plan 02\'s estimate must pair with plan 02\'s actuals — cross-pairing fabricates a sample '
+ 'and throws away the real signal');
});
test('counts every correctly-paired plan in a multi-plan phase', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
// Three plans in ONE phase, each cleanly 2x.
writePlan(tmpDir, '05-wave', '01', { estTokens: 40000, actTokens: 80000 });
writePlan(tmpDir, '05-wave', '02', { estTokens: 50000, actTokens: 100000 });
writePlan(tmpDir, '05-wave', '03', { estTokens: 60000, actTokens: 120000 });
const out = JSON.parse(runGsdTools('query estimate-calibrate', tmpDir).output);
assert.equal(out.sample_count, 3, 'all three plans must contribute — not just the first');
assert.equal(out.factor, 2);
assert.equal(out.applied, true, 'three samples in one phase must reach the minimum');
});
test('a plan with no matching summary contributes nothing', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
const dir = path.join(tmpDir, '.planning', 'phases', '06-partial');
fs.mkdirSync(dir, { recursive: true });
// 06-01 pairs; 06-02 is a plan with no summary (mid-execution).
fs.writeFileSync(path.join(dir, '06-01-PLAN.md'),
'---\nphase: 06-partial\nestimate:\n tokens: 100\n raw_tokens: 100\n tasks: 1\n confidence: low\nmust_haves:\n---\nx\n');
fs.writeFileSync(path.join(dir, '06-01-SUMMARY.md'),
'---\nphase: 06-partial\nactuals:\n tokens: 200\n tasks: 1\n commits: 1\n---\nx\n');
fs.writeFileSync(path.join(dir, '06-02-PLAN.md'),
'---\nphase: 06-partial\nestimate:\n tokens: 999999\n raw_tokens: 999999\n tasks: 1\n confidence: low\nmust_haves:\n---\nx\n');
const out = JSON.parse(runGsdTools('query estimate-calibrate', tmpDir).output);
assert.equal(out.sample_count, 1, 'an in-flight plan must not contribute a half-sample');
});
test('samples accumulate across BOTH plans and phases', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
writePlan(tmpDir, '01-a', '01', { estTokens: 100, actTokens: 200 });
writePlan(tmpDir, '01-a', '02', { estTokens: 100, actTokens: 200 });
writePlan(tmpDir, '02-b', '01', { estTokens: 100, actTokens: 200 });
const out = JSON.parse(runGsdTools('query estimate-calibrate', tmpDir).output);
assert.equal(out.sample_count, 3, 'two plans in phase 1 plus one in phase 2');
assert.equal(out.applied, true);
});
});
// ─── sentinel phases must not skew calibration (#3882, ADR-3473 §8.2) ──────
//
// collectCalibrationSamples enumerates .planning/phases with a raw readdirSync
// and treats every directory as a completed phase — it never routes through
// the phase-directory owner and never applies isSentinelPhaseId, so a sentinel
// directory (milestone 0 or 999 — SENTINEL_RANGES in src/phase-id.cts) whose
// PLAN/SUMMARY pair carries an estimate/actuals block contributes a phantom
// calibration sample.
//
// Governing constraint (.gsd/phase/feat-3882-enumerations/50-test-matrix.md
// row A1): computeCalibration is MEDIAN-based, so a single outlier sample does
// not move a 3-sample factor at all — asserting "the factor is unchanged" against
// exactly one sentinel would pass on the broken code for the wrong reason. The
// real, measured damage is the MIN_CALIBRATION_SAMPLES threshold crossing
// (applied flips false -> true, confidence low -> med on phantom evidence) and,
// once two sentinels are present, actual factor corruption. Every assertion
// below compares the WHOLE computed CalibrationResult object between a
// sentinel-free project and its sentinel-injected twin, reached the same way
// production reaches it (collectCalibrationSamples -> computeCalibration, the
// exact pair cmdEstimateCalibrate calls).
describe('sentinel phases must not skew calibration (#3882)', () => {
// Two genuine phases with DIFFERING ratios (1x and 2x), so A3 can pin each
// one's own unchanged value rather than two indistinguishable duplicates.
const REAL_PHASES = [
['01-alpha', 1000, 1000],
['02-beta', 2000, 4000],
];
// A deliberately extreme, fabricated ratio (50x) — the shape of what a
// sentinel's PLAN/SUMMARY pair would carry.
const SENTINEL_SAMPLE = { estTokens: 1000, actTokens: 50000 };
function buildRealOnlyProject() {
const tmpDir = createTempProject();
for (const [dir, estTokens, actTokens] of REAL_PHASES) {
writePhase(tmpDir, dir, { estTokens, actTokens });
}
return tmpDir;
}
/** Reached the way production reaches it — see cmdEstimateCalibrate above. */
function calibrationResultFor(tmpDir) {
return est.computeCalibration(estimateCli.collectCalibrationSamples(tmpDir));
}
test('A1a: sentinelPhaseDoesNotActivateCalibration', (t) => {
const realOnly = buildRealOnlyProject();
t.after(() => cleanup(realOnly));
const withSentinel = buildRealOnlyProject();
t.after(() => cleanup(withSentinel));
// milestone 999 — reserved icebox sentinel range (SENTINEL_RANGES).
writePhase(withSentinel, '999-icebox', SENTINEL_SAMPLE);
const realOnlyResult = calibrationResultFor(realOnly);
const withSentinelResult = calibrationResultFor(withSentinel);
assert.deepEqual(
withSentinelResult, realOnlyResult,
'a single sentinel phase must not change the computed calibration at all — '
+ `real-only=${JSON.stringify(realOnlyResult)} with-sentinel=${JSON.stringify(withSentinelResult)}`,
);
});
test('A1b: sentinelPhasesDoNotCorruptTheFactor', (t) => {
const realOnly = buildRealOnlyProject();
t.after(() => cleanup(realOnly));
const withTwoSentinels = buildRealOnlyProject();
t.after(() => cleanup(withTwoSentinels));
// milestone 999 and milestone 0 — both reserved sentinel ranges.
writePhase(withTwoSentinels, '999-icebox', SENTINEL_SAMPLE);
writePhase(withTwoSentinels, '0-backlog', SENTINEL_SAMPLE);
const realOnlyResult = calibrationResultFor(realOnly);
const withSentinelsResult = calibrationResultFor(withTwoSentinels);
assert.deepEqual(
withSentinelsResult, realOnlyResult,
'sentinel phases must not corrupt the calibration factor — '
+ `real-only=${JSON.stringify(realOnlyResult)} with-sentinels=${JSON.stringify(withSentinelsResult)}`,
);
});
test('A2: sentinelPhaseContributesNoCalibrationSample', (t) => {
const withTwoSentinels = buildRealOnlyProject();
t.after(() => cleanup(withTwoSentinels));
writePhase(withTwoSentinels, '999-icebox', SENTINEL_SAMPLE);
writePhase(withTwoSentinels, '0-backlog', SENTINEL_SAMPLE);
const samples = estimateCli.collectCalibrationSamples(withTwoSentinels);
const sentinelHits = samples.filter(
(s) => s.estimateTokens === SENTINEL_SAMPLE.estTokens && s.actualTokens === SENTINEL_SAMPLE.actTokens,
);
assert.equal(
sentinelHits.length, 0,
`the sentinel phases' sample must be absent from the returned list; got ${JSON.stringify(samples)}`,
);
});
test('A3: realPhasesStillContribute', (t) => {
const withTwoSentinels = buildRealOnlyProject();
t.after(() => cleanup(withTwoSentinels));
writePhase(withTwoSentinels, '999-icebox', SENTINEL_SAMPLE);
writePhase(withTwoSentinels, '0-backlog', SENTINEL_SAMPLE);
const samples = estimateCli.collectCalibrationSamples(withTwoSentinels);
const realSamples = samples.filter(
(s) => !(s.estimateTokens === SENTINEL_SAMPLE.estTokens && s.actualTokens === SENTINEL_SAMPLE.actTokens),
);
assert.deepEqual(
realSamples.sort((a, b) => a.estimateTokens - b.estimateTokens),
[
{ estimateTokens: 1000, actualTokens: 1000 },
{ estimateTokens: 2000, actualTokens: 4000 },
],
'the two genuine phases must still contribute their own, unchanged samples',
);
});
// A1a-A3 above only assert against the `collectCalibrationSamples` helper's
// return value. Per ADR-3180 Decision 4(b) / epic #3473 B7, a behavioral
// identity test must assert at the CONSUMER's output — the real `query
// estimate-calibrate` CLI JSON and the persisted calibration document it
// writes, not the internal sample array. #3372 named this exact surface
// (`src/estimate-cli.cts`) as one of four candidate sentinel-enumeration
// gaps; #3882 confirmed and fixed it. This closes the CLI-level gap.
test('CLI (#3372, #3882): query estimate-calibrate excludes sentinel phase directories from the emitted sample_count and the persisted document', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
writePhase(tmpDir, '01-alpha', { estTokens: 1000, actTokens: 1000 });
writePhase(tmpDir, '02-beta', { estTokens: 2000, actTokens: 4000 });
// milestone 999 — reserved icebox sentinel range (SENTINEL_RANGES).
writePhase(tmpDir, '999-icebox', { estTokens: 1000, actTokens: 50000 });
const r = runGsdTools('query estimate-calibrate', tmpDir);
assert.ok(r.success, `estimate-calibrate should succeed: ${r.error}`);
const out = JSON.parse(r.output);
assert.equal(
out.sample_count, 2,
`the sentinel phase directory must not be counted in the CLI's emitted sample_count; got ${JSON.stringify(out)}`,
);
const docPath = path.join(tmpDir, '.planning', 'estimation-calibration.json');
const persisted = est.parseCalibrationDocument(fs.readFileSync(docPath, 'utf8'));
assert.equal(
persisted.length, 2,
`the persisted calibration document must not carry the sentinel phase's sample; got ${JSON.stringify(persisted)}`,
);
});
});
// ─── PhasesUnreadableError / estimate_phases_unreadable (#3882, ADR-3473 §8.5,
// review finding #2) ─────────────────────────────────────────────────────
//
// `collectCalibrationSamples` now routes through `listMilestonePhaseDirs`
// (the #3882 fix above), which means an unreadable phases directory is no
// longer output-identical to a genuinely empty one — it surfaces as
// `scope: SCOPE.UNREADABLE`. `cmdEstimateCalibrate` converts that into a
// refusal (`PhasesUnreadableError`, `process.exit(1)`,
// `ERROR_REASON.ESTIMATE_PHASES_UNREADABLE`) instead of silently persisting
// a phantom empty calibration document — a real behavior change (previously
// silent-empty), disclosed and tested here rather than left as an
// undocumented side effect of the routing fix.
//
// H1 drives the real `cmdEstimateCalibrate` IN-PROCESS: `runGsdTools` spawns
// a real subprocess, and neither `fs.readdirSync` monkeypatching nor
// `process.exit` interception crosses that process boundary. The pattern
// below is the two established repo idioms composed, not invented: the
// process.exit-interception + `--json-errors`-mode capture from
// tests/config-get-default.test.cjs, and the `fs.readdirSync` method
// monkeypatch (never chmod, which root/CI bypasses — CLAUDE.md's
// cross-platform IO-failure-injection rule) already used in
// tests/phase-locator.test.cjs for `listAllPhaseDirs`'s own unreadable case.
describe('unreadable phases directory refuses calibration (#3882, ADR-3473 §8.5)', () => {
class _ExitSignal extends Error {
constructor(code, message) {
super(message ?? `process.exit(${code})`);
this.code = code;
}
}
/** Runs cmdEstimateCalibrate in-process with process.exit + stderr(fd 2) captured. */
function runCalibrateExpectError(tmpDir) {
const origExit = process.exit;
const origWriteSync = fs.writeSync;
io.setJsonErrorMode(true);
let exitCount = 0;
let exitCode;
let stderr = '';
fs.writeSync = (fd, ...rest) => {
if (fd !== 2) return origWriteSync.call(fs, fd, ...rest);
const [data, offset = 0, length] = rest;
const chunk = Buffer.isBuffer(data)
? data.subarray(offset, offset + (length ?? data.length - offset)).toString('utf8')
: String(data);
stderr += chunk;
return Buffer.byteLength(chunk);
};
const lastError = () => {
const parts = stderr.split('\n').filter(Boolean);
try { return JSON.parse(parts[parts.length - 1]); } catch { return {}; }
};
process.exit = (code) => {
exitCount++;
exitCode = code;
throw new _ExitSignal(code, lastError().message);
};
try {
estimateCli.cmdEstimateCalibrate(tmpDir, [], false);
} catch (e) {
if (!(e instanceof _ExitSignal)) throw e;
} finally {
process.exit = origExit;
fs.writeSync = origWriteSync;
io.setJsonErrorMode(false);
}
assert.ok(exitCode !== 0 && exitCode !== undefined, 'expected a non-zero exit code');
assert.equal(exitCount, 1, 'error() must fire exactly once (production process.exit terminates)');
return { status: exitCode, ...lastError() };
}
test('H1: unreadable phases directory exits non-zero with estimate_phases_unreadable', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
const phasesDir = path.join(tmpDir, '.planning', 'phases');
const originalReaddirSync = fs.readdirSync;
fs.readdirSync = (...args) => {
if (args[0] === phasesDir) {
const err = new Error('EACCES: permission denied, scandir');
err.code = 'EACCES';
throw err;
}
return originalReaddirSync.apply(fs, args);
};
let result;
try {
result = runCalibrateExpectError(tmpDir);
} finally {
fs.readdirSync = originalReaddirSync;
}
assert.equal(result.status, 1);
assert.equal(result.reason, io.ERROR_REASON.ESTIMATE_PHASES_UNREADABLE);
assert.ok(
!fs.existsSync(path.join(tmpDir, '.planning', 'estimation-calibration.json')),
'refusing the rebuild must not persist a phantom empty calibration document',
);
});
test('H2: a genuinely-empty phases directory still succeeds with empty calibration (boundary the H1 guard must not cross)', (t) => {
// Real CLI as a subprocess (runGsdTools) — this is the boundary case:
// .planning/phases/ EXISTS and is READABLE, but has zero entries. Must
// succeed, not be caught by the unreadable-directory refusal above.
// Overlaps the pre-existing "no phases at all is a clean no-op" test
// above; kept as its own named row because it pins THIS boundary
// specifically (see 50-test-matrix.md row H2), not incidentally.
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
const r = runGsdTools('query estimate-calibrate', tmpDir);
assert.ok(r.success, `a readable, genuinely-empty phases dir must succeed: ${r.error}`);
const out = JSON.parse(r.output);
assert.equal(out.sample_count, 0);
assert.equal(out.applied, false);
});
test('H3: a normal project with real phases is unaffected by the unreadable-dir guard', (t) => {
const tmpDir = createTempProject();
t.after(() => cleanup(tmpDir));
writePhase(tmpDir, '01-a', { estTokens: 100, actTokens: 200 });
writePhase(tmpDir, '02-b', { estTokens: 100, actTokens: 200 });
writePhase(tmpDir, '03-c', { estTokens: 100, actTokens: 200 });
const r = runGsdTools('query estimate-calibrate', tmpDir);
assert.ok(r.success, `a normal readable project must not be affected by the guard: ${r.error}`);
const out = JSON.parse(r.output);
assert.equal(out.sample_count, 3);
assert.equal(out.applied, true);
});
});